RE:NODE

Databases11 min read

Valkey persistence: AOF vs RDB snapshots explained

How Valkey saves data to disk: RDB snapshots, the append-only file, appendfsync choices, what each kind of crash loses, and how to recover a damaged file.

0 readers

Valkey keeps its data in memory and writes a copy to disk in one or both of two ways. An RDB snapshot is the whole dataset written to one compact file at intervals; an append-only file (AOF) logs every write as it happens and is replayed at start-up. With AOF on and appendfsync everysec - the sensible setting - a crash loses at most about a second of writes. With snapshots alone, a crash loses everything since the last snapshot, which by default can be up to an hour. Running both gives you the small loss window of AOF and a compact snapshot file to copy elsewhere. Neither is a backup on its own, because both live on the same disk as the server, and neither turns Valkey into a database that guarantees a committed write survives.

Two mechanisms in one picture#

write acknowledgedappendfsyncsave point reachedwhole datasetClientSET, HSET, INCRValkey memorythe real datasetAOF bufferin the serverForked childcopy-on-writeAppend-only filefsync every secondRDB snapshotdump.rdb
Where a write goes before it is safe

The client gets its OK as soon as the write is in memory. Everything to the right of that happens afterwards, which is the core fact about Valkey persistence: an acknowledged write is not yet on disk. How long it stays only in memory depends on the settings below.

RDB snapshots#

A snapshot is the entire dataset serialised into one file, dump.rdb by default, in the directory named by dir. Valkey takes one automatically when a save point is reached, and on a clean shutdown.

valkey.conf
save 3600 1 300 100 60 10000dbfilename dump.rdbrdbcompression yesrdbchecksum yesstop-writes-on-bgsave-error yes

The save line is pairs of seconds and changes: snapshot after 3600 seconds if at least one key changed, after 300 seconds if at least 100 changed, after 60 seconds if at least 10,000 changed. Those three pairs are the default. save "" disables automatic snapshots altogether.

How a snapshot is taken explains both its strength and its cost. Valkey calls fork(), and the child process writes the dataset to a temporary file and renames it over the old one when finished. The parent keeps serving clients throughout. The fork uses copy-on-write: parent and child share memory pages until the parent modifies one, at which point the kernel copies that page. A quiet dataset costs almost no extra memory to snapshot; a dataset receiving heavy writes during the snapshot can, in the worst case, need nearly twice its size, because every page touched is duplicated.

Strengths of RDB: a single compact file that is easy to copy off the machine; fast loading at start-up, much faster than replaying a log; and almost no cost between snapshots. The weakness: everything written since the last snapshot exists only in memory.

stop-writes-on-bgsave-error yes, the default, is worth understanding before it surprises you. If a background save fails - disk full, fork failed - Valkey refuses further writes with a MISCONF error, on the reasoning that you would rather know persistence has stopped than discover it after a crash. Fix the cause and it clears on the next successful save.

The commands: BGSAVE starts a background snapshot now; SAVE takes one in the foreground and blocks every client until it finishes, so do not use it on a live server; LASTSAVE returns the Unix time of the last successful save.

The append-only file#

AOF records every command that changes data, in the same protocol clients use, and appends it to a log. At start-up the server replays the log and arrives back at the same state.

valkey.conf
appendonly yesappendfilename "appendonly.aof"appenddirname "appendonlydir"appendfsync everysecno-appendfsync-on-rewrite noauto-aof-rewrite-percentage 100auto-aof-rewrite-min-size 64mbaof-use-rdb-preamble yesaof-load-truncated yes

The setting that decides how much you can lose is appendfsync. Writing to a file puts the data in the operating system's cache; only fsync forces it onto the disk.

appendfsyncWhen data reaches diskWorst-case lossCost
alwaysAfter every write commandAlmost nothingEvery write waits for the disk
everysecOnce a second, in a background threadAbout one secondSmall; the default
noWhenever the kernel decidesTens of secondsLeast

everysec is the right choice for almost everyone. On fast NVMe storage always is less painful than its reputation, but it turns every write into a disk round trip, and if you need that guarantee you probably want a database with real transactions instead. Under heavy load, if a background fsync is slow, everysec can briefly fall to about two seconds of exposure.

The log grows forever unless it is compacted. A rewrite builds a new, minimal log from the current dataset - a million INCR commands on one key become one SET - in a forked child, just like a snapshot. auto-aof-rewrite-percentage 100 triggers a rewrite when the log has doubled since the last one, and auto-aof-rewrite-min-size stops tiny logs from being rewritten constantly. BGREWRITEAOF starts one by hand.

Since Redis 7.0, and so in every Valkey release, the AOF is not one file but a directory, appendonlydir, holding a base file, one or more incremental files, and a manifest listing them. With aof-use-rdb-preamble yes (the default), the base file is written in RDB format, which makes rewrites and loading much faster; the incremental files hold commands since then. Treat the directory as a unit - copying one file out of it gives you a fragment, not a dataset.

Using both, and what loads at start-up#

Run both. The AOF gives you the one-second loss window; the RDB gives you a compact file for backups, and a faster restart if you ever need to start from it.

At start-up, if appendonly is yes, Valkey loads the AOF and ignores dump.rdb, because the log is more complete. If AOF is off, it loads the snapshot. One consequence catches people: turning appendonly yes on in the configuration file of a server that has data only in dump.rdb, then restarting, starts from an empty AOF and comes up empty. To enable AOF on a running server with existing data, use CONFIG SET appendonly yes while it runs, wait for the initial rewrite to finish (aof_rewrite_in_progress:0 in INFO persistence), and then make the configuration file match.

The opposite choice is legitimate too. A server that holds nothing but a cache can run with save "" and appendonly no: no fork, no disk writes, no MISCONF errors, and every restart starts empty. That is a reasonable trade when rebuilding the cache is cheap and the database behind it can take the cold start. It is the wrong trade the moment someone puts sessions or a job queue on the same server, which is how most shared instances end up holding data nobody meant to lose. Decide per server, write the decision down, and keep data with different durability needs on different servers if you can.

Loading takes time proportional to the data. While it happens, the server answers commands with a LOADING error; clients should retry it rather than treat it as fatal.

What each kind of stop loses#

The loss depends on how the process stopped, not just on the settings.

EventRDB onlyAOF everysec (with or without RDB)
Clean shutdown (SHUTDOWN, SIGTERM)Nothing, if save points are configuredNothing
Process killed (SIGKILL, out of memory)Everything since the last snapshotAbout the last second
Operating system crash or power lossEverything since the last snapshotAbout the last second, subject to the disk honouring fsync
Disk lost or corruptedEverything, without an off-machine copyEverything, without an off-machine copy
FLUSHALL run by mistakeRecoverable from an older snapshot copyThe FLUSHALL is in the log too - see below

That last row deserves a note. A FLUSHALL is a write like any other, so it goes into the AOF. If you catch it before the next rewrite, you can stop the server, remove the trailing FLUSHALL from the newest incremental file in appendonlydir, and start again. After a rewrite the data is gone from the log, and only a copy taken earlier will help. This is why persistence and backup are different subjects.

Recovering a damaged file#

Files get truncated - a crash in the middle of a write leaves half a command at the end of the AOF. With aof-load-truncated yes (the default), Valkey loads everything up to the broken last command, logs a warning, and starts. If the damage is in the middle of the file, it refuses to start, and you repair it by hand:

bash
$ valkey-check-aof appendonlydir/appendonly.aof.manifest$ valkey-check-aof --fix appendonlydir/appendonly.aof.manifest$ valkey-check-rdb dump.rdb

valkey-check-aof --fix truncates the log at the first invalid command, which loses everything after that point - copy the directory before running it. valkey-check-rdb reports whether a snapshot is readable. If the AOF is beyond repair and the snapshot is fine, you can start from the snapshot by disabling AOF, starting, then re-enabling it at runtime as described above.

Most recoveries, though, do not involve damaged files at all. They involve a server that started, loaded its AOF, and has the data it should. That is the normal case, and it is the point of the exercise.

Watching persistence#

Persistence fails quietly until a restart reveals it, so check it the way you check anything else. INFO persistence has everything:

bash
$ valkey-cli -h db.example.net -p 6380 --askpass INFO persistenceloading:0rdb_changes_since_last_save:184rdb_bgsave_in_progress:0rdb_last_save_time:1791449103rdb_last_bgsave_status:okaof_enabled:1aof_rewrite_in_progress:0aof_last_bgrewrite_status:okaof_last_write_status:ok

The fields worth alerting on: rdb_last_bgsave_status and aof_last_bgrewrite_status should be ok; aof_last_write_status should be ok; rdb_last_save_time should be recent if data is changing. rdb_changes_since_last_save climbing without a save means snapshots have stopped. In the server log, a line saying the fork failed or that a background save terminated with an error is the early warning that memory headroom has run out.

Persistence is not a backup#

Both files sit on the same disk as the server. A deleted server, a failed disk, a mistaken FLUSHALL that has since been rewritten into the AOF, or an application bug that overwrote half the keys all leave you with persistent copies of the wrong state. A backup is a copy somewhere else, from a time before the problem.

For Valkey that is straightforward, because a snapshot is one file:

  1. Trigger a fresh snapshot with BGSAVE, and wait for LASTSAVE to change.
  2. Copy dump.rdb off the machine - a different server, object storage, anywhere that does not share the failure.
  3. Keep several generations, so a problem noticed a week late is still recoverable.

Where the server permits it, valkey-cli --rdb backup.rdb fetches a snapshot over the network without touching the server's disk at all, which is convenient from a backup host. And restore one occasionally into a scratch instance - testing a restore before you need it explains why that step is the one that matters. Database backups and restores covers the wider schedule.

Then decide whether the data needs a backup at all. A cache does not: it can be rebuilt from the database. Sessions usually do not: losing them logs people out. Queues of pending jobs, counters you bill on, and anything with no other copy do.

Persistence on a RE:NODE Valkey server#

Valkey servers on the database line save to disk with both AOF and snapshots, so a restart - whether you press the button or the server is stopped at its memory limit and restarted clean, which is how the platform handles an out-of-memory condition - replays the log and comes back with the data it had, minus at most the last moment of writes. That memory-limit behaviour is also the practical reason to keep your dataset within the usable figure each plan lists, rather than the headline number: the snapshot and rewrite children need room. Backup slots come with every Valkey plan from 512 MB up; the 256 MB plan has none, so on that plan, if the data matters, take your own off-machine copies as above. Backups in the panel are restored with a button.

FAQ#

Should I use AOF or RDB?

Both, if the data matters at all. AOF with appendfsync everysec limits a crash to about a second of lost writes; RDB gives you a compact file to copy off the machine and a faster restart. If the server holds only a cache, snapshots alone, or no persistence, are reasonable.

Does Valkey lose data when it restarts?

Not on a clean restart with persistence enabled: it saves on shutdown and reloads on start. A forced stop loses whatever had not reached disk - about a second with AOF everysec, or everything since the last snapshot without AOF.

Why does my server refuse writes with a MISCONF error?

A background save failed and stop-writes-on-bgsave-error is on, so the server stops accepting writes until persistence works again. The cause is almost always a full disk or a fork that failed for lack of memory. The server log says which.

Is appendfsync always worth the cost?

Rarely. It makes every write wait for the disk to confirm, which cuts throughput sharply, and the data it protects is the last second before a crash. If losing that second is unacceptable, the data probably belongs in a transactional database such as PostgreSQL, with Valkey in front of it.

How big will the AOF get?

Roughly proportional to the dataset after each rewrite, growing between rewrites by every write made. With the default settings it is rewritten when it doubles, so expect it to stay between one and two times the size of the data, plus the snapshot file beside it. Leave disk space for both and for the rewrite's temporary copy.


Comments

Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.

0/2000