RE:NODE

Databases10 min read

Valkey memory: maxmemory, eviction policies and big keys

Size and tune Valkey memory: what maxmemory counts, choosing an eviction policy, finding big keys, compact encodings and reading the fragmentation ratio.

0 readers

Everything Valkey stores lives in memory, so memory is the limit you will actually meet. Two settings decide what happens when you meet it: maxmemory, the ceiling on the dataset, and maxmemory-policy, what to do at the ceiling. For a pure cache, set the ceiling and use allkeys-lru or allkeys-lfu, and Valkey quietly discards the least useful keys. For sessions, queues or anything you cannot lose, use noeviction, which refuses new writes with an OOM error instead of deleting data - and make sure you notice before it happens. The volatile-* policies sit in between and carry one trap: they only evict keys that have a TTL, so if none do, they behave exactly like noeviction. After the policy, the job is measurement: finding the keys that use the memory, keeping small structures in their compact encodings, and leaving enough headroom for the parts of memory that are not your data.

maxmemory: what it limits and what it does not#

valkey.conf
maxmemory 400mbmaxmemory-policy allkeys-lrumaxmemory-samples 5

maxmemory caps the memory used by the dataset, as reported in used_memory. When a write would take it over, the eviction policy runs. Its default is 0, meaning no limit on a 64-bit system - Valkey will grow until the operating system or container stops it, which is a much worse failure than any eviction policy. Always set it.

What it does not include matters as much as what it does:

  • Replication and AOF buffers are not counted towards the limit when deciding whether to evict.
  • Client buffers - query buffers, and output buffers for slow readers and pub/sub subscribers - are real memory that lives outside your keys.
  • Fragmentation - the gap between what the allocator holds and what the data uses - is outside used_memory altogether.
  • The fork for snapshots and AOF rewrites can need a lot of extra memory through copy-on-write while it runs, as described in Valkey persistence.

So the process's real footprint, used_memory_rss, is always above maxmemory, sometimes by a lot. A common starting point is maxmemory at 60 to 80 per cent of the memory available, lower if persistence is on and the write rate is high. CONFIG SET maxmemory 400mb changes it at runtime, where the server permits configuration changes.

Eviction policies#

PolicyEvictsRight for
noevictionNothing; writes fail with OOMSessions, queues, data you must not lose (the default)
allkeys-lruLeast recently used key, any keyA general cache
allkeys-lfuLeast frequently used key, any keyA cache with a stable hot set
allkeys-randomA random keyUniform access, rarely the best choice
volatile-lruLeast recently used, among keys with a TTLCache and durable data on one server
volatile-lfuLeast frequently used, among keys with a TTLThe same, frequency-based
volatile-randomA random key with a TTLRarely
volatile-ttlThe key with the nearest expiryWhen TTL reflects importance

How to choose:

  1. Is everything on this server disposable? Then allkeys-lru. Switch to allkeys-lfu if a small set of keys is read constantly and a scan or batch job tends to push them out - LFU remembers that a key is popular, where LRU only remembers that it was touched recently.
  2. Is anything on this server not disposable? Then noeviction, and monitor memory, because a full server refuses writes.
  3. Is it a mixture? The honest answer is two servers. The volatile-* policies let a mixed server evict only cache keys (which have TTLs) while keeping durable ones (which do not), but they depend on every cache write setting a TTL, forever, in every code path.

How eviction actually works#

Valkey does not keep a perfect list of keys ordered by last use; that would cost memory on every key. Instead, when it needs to evict, it samples a few keys and evicts the best candidate among them, keeping a small pool of good candidates between rounds. maxmemory-samples 5 is the default; raising it to 10 gets closer to true LRU at a little more CPU per eviction. In practice the approximation is good enough that you will not notice.

LFU uses a small logarithmic counter per key that rises with access and decays over time. lfu-log-factor (default 10) controls how quickly the counter saturates, and lfu-decay-time (default 1, in minutes) controls how fast it decays when the key is idle. OBJECT FREQ key shows a key's counter when an LFU policy is active; OBJECT IDLETIME key shows seconds since last access under the LRU policies.

Eviction happens when a write needs memory, inside that write. A large burst of writes into a full server therefore does eviction work on the write path, which shows up as latency. lazyfree-lazy-eviction yes hands the actual freeing of evicted values to a background thread, which helps when evicted values are large.

The counters to watch in INFO stats are evicted_keys, which climbs whenever the policy removes something, and expired_keys for keys removed by TTL. On a cache, some eviction is normal. A steadily rising rate means the working set no longer fits, and the hit ratio is falling with it.

What actually uses the memory#

Each key costs more than its name and value. There is the entry in the keyspace's hash table, the key string, an object header for the value, an expiry entry if it has a TTL, and allocator rounding on every allocation. For tiny values the overhead can exceed the data - a million keys holding a 4-byte counter use far more than 4 MB. Valkey 8.1 replaced the main hash table with a more compact design, which noticeably reduced per-key overhead, but small keys are still relatively expensive.

The big savings come from compact encodings. Small aggregates are stored as a single packed block rather than as a full hash table or skip list:

TypeCompact encodingStays compact while
HashlistpackAt most hash-max-listpack-entries (128) fields, values up to hash-max-listpack-value (64) bytes
Sorted setlistpackAt most zset-max-listpack-entries (128) members, up to 64 bytes each
Set of integersintsetAt most set-max-intset-entries (512) members
Small setlistpackAt most set-max-listpack-entries (128) members
Listquicklist of listpacksNode size set by list-max-listpack-size

OBJECT ENCODING key tells you which encoding a key is in. A hash that grows past the thresholds converts to a full hash table and does not convert back. Two practical consequences:

  • Group small values into hashes. One hash with a hundred small fields is far cheaper than a hundred separate keys, because it shares one key's overhead and sits in a listpack. Storing a user's profile as HSET user:42 name ... plan ... beats SET user:42:name and SET user:42:plan.
  • Keep fields small if you can. A single field over 64 bytes pushes the whole hash out of listpack encoding. Raising the thresholds trades a little CPU (listpacks are scanned linearly) for memory; values up to a few hundred are reasonable if memory is the constraint.

Finding big keys#

One oversized key - a list nobody trims, a set of every user who ever logged in, a cached blob of a whole catalogue - often explains a memory problem on its own. The CLI can find them without blocking the server, because it walks the keyspace with SCAN:

bash
$ valkey-cli -h db.example.net -p 6380 --askpass --bigkeys$ valkey-cli -h db.example.net -p 6380 --askpass --memkeys$ valkey-cli -h db.example.net -p 6380 --askpass MEMORY USAGE session:9f2c41 SAMPLES 0

--bigkeys reports the largest key of each type by element count; --memkeys reports by memory. MEMORY USAGE gives the bytes for one key, including overhead; SAMPLES 0 makes it measure every element of an aggregate rather than estimate from five.

Big keys hurt beyond their size. Reading one blocks the server while the reply is built and sent. Deleting one with DEL blocks while its memory is freed - use UNLINK, which frees in the background. And a big key can never be evicted partially: under memory pressure, it either all stays or all goes. Cap lists with LTRIM, give sets of identifiers a TTL or a rotation scheme, and split catalogue-sized values into per-item keys.

Fragmentation#

INFO memory reports both what the data uses and what the process holds:

bash
$ valkey-cli -h db.example.net -p 6380 --askpass INFO memory \    | grep -E '^(used_memory_human|used_memory_rss_human|mem_fragmentation_ratio|maxmemory_human|maxmemory_policy):'

mem_fragmentation_ratio is RSS divided by used_memory. Around 1.0 to 1.5 is healthy. Well above that means the allocator holds memory in pages that are only partly used, typically after many keys of varying sizes were deleted or expired. Below 1.0 means part of the process has been swapped out, which on a database that promises memory-speed answers is the worse problem. On a nearly empty server the ratio is meaningless - a few megabytes of baseline RSS over a tiny dataset produces a large number that signifies nothing.

Remedies, in order: MEMORY DOCTOR gives a plain-language diagnosis; activedefrag yes lets the server move values into fuller pages in the background, if it was built with jemalloc (the default on Linux), within the limits set by active-defrag-threshold-lower and active-defrag-ignore-bytes; MEMORY PURGE asks the allocator to release free pages. A restart reclaims everything, and with persistence on it is cheap, but it is a blunt tool.

Memory that is not your data#

Leave headroom above maxmemory for the things it does not count:

  • Client output buffers. A slow consumer, or a pub/sub subscriber that stops reading, accumulates replies in server memory. client-output-buffer-limit caps them - pubsub 32mb 8mb 60 by default, meaning a subscriber is disconnected at 32 MB, or after 60 seconds above 8 MB.
  • Copy-on-write during persistence. Every page written during a snapshot or AOF rewrite is duplicated for the child's sake.
  • Connections. Each client costs memory for its buffers. Thousands of idle connections from a leaky pool add up; CLIENT LIST and INFO clients show them.
  • Lua scripts and the scripting engine, small but non-zero.

Linux swap and the OOM killer covers what the operating system does when the total exceeds what is available, and it is never what you want.

Sizing: a worked example#

Estimation by arithmetic is unreliable; estimation by sampling is easy. Suppose an application will store 200,000 sessions. Load a realistic few thousand into a test server, then measure:

sample_size.py
import statistics, redisr = redis.Redis.from_url("redis://:pass@db.example.net:6380/0")sizes = []for key in r.scan_iter(match="session:*", count=1000):    sizes.append(r.memory_usage(key, samples=0))    if len(sizes) >= 2000:        breakprint(f"keys sampled: {len(sizes)}, mean bytes: {statistics.mean(sizes):.0f}")

If the mean comes out at around 600 bytes, 200,000 sessions need roughly 120 MB of dataset. Add a margin for growth and for the memory that is not your data - half again is a reasonable rule with persistence on - and you are looking at about 180 MB, which fits a server with around 200 MB usable, with little to spare. Measure again after launch with INFO memory, because real sessions are always bigger than test ones.

Sessions in Valkey covers keeping session values small, which is the cheapest memory you will ever save.

Memory on a RE:NODE Valkey server#

Valkey plans are sold by memory - 256 MB, 512 MB, 1 GB, 2 GB and 4 GB - and each lists its usable figure for data at roughly 80 per cent of that: about 200 MB, 400 MB, 800 MB, 1.6 GB and 3.2 GB. Size against the usable number. If the process itself reaches the plan's memory limit, the platform stops the container and restarts it clean rather than letting it swap; with AOF and snapshots on, the data comes back from disk, but connected clients see a short outage, and the crash watcher counts it. Check maxmemory and maxmemory-policy with INFO memory when the server is new, so you know which behaviour you will get when it fills, and move up a plan when evicted_keys or memory use says the working set has outgrown it.

FAQ#

What is the default eviction policy?

noeviction. When maxmemory is reached, write commands fail with an OOM error and nothing is deleted. That is safe for data you must keep and unhelpful for a cache, so set allkeys-lru or allkeys-lfu explicitly on a cache server.

LRU or LFU for a cache?

LRU is the safe default. LFU is better when a small set of keys is used constantly and occasional bulk reads - a report, a crawler - would otherwise push the hot keys out. If you cannot tell, use LRU and compare hit ratios after switching.

Why is the process using more memory than maxmemory?

Because maxmemory limits the dataset, not the process. Client buffers, fragmentation, and copy-on-write during snapshots and AOF rewrites all sit outside it. Leave a margin of a fifth to a half of the limit, depending on how write-heavy the server is.

How do I find which keys use the most memory?

Run valkey-cli --memkeys or --bigkeys, which scan the keyspace without blocking, and MEMORY USAGE key SAMPLES 0 on suspects. Look for unbounded lists and sets first; they are usually the culprit.

Will Valkey give memory back after I delete keys?

Partly. Freed memory is reused for new data straight away, but the allocator does not always return it to the operating system, so RSS can stay high and the fragmentation ratio rises. activedefrag and MEMORY PURGE help; a restart with persistence on resets it completely.


Comments

Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.

0/2000