The pattern almost every application should use with Valkey is cache-aside: read from the cache, and on a miss read from the database, store the result with a TTL, and return it. On writes, update the database and delete the cached key rather than updating it. Give every key an expiry, put a version in every key name, and protect expensive keys against stampedes - the moment a popular key expires and a hundred requests all rebuild it at once. Those five habits cover nearly all the ways a cache goes wrong: stale data that never refreshes, data that changes shape and breaks readers, and a cache that protects the database right up until the moment it is needed most.
This post walks through each of them with code, then covers how to tell whether the cache is earning its keep.
Cache-aside: the default pattern#
In cache-aside, the application owns the logic. The cache knows nothing about the database; the code checks one, falls back to the other, and fills the gap.
async function getProduct(id) { const key = `v2:product:${id}`; const hit = await valkey.get(key); if (hit !== null) return JSON.parse(hit); const product = await db.products.findById(id); // the expensive part if (product) { await valkey.set(key, JSON.stringify(product), "EX", 600); } return product;}Its virtues are the reason it is the default. If Valkey is down, the application still works, just slower, as long as the cache read is wrapped so that an error counts as a miss. Only data that someone actually asked for is cached, so memory goes to the hot set. And the cache can be flushed at any time without losing anything, because the database is the source of truth.
What to cache, in rough order of payoff:
- Expensive queries that cannot be indexed away - aggregations, reports, search results, anything with a
GROUP BYover a lot of rows. - Calls to other people's APIs, which are slow, rate-limited and sometimes billed per request.
- Rendered fragments - a page section, a serialised API response - where building the output costs more than fetching the data.
- Hot records read far more than they are written - a product, a configuration row, a user's permissions.
What not to cache: queries that are already fast with the right index. Caching a slow query that needs an index hides the problem until the cache is cold. Reading EXPLAIN ANALYZE is cheaper than a cache layer, and Redis: when you need it makes the wider case for doing without.
The other patterns, and why they are rarer#
| Pattern | How it works | When it fits |
|---|---|---|
| Cache-aside | App reads cache, falls back to database, fills cache | Almost always |
| Read-through | A cache library loads from the database on a miss | Same as cache-aside, wrapped in a library |
| Write-through | Every write goes to cache and database together | Data read immediately after it is written |
| Write-behind | Writes go to the cache, flushed to the database later | Counters and metrics where loss is acceptable |
| Refresh-ahead | Hot keys are rebuilt before they expire | A small set of very hot, expensive keys |
Write-through sounds attractive - the cache is never stale - but it caches everything written whether or not anyone reads it, and it still needs a TTL in case a write to one side fails. Write-behind makes Valkey the system of record for a while, which is fine for page-view counters flushed every minute and unacceptable for orders. Refresh-ahead appears again below as a stampede defence.
Choosing TTLs#
The TTL is the upper bound on how stale a cached value can be. Choose it by asking how long a wrong answer is tolerable, not by how long the value is likely to stay unchanged.
| Data | Typical TTL | Reasoning |
|---|---|---|
| Rendered page fragment for anonymous users | 30-300 seconds | Short staleness is invisible |
| Product details, article bodies | 5-60 minutes, plus invalidation on write | Invalidation handles edits; TTL is the backstop |
| Report or dashboard aggregate | 5-15 minutes | Users accept "as of a few minutes ago" |
| Third-party API response | As long as their terms and freshness allow | Saves money and rate limit |
| Permissions and roles | 1-5 minutes, plus invalidation | A revoked permission must not linger |
| Configuration and lookup tables | 1 hour or more | Rarely changes; invalidate on deploy |
Two refinements make TTLs behave better under load:
- Add jitter. If a batch job fills ten thousand keys at once with
EX 3600, they all expire in the same second an hour later, and the database takes ten thousand misses at once. Add a random few per cent:EX 3600 + random(0, 300). - Never cache without one. A key with no expiry is a promise that invalidation will always work. It will not - a missed code path, a manual database fix, a failed delete - and the key will serve the wrong value until somebody flushes the cache by hand. Even data you invalidate carefully deserves a TTL as a backstop.
Watch for one trap: a plain SET without EX on an existing key removes its expiry. Every write path must pass the TTL, or use KEEPTTL.
Invalidation: delete, do not update#
When the underlying data changes, the cached copy must go. The robust approach is to update the database first, then delete the key, and let the next read rebuild it.
async function updateProduct(id, changes) { await db.products.update(id, changes); await valkey.unlink(`v2:product:${id}`);}Why delete rather than write the new value into the cache? Because two concurrent writers can finish in the opposite order in the database and in the cache, leaving the cache permanently holding the older value. Deleting has no ordering problem: whichever delete arrives last, the next read loads the current row.
There is still one narrow race. A reader misses, loads the old row, gets delayed; a writer updates the row and deletes the key; the delayed reader then writes the old row into the cache. The TTL bounds how long that survives, which is one more reason every key needs one. Where that window is unacceptable, delete the key a second time a short delay after the write, or keep the value out of the cache entirely.
Invalidating groups of keys is where people reach for KEYS product:* and take the server down. Better options:
- Version prefixes. Put a version in the key (
v2:product:...). Changing the cached shape is a deploy that bumps the version; the old keys simply expire. - Generation counters. Keep a counter per group (
gen:catalogue) and include its value in the key names. Invalidating the whole group is oneINCR; the old keys are never read again and age out on their TTLs. - Tag sets. Add each cache key to a set per tag (
SADD tag:category:7 v2:product:42). To invalidate the tag, read the set's members,UNLINKthem, and delete the set. Give the set a TTL as well.
Use UNLINK rather than DEL for anything large: it removes the key immediately and frees its memory in the background, so deleting a big value does not stall other clients.
Stampedes: when the cache fails exactly when it matters#
A stampede, also called a dogpile or thundering herd, happens when a popular key expires and many requests miss at the same moment. Each one runs the expensive query, each one writes the result, and the database takes the full load of every request in that window - often at peak traffic, which is exactly when the key was popular. A cold start after a cache flush or restart is the same thing across every key at once.
Three defences, from simplest to most robust.
A rebuild lock. The first request to miss takes a short lock with SET ... NX PX, rebuilds, and releases it. Others wait briefly and re-read the cache, or serve a stale copy if one exists.
import json, time, uuidUNLOCK = valkey.register_script("""if redis.call('GET', KEYS[1]) == ARGV[1] then return redis.call('DEL', KEYS[1])endreturn 0""")def get_report(team_id): key, lock = f"v1:report:{team_id}", f"lock:v1:report:{team_id}" for _ in range(50): hit = valkey.get(key) if hit is not None: return json.loads(hit) token = str(uuid.uuid4()) if valkey.set(lock, token, nx=True, px=10_000): try: report = build_report(team_id) # the expensive part valkey.set(key, json.dumps(report), ex=900) return report finally: UNLOCK(keys=[lock], args=[token]) time.sleep(0.1) # someone else is rebuilding return build_report(team_id) # give up waitingThe lock has an expiry (PX 10000) so a crashed worker cannot hold it forever, and it is released with a script that checks the token, so a slow worker whose lock already expired cannot delete a lock someone else now holds.
Stale-while-revalidate. Store the value with a long hard TTL, and keep a soft expiry time inside the value. A reader who finds the soft time passed still returns the stale value at once, and triggers a rebuild in the background - guarded by the same lock so only one runs. Users never wait, and the database sees one rebuild per key per interval.
Probabilistic early refresh. Each reader, on a hit, rolls a die weighted by how close the key is to expiry and how long it takes to rebuild, and occasionally refreshes early. The usual formula (from the "XFetch" paper) rebuilds when now - rebuild_time * beta * ln(random()) >= expiry, with beta around 1. The effect is that one request, statistically, refreshes the key shortly before it expires, with no lock at all. It suits a small set of very hot keys.
Negative caching and errors#
If a lookup finds nothing - a product ID that does not exist, a user with no avatar - cache that too. Otherwise every request for a missing thing goes to the database, and a scraper walking through IDs, or an attacker doing it deliberately, bypasses the cache entirely. Store a sentinel such as an empty string or {"missing":true} with a shorter TTL than real values, so a newly created record appears quickly.
Errors are different. Do not cache a failed call to a third-party API as if it were an answer - or if you must, to stop hammering a service that is down, cache it for seconds, not minutes, and distinguish it clearly from a real empty result.
Values: size and serialisation#
How you store the value affects memory and speed more than people expect.
- JSON is readable and universal, and the right default. MessagePack or a similar binary format is smaller and faster to parse, at the cost of readability in
valkey-cli. - Compress large values. A 200 KB JSON document compresses to a fraction of that with gzip or zstd, and the CPU time is usually smaller than the network time saved. Below a few kilobytes, it is not worth it.
- Avoid huge values. A multi-megabyte value takes milliseconds to send, and while the server sends it, it is not serving anyone else. Split it, or cache the parts that are actually read.
- Use a hash for records that are read partially.
HGET product:42 pricereads one field without transferring the whole object; small hashes are also stored compactly.
Valkey memory and eviction policies covers what happens when the cache fills, and why the eviction policy has to match the fact that this data is disposable.
Measuring whether the cache works#
A cache that is not measured is a guess. Valkey keeps the basic counters itself:
$ valkey-cli -h db.example.net -p 6380 --askpass INFO stats \ | grep -E 'keyspace_hits|keyspace_misses|evicted_keys|expired_keys'keyspace_hits:1843920keyspace_misses:211347evicted_keys:0expired_keys:96112The hit ratio is hits divided by hits plus misses - here about 90 per cent. Server-wide numbers mix every kind of key together, so also count hits and misses per key prefix in the application, where you can see that product pages hit 98 per cent and search results 40 per cent. A low ratio means TTLs that are too short, keys that are too specific (a timestamp or a session ID in a key meant to be shared), or data that is simply not reused. Rising evicted_keys means the cache is too small for its working set. Monitoring that tells you something covers turning these into alerts.
Measure the database too. The point of the cache is the load it removes; if query rate and latency do not drop when the cache is added, it is not doing its job.
Caching with Valkey on RE:NODE#
The Valkey line suits exactly this job. The smallest plan has 256 MB of memory, of which the plan lists roughly 200 MB as usable - plenty for a cache in front of one application - and plans go up to 4 GB. Each server is password-protected and reached on the host and port shown in the panel; there is no proxy slot, so the application connects directly. Because each server saves to disk with AOF and snapshots, a restart brings the cache back warm rather than empty, which removes the cold-start stampede from planned restarts. Treat that as a convenience rather than a guarantee: the cache is still disposable data, and the code must work when it is empty.
FAQ#
Should I update the cache or delete it when data changes?
Delete it. Updating can leave the older value in the cache when two writes finish in different orders, while deleting always leads the next read to load the current data. Keep a TTL on every key as a backstop for the cases invalidation misses.
What TTL should I use?
The longest staleness your users can tolerate for that data, not a guess at how often it changes. Seconds for page fragments, minutes for product data and aggregates, longer for configuration - and always with invalidation on write for anything users edit. Add a little random jitter so keys filled together do not expire together.
How do I clear all cached keys for one feature?
Do not use KEYS. Put a version or generation number in the key names and increment it, so old keys are never read again and expire on their own. If you need immediate removal, keep a set of the keys per tag and UNLINK its members.
What is a cache stampede?
Many requests missing the same expired key at the same moment, all running the expensive rebuild together. Prevent it with a short rebuild lock using SET NX PX, by serving stale values while one request refreshes, or by refreshing hot keys slightly before they expire.
What hit ratio should a cache have?
It depends on the data, but below about 80 per cent for a general-purpose cache is worth investigating. Look at hit ratios per key prefix in the application; a server-wide average hides the one key type that never hits.




Comments
Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.