RE:NODE

Databases12 min read

Valkey data types explained: which one to use and when

Strings, hashes, lists, sets, sorted sets, streams, bitmaps and HyperLogLog in Valkey: commands, memory cost, encodings and the right type for each job.

0 readers

Valkey is not a key-value store in the sense of "store a blob under a name". Each key holds a typed structure, and the commands operate on the structure in place: increment a field of a hash, push to the end of a list, add a member to a set, read the top ten of a sorted set. Choosing the right type is most of the design work, because it decides how many round trips a feature takes, whether an update is atomic, and how much memory the data uses.

The short version: strings for single values and counters, hashes for objects, lists for queues and recent-items, sets for membership and uniqueness, sorted sets for anything ranked or ordered by time, streams for event logs. Bitmaps, HyperLogLog and geospatial indexes cover three specialised jobs very cheaply. The rest of this post goes through each with the commands that matter, the memory behaviour, and the mistakes that turn a fast server into a slow one.

How Valkey stores a key#

Every key has a name, one type and an optional expiry. Asking for the wrong type is an error, not a conversion:

code
SET page:views 10LPUSH page:views 1(error) WRONGTYPE Operation against a key holding the wrong kind of valueTYPE page:viewsstring

Expiry belongs to the whole key. Until recently you could not expire one field of a hash or one member of a set; Valkey 9.0 added per-field expiry for hashes (covered below), but other types still expire as a unit. That is a design constraint worth knowing early: if items need individual lifetimes, they usually need individual keys or a sorted set scored by expiry time.

Key names are free-form strings. The convention that makes a keyspace manageable is colon-separated segments from general to specific - app:user:4182:profile - with a version segment where the shape of the value might change. Long key names cost memory on every key, so app:u:4182 is defensible when you have tens of millions; for most applications readability wins.

Two commands tell you what a key really is and what it costs:

code
OBJECT ENCODING user:4182"listpack"MEMORY USAGE user:4182(integer) 120

The encoding matters because small collections are stored in a compact form - a listpack or an intset - that uses a fraction of the memory of the general form. Valkey converts to the general form automatically when a collection crosses a size threshold. More on that in the memory section.

Strings and counters#

The basic type, despite the name, is a binary-safe byte sequence up to 512 MB. It holds text, JSON, serialised objects, images if you insist, and integers or floats that the server can do arithmetic on.

CommandDoes
SET key value EX 300 NXSet with a 300-second TTL, only if absent
GET key, MGET k1 k2Read one or many
INCR, INCRBY, INCRBYFLOATAtomic arithmetic
GETDEL, GETEXRead and delete, or read and change expiry
APPEND, STRLEN, GETRANGEByte-level operations

SET with NX and EX is the basis of a simple lock and of idempotency guards: "only the first request with this id proceeds". INCR is the basis of counters and rate limits (rate limiting with Valkey builds on it). For caching, a string holding serialised JSON is the default choice and usually the right one; Valkey caching patterns goes into TTLs and invalidation.

The common mistake is storing a large JSON document as a string and then rewriting all of it to change one field. If you update parts of an object independently, it wants to be a hash.

Hashes for objects#

A hash is a map of fields to string values under one key - the natural shape for a user record, a session, a product, a configuration block.

code
HSET user:4182 name "Nino" plan "pro" logins 17HGET user:4182 planHINCRBY user:4182 logins 1HMGET user:4182 name planHGETALL user:4182HDEL user:4182 plan

Each field can be read and updated independently, HINCRBY is atomic per field, and small hashes are stored very compactly. Many small hashes are one of the most memory-efficient ways to hold object data in Valkey.

Hash values are flat strings. Nested data either goes in as serialised JSON in one field, or gets flattened into field names such as address.city. If you need to query inside nested documents, you are describing a document database, not a hash - see which database to use.

Valkey 9.0 added expiry for individual hash fields: commands such as HSETEX, HGETEX, HEXPIRE, HTTL and HPERSIST set, read and clear a TTL on one field. That makes a single hash usable for things like a set of per-device tokens that each expire on their own. Check the server version with INFO server (Valkey reports a valkey_version field alongside a redis_version kept for client compatibility) and the exact syntax in the command reference before relying on it - older servers return an unknown-command error.

Lists for queues and recent items#

A list is an ordered sequence of strings with fast operations at both ends.

code
LPUSH events:recent "login 4182"LTRIM events:recent 0 99LRANGE events:recent 0 9RPUSH jobs '{"task":"email","user":4182}'BLMOVE jobs jobs:processing LEFT RIGHT 5

Pushing and popping at either end is constant time. Reading by index from the middle of a long list is not: LINDEX and LSET walk the list. The two common uses fit that shape:

  • A capped recent-items list. LPUSH then LTRIM 0 99 keeps the newest hundred items, in one round trip if pipelined. Activity feeds, last-seen logs, recent searches.
  • A simple queue. Producers RPUSH, consumers block on BLPOP or, better, BLMOVE, which atomically moves the item to a processing list so a crashed consumer does not lose it. This is the foundation that libraries like RQ build on; for anything serious use a library - job queues with Valkey covers them.

Do not use a list for membership tests ("is 4182 in this list?") - that is a linear scan. That is what sets are for.

Sets and sorted sets#

A set is an unordered collection of unique strings.

code
SADD online:users 4182 5530 7781SISMEMBER online:users 4182SCARD online:usersSINTER tags:valkey tags:tutorialSRANDMEMBER giveaway:entrants 3

Adding and membership tests are constant time, and duplicates are impossible by construction. Uses: online users, tags, "has this user already voted", deduplicating a crawl, unique visitors when exact counts matter and the numbers are small. SINTER, SUNION and SDIFF compute set algebra on the server, which beats fetching two lists into the application.

A sorted set gives each unique member a floating-point score and keeps members ordered by it. It is the most versatile type Valkey has.

code
ZADD leaderboard 3120 "nino" 2875 "giorgi" 4010 "ana"ZINCRBY leaderboard 50 "giorgi"ZRANGE leaderboard 0 9 REV WITHSCORESZREVRANK leaderboard "nino"ZRANGE delayed 0 1791468000000 BYSCORE LIMIT 0 100ZREMRANGEBYSCORE sessions:active 0 1791460000000

Adds, removes and rank lookups are logarithmic; range reads are logarithmic plus the size of the range. Sorted sets handle:

  • Leaderboards and rankings, where ZINCRBY and ZREVRANK are exactly the operations needed.
  • Time-ordered indexes, with a timestamp as the score: "jobs due before now", "sessions inactive since an hour ago", "posts from the last day".
  • Delayed jobs and scheduling, the same thing seen from the other side.
  • Sliding-window rate limits, recording each request's time.
  • Secondary indexes you maintain yourself: "products by price" as a sorted set of product ids scored by price.

Keep the large commands in mind. SMEMBERS and ZRANGE 0 -1 on a set of a million members return a million members in one blocking call; use SSCAN and ZSCAN, or bounded ranges.

Streams, bitmaps, HyperLogLog and geospatial#

Streams are an append-only log with ids, consumer groups and acknowledgements - the durable alternative to pub/sub. They have enough depth for their own post: Valkey pub/sub and streams.

Bitmaps are not a separate type but string commands that treat a string as an array of bits.

code
SETBIT active:2026-10-08 4182 1GETBIT active:2026-10-08 4182BITCOUNT active:2026-10-08BITOP AND active:both active:2026-10-07 active:2026-10-08

With numeric user ids as offsets, one bit per user records "active today". A million users cost about 125 KB per day, and BITOP answers "active on both days" without touching the application. The catch is that the string is as long as the highest offset: a single user with id 4,000,000,000 allocates 500 MB. Use bitmaps only with dense, small integer ids.

HyperLogLog estimates the number of distinct items in a stream of values using at most about 12 KB per key, with a standard error of 0.81 per cent.

code
PFADD visitors:2026-10-08 "203.0.113.5" "198.51.100.7"PFCOUNT visitors:2026-10-08PFMERGE visitors:week visitors:2026-10-02 visitors:2026-10-08

Unique visitors, unique search terms, distinct devices - anywhere an approximate count is enough and storing every value would be too much. You cannot list the members back out; it only counts.

Geospatial indexes are sorted sets with coordinates encoded into the score.

code
GEOADD stores 44.7930 41.7151 "tbilisi-1" 13.4050 52.5200 "berlin-1"GEOSEARCH stores FROMLONLAT 44.80 41.71 BYRADIUS 5 km ASC COUNT 5 WITHDISTGEODIST stores "tbilisi-1" "berlin-1" km

Note the order: longitude first, then latitude. Getting that backwards puts your Tbilisi shop in the Indian Ocean.

Some things people expect are not part of core Valkey: JSON documents with path queries, full-text search, Bloom filters and time series come from modules, and a server without those modules loaded does not have the commands. Do not design around them unless you have confirmed they exist on the server you will use - MODULE LIST shows what is loaded, if your user is allowed to run it.

Memory, encodings and big keys#

Every key carries overhead - the key name, an internal object header, an expiry entry if set - in the region of tens of bytes before the value. A million tiny keys is therefore not a tiny dataset. Grouping related small values into one hash often halves memory, because a small hash is stored as a compact listpack.

TypeCompact formThresholds (defaults)
Hashlistpackup to 128 fields, values up to 64 bytes
Set of integersintsetup to 512 members
Setlistpackup to 128 members, values up to 64 bytes
Sorted setlistpackup to 128 members, values up to 64 bytes
Listlistpack nodes in a quicklistnode size set by list-max-listpack-size

The thresholds are configuration settings (hash-max-listpack-entries, hash-max-listpack-value, set-max-intset-entries, set-max-listpack-entries, zset-max-listpack-entries and the rest). Past them, the structure converts to its general form, which is faster for large collections and uses noticeably more memory per element. You rarely need to change the defaults; you do need to know that OBJECT ENCODING explains why one hash is ten times the size of another.

The bigger operational risk is the big key: one collection with millions of members, or one string of hundreds of megabytes. Deleting it with DEL blocks the server while it frees memory (use UNLINK, which frees in the background). Reading it whole blocks the server and floods the network. Find them with:

bash
$ valkey-cli -h 203.0.113.20 -p 6380 --askpass --bigkeys$ valkey-cli -h 203.0.113.20 -p 6380 --askpass --memkeys

Both sample the keyspace with SCAN, so they are safe on a live server. Valkey memory and eviction policies covers what happens when memory runs out.

On RE:NODE, Valkey plans run from 256 MB to 4 GB of memory, which is the number to size against: everything above lives in RAM, and the disk holds the AOF and snapshot copies.

Choosing a type: a worked example#

A small game community site wants: player profiles, a weekly leaderboard, "who is online", a feed of the last fifty events, a daily active-player count, and per-player rate limits on chat.

FeatureTypeKeyWhy
ProfileHashp:4182Fields updated independently, compact
Weekly leaderboardSorted setlb:2026-w41ZINCRBY and ZRANGE ... REV, new key each week with a TTL
Online nowSorted setonlineScore = last-seen time; ZREMRANGEBYSCORE drops the stale
Recent eventsListfeedLPUSH plus LTRIM 0 49
Daily activesHyperLogLogdau:2026-10-0812 KB however many players
Chat rate limitString counterrl:chat:4182:<minute>INCR with an expiry

Two choices there are worth noticing. "Online now" uses a sorted set rather than a set, because a set cannot drop members who stopped sending heartbeats - a sorted set scored by time can, in one command. And the leaderboard gets a new key each week with a TTL, so last week's board remains readable until it expires, and nothing has to reset scores at midnight on Sunday.

Updating several types at once

Real features touch more than one key. When a player finishes a match, the site increments their score in the leaderboard, bumps a field in their profile and pushes an event to the feed. Three ways to send that, with different guarantees:

  • A pipeline sends all three commands in one network round trip and reads the replies together. It is the cheapest option and the right default when the commands do not depend on each other. It is not atomic: another client's command can run between them.
  • `MULTI` and `EXEC` queue the commands and run them as one uninterruptible block. Nothing else runs in between. There is no rollback, though - if one command fails, for example with a WRONGTYPE error, the others still apply.
  • A Lua script runs atomically and can branch on what it reads: "add to the leaderboard only if the match id has not been recorded yet". Use it when a decision depends on current values.
code
MULTIZINCRBY lb:2026-w41 120 "nino"HINCRBY p:4182 matches 1LPUSH feed "nino won on Dust"LTRIM feed 0 49EXEC

WATCH adds optimistic locking to MULTI: watch a key, read it, and if anybody changes it before your EXEC, the transaction is discarded and you retry. It works, but under contention a short Lua script is usually simpler. For counters and leaderboards, single commands such as INCR and ZINCRBY are already atomic on their own, which is the main reason to pick the type that has the operation you need built in.

FAQ#

Can I store JSON in Valkey?

Yes, as a string, and that is the usual way. You read and write the whole document. If you update individual fields often, use a hash instead. Path queries into JSON need the JSON module, which is not part of core Valkey.

What is the maximum size of a value?

A string can be up to 512 MB, and a collection can hold billions of elements in principle. In practice keep values small: anything over a few megabytes is slow to transfer and blocks the server while it is read or written.

Hash or many string keys?

For an object with several fields, a hash: one key's overhead, compact encoding, and atomic field updates. Separate string keys are right when each value needs its own TTL, on servers without per-field hash expiry.

Why did a command block the whole server?

Commands run one at a time, so an expensive one - KEYS, SMEMBERS on a huge set, DEL on a huge key, ZRANGE 0 -1 on a big sorted set - makes every other client wait. Use the SCAN family and bounded ranges, and UNLINK instead of DEL.

Do numbers stored as strings use more memory?

Values that are integers are stored in a compact integer form, so SET counter 12345 costs little more than the key itself. Floats and numbers with leading zeros are stored as text.


Comments

Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.

0/2000