A job queue moves slow work out of the request: the web process writes a small record saying "send this email" or "resize this image", returns to the user at once, and a separate worker process picks the record up and does the work. Valkey is the most common place to keep that record, because every popular queue library - BullMQ on Node, Celery and RQ on Python, Laravel's queue on PHP, Sidekiq on Ruby - was built on the Redis protocol, and Valkey speaks it unchanged.
The libraries do the hard parts well: retries, delays, backoff, dead letters, concurrency. What they cannot do is protect you from a Valkey configured as a cache. A queue on a server that evicts keys under memory pressure loses jobs silently, and a queue on a server that does not persist loses every pending job on restart. This post sets up each library and then goes through the server settings, the failure modes and the sizing that make a queue trustworthy.
When a queue is the right tool#
A queue earns its place when work is slow, can fail and be retried, and does not need to finish before the user sees a response. The usual candidates:
- Sending email, SMS and push notifications.
- Generating thumbnails, PDFs, exports and reports.
- Calling third-party APIs that are slow or rate limited.
- Webhooks you receive and want to acknowledge immediately, then process.
- Scheduled and repeating work: nightly cleanups, digests, retrying payments.
What it is not for is work the user is waiting on. If the next page needs the result, a queue just adds latency and a polling loop.
Before adding Valkey for this, check whether your database already does the job. A table with SELECT ... FOR UPDATE SKIP LOCKED is a correct queue in PostgreSQL and MySQL 8, and it has one property no Valkey queue has: enqueueing the job and committing the data that caused it happen in one transaction. Background jobs on a small server walks through that option. Valkey wins when job volume is high, when you want the features of a mature library rather than writing your own, or when you already run Valkey for sessions and cache.
Server settings a queue depends on#
Three settings on the Valkey server decide whether jobs survive. Check them before writing any code.
| Setting | Queue needs | Why |
|---|---|---|
maxmemory-policy | noeviction | Any other policy can delete job keys when memory is full |
appendonly | yes | Pending jobs survive a restart |
appendfsync | everysec | At most about a second of enqueues lost on a hard stop |
maxmemory | Set, with headroom | A backlog must fit in memory |
Eviction is the one that loses jobs. Under allkeys-lru a full Valkey deletes the least recently used keys, and a job that has been waiting longest in a delayed set is exactly that. BullMQ checks this at startup and logs a warning when the policy is not noeviction; the others do not check at all. With noeviction, a full server rejects new jobs with an error, which your code sees and can alert on. That is the behaviour you want: a loud failure instead of a quiet one. Valkey memory and eviction policies covers the policies in detail.
$ valkey-cli -h 203.0.113.20 -p 6380 --askpass INFO memory | grep -E 'maxmemory|used_memory_human'$ valkey-cli -h 203.0.113.20 -p 6380 --askpass INFO persistence | grep -E 'aof_enabled|rdb_last'If the same Valkey holds a cache that should evict and a queue that must not, the honest answer is two instances. A small one for the queue costs less than debugging lost jobs.
On RE:NODE a Valkey plan saves to disk with both AOF and snapshots, and comes password-protected with the credentials generated per server. It runs as a separate server reached on its own host and port, so workers on an app plan and a web process elsewhere all connect to the same queue.
BullMQ on Node.js#
BullMQ is the current queue library for Node, the successor to Bull. It uses ioredis underneath and stores each queue as a set of keys under a bull:<queue>: prefix.
$ npm install bullmqimport { Queue } from "bullmq";export const connection = { host: process.env.VALKEY_HOST, port: Number(process.env.VALKEY_PORT), password: process.env.VALKEY_PASSWORD,};export const emails = new Queue("emails", { connection, defaultJobOptions: { attempts: 5, backoff: { type: "exponential", delay: 10_000 }, removeOnComplete: { age: 3600, count: 1000 }, removeOnFail: { age: 7 * 24 * 3600 }, },});// in the web processawait emails.add("welcome", { userId: 42 }, { jobId: `welcome:42` });import { Worker } from "bullmq";import { connection } from "./queue.js";const worker = new Worker("emails", async (job) => { await sendWelcomeEmail(job.data.userId);}, { connection: { ...connection, maxRetriesPerRequest: null }, concurrency: 5 });worker.on("failed", (job, err) => console.error(job?.id, err.message));process.on("SIGTERM", async () => { await worker.close(); process.exit(0); });The lines that matter:
- `maxRetriesPerRequest: null` on the worker's connection. BullMQ's workers use blocking commands that wait indefinitely;
iorediswith a retry limit gives up on them and breaks the worker. BullMQ sets this itself when it creates the connection, and warns if you pass a client configured otherwise. - `removeOnComplete` and `removeOnFail`. By default BullMQ keeps every finished job. On a busy queue that is the most common way a Valkey fills up: not pending work, but a history nobody reads. Keep a bounded window.
- `jobId`. A custom id makes enqueueing idempotent: adding a job whose id already exists does nothing. That is how you avoid sending the welcome email twice when a request is retried.
- `worker.close()` on `SIGTERM`. It stops taking new jobs and waits for the running ones, so a deploy does not interrupt a job halfway. Graceful shutdown and health checks explains why that signal handler matters on every platform.
Repeating jobs - "every night at 03:00" - are built in as job schedulers in current BullMQ versions, with older code using the repeat option on add. Use one, not both, and check the docs for the version you have installed.
Celery and RQ on Python#
Celery uses Valkey as a broker (where messages wait) and optionally as a result backend (where return values are kept).
import osfrom celery import Celeryurl = os.environ["VALKEY_URL"] # redis://:password@host:port/0app = Celery("shop", broker=url, backend=url.replace("/0", "/1"))app.conf.update( task_acks_late=True, worker_prefetch_multiplier=1, task_reject_on_worker_lost=True, result_expires=3600, broker_transport_options={"visibility_timeout": 3600},)@app.task(bind=True, autoretry_for=(ConnectionError,), retry_backoff=True, max_retries=5)def send_welcome(self, user_id): ...$ celery -A tasks worker --loglevel=INFO --concurrency=4$ celery -A tasks beat --loglevel=INFOThe setting to understand is visibility_timeout. With the Redis-protocol transport, a message taken by a worker is not deleted; it is hidden until acknowledged. If the worker has not acknowledged it within the visibility timeout - an hour by default - the message is delivered again to another worker. A task that legitimately runs longer than that timeout therefore runs twice, or many times. Set the timeout above your longest task, and with task_acks_late=True make tasks safe to run twice anyway. result_expires matters for the same reason as BullMQ's removeOnComplete: results you never read pile up as keys.
RQ is the simpler choice when you do not need Celery's routing, chords and canvas.
from redis import Redisfrom rq import Queue, Retryq = Queue("default", connection=Redis.from_url(os.environ["VALKEY_URL"]))job = q.enqueue(send_welcome, 42, retry=Retry(max=3, interval=[10, 60, 300]), job_timeout=600)$ rq worker --url "$VALKEY_URL" high default lowQueue names are listed in priority order: the worker drains high before looking at default. Failed jobs go to a failed job registry you can inspect and requeue with rq info and the RQ dashboard. Scheduled jobs (enqueue_in, enqueue_at) need the worker started with --with-scheduler.
Laravel queues#
Laravel's queue system needs only configuration to run on Valkey.
QUEUE_CONNECTION=redisREDIS_HOST=203.0.113.20REDIS_PORT=6380REDIS_PASSWORD=a-long-generated-password$ php artisan queue:work redis --queue=high,default --tries=3 --backoff=10 --max-time=3600In config/queue.php, the redis connection has a retry_after value, 90 seconds by default. It plays the same role as Celery's visibility timeout: a job reserved for longer than retry_after is released back onto the queue. Your longest job's $timeout must be a few seconds shorter than retry_after, or long jobs run twice. Laravel's documentation states this directly, and it is still the most common Laravel queue bug.
--max-time makes the worker exit cleanly after an hour so it is restarted fresh by whatever supervises it, which bounds memory leaks in long-lived PHP processes. After every deploy run php artisan queue:restart, because a running worker keeps the old code in memory. Laravel Horizon adds a dashboard and auto-balancing on top of the same Redis-protocol queues. Deploying Laravel to production covers supervising the worker.
Retries, idempotency and dead letters#
Every queue above delivers at least once. A worker can finish a job and die before acknowledging it, and the job will run again. Design for that rather than against it.
- Make jobs idempotent. "Mark invoice 512 as paid" is safe to repeat. "Add 10 to the balance" is not. Store a processed marker - a database unique constraint on an event id is the strongest - and check it first.
- Put ids in the payload, not objects. A job carrying a whole user record runs with stale data if it waits an hour. A job carrying
userId: 42loads the current state. - Retry with backoff, then stop. Exponential backoff with a cap of five or so attempts handles transient failures. A job that fails five times is not going to succeed on the sixth; it needs a human.
- Keep failed jobs where someone will look. BullMQ's failed set, RQ's failed registry and Laravel's
failed_jobstable are dead-letter queues. Alert on their size, not on individual failures. - Separate slow and fast queues. If a nightly export and password-reset emails share a queue and two workers, the export blocks the emails. Give urgent work its own queue and workers.
Sizing Valkey for a queue#
A queue's memory use is the backlog, plus any history you keep. A typical job record - payload, options, timestamps and a little library bookkeeping - is somewhere between a few hundred bytes and a few kilobytes. A backlog of 100,000 jobs at 2 KB is about 200 MB, which is why the steady-state size matters less than the worst case: what happens when workers are down for an hour during an incident and the web process keeps enqueueing.
| Situation | What uses memory | Rough guide |
|---|---|---|
| Small app, jobs processed in seconds | Little; backlog near zero | 256 MB is plenty |
| Steady traffic, completed jobs kept for an hour | Completed-job history | Count jobs per hour times size |
| Large payloads (HTML, base64 files) | The payloads | Store the file elsewhere, enqueue a reference |
| Workers stopped for an incident | The whole backlog | Size for an hour of enqueues |
Never put files in a job. Upload the image to disk or object storage and enqueue its path. Measure instead of estimating: INFO memory for the total, MEMORY USAGE on a sample job key, and the library's own counts (getJobCounts() in BullMQ, rq info, celery inspect).
CPU is rarely the limit for a queue on Valkey; a single small instance handles thousands of enqueues per second. What workers need is CPU and memory of their own, on whatever runs them.
What to watch
A queue fails slowly and then all at once, so watch the numbers that move first:
- Backlog size - waiting jobs per queue. A backlog that grows during the day and drains at night is fine; one that only grows means workers cannot keep up.
- Age of the oldest waiting job. More useful than the count: 10,000 jobs that take a millisecond each are nothing, ten jobs waiting for an hour are an incident.
- Failed jobs per hour, alerting on a rise rather than on individual failures.
- Valkey memory against its limit, from
INFO memory, with an alert well before thenoevictionwall. - Connected clients, from
INFO clients. A worker fleet that leaks connections ends at the server's client limit.
Each library has a dashboard that shows most of this - Bull Board or Taskforce for BullMQ, Flower for Celery, rq-dashboard for RQ, Horizon for Laravel - but a dashboard nobody looks at is not monitoring. Export the two or three numbers that matter to whatever already alerts you; monitoring that tells you something is about choosing them.
Workers also need supervising. A worker process that exits - an unhandled exception, an out-of-memory stop, a deploy - must be restarted automatically, or the queue quietly stops draining while the web process keeps adding to it. On a VDS that is a systemd unit with Restart=always; on a hosting panel, it is the server's own restart behaviour. Either way, check after every deploy that workers came back.
Troubleshooting#
Jobs disappear without running or failing. Eviction, or a restart without persistence. Check evicted_keys in INFO stats and the eviction policy.
The same job runs twice. A visibility timeout or retry_after shorter than the job, or a worker that died before acknowledging. Raise the timeout and make the job idempotent.
Workers stop picking up jobs after a network blip. The connection dropped and the client gave up. In BullMQ check maxRetriesPerRequest: null; elsewhere make sure the worker is supervised and restarted on exit.
Valkey memory grows steadily with no backlog. Completed jobs and results being kept forever. Set removeOnComplete, result_expires or the equivalent, then let the old ones expire.
`OOM command not allowed when used memory > 'maxmemory'`. The noeviction policy doing its job. The queue is full: workers are down or too slow, or history is filling memory.
FAQ#
Can Valkey replace RabbitMQ or Kafka?
For a single application's background jobs, yes, and with less to operate. RabbitMQ is better at complex routing between many services, and Kafka at retaining large event logs for replay. If you are choosing between them for an app's email and thumbnail jobs, a Valkey queue is the proportionate choice.
Do BullMQ, Celery and RQ work with Valkey unchanged?
They speak the Redis protocol and use standard commands and Lua scripting, which Valkey implements. You connect with the same redis:// URL and the same client libraries. Test your version against your server before going live, as you would with any infrastructure change.
Should the queue and the cache share one Valkey?
Better not. A cache wants an evicting policy and a queue must never evict. On one instance you have to choose, and either choice is wrong for one of them. Two small instances solve it.
How many workers should I run?
Enough to keep the backlog near zero at peak, and no more than the work can use. For I/O-bound jobs such as sending email, one process with concurrency 5 to 20 goes a long way. For CPU-bound jobs such as image processing, one worker per core.
What happens to scheduled jobs if Valkey restarts?
With persistence enabled they are on disk and reload with everything else; a delayed job whose time passed during the restart runs when workers reconnect. Without persistence, every pending and scheduled job is gone.




Comments
Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.