RE:NODE

App hosting12 min read

.NET worker services and background jobs: hosted to Hangfire

Run background work in .NET properly: BackgroundService, PeriodicTimer, in-process queues with Channels, Quartz.NET, Hangfire, error handling and clean shutdown.

0 readers

Background work in .NET starts with BackgroundService: a class with one method, ExecuteAsync, that the Generic Host starts with the app and cancels when the app stops. For "every five minutes, do this", a BackgroundService with a PeriodicTimer is all you need. For "do this soon, without making the HTTP request wait", a bounded Channel feeding a BackgroundService is all you need. You reach for a library when the work must survive a restart or follow a calendar: Quartz.NET for cron-style schedules, Hangfire for jobs persisted in a database with retries and a dashboard. Whichever you use, the two things that decide whether it works in production are the same - what happens when a job throws, and what happens when the process is told to stop halfway through one.

This post builds each of those, in order of how much machinery they need.

The hosting model#

Every modern .NET app - web, worker or console with the host - runs on the Generic Host, and the host runs hosted services. IHostedService is the interface: StartAsync when the app starts, StopAsync when it stops. BackgroundService is the base class almost everyone uses instead, because it turns that into one long-running method with a cancellation token.

bash
$ dotnet new worker -n Reports.Worker
Program.cs
var builder = Host.CreateApplicationBuilder(args);builder.Services.AddHostedService<CleanupWorker>();builder.Build().Run();

In an ASP.NET Core app it is the same line, builder.Services.AddHostedService<CleanupWorker>(), and the worker runs inside the web process alongside the request pipeline. That is not a compromise. For most small apps it is the right place, because the worker shares configuration, logging and services with the web app, and there is only one process to deploy.

Hosted services start in the order they are registered, and in older versions the start was sequential: code in ExecuteAsync before its first await ran on the startup path and delayed everything registered after it. .NET 10 changed BackgroundService to run all of ExecuteAsync in the background. On earlier versions, keep the synchronous part short, or start the method with await Task.Yield(). .NET 8 also added HostOptions.ServicesStartConcurrently and IHostedLifecycleService, with StartingAsync, StartedAsync, StoppingAsync and StoppedAsync hooks for services that need finer control.

A worker that runs on a timer#

The classic shape is a loop with a PeriodicTimer:

CleanupWorker.cs
public sealed class CleanupWorker(    IServiceScopeFactory scopes,    ILogger<CleanupWorker> log) : BackgroundService{    protected override async Task ExecuteAsync(CancellationToken stoppingToken)    {        using var timer = new PeriodicTimer(TimeSpan.FromMinutes(5));        do        {            try            {                await using var scope = scopes.CreateAsyncScope();                var db = scope.ServiceProvider.GetRequiredService<ShopDb>();                var removed = await db.Carts                    .Where(c => c.UpdatedAt < DateTime.UtcNow.AddDays(-7))                    .ExecuteDeleteAsync(stoppingToken);                log.LogInformation("Removed {Count} stale carts", removed);            }            catch (Exception ex) when (ex is not OperationCanceledException)            {                log.LogError(ex, "Cart cleanup failed");            }        }        while (await timer.WaitForNextTickAsync(stoppingToken));    }}

The details that matter:

  • Scopes. A hosted service is a singleton. Scoped services such as an EF Core DbContext cannot be injected into it directly - the host refuses at startup - and should not be held for its lifetime anyway, because the change tracker grows forever. Create a scope per unit of work with IServiceScopeFactory, as above.
  • `PeriodicTimer` does not overlap. If one run takes longer than the period, the next tick waits; runs never pile up on top of each other. That is usually what you want, and it is not true of System.Threading.Timer, which fires on schedule whether the last callback finished or not.
  • `WaitForNextTickAsync` throws on cancellation. When the host stops, the token is cancelled and the wait throws OperationCanceledException, which ends the loop. BackgroundService treats that as a normal stop.
  • Catch per iteration. The try inside the loop means one failed run is logged and the next one still happens. Without it, the first exception ends the worker - see below.

This covers a surprising amount: cache warming, cleanup, polling an external API, sending queued email, recalculating a leaderboard. Background jobs on a small server discusses which of these belong in the app at all.

Queues inside the process with Channels#

For work triggered by a request - send the welcome email, resize the upload, call the slow webhook - the request should return immediately and the work should happen behind it. System.Threading.Channels is the built-in, allocation-light queue for that:

csharp
public sealed class EmailQueue{    private readonly Channel<int> _channel = Channel.CreateBounded<int>(        new BoundedChannelOptions(500) { FullMode = BoundedChannelFullMode.Wait });    public ValueTask EnqueueAsync(int userId, CancellationToken ct) =>        _channel.Writer.WriteAsync(userId, ct);    public IAsyncEnumerable<int> ReadAllAsync(CancellationToken ct) =>        _channel.Reader.ReadAllAsync(ct);}public sealed class EmailWorker(EmailQueue queue, IServiceScopeFactory scopes,    ILogger<EmailWorker> log) : BackgroundService{    protected override async Task ExecuteAsync(CancellationToken stoppingToken)    {        await foreach (var userId in queue.ReadAllAsync(stoppingToken))        {            try            {                await using var scope = scopes.CreateAsyncScope();                var sender = scope.ServiceProvider.GetRequiredService<IWelcomeMailer>();                await sender.SendAsync(userId, stoppingToken);            }            catch (Exception ex) when (ex is not OperationCanceledException)            {                log.LogError(ex, "Welcome mail for {UserId} failed", userId);            }        }    }}

Register EmailQueue as a singleton and the worker as a hosted service; an endpoint calls EnqueueAsync and returns.

Make the channel bounded. An unbounded channel accepts work faster than you can process it until memory runs out; a bounded one with FullMode = Wait slows producers down instead, which is the honest behaviour. Queue small things - an ID, not an entity or a file - so the queue's memory stays predictable.

Be clear about the weakness: an in-memory queue dies with the process. Whatever is waiting in it when the app restarts or crashes is gone. For a welcome email that is often acceptable; for a payment confirmation it is not. That line - "can I lose this?" - is the point at which you need persistence, which means Hangfire, a database table you poll, or a queue server. Job queues with Valkey covers the queue server option.

Errors, and what kills the host#

Since .NET 6, an exception that escapes ExecuteAsync stops the whole host. The setting is HostOptions.BackgroundServiceExceptionBehavior, and its default is StopHost:

csharp
builder.Services.Configure<HostOptions>(o =>{    o.BackgroundServiceExceptionBehavior = BackgroundServiceExceptionBehavior.StopHost;});

That default is right, and the alternative, Ignore, is usually wrong. With Ignore, the worker silently stops while the app keeps serving requests, and nothing tells you that cleanup stopped running three weeks ago. With StopHost, the process exits, the exception is logged, and the platform restarts it - visible, and self-healing if the failure was transient.

The real fix is the try/catch per iteration shown above: a single bad record should cost one iteration, not the process. Reserve process-ending failures for things that are actually fatal, such as configuration that is missing.

Catching everything has its own failure mode: a worker that fails on every iteration, logs an error each time, and never stops. Nobody reads the log, and the job has effectively been dead for weeks. Record the time of the last successful run in a singleton, and expose it through a health check that reports unhealthy when the last success is older than, say, three periods. Then a monitoring check, or simply the health endpoint you already call after deploys, tells you the job is stuck before a customer does. Monitoring that tells you something covers what to alert on and what to leave alone.

On RE:NODE, a process that exits is restarted, and the crash watcher counts restarts nobody asked for: three in an hour raise a warning on the server page and open a ticket automatically, and six lead to suspension. A worker that throws on every start therefore gets noticed quickly - but it is better to catch per item and log than to rely on that.

Shutdown: what happens on stop and deploy#

When the host is asked to stop - a deploy, a restart, a stop from the panel - it cancels stoppingToken and waits for every hosted service to finish, up to HostOptions.ShutdownTimeout. The default has been 30 seconds since .NET 6. After that, the host stops waiting and the process exits, mid-job if necessary.

So every job must cope with being interrupted:

  1. Pass the token everywhere. Database calls, HTTP calls and delays that receive stoppingToken end promptly when it is cancelled. A loop that ignores it runs until the timeout and is then cut off anyway.
  2. Make work idempotent. A job interrupted halfway will run again. Sending the same email twice, or charging twice, must be impossible by design: record what was done, and check before doing it.
  3. Keep units small. A job that processes 10,000 rows in one transaction loses all of it on interruption. Batches of a few hundred, each committed, lose at most one batch.

If you start the process yourself, the start command should hand the stop signal to the .NET process directly - exec dotnet App.dll rather than a shell wrapper - or the host never hears it and cannot shut down cleanly. Graceful shutdown and health checks covers the signal side in general.

Quartz.NET for schedules#

When work must happen at a time - 03:00 every night, the first Monday of the month - rather than on an interval, use a scheduler. Quartz.NET is the established one:

bash
$ dotnet add package Quartz.Extensions.Hosting
csharp
builder.Services.AddQuartz(q =>{    var key = new JobKey("nightly-report");    q.AddJob<NightlyReportJob>(o => o.WithIdentity(key));    q.AddTrigger(t => t.ForJob(key).WithCronSchedule("0 0 3 * * ?"));});builder.Services.AddQuartzHostedService(o => o.WaitForJobsToComplete = true);[DisallowConcurrentExecution]public sealed class NightlyReportJob(ShopDb db) : IJob{    public async Task Execute(IJobExecutionContext context)    {        // build and send the report, honouring context.CancellationToken    }}

Note the cron expression. Quartz cron has a seconds field first and requires ? in one of the day fields, so 0 0 3 * * ? is "03:00:00 every day" - not the five-field Unix syntax most people know. Cron expressions explained covers the Unix form; read Quartz's own documentation for its variant before you trust a schedule.

[DisallowConcurrentExecution] stops a slow run overlapping the next. WaitForJobsToComplete makes shutdown wait for running jobs, within the host's shutdown timeout. Schedules are evaluated in the server's time zone unless you set one on the trigger with InTimeZone, which matters twice a year when the clocks change.

By default Quartz keeps its schedule in memory, so a run that should have happened while the app was down is simply missed. If missed runs matter, Quartz supports a database job store with misfire handling; at that point, compare it with Hangfire.

Hangfire for persistent jobs#

Hangfire stores jobs in a database, so they survive restarts, are retried automatically on failure, and can be inspected and re-run from a web dashboard:

csharp
builder.Services.AddHangfire(c => c.UseSqlServerStorage(    builder.Configuration.GetConnectionString("Hangfire")));builder.Services.AddHangfireServer();var app = builder.Build();app.UseHangfireDashboard("/jobs");// enqueue from anywhere, via IBackgroundJobClientjobs.Enqueue<IWelcomeMailer>(m => m.SendAsync(userId, CancellationToken.None));// recurring, via IRecurringJobManagerrecurring.AddOrUpdate<NightlyReport>("nightly-report", r => r.RunAsync(), Cron.Daily());

How it works: Enqueue serialises the method call and its arguments into the storage, and the server component polls storage and runs them. Consequences:

  • Arguments must be small and serialisable. Pass an ID, not an entity; the job loads fresh data when it runs.
  • Failed jobs retry automatically, ten times by default with increasing delays. That is why jobs must be idempotent.
  • Storage is a real dependency. SQL Server storage is maintained by the Hangfire authors; PostgreSQL and MySQL storage come from community packages; Redis storage is part of the commercial Pro edition. Check the package for your database before you choose.
  • The dashboard is local-only by default. Requests from anywhere but localhost are refused until you add an authorisation filter. Do not "fix" that by allowing everyone; the dashboard can delete and re-run jobs.

Hangfire's core is open source under the LGPL; batches and some other features are in the paid Pro edition.

Choosing, and where the work runs#

NeedUse
Every N minutes, losing a run on restart is fineBackgroundService + PeriodicTimer
Fire-and-forget from a request, loss acceptableChannel + BackgroundService
At set times, calendar-styleQuartz.NET
Must survive restarts, retries, visibilityHangfire, or a queue on Valkey

Where the work runs is the other decision. A single server runs one process, so on one app plan the background work lives inside the web app as hosted services - which is fine for most apps. Split it into a separate worker process, on a server of its own, when background work starts to compete with requests for CPU or memory: on RE:NODE CPU is a hard throttle to the share bought, so a heavy nightly job inside the web process slows the site while it runs. Hangfire and database-backed queues make that split easy later, because the web app and the worker only share a database. App plans include two database slots, which is enough for the app's data and a job store.

A worked example: a small shop on one app plan. Stale carts are cleaned every hour by a BackgroundService with a PeriodicTimer - losing one run to a restart costs nothing. Order confirmation emails go through Hangfire with SQL Server or PostgreSQL storage, because a lost confirmation is a support ticket and Hangfire retries it if the mail relay is briefly down. The nightly sales report is a Hangfire recurring job rather than a separate Quartz setup, since Hangfire is already there and one scheduler is easier to reason about than two. Everything runs in the web process. If the report ever grows heavy enough to slow the site at 03:00, the Hangfire server moves to a second plan unchanged and the web app simply stops calling AddHangfireServer().

The panel's own Schedules tab is a different tool: it runs ordered tasks on a cron expression - a console command, a backup or a power action. It is the right place for a nightly restart or a backup before a risky job, not for application logic. Scheduled tasks worth having lists the useful ones.

FAQ#

Should background jobs run in the web app or a separate process?

In the web app until they compete with it for CPU or memory. One process is simpler to deploy and monitor. Split when a job's load visibly slows requests, and use persistent jobs so either process can be restarted independently.

Why does my BackgroundService stop running without any error?

Either it threw once with BackgroundServiceExceptionBehavior set to Ignore, or the loop exited normally - often a while condition that became false. Log at the start and end of ExecuteAsync so you can see when it stops.

Can I inject a DbContext into a BackgroundService?

Not directly; it is scoped and the service is a singleton. Inject IServiceScopeFactory and create a scope per unit of work.

How do I run a job exactly once across several instances?

An in-process timer runs once per instance. Use Hangfire or Quartz with a shared database store, both of which coordinate between servers, or take a lock in the database before running.

Is Task.Run in a controller a background job?

No. It runs on the thread pool with no shutdown handling, no error reporting and no scope, and it is lost on restart. Put the work on a queue processed by a hosted service instead. .NET memory and garbage collection explains why work that captures request objects also holds memory longer than you expect.


Comments

Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.

0/2000