RE:NODE

App hosting12 min read

ASP.NET Core performance on small servers

Make ASP.NET Core fast on 1-2 vCPU and a few GB of RAM: Kestrel limits, response compression, output caching, thread pool starvation and allocations.

0 readers

ASP.NET Core is one of the fastest web stacks there is, and on a small server it is still easy to make slow. The framework is rarely the problem. On half a core and a gigabyte of memory, what costs you is blocking on async code, allocating megabytes per request, sending uncompressed JSON, recomputing the same response a thousand times a minute, and letting an unbounded queue of requests pile up instead of refusing the excess. Fix those five and a single small instance handles more traffic than most projects ever see. This post goes through each one with the settings, their defaults, and how to tell which one is actually hurting you before you change anything.

What a small server changes#

A .NET app sized on a developer laptop meets two hard limits in a container: a CPU quota and a memory limit. Both change how the runtime behaves.

The runtime reads the CPU quota from the container's cgroup and sets Environment.ProcessorCount from it, rounded up. On a plan with half a vCPU, the app believes it has one processor, which affects how many garbage collector heaps it creates, how many thread pool threads it starts with, and how much parallelism libraries choose by default. The quota is still a hard throttle: once the process has used its share of a scheduling period, it waits for the next one, which shows up as latency spikes rather than as an error.

Memory is the other ceiling. The garbage collector reads the container's limit and sizes its heap budget against it, so the app will not grow without bound, but every allocation still costs CPU time later in collections. On RE:NODE, a server that reaches its memory limit is stopped by the kernel and restarted clean rather than left to swap, so a memory spike is an outage, not a slowdown. .NET memory and garbage collection covers the collector settings in depth; this post covers what the application does to create the pressure.

Plan sizeWhat it comfortably runsWhere it runs out first
0.5 vCPU, 1 GBAn API or small site, tens of requests per secondCPU during bursts and GC
1 vCPU, 2 GBA busy API, Blazor Server for a few dozen usersMemory per connection, CPU in serialisation
1.5-2 vCPU, 4 GBA product with real traffic and background jobsDatabase round trips, not the app
3 vCPU, 8 GBSeveral apps or heavy in-process cachingUsually you have a design problem by now

Those rows are rough. An app that renders server-side pages with large models, or holds big caches in memory, moves up a row; a thin JSON API over a well-indexed database moves down one.

Measure before you tune#

Every tuning step below has a cost, and most performance work on small apps is wasted on the wrong layer. Get three numbers first:

  1. Where does the time go in one slow request? Log elapsed time per request (Serilog's request logging does this in one line) and, for the slow ones, how long the database part took. EF Core logs each command's duration at Information under Microsoft.EntityFrameworkCore.Database.Command; turn it on briefly.
  2. Is the CPU at its limit under load? The panel's CPU graph against the plan's limit answers this. Flat at the ceiling means CPU-bound; low CPU with high latency means you are waiting on something - the database, an external API, or the thread pool.
  3. How does it behave under concurrency? Run a load test from another machine with k6, bombardier or wrk at a realistic request mix, and watch latency percentiles, not the average.

On your own machine, dotnet-counters gives the runtime's view while a load test runs:

bash
$ dotnet tool install --global dotnet-counters$ dotnet-counters monitor --name MyApp System.Runtime Microsoft.AspNetCore.Hosting

The counters that matter are the thread pool queue length, the thread pool thread count, the GC heap size and the time spent in GC, and requests per second. A growing thread pool queue with low CPU is the signature of the most common .NET performance bug, which comes next. The exact counter names changed between .NET 8 and .NET 9, when the runtime moved to a new System.Runtime meter, so read them from the tool's output rather than from an old blog post.

Thread pool starvation and sync-over-async#

ASP.NET Core serves requests on thread pool threads. An async method that awaits I/O hands its thread back while it waits, so a handful of threads can serve thousands of concurrent requests. Code that blocks on async work - .Result, .Wait(), .GetAwaiter().GetResult(), or a synchronous database or HTTP call - holds a thread for the whole wait.

csharp
// Blocks a thread pool thread for the full round tripvar user = db.Users.FirstOrDefaultAsync(u => u.Id == id).Result;// Hands the thread back while waitingvar user = await db.Users.FirstOrDefaultAsync(u => u.Id == id);

When enough threads are blocked, new requests queue. The pool does add threads, but deliberately slowly - its injection heuristic is designed not to overreact, and adding a few threads per second is nowhere near enough when a burst arrives. The result is the classic starvation pattern: CPU low, latency climbing into seconds, then timeouts, then recovery once the burst passes. On a one-core plan the pool starts with very few threads, so it takes surprisingly little blocking to trigger.

The fix is to make the path async all the way down. The usual culprits:

  • Synchronous EF Core calls (ToList(), SaveChanges(), First()) in request handlers. Use the Async versions.
  • HttpClient calls wrapped in .Result because the caller was not async. Make the caller async.
  • Synchronous reads of the request body. Kestrel rejects these by default (AllowSynchronousIO is false) with an InvalidOperationException, and turning that setting on to silence the error is how you reintroduce starvation.
  • Locks around async work. Use SemaphoreSlim.WaitAsync() rather than lock, which cannot contain an await anyway.

ThreadPool.SetMinThreads raises the starting number of threads and is a legitimate stopgap while you remove the blocking calls, but it is a stopgap: more blocked threads still means more memory and more context switching on a single core.

Kestrel limits that protect a small server#

Kestrel's defaults are generous because they are written for servers of every size. A small server benefits from saying no early instead of accepting work it cannot finish.

SettingDefaultWhy you might change it
MaxConcurrentConnectionsunlimitedCap open connections so a flood queues at the proxy, not in your memory
MaxConcurrentUpgradedConnectionsunlimitedSame for WebSockets and SignalR
MaxRequestBodySize30,000,000 bytesLower it for an API that never accepts uploads
KeepAliveTimeout130 secondsShorter frees idle connections sooner
RequestHeadersTimeout30 secondsShorter limits slow-header attacks
MinRequestBodyDataRate240 bytes/s after 5 s graceDrops clients trickling a body to hold a connection
MaxRequestHeadersTotalSize32 KBRarely needs changing
csharp
builder.WebHost.ConfigureKestrel(options =>{    options.Limits.MaxConcurrentConnections = 500;    options.Limits.MaxConcurrentUpgradedConnections = 200;    options.Limits.MaxRequestBodySize = 2 * 1024 * 1024;    options.Limits.KeepAliveTimeout = TimeSpan.FromSeconds(60);});

The same values can live in configuration under Kestrel:Limits, which lets you change them with an environment variable such as Kestrel__Limits__MaxConcurrentConnections=500 instead of a rebuild. An individual endpoint that does accept uploads can raise the body limit for itself with [RequestSizeLimit] or the IHttpMaxRequestBodySizeFeature.

Connection limits do not limit work in progress, though. For that, the rate limiting middleware (in the box since .NET 7) has a concurrency limiter, which is the most useful single protection for a small instance: it lets a fixed number of requests run at once, queues a few more, and rejects the rest with 503 instead of letting latency grow without bound.

csharp
builder.Services.AddRateLimiter(options =>{    options.RejectionStatusCode = StatusCodes.Status503ServiceUnavailable;    options.GlobalLimiter = PartitionedRateLimiter.Create<HttpContext, string>(_ =>        RateLimitPartition.GetConcurrencyLimiter("global", _ =>            new ConcurrencyLimiterOptions { PermitLimit = 64, QueueLimit = 32 }));});app.UseRateLimiter();

Per-client limits (fixed window, sliding window, token bucket) live in the same middleware and are the right tool against one noisy client, as long as forwarded headers are configured so the partition key is the real client address. Rate limits and abuse has the reasoning for picking numbers.

Response compression#

JSON and HTML compress to a fraction of their size, and on a small server bandwidth is rarely the constraint - latency for the client is. Compression trades a little CPU for much less time on the wire.

csharp
builder.Services.AddResponseCompression(options =>{    options.EnableForHttps = true;    options.Providers.Add<BrotliCompressionProvider>();    options.Providers.Add<GzipCompressionProvider>();});builder.Services.Configure<BrotliCompressionProviderOptions>(o =>    o.Level = CompressionLevel.Fastest);app.UseResponseCompression();

Three cautions. EnableForHttps is off by default because compressing responses that mix secrets with attacker-controlled input over TLS is the setting for the CRIME and BREACH family of attacks; it is safe for public JSON and pages without secrets in them, and you should think before enabling it on responses that carry tokens. CompressionLevel.Fastest is the right level for dynamic responses on a small CPU; the optimal level can cost several times the CPU for a few percent less output. And if a reverse proxy in front of the app already compresses, do not compress twice: it wastes CPU and the proxy will usually pass your compressed body through unchanged anyway. Check the response headers - Content-Encoding: br or gzip - with compression disabled in the app, and you will know whether the proxy is doing it.

Static assets are better compressed once at build time than on every request. MapStaticAssets() in .NET 9 and later does this for files known at build time, serving precompressed versions with fingerprinted names and long cache lifetimes.

Output caching and response caching#

The cheapest request is the one you do not compute. ASP.NET Core has two caching middlewares and they are often confused:

  • Response caching (AddResponseCaching) follows HTTP caching headers. It caches only what the headers allow, and a browser sending Cache-Control: no-cache bypasses it. It is mostly useful for setting headers correctly for downstream caches.
  • Output caching (AddOutputCache, .NET 7 and later) is a server-side cache you control. The client cannot bypass it, entries can be tagged and evicted, and concurrent requests for the same uncached entry are collapsed into one computation, which protects a small server from a stampede.
csharp
builder.Services.AddOutputCache(options =>{    options.AddBasePolicy(b => b.Expire(TimeSpan.FromSeconds(10)));    options.AddPolicy("Products", b => b.Expire(TimeSpan.FromMinutes(5)).Tag("products"));});app.UseOutputCache();app.MapGet("/products", GetProducts).CacheOutput("Products");app.MapPost("/products", async (Product p, IOutputCacheStore cache, CancellationToken ct) =>{    // ... save    await cache.EvictByTagAsync("products", ct);});

By default, output caching stores only GET and HEAD responses with status 200, and skips requests that are authenticated and responses that set cookies, which is what keeps one user's page from being served to another. The default store is in memory with a 100 MB size limit; on a 1 GB plan, set SizeLimit lower. For several instances sharing one cache, .NET 8 added a Redis-protocol store (Microsoft.AspNetCore.OutputCaching.StackExchangeRedis, registered with AddStackExchangeRedisOutputCache), which works against Valkey. Valkey caching patterns covers what to cache and how to expire it.

Even ten seconds of output caching on a busy endpoint turns hundreds of database queries a minute into six. That is usually the biggest single win available, and it costs a few lines.

Allocations, serialisation and the database#

On a small server the garbage collector competes with your requests for the same CPU, so allocation rate is throughput. The common sources, roughly in order of how often they matter:

  • Loading more than you need from the database. ToListAsync() on an unfiltered query, or entities with every column when the page shows three. Project with Select into a small record, page with Skip and Take, and use AsNoTracking() for reads - change tracking keeps a snapshot of every entity it loads.
  • N+1 queries. A loop that lazily loads a relation per row turns one request into a hundred round trips. Use Include or a projection, and read the SQL EF Core actually generates.
  • Buffering large responses. Returning a List<T> of fifty thousand rows builds the whole list and then the whole JSON in memory. Page it, or return an IAsyncEnumerable<T>, which System.Text.Json streams.
  • New `HttpClient` per request. Use IHttpClientFactory or a single long-lived client. Creating clients per request exhausts sockets and allocates for nothing.
  • Reflection-based serialisation on hot paths. System.Text.Json source generation (JsonSerializerContext) removes the reflection and some allocation, and is required for trimmed or Native AOT builds anyway.
  • String building in loops and large `byte[]` buffers. StringBuilder, ArrayPool<T>.Shared and Span<T> exist for this.

The database is usually the slowest component in a small deployment, and nothing in this post helps a query missing an index. A connection pool sized far beyond what the database can serve also hurts more than it helps; connection pools and limits explains why the right pool is smaller than people expect.

Startup, the GC mode and publishing options#

A few build and runtime settings matter specifically on small machines:

MyApp.csproj
<PropertyGroup>  <InvariantGlobalization>true</InvariantGlobalization>  <TieredPGO>true</TieredPGO>  <PublishReadyToRun>true</PublishReadyToRun></PropertyGroup>

InvariantGlobalization drops the ICU culture data, saving memory and startup time, at the cost of culture-specific formatting and sorting - fine for most APIs, wrong for an app that formats dates and currencies per user. TieredPGO is already on by default from .NET 8, so listing it only documents intent. PublishReadyToRun precompiles code so the first requests after a restart are not slowed by the JIT, which matters when a deploy restarts the app under traffic; it needs a runtime identifier at publish time and makes the output larger. dotnet publish and runtime options covers RIDs, self-contained builds and trimming.

ASP.NET Core apps use Server GC by default. With one visible processor the runtime uses workstation GC anyway, and from .NET 9 Server GC enables DATAS, which adapts the heap count to the load and keeps memory much lower than older Server GC did. If an app on a 2-core plan sits at a higher baseline memory than you can afford, <ServerGarbageCollector>false</ServerGarbageCollector> trades some throughput for a smaller footprint; measure both before deciding.

FAQ#

How many requests per second can a small ASP.NET Core server handle?

For simple JSON endpoints that hit a cache or a fast query, thousands on one core. For endpoints that query a database each time, the database sets the number, typically hundreds. The framework is rarely the limit; your slowest dependency and your allocation rate are.

Is Kestrel safe to expose without a reverse proxy?

Kestrel is a supported edge server and handles TLS and HTTP/2 itself. A proxy in front is still useful for certificates, a stable domain and header handling, and it absorbs slow clients before they reach your app. If you expose Kestrel directly, set the limits in this post deliberately.

Should I enable response compression in the app or at the proxy?

At one of them, not both. If the proxy in front of you compresses, leave it there. If it does not, enable it in the app with Brotli and gzip at the fastest level, and think about EnableForHttps for responses that contain secrets.

Why does my app get slow under load while CPU stays low?

Almost always thread pool starvation from blocking on async code, or waiting on a saturated database connection pool. A growing thread pool queue in dotnet-counters confirms the first; slow query logs or pool timeout errors confirm the second.

Does Native AOT make sense on a small server?

It gives faster startup and a smaller memory footprint, but not every library supports it, and in .NET 8 and later only part of ASP.NET Core is compatible - minimal APIs and gRPC, not MVC or Blazor Server. For a small, self-contained API it can be worth it; for a typical app with EF Core and MVC, ReadyToRun plus the fixes above is the better trade.


Comments

Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.

0/2000