A .NET app's memory use is mostly a decision made by the garbage collector, not by your code. ASP.NET Core apps run Server GC by default, which trades memory for throughput by keeping a separate heap per CPU and collecting less often; console apps and worker services run Workstation GC, which collects more often and stays smaller. Inside a container, the runtime reads the memory limit and caps the managed heap at 75% of it by default, and since .NET 9 Server GC adapts its heap count to the app's actual size (DATAS). So a small API that "uses 400 MB" on a 4 GB plan is often just a GC that has been given room and is using it - and the same app on a 1 GB plan will run in far less.
This post explains what the GC does with the memory it is given, which settings change that, how to read the numbers that actually matter, and how to tell a leak from a heap that is simply large.
Where a .NET process's memory goes#
The memory graph shows the whole process. The managed heap - the objects your code allocates - is the largest part, but not all of it:
- The GC heap, split into generations. New objects go into generation 0; survivors are promoted to 1 and then 2. Objects of 85,000 bytes or more go straight to the large object heap (LOH), which is collected with generation 2 and not compacted by default. Pinned objects have their own heap since .NET 5.
- Committed but empty heap space. After a collection the GC keeps memory it expects to need again rather than returning it to the operating system immediately. This is the main reason the graph does not fall after a burst of traffic.
- JIT-compiled code and runtime data: type metadata, method tables, compiled code for every method that has run. A large app with many dependencies carries tens of megabytes of this regardless of traffic.
- Native memory: anything allocated outside the GC - database drivers with native parts, image libraries, compression buffers, SQLite.
- Thread stacks, one per thread. Usually small in total, unless something creates hundreds of threads.
Two different failures follow from this. When the managed heap reaches its limit, the GC throws OutOfMemoryException inside your app. When the whole process reaches the container's limit, the kernel stops it from outside, with no exception and no last log line. On RE:NODE, a container that reaches its memory limit is stopped and restarted clean rather than left to swap, so the second case shows up as a restart. Linux swap and the OOM killer explains the kernel side.
Workstation and Server GC#
.NET has two GC flavours, and which one you get depends on the project type, not on the server:
| Workstation GC | Server GC | |
|---|---|---|
| Default for | Console apps, worker services | ASP.NET Core (Web SDK) projects |
| Heaps | One | One per logical CPU (adapted by DATAS) |
| GC threads | Collects on the triggering thread | A dedicated thread per heap |
| Gen 0 budget | Small | Large |
| Memory use | Lower | Higher |
| Throughput | Lower under load | Higher under load |
Both run in concurrent (background) mode by default, which lets generation 2 collections happen alongside your code instead of stopping it.
Server GC's larger budgets mean it allocates for longer between collections and holds more memory while it does. On a machine with 16 cores and no limit, a Server GC process can sit at several hundred megabytes doing almost nothing. That is the source of the folklore that ASP.NET Core is memory-hungry. The runtime is optimising for throughput because it was told it could.
Two rules about the defaults that matter on small plans:
- One CPU means Workstation GC. If the runtime sees a single processor, it uses Workstation GC whatever the setting says. A plan with half a vCPU usually presents as one processor (see below), so a small ASP.NET Core app on the smallest plan is already running the lean collector.
- You can choose explicitly. In the project file, or as an environment variable:
<PropertyGroup> <ServerGarbageCollection>false</ServerGarbageCollection> <ConcurrentGarbageCollection>true</ConcurrentGarbageCollection></PropertyGroup>DOTNET_gcServer=0Switching a web app to Workstation GC on a two-core plan reduces memory noticeably and costs some throughput at high request rates. For an app serving a few requests a second, you will not see the throughput difference; you will see the memory difference.
DATAS: Server GC that sizes itself#
Dynamic Adaptation To Application Sizes (DATAS) changes Server GC from "one heap per core, always" to "as many heaps as this app's load justifies". It starts with one heap and adds more as allocation pressure grows, and it adjusts the generation 0 budget to the size of the live data rather than to the core count.
It was opt-in in .NET 8 and is on by default with Server GC from .NET 9. If you are on .NET 9 or 10 and your app uses Server GC, you already have it. The practical effect is that a quiet ASP.NET Core app on a machine with many cores no longer reserves memory as if it were under full load.
You can turn it off, which is occasionally right for a throughput-critical service that is always busy:
DOTNET_GCDynamicAdaptationMode=0The MSBuild equivalent is <GarbageCollectionAdaptationMode>0</GarbageCollectionAdaptationMode>. On .NET 8, set it to 1 to opt in. For small and medium apps on modest plans, leave DATAS on.
What the runtime sees inside a container#
The runtime reads the container's cgroup limits, both cgroup v1 and v2, and adjusts two things.
The memory limit. When the container has a memory limit, the GC sets a hard limit on the managed heap of 75% of it (or 20 MB, whichever is larger). The remaining quarter is left for native memory, code, stacks and everything else the process needs. On a 1 GB plan the managed heap can grow to about 768 MB; on 4 GB, about 3 GB.
The CPU count. Environment.ProcessorCount reflects the container's CPU quota rounded up to a whole number. A limit of half a CPU appears as 1; one and a half appears as 2. That number decides how many heaps Server GC may create, how the thread pool sizes itself, and how many things the runtime does in parallel. On RE:NODE the CPU limit is a hard throttle to the share you bought, so the runtime's view of "2 processors" on a 1.5 vCPU plan is right about the count but generous about the time: the process gets 150% of one core's time across those threads, no more. You can override the count with DOTNET_PROCESSOR_COUNT if the default leads to too many heaps or threads.
The consequence for sizing is that the same app uses different amounts of memory on different plans, because the GC is given different room. Do not read a memory figure from a large machine and assume you need that much.
The settings that change memory use#
| Setting | runtimeconfig / MSBuild | Environment variable | Effect |
|---|---|---|---|
| Server GC | ServerGarbageCollection | DOTNET_gcServer | 0 for workstation, 1 for server |
| Concurrent GC | ConcurrentGarbageCollection | DOTNET_gcConcurrent | Background gen 2 collections |
| Heap hard limit | System.GC.HeapHardLimit | DOTNET_GCHeapHardLimit | Absolute cap on the managed heap |
| Heap limit percent | System.GC.HeapHardLimitPercent | DOTNET_GCHeapHardLimitPercent | Cap as a share of the limit |
| Conserve memory | System.GC.ConserveMemory | DOTNET_GCConserveMemory | 0-9; compacts more to save memory |
| DATAS | GarbageCollectionAdaptationMode | DOTNET_GCDynamicAdaptationMode | 1 on, 0 off |
When to touch the heap limit: if your app uses significant native memory - an image processing library, a large SQLite cache, an embedded engine - the default 25% headroom may not be enough, and the process is killed by the container limit before the GC ever feels pressure. Lowering the managed limit to 60% gives native code more room and makes the GC collect harder sooner. The opposite case, raising it to 85% or more, is only safe for apps with almost no native allocation.
GCConserveMemory is for apps with heavy LOH fragmentation: large buffers allocated and freed repeatedly leave holes the GC does not compact. A value around 5 makes it compact the LOH when fragmentation is high, at some CPU cost.
Reading the numbers that matter#
The process working set (the number on the graph) tells you whether you will hit the limit. It does not tell you why. For that, ask the GC:
app.MapGet("/debug/memory", () =>{ var info = GC.GetGCMemoryInfo(); return new { heapMB = info.HeapSizeBytes / 1_048_576, committedMB = info.TotalCommittedBytes / 1_048_576, fragmentedMB = info.FragmentedBytes / 1_048_576, limitMB = info.TotalAvailableMemoryBytes / 1_048_576, workingSetMB = Environment.WorkingSet / 1_048_576, gen0 = GC.CollectionCount(0), gen1 = GC.CollectionCount(1), gen2 = GC.CollectionCount(2), serverGc = System.Runtime.GCSettings.IsServerGC, cpus = Environment.ProcessorCount };});Put it behind authentication, or remove it after use. How to read it:
limitMBis the memory the GC believes it has. If it shows the host's whole memory rather than your plan, the runtime did not detect the container limit, and nothing else in this post applies until it does.heapMBagainstworkingSetMBtells you whether the growth is managed (your objects) or native (something outside the GC).committedMBwell aboveheapMBis memory the GC is holding for reuse. That is not a leak.gen2climbing quickly under steady load means objects are surviving long enough to be promoted and then dying - often a cache with no size limit, or per-request objects captured by something long-lived.
For a deeper look, the dotnet-counters and dotnet-gcdump diagnostic tools read the same data live and capture a heap snapshot you can open in Visual Studio or PerfView. They need to run where the process runs, which is easy on a VDS and depends on the environment elsewhere. Since .NET 9, the runtime also publishes these as built-in metrics on the System.Runtime meter, so an OpenTelemetry exporter can send them wherever you keep metrics. Reading a server load graph covers the panel's graph.
Leaks, and things that look like leaks#
A real leak is memory that grows without bound under steady load and is never collected because something still references it. In .NET the usual culprits are:
- Unbounded caches. A static
DictionaryorConcurrentDictionarythat only ever gains entries.IMemoryCacheis unbounded unless you setSizeLimitand give every entry aSize. - Event handlers on long-lived objects. Subscribing a short-lived object to an event on a singleton keeps the short-lived object alive forever unless it unsubscribes.
- Disposables resolved from the root provider. A transient
IDisposableresolved from the application's root service provider - instead of a scope - is tracked until the app shuts down. Resolve it from a scope, especially in background services. - A long-lived `DbContext`. EF Core's change tracker holds every entity it has loaded. A context kept alive in a singleton or a background loop grows until it is disposed. Use a scope per unit of work, and
AsNoTracking()for reads. - Buffered large requests and responses. Reading a whole upload into a
MemoryStreamputs it on the LOH; a few concurrent uploads of 50 MB each is a few hundred megabytes. Stream to disk or to object storage instead.
Things that look like leaks and are not: memory that rises after start and plateaus (the JIT, caches warming, the GC settling on a budget); memory that does not fall after a traffic spike (committed space retained for reuse); a sawtooth that returns to the same floor (normal collection). A leak's floor rises. Watch the low points of the graph over hours, not the peaks over minutes. Game server memory leaks describes the same reasoning for a different kind of process.
Sizing a plan for a .NET app#
These are working figures for a framework-dependent app on .NET 8 to 10 under modest load. Measure your own; dependencies change them more than anything else.
| App | Typical working set | Plan to start on |
|---|---|---|
| Discord bot or worker service | 60-150 MB | 1 GB |
| Minimal API, no ORM | 50-120 MB | 1 GB |
| API with EF Core, light traffic | 120-300 MB | 1-2 GB |
| Razor Pages or MVC site | 150-400 MB | 2 GB |
| Blazor Server, dozens of users | 300 MB plus per-user state | 2-4 GB |
Add room for builds if you compile on the server: the SDK and compiler need more memory than most small apps do at runtime. Deploy an ASP.NET Core app from GitHub covers moving the build to CI, and Blazor Server vs WebAssembly explains why Blazor Server scales with users rather than requests.
Upgrade when the floor of the memory graph, not the peak, is above about 70% of the limit, or when the app restarts at the limit. On RE:NODE a restart at the memory limit counts towards the crash watcher - three in an hour raise a warning and open a ticket - so a leak does not go unnoticed for long. Moving up a plan changes the limit on the existing server rather than rebuilding it.
FAQ#
Why does my ASP.NET Core app use more memory on a bigger plan?
Because the GC is told it has more memory and more CPUs, and sizes its budgets accordingly. The app is not using more because it needs it; it is collecting less often because it can. On a smaller limit it will collect more often and stay smaller.
Should I call GC.Collect() to free memory?
No. Forced collections pause the app, promote objects that were about to die, and usually leave memory where it was a minute later. If memory is a problem, fix what holds it or set a heap limit.
Is Workstation GC slower for a web app?
Under heavy concurrent load, yes. For an app on one or two CPUs serving modest traffic, the difference is hard to measure, and the memory saving is real. On a single-CPU container you are already running Workstation GC.
What happens when the managed heap hits its hard limit?
The GC collects as hard as it can and, if that is not enough, throws OutOfMemoryException at the allocation that failed. If the whole process hits the container limit first, the kernel stops it with no exception at all.
Does Native AOT use less memory?
At idle, yes, because there is no JIT and less runtime metadata. Under load the GC behaves the same. dotnet publish options covers what Native AOT costs in compatibility.




Comments
Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.