A managed runtime like .NET comes with a garbage collector that promises to handle memory so you don’t have to. For a typical line-of-business app, that deal works just fine. But when you’re building systems where response time has to stay below a hard ceiling—trading engines, real-time bidding platforms, multiplayer game servers, high-frequency telemetry ingestion—the GC can turn into your biggest liability. The real problem isn’t the small, frequent Gen 0 or Gen 1 collections. It’s the full, blocking Gen 2 collection that tears up your latency guarantees.
I’ve spent years profiling production .NET applications, often under punishing load, and one pattern keeps surfacing: Gen 2 collections introduce pauses that are orders of magnitude bigger than a typical request budget. Knowing exactly why this happens—and what you can realistically do about it—is what separates a system that bends from one that snaps when the pressure peaks.

The Generational Hypothesis and Its Limits
The .NET GC is a generational collector. It’s built on a simple observation: most objects die young. Gen 0 is where newly allocated objects land, and it gets cleaned up often with a fast, stop-the-world pause—usually microseconds to a few milliseconds. Objects that make it through a Gen 0 collection get promoted to Gen 1, a kind of staging area before the long-lived Gen 2. Gen 1 collections are still pretty quick. Gen 2, though, is the attic. That’s where objects that have survived multiple rounds live—static data, long-lived caches, pooled buffers, anything that sticks around for the application’s lifetime.
When a Gen 2 collection kicks in, the GC has to walk the entire managed heap, across all generations. In workstation GC mode, this is a fully blocking affair: every application thread gets suspended until the collection finishes. Server GC mode runs the collection on dedicated GC threads, but managed threads still get paused during certain phases. The pause time grows with the size of the Gen 2 heap and the tangle of your object graph. With a heap that’s several gigabytes, a full Gen 2 collection can easily eat up hundreds of milliseconds, sometimes more than a second. For an API that has to answer in under 50 milliseconds, that’s a lifetime.
What Triggers a Gen 2 Collection
The GC fires a Gen 2 collection under a handful of conditions:
- Gen 2 budget exhaustion: Each generation has a budget, a threshold that, when crossed, triggers a collection. For Gen 2, that budget adjusts itself over time, but if the heap keeps ballooning, a full collection becomes a matter of when, not if.
- Large object heap (LOH) allocation pressure: Objects over 85,000 bytes go straight to the LOH, which only gets cleaned during Gen 2 collections. So if you’re churning through big arrays or strings, you’re forcing frequent full collections.
- Explicit calls to
GC.Collect(): Any code that demands a full collection—including a few libraries that quietly callGC.Collect()behind the scenes—short-circuits the generational logic and slaps you with a full pause. - Low system memory: When the OS reports memory pressure, the GC might run a full collection to trim the working set, even if Gen 2 isn’t really full yet.
Each of these triggers can fire at seemingly random moments from the viewpoint of an incoming request. The pause lands right in the middle of something time-sensitive, blowing out your 99th percentile latency and carving a sawtooth pattern into your response-time graphs.

How Gen 2 Pauses Manifest in Production
The damage from a Gen 2 collection isn’t just the raw pause time. When every thread gets suspended, requests pile up. The moment threads wake up, they have to chew through that backlog, which spikes CPU and drags out latency even further. That cascade often turns a 200-millisecond GC pause into a multi-second brownout.
Think about a trading system handling 10,000 orders per second. A 500-millisecond Gen 2 collection stops the world. In that half-second, 5,000 orders stack up. When processing kicks back in, the system scrambles to catch up, but that sudden burst of work can trigger yet more allocations—and maybe another collection sooner than you’d like. This feedback loop can destabilize the whole service.
One investigation I remember: a market data gateway kept failing at regular intervals. We traced it to a Gen 2 collection caused by a logging library that was allocating big temporary byte arrays for serialization. Those arrays landed on the LOH. After minutes of steady traffic, the LOH filled, forcing a full GC. The pause blew past the heartbeat timeout for downstream consumers, and they dropped the connection. The fix wasn’t some clever GC tuning knob—it was ripping out those large allocations entirely.
The Diagnostic Trail
To pin down Gen 2 latency problems, you need to line up GC events with your application’s response times. Event Tracing for Windows (ETW) and tools like PerfView or dotnet-trace give you Microsoft-Windows-DotNETRuntime events, including GC/Start and GC/Stop with the generation depth. I usually grab a trace during a window of elevated tail latency and hunt for GC/Stop events where Generation is 2 and the duration matches the latency spike.
Here’s a quick dotnet-trace command to pull GC events:
dotnet-trace collect --process-id [PID] --providers Microsoft-Windows-DotNETRuntime:GC:4
You can crack open that trace in PerfView and see the exact pause times plus the reason for each collection. With server GC, you might see concurrent Gen 2 collections (background GC), but even those have a brief stop-the-world phase that can still hurt. The metric that really matters is the suspend time for managed threads, not just the total GC wall-clock duration.
Strategies to Eliminate Gen 2 Pauses
Given how brutal the problem is, the real goal isn’t just to shrink Gen 2 pause times. It’s to prevent Gen 2 collections from happening at all during the moments you care about. There are a few paths, and they all come with trade-offs.
1. Reduce Heap Size and Allocation Rate
The most straightforward move is to keep the Gen 2 heap small and stop doing the kind of allocations that promote objects. That means aggressive object pooling, especially for anything large. ArrayPool<T> is a workhorse here. Instead of allocating a fresh byte array for every request, you rent one from the pool and return it when you’re done. This cuts both Gen 0 pressure and LOH fragmentation.
For classes, think about object pooling with a custom pool or ObjectPool<T> from Microsoft.Extensions.ObjectPool. The cost of resetting an object is almost always dwarfed by the cost of a Gen 2 collection. Watch out for string allocations—they’re immutable and often drift into Gen 2 if they hang around through a few collections. Reach for StringBuilder with a pooled backing array or Span<char> on the stack when you can.
2. Tune GC Mode and Latency Mode
.NET gives you two GC modes that matter for latency: workstation GC with sustained low latency and server GC with background collections. Workstation GC runs a single heap and is tuned for interactive apps. Flipping GCSettings.LatencyMode to GCLatencyMode.SustainedLowLatency or GCLatencyMode.LowLatency (the latter for short bursts) tells the GC to hold off on Gen 2 collections as long as it can. This works, but it’s a gamble: if memory pressure builds, the GC eventually forces a full collection, and the pause might be even worse.
Server GC spreads the load across multiple heaps (one per CPU) with dedicated GC threads. It usually pushes higher throughput and is the default for ASP.NET Core apps. Background Gen 2 collections let the application keep running while the GC scans the heap, but a short stop-the-world phase remains. Some teams building really low-latency systems turn off background GC and instead schedule manual collections during quiet periods—though that takes real discipline.
3. Segment the Workload with Out-of-Process Caching
If your Gen 2 heap is swollen because of a giant cache, moving that cache out of process can lighten the load. A distributed cache like Redis, or a separate caching service, pulls those long-lived objects out of the managed heap. What’s left inside your application process are mostly short-lived, request-scoped objects that rarely make it to Gen 2. You’ll see this pattern a lot in microservice architectures where state gets externalized.
For in-memory caches that have to stay in-process, lean on MemoryCache with size limits and compaction callbacks to evict entries before they balloon the heap. Or, use a library that parks cached objects in unmanaged memory, sidestepping the GC entirely—though that definitely adds complexity.
4. Avoid LOH Allocations
The LOH is a quiet assassin. Every LOH allocation nudges you closer to a Gen 2 collection. Starting with .NET Core 3.0, the GC can compact the LOH on demand, but that compaction is itself a full blocking operation. The smarter path is to avoid LOH allocations altogether. That means:
- Breaking large collections into smaller arrays (under 85,000 bytes each).
- Using
Span<T>andMemory<T>to carve up existing buffers instead of allocating new ones. - Leaning on
ArrayPool<T>for temporary large buffers. - Steering clear of huge string concatenations that produce a string over 85,000 characters.
One team I worked with wiped out 90% of their Gen 2 collections by swapping a single large memory stream allocation per request for a recycled buffer from a pool. The latency improvement was instant and hard to ignore.

When You Cannot Eliminate Gen 2 Collections
Some applications just can’t dodge Gen 2 collections—systems that have to keep large, shifting in-memory data structures, for instance. Then the play becomes making those pauses predictable and squeezing them inside your budget.
GC notification API: .NET exposes GC.RegisterForFullGCNotification, which can tip you off when a Gen 2 collection is on its way. You can then throttle incoming requests, shed load, or steer traffic to another node. This is a bit fragile because the notification timing isn’t exact, but it can hold up for batch processing systems.
Manual collections during quiescence: If your application has natural lulls—say, a market data system that goes quiet between trading sessions—you can force a full collection during that window with GC.Collect(2, GCCollectionMode.Forced, true). This resets the Gen 2 budget and pushes the next collection further out. The trick is making absolutely sure the window is really idle; otherwise, you’re creating the exact pause you wanted to avoid.
Process recycling: A blunt but effective technique is to recycle the process before Gen 2 gets too large. In Kubernetes or similar orchestrators, you can set a lifetime limit on pods. If your application can handle a graceful shutdown in a few seconds, a rolling restart keeps heap sizes in check and prevents full collections during peak traffic.
Measuring the Impact
Whatever you try, you have to validate it with production metrics. I instrument every latency-sensitive application with a GC.CollectionCount counter per generation and a histogram of GC.GetTotalPauseDuration() sampled periodically. In .NET 5 and later, the System.Runtime runtime counters exposed through dotnet-counters or Prometheus give you:
gc-pause-timeby generationgc-heap-sizeby generationalloc-ratein bytes per second
Plot those alongside request latency percentiles, and you can see exactly how much of your tail latency traces back to GC pauses. The aim is to drive Gen 2 pause time down near zero during normal operations, with any leftover pauses happening only during off-peak hours.
FAQ
Why does a Gen 2 collection pause all threads?
The GC needs a consistent view of the object graph while it compacts the heap and updates references. To stop threads from changing references mid-collection, the runtime suspends all managed threads at a safe point. With workstation GC, that suspension lasts the whole collection. Server GC with background collection lets threads run during most of the work, but a short suspension is still required at the end.
Can I use GC.TryStartNoGCRegion to prevent Gen 2 collections?
This API, available in .NET Framework and .NET Core, tries to block any GC for a specified amount of allocated memory. It’s meant for very short, critical sections where any GC pause would be a disaster. But it’s narrow: the size has to be small (usually a few megabytes), and if you blow past the budget, the API either fails or forces a full collection. It’s not a general-purpose shield against Gen 2.
What is the difference between workstation and server GC for latency?
Workstation GC uses one heap and aims to minimize pause time for interactive apps. It can use sustained low latency mode to delay Gen 2 collections. Server GC uses multiple heaps and is tuned for throughput, which often means better overall performance but potentially longer individual pauses. For latency-sensitive services, it’s worth testing workstation GC with low latency mode, though server GC is the ASP.NET default and frequently delivers better tail latency in practice when dialed in right.
How do I identify which objects are causing Gen 2 collections?
Grab a memory profiler like dotMemory or PerfView and capture a heap snapshot before and after a Gen 2 collection. Focus on the objects that survived to Gen 2—those are your long-lived allocations. Keep a close eye on the LOH and pinned handles. PerfView’s GC Heap Analyzer can show you the root paths keeping things alive. The fix is usually to shorten the lifetime of those objects or yank them out of the GC heap.