How to Use PerfView Like a Senior Engineer

How to Use PerfView Like a Senior Engineer

PerfView is the lens. When a production service starts burning CPU or leaking memory, I don’t open a memory profiler first. I open PerfView. The raw ETW (Event Tracing for Windows) data it captures separates what you think is happening from what the runtime actually did. This isn’t a getting-started guide. I’m writing for engineers who have clicked around the tool but still feel like they’re fumbling when an incident wakes them at 3 a.m. By the end, the goal is simple: stop guessing, start proving.

Close-up of performance profiling charts on a monitor
PerfView captures high-resolution ETW traces that reveal thread-level timing and allocation patterns.

Setting Up for Production-Grade Collection

Nobody gets excited about trace collection setup. But senior engineers treat it like a biopsy, not a hammer. You want exactly the right data, minimal overhead, and a clean detach. If you start the collector without a plan, you’ll either miss the problem or drown in noise.

Choosing the Right Collection Mode

PerfView gives you three main modes: Collect, Heap Snapshot, and Thread Time. For a CPU-bound issue, Thread Time with the “/CpuSample” flag is the boring, correct choice. For memory investigations, a Heap Snapshot with /ForceGC gives you a steady baseline to compare against. There’s a subtler option tucked into the Collect menu: CPU Sample with GC Only Events. The overhead it adds is tiny, but it still exposes GC pauses in the trace. If you’re hunting a managed memory leak, don’t skip /GCOnly. It shrinks the trace an order of magnitude while keeping allocation ticks intact.

Preparing the Target Process

Before you launch the collector, set up your symbol paths. Use _NT_SYMBOL_PATH to point at your private symbol server or a local cache. I’ve watched a single missing PDB stretch a 10-minute investigation into two hours of dead ends. Also, make sure no other profiler is latched onto the process. ETW sessions conflict, and PerfView will fail silently sometimes. Run logman query -ets and kill any leftover sessions from earlier debugging attempts.

Collecting the Trace with Precision

Clicking through the GUI is fine when you’re poking at a dev box. But production incidents demand a scripted response. Command-line collection is the only way to guarantee you capture the same thing every time.

Command-Line Invocation for CPU Sampling

PerfView.exe /DataFile:HighCPU.etl /AcceptEula /CpuSample /MaxCollectSec:60 /NoView /Process:1234 /StopOnMaxCollectSec /Zip:true collect

That one-liner attaches to process 1234, samples CPU for 60 seconds, stops on its own, and packages the result into a compressed ETL.ZIP file. /NoView keeps the GUI from popping up—irrelevant on a headless server, but you’d be surprised how many people forget it. /Zip:true often squeezes a 600 MB trace down to 40 MB, which matters when you’re pulling it over a throttled VPN.

Adjusting Provider Keywords for Managed Code

The default .NET provider keyword set is moderate. For deeper IL-level work, add /Providers:Microsoft-Windows-DotNETRuntime:0x1F000080018:5. That keyword mask enables GC, JIT, loader, and exception events at informational level 5. It gives you thread time attribution and exception callbacks without filling the trace with verbose ETW chatter. It’s the minimal set I reach for in almost every managed-code investigation.

Engineer analyzing a detailed performance trace on a laptop
Analyzing a collected ETL trace requires methodical filtering to isolate hot paths.

Analyzing CPU Spikes Like a Surgeon

This is where most engineers lose an hour. They open the CPU stacks view and start scrolling. Thousands of methods scroll past, and nothing stands out. Senior engineers filter aggressively from the first click.

Filtering by Time Range and Thread

Go to the Events view. You’ll see the CPU utilization graph. Drag-select the spike—the exact region where CPU was pegged—right-click, and choose Filter to Selected Time Range. You’ve just thrown out the quiet periods that dilute the signal. If the spike is tied to a specific thread pool thread, filter further by ThreadID. Now the call tree shows what the CPU was executing during the incident window, not an aggregate of the entire trace.

Reading the CallTree with the “By Name” Tab

Switch to the CallTree tab and pick By Name. This collapses methods by assembly and namespace. Immediately you’ll see whether the hot path sits in your code or a third-party library. Focus on the Exc % column—exclusive CPU percentage. That’s time spent directly inside the method, not in its children. A method with high inclusive time and low exclusive time is just a dispatcher; the real work is further down. Triage from the highest exclusive percentage downward.

Spotting Async Continuations

Async code makes a mess of call stacks. A Task.Delay continuation might show up under a completely unrelated thread. Use the Thread Time (with Start Stop Activities) view. It reconstructs the logical activity chain using System.Threading.Tasks.TplEventSource events. Without it, you can misattribute 30% of CPU to a thread pool dispatch loop and never see the actual business logic sitting behind an await.

Diagnosing Managed Memory Leaks

Heap size alone is a distraction. A leak isn’t just about how much memory an object uses; it’s about rootedness. An object is a leak only if it’s reachable from a static root or an active thread and shouldn’t be.

Taking and Comparing Snapshots

Take a baseline snapshot with /HeapSnapshot /ForceGC. Exercise the application. Take a second snapshot. Open both in PerfView and use the Diff feature from the Heap Stacks view. It highlights object types whose instance count grew between the two points. Ignore transient types—byte arrays from I/O buffers come and go. You’re looking for types with a steady, monotonic increase.

Tracing Roots with the “Ref Tree”

Found a suspicious type? Open the References view and expand the Ref Tree. Start from the Static vars root or the Thread root and walk the reference chain. What you’re looking for is the unexpected holder—maybe a static ConcurrentDictionary that accumulates entries with no eviction policy. The reference graph is directed; PerfView renders it as a tree by picking one path per object, so verify a few times so you’re not staring at a coincidental reference.

Server room with blinking lights symbolizing memory pressure
Heap snapshot analysis reveals root causes that heap size alone cannot explain.

Correlating GC Pauses with Application Latency

GC pauses are the silent latency killers in managed applications. PerfView’s GCStats view logs every garbage collection during the trace, including pause duration and generation.

Interpreting GCStats Output

Open GCStats and sort by Pause MSec. You’re hunting for full blocking Gen 2 collections with pauses above 100 ms. These often point to Large Object Heap pressure. If you see frequent Gen 2 collections with high pause times, drill into the Heap Snapshot you took during the trace and examine the LOH for arrays larger than 85,000 bytes that are pinned or have long lifetimes.

Using the “GC Reason” Column

The Reason column tells you why the GC fired. AllocSmall means a small-object allocation hit the budget. Induced means someone called GC.Collect explicitly. An excessive number of induced GCs in production is a code smell—usually a well-intentioned attempt to manage memory that backfires by promoting objects to higher generations too early. Trace the call stack of the induced GC through the Any Stacks view to find the method responsible.

Advanced Analysis with Custom Views

PerfView’s built-in views handle 90% of what you’ll encounter. The remaining 10% forces you to write custom queries over the ETW data. That’s the moment the tool stops being a profiler and becomes a data platform.

Using the PerfViewExcel Extensions

Export trace data to XML and pull it into Excel with the PerfViewData.xlsx template. It lets you build pivot tables that aggregate CPU time by namespace, module, and time window simultaneously—something the native views won’t do. For example, pivot on Module!Namespace over 10-second intervals to catch a background service whose CPU consumption climbs gradually. A static call tree would average that rise into invisibility.

Querying with LINQPad and TraceEvent

For programmatic access, reference Microsoft.Diagnostics.Tracing.TraceEvent in a LINQPad script. Open an ETL file and write LINQ queries against the events:

var source = new ETWTraceEventSource("trace.etl");
var clr = new TraceEventParser(source);
clr.GCStart += (data) => {
    if (data.Depth >= 2) Console.WriteLine($"Gen2 GC at {data.TimeStamp}");
};
source.Process();

That script prints a timestamp for every Gen 2 GC. Extend it with HTTP request events from Microsoft-Windows-HttpService and you can build a latency impact report the GUI can’t generate. I keep a library of these scripts because a one-off investigation today becomes a repeatable diagnostic tomorrow.

Automating Trace Collection for CI/CD Pipelines

PerfView isn’t just a desktop tool. It can sit inside a performance regression suite. A nightly build that runs a benchmark and collects a PerfView trace catches CPU regressions before they ever hit production.

Headless Trace Collection in a Test Runner

Wrap your test in a batch script that starts PerfView collection beforehand and stops it afterward. The /StopOnPerfCounter flag lets you stop collection automatically when a performance counter crosses a threshold:

start /B PerfView.exe /DataFile:Benchmark.etl /CpuSample /StopOnPerfCounter:Process("%ProcessorTime","MyApp")>50 /NoView collect

That command watches the “%ProcessorTime” counter for “MyApp” and stops collection the moment it drops below 50%—meaning the benchmark ended. Pair it with a script that triggers PerfView’s HeapSnapshotFromProcess action to capture final memory state. The resulting ETL and GC dump files become build artifacts you can diff against the previous build.

Common Pitfalls and How to Avoid Them

Even people who’ve used PerfView for years fall into traps that waste time or produce garbage data. Here are three I see in post-mortems over and over.

Collecting Too Much Data

Enable every ETW provider at verbose level and you’ll get a trace so big PerfView can’t open it—or it takes 20 minutes to load. The temptation to “capture everything just in case” is strong, but a trace that crashes your analysis tool has zero value. Start with the minimal provider set, look at the data, and add providers only if gaps appear.

Ignoring Symbol Loading Failures

When PerfView shows ?!? instead of method names, stop. Don’t proceed. Fix the symbol path, reload the trace, and verify your own assemblies resolve. If third-party assemblies won’t resolve, try /NGenSymbols to include precompiled native images. A call stack full of question marks isn’t a basis for any conclusion.

Misreading Inclusive vs. Exclusive Time

Inclusive time lies. A method that calls Thread.Sleep can show high inclusive CPU if the sampling interval catches the sleep, but exclusive time will hover near zero. When you’re dealing with I/O-bound or waiting operations, cross-reference the Wall Clock Time view. CPU time and wall clock time diverge sharply in those scenarios.

FAQ

Why does PerfView show “Broken Stacks” in my trace?

Broken stacks happen when ETW can’t unwind the call stack. The usual suspects are missing symbols, optimized native code without frame pointers, or a corrupted stack. Fix it by ensuring your symbol path includes Microsoft public symbols (srv*C:\Symbols*https://msdl.microsoft.com/download/symbols) and enable /NGenSymbols if you’re using NGEN’d assemblies. In some cases, disabling “Suppress JIT Optimization” in your build and collecting with /ClrStackWalk improves stack resolution.

Can PerfView collect traces on Linux?

PerfView itself is a Windows application and relies on ETW, which is Windows-specific. But the trace analysis engine (TraceEvent) can open traces collected on Linux with perfcollect or dotnet-trace, as long as they’re converted to ETL format. For .NET Core applications on Linux, collect with dotnet-trace, then analyze the .nettrace file using PerfView on a Windows machine, or run dotnet-trace convert to change the format.

How do I reduce the size of a PerfView trace without losing critical data?

The most effective method is to limit ETW providers and keywords to exactly what you need. For CPU analysis, the default /CpuSample provider set is already minimal. For memory, use /GCOnly instead of a full heap snapshot unless you need object references. Also, use /MaxCollectSec to limit duration and /BufferSizeMB to set a smaller buffer (e.g., 64 MB) if the system is memory-constrained. Post-collection, /Zip:true compresses the trace significantly without losing fidelity.

Debugging .NET Exceptions That Disappear Into the Void

Debugging .NET Exceptions That Disappear Into the Void

An exception handling contract in .NET seems simple on paper. Something goes sideways, an exception object gets born, the runtime fills in a stack trace, and the thing gets thrown. You catch it, log it, and either recover or fail without trashing the process. Nice and tidy. Except sometimes the contract just breaks. The exception evaporates. No log line. No dump. No crash. Just a process that quietly stops doing whatever it was supposed to do. I call these exceptions that disappear into the void, and chasing them down takes a methodical, low-level approach that goes miles past ordinary try-catch blocks.

I want to walk you through the technical reasons why .NET exceptions can get swallowed without a whisper, the runtime internals that make it possible, and the debugging moves—WinDbg, SOS, PerfView—that let you resurrect these lost errors. If you work on high-reliability systems, distributed architectures, or codebases with layers of exception-handling anti-patterns, you’ll hit this problem sooner or later. So let’s look at the root causes and the exact steps I use to fix them.

Abstract digital debugging visualization with glowing lines and code fragments
The invisible failures in your .NET application often leave no trace in standard logs.

The Anatomy of a Vanished Exception

To understand why an exception can vanish, you need a clear picture of what happens when one gets thrown. The CLR walks the stack, hunting for a matching catch block. If it finds one, control transfers. If it doesn’t, the exception becomes unhandled. The default behavior depends on the .NET version and the application model—maybe the AppDomain.UnhandledException event fires, maybe the Windows Event Log gets a write, maybe the process just terminates. But between “thrown” and “handled” (or “unhandled”) there are a handful of dark corners where an exception can get absorbed silently.

Empty Catch Blocks: The Classic Culprit

The most obvious cause is the empty catch block, often written with good intentions that lead to awful results:

try
{
    RiskyOperation();
}
catch (Exception)
{
    // Swallow silently
}

This pattern shows up everywhere in codebases where developers were afraid of crashes, wanted to keep a loop spinning, or just had no idea what to do with the exception. The outcome is a silent failure that can leave the application in a weird, inconsistent state. Static analysis tools—Roslyn-based analyzers or SonarQube—catch a lot of these, but not all of them. Especially when the catch block has a comment like // Logged elsewhere that turns out to be wishful thinking.

Thread Pool and Task Parallel Library Swallowing

The ThreadPool and the Task Parallel Library are common sources of lost exceptions. When you queue work to the thread pool with ThreadPool.QueueUserWorkItem, an unhandled exception inside the callback usually won’t crash the process on .NET Framework. The thread pool’s internal plumbing catches the exception and tosses it aside. The thread goes back into the pool, ready for the next work item, and you never get a hint that anything failed.

The TPL has its own version of this problem. An unobserved Task exception doesn’t surface right away. Before .NET 4.5, these exceptions got caught and stored until the task was garbage collected; then the TaskScheduler.UnobservedTaskException event fired. Starting with .NET 4.5, the default changed: unobserved task exceptions are silently swallowed and the event doesn’t even get raised unless you flip a configuration switch to enable it.

Code visualization with multiple parallel execution paths and one highlighted as failing quietly
Parallel execution paths can mask exceptions that occur off the main thread.

Async Void Methods: The Black Hole of Exceptions

If you’ve ever written an async void method—maybe as an event handler—you’ve built a direct pipe into the void. When an exception gets thrown inside an async void method, no Task captures it. Instead, the exception is posted straight to the SynchronizationContext that was active when the method started. If that context is the UI thread’s context (WPF or WinForms), the exception might crash the application. But if the context is the thread pool’s default context—or no context at all—the exception gets raised on a thread pool thread and then swallowed. The CLR has no way to propagate it back to the caller because no caller is waiting on a task.

This is exactly why the rule “avoid async void” exists. The only reasonable use is top-level event handlers, and even then you have to wrap the whole body in a try-catch that logs the exception and does something sensible to recover.

Finalizers and Dispose: Silent Failures in Cleanup

Exceptions thrown inside finalizers (destructors) are a guaranteed recipe for making an exception disappear. The CLR treats an exception from a finalizer as a catastrophic failure of the finalization process and aborts whatever the finalizer thread was doing. It does not propagate the exception. In .NET Framework, the runtime catches the exception and the finalizer thread moves on to the next object—behavior that can hide resource leaks or corrupted state. In .NET Core and .NET 5+, the process ends instead, which is a safer default but still surprising when you expected a log entry.

Exceptions inside Dispose methods can get swallowed too, especially when dispose runs from a using block while another exception is already in flight. The using statement is just syntactic sugar for a try-finally block. If an exception happens in the try block and another one fires in the finally block during dispose, the original exception disappears unless you explicitly handle the dispose exception.

Diagnosing the Void: Tools and Techniques

When an exception disappears, standard logging and crash dumps won’t save you because there’s no crash and no log. You need to intercept the exception at the runtime level or reconstruct the failure from side effects. Here are the techniques I lean on in production and dev environments.

First-Chance Exception Handling in WinDbg

WinDbg can break on first-chance exceptions—meaning it stops as soon as an exception is thrown, before any catch blocks run. That’s the most direct way to see exceptions that would otherwise get swallowed. Attach WinDbg to your process and use these commands:

.loadby sos clr
sxe clr
sxi clr

Or, more precisely, you can break on a specific exception type:

sxe -c "!PrintException; gc" System.NullReferenceException

That tells the debugger to break when a NullReferenceException gets thrown, print the exception object with SOS’s !PrintException, and then continue execution (gc). You can tweak the command to log to a file, capture a mini-dump, or just watch. This technique is a lifesaver for catching swallowed exceptions in thread pool work items, async void methods, and finalizers.

When you can’t attach a debugger to a production process, reach for EventPipe or dotnet-trace to capture exception events. The Microsoft-Windows-DotNETRuntime provider emits an ExceptionThrown_V1 event for every managed exception. Collect those traces and sift through them offline with PerfView.

Using PerfView to Capture Silent Exceptions

PerfView is a performance analysis tool that also grabs ETW (Event Tracing for Windows) events from the .NET runtime. To collect all managed exceptions—including the ones that get caught and discarded—run this from the command line:

PerfView.exe /Providers=*Microsoft-Windows-DotNETRuntime:ExceptionKeyword collect

After you’ve got the trace, open it in PerfView and head to the “Events” view. Filter for Exception events. Each record gives you the exception type, message, the stack trace at the point of throw, and—if the exception was caught—the handler’s stack. By looking for exceptions that have no matching catch event (or whose catch lives in a system assembly like System.Threading.ThreadPoolWorkQueue), you can spot the swallowed ones.

Performance analysis dashboard with exception graphs and event timelines
ETW traces reveal the complete lifecycle of every exception, caught or uncaught.

Configuring the Runtime to Surface Hidden Exceptions

Several runtime configuration knobs can force exceptions to surface instead of getting quietly dropped. In .NET Framework, you can enable legacyUnhandledExceptionPolicy in your app.config to make unhandled exceptions on thread pool threads crash the process:

<configuration>
  <runtime>
    <legacyUnhandledExceptionPolicy enabled="1"/>
  </runtime>
</configuration>

In modern .NET, subscribe to TaskScheduler.UnobservedTaskException and configure it to crash the process. You can also set ThrowUnobservedTaskExceptions in the runtime configuration to get the pre-.NET 4.5 behavior back. For async void exceptions, there’s no global setting; you must wrap each async void method body in a try-catch block.

Proactive Prevention: Writing Void-Resistant Code

Debugging is reactive by nature. A better path is designing your code so exceptions can’t disappear in the first place. That takes discipline in exception handling, async patterns, and resource management.

Centralize Exception Logging with Global Handlers

Register handlers for AppDomain.UnhandledException, TaskScheduler.UnobservedTaskException, and Application.DispatcherUnhandledException (WPF) or Application.ThreadException (WinForms). These handlers need to log the exception and, depending on severity, either attempt recovery or shut things down. Don’t try to keep running normally after an unhandled exception unless you fully understand the corrupted state that might be lurking.

Enforce Catch-Block Standards with Roslyn Analyzers

Use Roslyn analyzers to lock down rules like:

  • Every catch block must log the exception (at minimum).
  • Empty catch blocks are forbidden.
  • Catching Exception requires a justification, either via a comment or an attribute.

That catches most swallowed-exception anti-patterns at build time. Pair it with code review checklists that explicitly look for missing exception propagation.

Adopt Async Task Over Async Void

Treat async void as a code smell. Swap it for async Task whenever you can. For event handlers that have to be void, use a safe wrapper that invokes the async handler and catches exceptions:

public async void OnButtonClick(object sender, EventArgs e)
{
    try
    {
        await HandleClickAsync();
    }
    catch (Exception ex)
    {
        Logger.Log(ex);
        // Show user-friendly message or recover
    }
}

That guarantees no exception from the async workflow can escape into the void.

Guard Finalizers and Dispose

Never let an exception propagate out of a finalizer. Wrap the whole finalizer body in a try-catch that logs and swallows (there’s no better option at that point). For Dispose methods, follow the standard pattern with a disposing flag and avoid throwing exceptions from Dispose entirely. If cleanup fails, log it and move on; the object is getting discarded anyway.

FAQ

Why do exceptions in Task.Run sometimes get lost?

Exceptions thrown inside a Task.Run delegate are captured by the returned Task object. If that task isn’t awaited, isn’t stored, and isn’t observed via Wait() or Result, the exception becomes an unobserved task exception. In .NET 4.5 and later, unobserved task exceptions are silently swallowed by default. Always await or explicitly handle tasks you create.

How can I find empty catch blocks in a large codebase?

Static analysis tools like SonarQube, the built-in Roslyn analyzers in Visual Studio (rule CA1031), or a custom analyzer that flags catch blocks with empty bodies or only a comment all work. You can also run regex searches across the codebase for patterns like catch\s*\([^)]*\)\s*\{\s*\}.

Does .NET Core handle swallowed exceptions differently from .NET Framework?

Yes. In .NET Core and .NET 5+, unhandled exceptions on thread pool threads and in finalizers are more likely to crash the process, which is a safer default. Unobserved task exceptions are still swallowed by default, but you can configure the runtime to throw them. The behavior of async void exceptions stays the same—they get posted to the SynchronizationContext and may be swallowed if the context doesn’t handle them.

What is the best tool for catching exceptions in production without a debugger?

For production diagnostics, use dotnet-trace or PerfView to collect ETW events. The Microsoft-Windows-DotNETRuntime provider with the ExceptionKeyword captures every managed exception with stack traces. It’s low-overhead and doesn’t need a debugger attached. You can also use Application Insights or OpenTelemetry with exception tracking turned on.

Debugging silent failures is one of the hardest corners of .NET development. Once you understand the runtime internals and get comfortable with first-chance exception tools, you can drag those lost exceptions back into the light. Stay precise, trust the debugger, and never let an exception escape your scrutiny.

Inside PerfView: Diagnosing Real-World .NET Performance Problems Like a Seasoned Engineer

When a production incident hits and the CPU graph jumps past 95%, junior engineers open Task Manager, guess at the problem, and restart the service. Senior engineers reach for PerfView. Not the GUI-first, click-through version — the one you drive from the console, with custom collection flags, symbol servers configured ahead of time, and a working mental model of the CLR’s runtime internals. This is that PerfView.

Close-up of server rack LED indicators blinking in a dark data center
Production diagnostics demand precision — PerfView delivers it when configured correctly.

Why the Default Collection Is a Trap

Double-clicking PerfView.exe and hitting “Collect” with whatever providers are ticked by default — that’s the first mistake I spot even in mid-level engineers. The default kernel and CLR providers pull in so much noise that the .etl file balloons to gigabytes, and the trace processing time outlasts the incident window. You have to be specific about what you ask for.

Start from the command line. Always. The GUI is fine for an initial poke around a trace, not for collection. Use the /DataFile and /Providers flags to narrow the scope. For a high-CPU investigation, I typically disable kernel events beyond Profile and ContextSwitch, and I specify CLR providers at the Informational level only for the modules I actually suspect.

PerfView.exe /DataFile:HighCPU.etl /BufferSizeMB:256 /Providers:"Microsoft-Windows-DotNETRuntime:0x1:5,Microsoft-Windows-DotNETRuntimeRundown:0x1:5" collect

That provider GUID specifies only GC, JIT, and Exception events at the verbose level. No loader callbacks, no interop stubs, no assembly resolve noise. The trace stays small, focused, and parseable in under 30 seconds.

Symbol Resolution: The Divide Between Guesswork and Certainty

If your call stacks show [Unknown] or generic mscorlib.ni.dll entries, you’re flying blind. Senior engineers set up a symbol server path before collection, not after. The /SymbolPath parameter tells PerfView where to find PDBs, and it must include both your build server’s symbol store and the Microsoft public symbol server.

I keep a small script that maps a network drive to the latest release symbols and appends the Microsoft URL. The syntax is semicolon-delimited:

/SymbolPath:\"\\build-server\symbols;https://msdl.microsoft.com/download/symbols"

When the trace loads, every frame resolves. You see the exact method name, the source file, and the line number. No guesswork. This alone separates a useful trace from a waste of disk space.

Filtering CPU Stacks to the Thread That Matters

Open the trace in the GUI, navigate to the CPU Stacks view, and immediately group by process and thread. The default flame graph aggregates all threads, which hides the one thread spinning at 100%. Right-click the thread column, select Filter to Selection, and now the flame graph shows only that thread’s activity.

Look for flat tops — methods with no children. Those are the CPU consumers. Often you will see a while loop inside an IDisposable implementation that never yields. The call stack tells the story: MyService.Dispose() → SpinWait.SpinOnce() → kernel wait. That is a bug, not a feature.

Software engineer analyzing code on multiple monitors with debugging tools open
Interpreting a PerfView flame graph requires understanding the runtime’s execution model.

GC Pauses: The Silent Latency Killer

High memory traffic doesn’t always show up as CPU spikes. It shows up as GC pause time. In PerfView, the GCStats view gives you the total pause duration per generation. If gen-2 collections are taking 200ms on a server that should respond in under 50ms, you have a real problem.

Drill into GC Heap Alloc Ignore Free to see which types allocate the most. Sort by Alloc MB. I once found a logging wrapper that allocated a new StringBuilder on every call, totaling 80MB per minute under load. The fix was a pooled instance. No profiler needed — just PerfView and a careful read of the allocation graph.

Cross-Referencing with ETW Events

The Events view is underused. Filter to Microsoft-Windows-DotNETRuntime/GC/Start and /Stop events, and export the timestamps. Plot them in Excel alongside your latency metrics. The correlation is often striking: every spike in request duration aligns precisely with a gen-2 collection. Then you know the fix isn’t in the code path itself but in the object lifetime management.

Contention and Blocking: The Lock That Broke Production

PerfView captures Monitor.Enter contention by default if you include the right CLR keywords. In the .NET ThreadPool view, look for threads in the Waiting state with a Monitor wait reason. The stack trace shows the lock location.

I diagnosed a four-thread deadlock in under ten minutes using this technique. The classic pattern: Thread A holds lock X and waits for lock Y; Thread B holds lock Y and waits for lock X. PerfView’s Thread Time view, sorted by wait time, gave me the exact four stacks. No magic, just methodical filtering.

Memory Dumps vs. Traces: When to Switch Tools

PerfView is not a memory dump analyzer. If the issue is a slow leak over hours, you need a series of dumps compared in dotMemory or WinDbg+SOS. PerfView excels at active analysis — what is happening right now, in the running process. Use it to capture the burst, then switch to a dump for the snapshot.

I combine them: start a PerfView collection, wait for the spike, stop it, then take a dump with procdump immediately after. The trace shows the dynamic behavior; the dump shows the state. Together, they form a complete picture.

Server hardware with glowing indicators and detailed circuit boards
A well-configured PerfView trace captures the exact moment of failure without disrupting the production service.

Automating PerfView in CI/CD for Regression Detection

Senior engineers don’t wait for production. They integrate PerfView into the performance test pipeline. A short five-second trace during a load test can capture GC pause times and allocation rates automatically. Write a script that invokes PerfView.exe with the /NoGui and /AcceptEula flags, parses the XML output, and fails the build if gen-2 pauses exceed a threshold.

PerfView.exe /NoGui /AcceptEula /DataFile:LoadTest.etl /Providers:... collect
PerfView.exe /NoGui /AcceptEula /DataFile:LoadTest.etl GcStats > GcStats.txt

Parse GcStats.txt for “Gen2 Pause Time Mean” and assert. This catches allocation regressions before they reach staging, let alone production.

FAQ

Why do my call stacks show [Broken] even with symbols configured?

[Broken] typically indicates a mismatch between the trace’s recorded module version and the PDB version. Make sure the symbol path points to the exact build artifacts deployed to the profiled machine. Also verify that the _NT_SYMBOL_PATH environment variable isn’t overriding your PerfView argument.

How do I reduce the trace file size without losing critical events?

Use provider filter keywords aggressively. For GC-only analysis, specify 0x1 for GC events instead of the default 0x1F. Increase the circular buffer size (/BufferSizeMB) to 1024 and set /CircularMB to a reasonable limit. This drops old events when the buffer fills, keeping the file size constant while preserving the most recent activity.

Can PerfView collect on a production server with minimal overhead?

Yes, if configured carefully. Use the /KernelEvents=Default flag to limit kernel providers, set CLR providers to 0x1:4 (Informational level for GC only), and reduce the sampling interval with /SampleIntervalMS=100. This lowers CPU overhead to under 2% in typical scenarios, making it safe for short production collections.

What is the difference between CPU Samples and CPU Stacks?

CPU Samples shows a flat list of methods with sample counts, useful for ranking hot methods. CPU Stacks shows a hierarchical call tree, allowing you to trace the exact path to the hot method. Always start with CPU Stacks grouped by thread to understand the context, then switch to CPU Samples for a quantitative summary.

Mastering PerfView for Production Diagnostics: A Senior Engineer’s Guide

Why Most Engineers Under-Use PerfView

PerfView is not simply a performance profiler. It is a forensic toolkit for .NET production incidents, designed by the .NET runtime team to expose what is happening at the memory, thread, and CPU level. Many developers, even those with several years of .NET experience, reach for PerfView only when they need a quick CPU trace or a managed memory dump analysis. The tool is capable of far more than that. When I mentor engineers, I do not teach them just the commands. I teach them a systematic approach: how to capture data without destabilizing a live server, how to interpret the output, and how to present findings that leave no room for speculation.

This guide is not a beginner tutorial. I assume you already understand basic performance concepts, ETW events, and the difference between managed and native allocations. What I will show you is the workflow I use when I get a call at 3 a.m. because a production service is consuming 100% CPU or the GC pause time has jumped from 200 ms to 3 seconds. I will show you how to use PerfView the way a senior engineer uses it: with precise commands, a critical eye on overhead, and an understanding of the underlying runtime mechanics.

Server rack with blinking lights representing production environment

Building a Reliable Collection Strategy

Before you look at a single stack trace, you must plan the collection. The biggest mistake I see is engineers running high-overhead traces on a production server during peak load and then being surprised when the service tips over. PerfView can collect kernel events, .NET runtime events, and sample call stacks with minimal impact—if you configure it correctly.

Choosing the Right Provider Set

The default collection is often too broad or too narrow. For CPU investigations, I start with CPU Samples and enable only the Microsoft-Windows-DotNETRuntime provider with the GC and Contention keywords if I suspect lock issues. Avoid enabling the JIT keyword unless you specifically need to see method compilation overhead. The JIT events are verbose and can add measurable pressure on the ETW session. The exact command I use looks like this:

PerfView.exe /DataFile:HighCpuTrace.etl /BufferSizeMB:256 /CircularMB:2000 /Providers:Microsoft-Windows-DotNETRuntime:0x1:5 collect

The keywords 0x1 give you GC events only. The 5 is the verbosity level, which filters out unnecessary payload. If I need thread time or context switch data, I also add the kernel provider with Loader and ThreadTime flags, but I always measure the impact in a staging environment first. A senior engineer knows that a trace that is too large becomes unreadable anyway.

Managing Trace Size and Duration

Circular buffers are essential for capturing intermittent problems. I set /CircularMB to at least 2000 for a service handling hundreds of requests per second. The buffer wraps; when you stop the trace, you get the last N megabytes of events. This lets you leave PerfView running for hours with negligible overhead, then trigger a stop when the issue occurs. For high-memory services, I pair this with a low /BufferSizeMB to avoid ETW session allocation failures. The balance is always between event loss and system impact.

For heap snapshots, the rules change. A full heap dump with GCCollectOnly or a forced GC via PerfView’s HeapSnapshot provider will block managed threads. I never run this on a production node without first draining traffic. The trace itself must be short—30 seconds is usually sufficient—and I warn the on-call team that a pause will occur. There is no magic: a heap snapshot triggers a full blocking GC, and that is a design decision, not a bug.

Close-up of code on a monitor during debugging session

Analyzing CPU Traces with Surgical Precision

Once the ETL file is on your machine, the real work begins. I open the trace in PerfView and immediately go to the CPU Stacks view. The default grouping by module is nearly useless for managed code. I switch to the By Name grouping and set the time range to exclude the startup and shutdown periods. What I want is a flat cost table that shows me inclusive and exclusive CPU time for each method, filtered to the time window where the problem was observed.

Interpreting the Flame Graph

PerfView’s flame graph is not a toy. I expand nodes from the bottom up, looking for broad plateaus where one method dominates the sample count. If I see System.String.Concat taking 40% of samples inside a request handler, I do not immediately blame string concatenation. I look at the calling context. Often, the real problem is a loop that builds a large string for logging or serialization, and the fix is to use a StringBuilder or restructure the log pipeline. The flame graph shows you the relationship; the flat list shows you the cost. You need both.

A senior engineer also knows when the trace itself is lying. ETW sampling is statistical. If you have many short-lived threads, a method that runs for 1 ms but is called millions of times might be under-sampled. I always cross-reference PerfView CPU data with dotnet-counters or Windows Performance Monitor metrics collected during the same window. The CPU percentage reported by the OS should roughly align with the trace’s sampled busy time. If they diverge significantly, your sampling interval is too coarse, or kernel overhead is masking the signal.

Spotting Hidden Bottlenecks: GC and JIT

CPU traces often hide GC pauses because the GC runs on dedicated threads that may not be sampled in the same way. I always open the GCStats view immediately after a CPU analysis. A sudden increase in % Time in GC from 5% to 45% tells you more than any stack trace. If Gen 2 collections are frequent, I then take a heap snapshot to find the pinning or large object heap fragmentation causing them. Do not tune code if the GC is the bottleneck. Tune allocations.

Similarly, if I see high CPU in clr!ThePreStub or JIT-related stacks, it means methods are being compiled at runtime under load. The fix is not always to enable background JIT; sometimes you need to run a warm-up script or use ReadyToRun images. PerfView’s JITStats view gives you the exact method names and compilation times. This is data you can take to the build team.

Engineer analyzing performance graphs on dual monitors

Memory Diagnostics Without Guessing

Memory leaks are the most misdiagnosed problems in .NET. A senior engineer does not guess. We collect a heap snapshot with PerfView using the /GCCollectOnly flag, which forces a full GC before the snapshot so we see only live objects. Then we open the Heap Snapshot view and sort by Inclusive Size.

Rooting Paths, Not Just Object Counts

The object list is distracting. What matters is the reference graph. I select a suspicious type—say, System.Byte[] taking 800 MB—and click Open in GC Heap Explorer. From there, I pick a random large instance and trace its roots back to a static field, an event handler, or a pinned Gen 2 segment. The tool shows me the full chain: StaticVar -> List -> Byte[]. That is the leak. Fixing it means breaking that chain, usually by nulling out the static reference after use or unsubscribing from an event.

For managed memory, I ignore shallow size. A 24-byte object that holds a reference to a 2 MB array is the real problem. PerfView’s Inclusive Size column accounts for this. I also filter by generation. Objects in Gen 2 that are not supposed to be there—configuration data, cached responses—indicate a leak that survives collections. The Gen 2 Object Count trend over multiple snapshots is a dead giveaway.

Native Memory and VirtualAlloc

Managed memory is only part of the story. When the private working set grows but managed heap size is stable, I switch to the Native Memory view. PerfView can show you allocations from VirtualAlloc by call stack, provided you collected kernel events. I look for stacks ending in System.Net.Http or Socket buffers, which often point to pinned object arrays that the GC cannot move. The fix might be to pool buffers or reduce the send/receive buffer sizes at the socket level. PerfView gives you the exact allocation size and count, so you can calculate the potential savings.

Advanced Diffing and Baseline Comparisons

One trace is a story; two traces are evidence. When I suspect a regression, I capture a baseline trace from a known-good build and a comparison trace from the suspect build, using identical collection settings and load patterns. PerfView’s Diff feature lets me subtract one trace from another, showing which methods increased in CPU time or which types grew in memory.

Automating Comparisons with the Command Line

The GUI is fine for exploration, but for repeatable analysis, I script PerfView. The command PerfView.exe /Diff baseline.etl regression.etl /Out:diff.xml produces an XML report I can parse in a CI pipeline. I focus on methods with a delta greater than 5% and an absolute inclusive time over 100 ms. These thresholds filter out noise. The report becomes a checklist for the developer who introduced the change. No opinions, just data.

For memory, the diff shows you which types increased in count and total size. I once caught a leak of SemaphoreSlim objects this way. The diff showed 50,000 more instances in the regression trace than in the baseline. The root cause was a missing Dispose in a fire-and-forget task. Without the diff, the object count looked normal because the process had been running for days and the absolute numbers were large. The baseline made the leak obvious.

Production-Safe Practices and Team Integration

Collecting data is the technical part. Acting on it without causing an outage is the engineering part. I never run PerfView on a production server without a pre-agreed rollback plan. The collection commands are stored in a runbook, reviewed by the ops team, and tested on a canary node. The trace files are pulled immediately after collection and analyzed offline. Leaving an ETW session running for days because someone forgot to stop it is a common and dangerous mistake.

I also insist that every PerfView analysis ends with a concise report: the trace duration, the collection flags, the key findings, and the recommended change. Screenshots of the flame graph are not enough. The report must link to the exact stack frames and object types, so any engineer can reproduce the analysis. When I train teams, I make them write this report before they touch any code. It forces clarity.

FAQ

What is the difference between CPU Samples and CPU Stacks in PerfView?

CPU Samples collects ETW sampling events at a configurable interval (default 1 ms) and shows you a flat list of methods with inclusive and exclusive time. CPU Stacks is the flame-graph visualization of the same data, showing call hierarchies. Use the flat list for cost ranking and the flame graph for understanding the call path that led to the cost. They are two views of the same underlying trace data.

How can I collect a trace without blocking the finalizer thread?

When you enable the GCCollectOnly provider for a heap snapshot, PerfView triggers a full blocking garbage collection across all generations. This suspends managed threads, including the finalizer, until the collection completes. To avoid blocking, do not use GCCollectOnly in production. Instead, use the HeapSnapshot provider without the forced GC flag, which takes a snapshot of the current heap state without inducing a collection. The trade-off is that you will see dead objects that have not yet been collected, which can complicate the analysis.

Why does my PerfView trace show high CPU in ntoskrnl.exe and not in my managed code?

High kernel CPU (ntoskrnl.exe or ntdll.dll) typically indicates that your application is spending time in Windows system calls rather than managed code. Common causes include excessive context switching due to thread contention, heavy file I/O, or socket operations that are not async. Use the Thread Time and Context Switch views in PerfView to correlate the kernel stacks with your managed threads. You may find that a seemingly innocent File.ReadAllBytes call is triggering synchronous I/O that burns kernel CPU.

Can PerfView analyze dumps from Linux containers?

PerfView can open managed memory dumps (`.dmp` files) collected on Linux via `dotnet-dump` or `createdump`, provided the dump is in the Windows minidump format. You must have the correct version of `mscordaccore.dll` and `mscordbi.dll` from the target runtime accessible to PerfView. The tool will prompt you for the DAC path if it cannot locate them automatically. Symbol resolution may require manual configuration of `_NT_SYMBOL_PATH`. CPU traces (ETL files) are Windows-only; for Linux, use `perf` or `dotnet-trace` and convert the format if needed.

Why Gen 2 Garbage Collections Kill Latency-Sensitive Apps

A managed runtime like .NET comes with a garbage collector that promises to handle memory so you don’t have to. For a typical line-of-business app, that deal works just fine. But when you’re building systems where response time has to stay below a hard ceiling—trading engines, real-time bidding platforms, multiplayer game servers, high-frequency telemetry ingestion—the GC can turn into your biggest liability. The real problem isn’t the small, frequent Gen 0 or Gen 1 collections. It’s the full, blocking Gen 2 collection that tears up your latency guarantees.

I’ve spent years profiling production .NET applications, often under punishing load, and one pattern keeps surfacing: Gen 2 collections introduce pauses that are orders of magnitude bigger than a typical request budget. Knowing exactly why this happens—and what you can realistically do about it—is what separates a system that bends from one that snaps when the pressure peaks.

Server racks in a data center with blinking lights, representing high-performance computing environments where latency is critical
Latency-sensitive systems demand predictable execution, not just high throughput.

The Generational Hypothesis and Its Limits

The .NET GC is a generational collector. It’s built on a simple observation: most objects die young. Gen 0 is where newly allocated objects land, and it gets cleaned up often with a fast, stop-the-world pause—usually microseconds to a few milliseconds. Objects that make it through a Gen 0 collection get promoted to Gen 1, a kind of staging area before the long-lived Gen 2. Gen 1 collections are still pretty quick. Gen 2, though, is the attic. That’s where objects that have survived multiple rounds live—static data, long-lived caches, pooled buffers, anything that sticks around for the application’s lifetime.

When a Gen 2 collection kicks in, the GC has to walk the entire managed heap, across all generations. In workstation GC mode, this is a fully blocking affair: every application thread gets suspended until the collection finishes. Server GC mode runs the collection on dedicated GC threads, but managed threads still get paused during certain phases. The pause time grows with the size of the Gen 2 heap and the tangle of your object graph. With a heap that’s several gigabytes, a full Gen 2 collection can easily eat up hundreds of milliseconds, sometimes more than a second. For an API that has to answer in under 50 milliseconds, that’s a lifetime.

What Triggers a Gen 2 Collection

The GC fires a Gen 2 collection under a handful of conditions:

  • Gen 2 budget exhaustion: Each generation has a budget, a threshold that, when crossed, triggers a collection. For Gen 2, that budget adjusts itself over time, but if the heap keeps ballooning, a full collection becomes a matter of when, not if.
  • Large object heap (LOH) allocation pressure: Objects over 85,000 bytes go straight to the LOH, which only gets cleaned during Gen 2 collections. So if you’re churning through big arrays or strings, you’re forcing frequent full collections.
  • Explicit calls to GC.Collect(): Any code that demands a full collection—including a few libraries that quietly call GC.Collect() behind the scenes—short-circuits the generational logic and slaps you with a full pause.
  • Low system memory: When the OS reports memory pressure, the GC might run a full collection to trim the working set, even if Gen 2 isn’t really full yet.

Each of these triggers can fire at seemingly random moments from the viewpoint of an incoming request. The pause lands right in the middle of something time-sensitive, blowing out your 99th percentile latency and carving a sawtooth pattern into your response-time graphs.

Close-up of a performance monitoring dashboard showing latency spikes and sawtooth patterns
A single Gen 2 collection can push tail latency far beyond acceptable thresholds.

How Gen 2 Pauses Manifest in Production

The damage from a Gen 2 collection isn’t just the raw pause time. When every thread gets suspended, requests pile up. The moment threads wake up, they have to chew through that backlog, which spikes CPU and drags out latency even further. That cascade often turns a 200-millisecond GC pause into a multi-second brownout.

Think about a trading system handling 10,000 orders per second. A 500-millisecond Gen 2 collection stops the world. In that half-second, 5,000 orders stack up. When processing kicks back in, the system scrambles to catch up, but that sudden burst of work can trigger yet more allocations—and maybe another collection sooner than you’d like. This feedback loop can destabilize the whole service.

One investigation I remember: a market data gateway kept failing at regular intervals. We traced it to a Gen 2 collection caused by a logging library that was allocating big temporary byte arrays for serialization. Those arrays landed on the LOH. After minutes of steady traffic, the LOH filled, forcing a full GC. The pause blew past the heartbeat timeout for downstream consumers, and they dropped the connection. The fix wasn’t some clever GC tuning knob—it was ripping out those large allocations entirely.

The Diagnostic Trail

To pin down Gen 2 latency problems, you need to line up GC events with your application’s response times. Event Tracing for Windows (ETW) and tools like PerfView or dotnet-trace give you Microsoft-Windows-DotNETRuntime events, including GC/Start and GC/Stop with the generation depth. I usually grab a trace during a window of elevated tail latency and hunt for GC/Stop events where Generation is 2 and the duration matches the latency spike.

Here’s a quick dotnet-trace command to pull GC events:

dotnet-trace collect --process-id [PID] --providers Microsoft-Windows-DotNETRuntime:GC:4

You can crack open that trace in PerfView and see the exact pause times plus the reason for each collection. With server GC, you might see concurrent Gen 2 collections (background GC), but even those have a brief stop-the-world phase that can still hurt. The metric that really matters is the suspend time for managed threads, not just the total GC wall-clock duration.

Strategies to Eliminate Gen 2 Pauses

Given how brutal the problem is, the real goal isn’t just to shrink Gen 2 pause times. It’s to prevent Gen 2 collections from happening at all during the moments you care about. There are a few paths, and they all come with trade-offs.

1. Reduce Heap Size and Allocation Rate

The most straightforward move is to keep the Gen 2 heap small and stop doing the kind of allocations that promote objects. That means aggressive object pooling, especially for anything large. ArrayPool<T> is a workhorse here. Instead of allocating a fresh byte array for every request, you rent one from the pool and return it when you’re done. This cuts both Gen 0 pressure and LOH fragmentation.

For classes, think about object pooling with a custom pool or ObjectPool<T> from Microsoft.Extensions.ObjectPool. The cost of resetting an object is almost always dwarfed by the cost of a Gen 2 collection. Watch out for string allocations—they’re immutable and often drift into Gen 2 if they hang around through a few collections. Reach for StringBuilder with a pooled backing array or Span<char> on the stack when you can.

2. Tune GC Mode and Latency Mode

.NET gives you two GC modes that matter for latency: workstation GC with sustained low latency and server GC with background collections. Workstation GC runs a single heap and is tuned for interactive apps. Flipping GCSettings.LatencyMode to GCLatencyMode.SustainedLowLatency or GCLatencyMode.LowLatency (the latter for short bursts) tells the GC to hold off on Gen 2 collections as long as it can. This works, but it’s a gamble: if memory pressure builds, the GC eventually forces a full collection, and the pause might be even worse.

Server GC spreads the load across multiple heaps (one per CPU) with dedicated GC threads. It usually pushes higher throughput and is the default for ASP.NET Core apps. Background Gen 2 collections let the application keep running while the GC scans the heap, but a short stop-the-world phase remains. Some teams building really low-latency systems turn off background GC and instead schedule manual collections during quiet periods—though that takes real discipline.

3. Segment the Workload with Out-of-Process Caching

If your Gen 2 heap is swollen because of a giant cache, moving that cache out of process can lighten the load. A distributed cache like Redis, or a separate caching service, pulls those long-lived objects out of the managed heap. What’s left inside your application process are mostly short-lived, request-scoped objects that rarely make it to Gen 2. You’ll see this pattern a lot in microservice architectures where state gets externalized.

For in-memory caches that have to stay in-process, lean on MemoryCache with size limits and compaction callbacks to evict entries before they balloon the heap. Or, use a library that parks cached objects in unmanaged memory, sidestepping the GC entirely—though that definitely adds complexity.

4. Avoid LOH Allocations

The LOH is a quiet assassin. Every LOH allocation nudges you closer to a Gen 2 collection. Starting with .NET Core 3.0, the GC can compact the LOH on demand, but that compaction is itself a full blocking operation. The smarter path is to avoid LOH allocations altogether. That means:

  • Breaking large collections into smaller arrays (under 85,000 bytes each).
  • Using Span<T> and Memory<T> to carve up existing buffers instead of allocating new ones.
  • Leaning on ArrayPool<T> for temporary large buffers.
  • Steering clear of huge string concatenations that produce a string over 85,000 characters.

One team I worked with wiped out 90% of their Gen 2 collections by swapping a single large memory stream allocation per request for a recycled buffer from a pool. The latency improvement was instant and hard to ignore.

A programmer analyzing code on multiple monitors, highlighting memory diagnostics and performance counters
Profiling and careful allocation discipline are essential to controlling GC pauses.

When You Cannot Eliminate Gen 2 Collections

Some applications just can’t dodge Gen 2 collections—systems that have to keep large, shifting in-memory data structures, for instance. Then the play becomes making those pauses predictable and squeezing them inside your budget.

GC notification API: .NET exposes GC.RegisterForFullGCNotification, which can tip you off when a Gen 2 collection is on its way. You can then throttle incoming requests, shed load, or steer traffic to another node. This is a bit fragile because the notification timing isn’t exact, but it can hold up for batch processing systems.

Manual collections during quiescence: If your application has natural lulls—say, a market data system that goes quiet between trading sessions—you can force a full collection during that window with GC.Collect(2, GCCollectionMode.Forced, true). This resets the Gen 2 budget and pushes the next collection further out. The trick is making absolutely sure the window is really idle; otherwise, you’re creating the exact pause you wanted to avoid.

Process recycling: A blunt but effective technique is to recycle the process before Gen 2 gets too large. In Kubernetes or similar orchestrators, you can set a lifetime limit on pods. If your application can handle a graceful shutdown in a few seconds, a rolling restart keeps heap sizes in check and prevents full collections during peak traffic.

Measuring the Impact

Whatever you try, you have to validate it with production metrics. I instrument every latency-sensitive application with a GC.CollectionCount counter per generation and a histogram of GC.GetTotalPauseDuration() sampled periodically. In .NET 5 and later, the System.Runtime runtime counters exposed through dotnet-counters or Prometheus give you:

  • gc-pause-time by generation
  • gc-heap-size by generation
  • alloc-rate in bytes per second

Plot those alongside request latency percentiles, and you can see exactly how much of your tail latency traces back to GC pauses. The aim is to drive Gen 2 pause time down near zero during normal operations, with any leftover pauses happening only during off-peak hours.

FAQ

Why does a Gen 2 collection pause all threads?

The GC needs a consistent view of the object graph while it compacts the heap and updates references. To stop threads from changing references mid-collection, the runtime suspends all managed threads at a safe point. With workstation GC, that suspension lasts the whole collection. Server GC with background collection lets threads run during most of the work, but a short suspension is still required at the end.

Can I use GC.TryStartNoGCRegion to prevent Gen 2 collections?

This API, available in .NET Framework and .NET Core, tries to block any GC for a specified amount of allocated memory. It’s meant for very short, critical sections where any GC pause would be a disaster. But it’s narrow: the size has to be small (usually a few megabytes), and if you blow past the budget, the API either fails or forces a full collection. It’s not a general-purpose shield against Gen 2.

What is the difference between workstation and server GC for latency?

Workstation GC uses one heap and aims to minimize pause time for interactive apps. It can use sustained low latency mode to delay Gen 2 collections. Server GC uses multiple heaps and is tuned for throughput, which often means better overall performance but potentially longer individual pauses. For latency-sensitive services, it’s worth testing workstation GC with low latency mode, though server GC is the ASP.NET default and frequently delivers better tail latency in practice when dialed in right.

How do I identify which objects are causing Gen 2 collections?

Grab a memory profiler like dotMemory or PerfView and capture a heap snapshot before and after a Gen 2 collection. Focus on the objects that survived to Gen 2—those are your long-lived allocations. Keep a close eye on the LOH and pinned handles. PerfView’s GC Heap Analyzer can show you the root paths keeping things alive. The fix is usually to shorten the lifetime of those objects or yank them out of the GC heap.

Understanding Finalizer Queue Backlogs and Their Impact

Introduction to Finalizer Queues and the FReachable Mechanism

Inside .NET’s managed runtime, the garbage collector (GC) usually calls the shots on object lifetimes. When an object drops off the application’s root references, the GC reclaims that memory. But if a class overrides Finalize — or uses the C# destructor syntax ~ClassName() — the story changes. The GC can’t free the object right away. Instead, it pushes the object into a special internal structure: the finalizer queue. That queue is just a linked list of objects waiting for finalization, and it adds a whole extra act to memory cleanup that can quietly throttle throughput or even destabilize a process.

A single runtime thread, widely called the FReachable thread, handles the finalization. After the GC marks an object and places it in the finalizer queue, this thread dequeues objects one at a time and runs their finalizers. The split between marking and execution is intentional — it stops the GC sweep from getting stuck on arbitrary user code. But it also sets up a classic producer-consumer problem: the GC can produce finalizable candidates far quicker than a single thread can eat through them, and that’s when a backlog starts.

Abstract digital network representing queue processing

How Finalizer Queue Backlogs Develop

A backlog builds when you allocate finalizable objects (and the GC promotes them) faster than finalization can keep up. And remember: every object sitting in that queue holds its own memory plus any graph of otherwise unreachable objects it still references. Those objects sit in generation 2 — or on the large object heap if big enough — until the finalizer finishes. So your live memory footprint swells. The FReachable thread works through finalizers one after another. A single slow finalizer — maybe one flushing a big file buffer or closing a remote connection — stalls the whole line.

Take a hot allocation loop that pumps out thousands of finalizable objects each second. Even if each finalizer only burns a few microseconds, the cumulative lag piles up and the queue grows. Memory profilers often tip you off with a climbing Finalization Survivors metric, or a widening gap between total promoted finalizable objects and finalized objects. At worst, you get an effective memory leak: the GC simply can’t reclaim the objects until their finalizers run.

Root Causes of Slow Finalization

A few design habits choke finalizer throughput. Blocking calls inside finalizers — synchronous I/O, lock grabs — sit right at the top. Since the FReachable thread is a single point of execution, any delay hits every pending finalizer downstream. Finalizers that call into other managed objects can accidentally resurrect those objects, which confuses the GC’s accounting. Then there are the finalizers that try to do too much: complex calculations, thread-static data access, you name it. All common sources of slowness.

One less obvious contributor is Large Object Heap (LOH) fragmentation. Finalizable objects on the LOH stay indirectly pinned by the finalizer queue, blocking compaction and making fragmentation worse. Over time, that forces the GC into more frequent full collections. More full collections mean more objects marked as finalizable, and that feedback loop accelerates the backlog.

Impact on Application Performance and Stability

Performance monitoring dashboard with graphs

The first thing you notice with a finalizer queue backlog is memory pressure. As the queue expands, the heap hangs on to objects that should be dead, driving up the memory footprint. If your process is already near its limits, out-of-memory exceptions loom. That pressure also nudges the GC into more gen2 collections — fully stop-the-world events that suspend all managed threads. The total pause time eats into throughput and can bust latency budgets in interactive or real-time applications.

Memory isn’t the only victim. Under a heavy backlog, the FReachable thread’s CPU use starts to stand out. While finalizers execute, they compete with application threads for CPU time, and the GC itself might wait on finalizer completion during certain phases. In server workloads you see longer request latencies and weaker scalability. Tools like PerfView or dotnet-counters can expose high “% Time in GC” numbers and elevated finalizer counts — a solid hint that a backlog is in play.

Diagnosing Finalizer Queue Backlogs

Getting a clear diagnosis means lining up a few telemetry sources. Memory dumps, analyzed with SOS or ClrMD, let you peek straight into the finalizer queue. In WinDbg, !FinalizeQueue shows the count of objects ready for finalization and those awaiting cleanup. A number that stays stubbornly high, even after forced GC collections, points to a backlog. Application-level logging of GC.GetTotalMemory and finalizer execution timestamps can add useful context to that low-level data.

Watching the .NET CLR Memory\Finalization Survivors performance counter gives a near-real-time look at objects that survive because of finalization. If that line keeps climbing, finalizers aren’t keeping pace. Pair it with the \Finalization Promoted from Gen 0 counter to gauge the inflow rate. When the inflow consistently outruns processing, the backlog is baked in.

Mitigation Strategies and Best Practices

The most straightforward way to avoid finalizer queue backlogs is to cut back on finalizers altogether. The IDisposable pattern and the using statement give you deterministic cleanup without ever touching the GC’s finalization machinery. For types that wrap native resources, SafeHandle derivatives handle cleanup and remove the need for a hand-rolled finalizer. When you absolutely can’t avoid a finalizer, keep its code minimal and never async: release a native handle, unhook from events, set a flag to block re-entry. That’s it.

In older codebases where tearing out finalizers isn’t on the table, consider offloading heavy cleanup to a dedicated background thread. The finalizer can post a work item to a custom queue and signal a worker thread, letting the FReachable thread get back to business fast. Be careful, though — this adds synchronization overhead and might delay resource release. Another practical step: call GC.SuppressFinalize after deterministic cleanup. That removes the object from the finalizer queue entirely.

Monitoring in Production

Keeping an eye on finalizer-related metrics in production catches problems before they spiral. Set alerts on the finalization survivors counter and track how often gen2 collections fire. In cloud environments, correlate those numbers with memory usage and CPU load to spot degradation trends. Automated memory dump collection, triggered by sustained high finalizer counts, can preserve the queue state for later post-mortem work.

Server rack indicating infrastructure monitoring

FAQ

What is the difference between the finalizer queue and the f-reachable queue?

People often use the terms as synonyms, but the runtime separates them. The finalizer queue holds objects the GC has marked as unreachable and needing finalization. The FReachable thread reads from that queue and moves objects into a distinct “f-reachable” state while executing their finalizers. The main distinction: the finalizer queue is the work source for the FReachable thread, while the f-reachable state covers objects currently or recently finalized.

Can a finalizer queue backlog cause the GC to block application threads?

Yes, but indirectly. The GC doesn’t normally suspend application threads just for finalization. Yet a large backlog increases memory pressure and drives up gen2 collection frequency. Gen2 collections block all managed threads, so they do suspend your app. And if the process runs out of memory because the GC can’t reclaim finalized objects, application threads will be blocked during GC attempts or may hit out-of-memory exceptions.

How can I determine the optimal number of finalizable objects in my application?

For most types you write, the optimal number is zero. Every finalizable object costs memory and GC cycles. Use finalizers only for types that directly own native resources and can’t lean on SafeHandle. Profile your application under realistic load to measure finalizer throughput: if the finalizer queue count keeps growing over time, you’ve blown past the sustainable rate. Aim for a steady state where the count hovers near zero after a forced full GC.

Understanding Finalizer Queue Backlogs and Their Impact on .NET Memory

Memory cleanup in a managed runtime like .NET usually feels invisible. The garbage collector hums along, sweeping through heaps and reclaiming objects that have fallen out of scope. But there’s a less obvious side to this process—the finalizer queue. When objects with finalizers stack up faster than the runtime can clear them, you get a backlog. That backlog doesn’t just sit there; it warps memory usage, starves resources, and drags down performance in ways that are easy to overlook until they bite you.

This article picks apart the finalizer queue backlog: how it forms, how you can spot it, and what you need to do to stop it from undermining your application’s stability.

Close-up of a computer motherboard with intricate circuit traces

How Finalization Actually Works in .NET

When an object overrides Finalize()—or uses a C# destructor—the GC won’t reclaim its memory the moment it becomes unreachable. Instead, the object gets shuffled into a dedicated structure called the finalizer queue. A separate thread, the finalizer thread, plucks objects from this queue one at a time and runs their finalizers. Only afterward, in a later GC cycle, does the memory get freed.

So you’re looking at a two-stage reclamation: first the object is marked for finalization, then—after its finalizer has run—the memory can actually be collected. That separation is there for a reason. It lets cleanup logic (releasing native handles, for instance) execute safely outside the main GC pause. But it also creates a natural choke point. The finalizer queue runs on a single thread, and every object waiting for finalization sits on the heap, holding onto memory, until both stages are done.

The Structure of the Finalizer Queue

Inside the runtime, the finalizer queue is a linked list of objects that registered a finalizer. During a collection, the GC identifies dead objects with finalizers and adds them to a sub-list called the “freachable” queue—these are the ones actually ready to be finalized. The finalizer thread walks through that freachable list, calling finalizers in order. After execution, the object disappears from the freachable queue and becomes a normal candidate for collection.

But here’s the catch: the finalizer thread runs asynchronously. An object can idle in the freachable queue for an unpredictable stretch. If dead objects pour into the queue faster than the finalizer thread can handle them, the backlog builds. And that backlog isn’t just a counter—it’s live memory you can’t touch until each finalizer completes.

Server racks with blinking lights in a data center

What Causes a Finalizer Queue Backlog

Backlogs don’t show up in applications with a handful of finalizable objects. They emerge when the volume—or the behavior—of finalizable objects overwhelms the single-threaded cleanup model.

High Allocation Rate of Finalizable Objects

Applications that constantly allocate types with finalizers—database connections, file streams, custom wrappers around native resources—can flood the queue. Every allocation that eventually becomes unreachable adds another entry. In high-throughput systems, even well-written finalizers can’t keep up if the allocation rate spikes beyond what a single thread can process.

Slow or Blocking Finalizers

One badly behaved finalizer can stall the entire queue. If a finalizer does lengthy I/O, grabs a lock, or calls into unmanaged code with high latency, it holds up every other entry behind it. Remember, the finalizer thread is shared across all finalizable objects in the process. A single slowpoke turns a modest allocation rate into a serious backlog.

Thread Starvation or Priority Issues

The finalizer thread normally runs at a high priority, but certain runtime configurations or host processes can interfere. If it gets starved of CPU time—maybe because of heavy user thread activity or aggressive thread pool tuning—the queue drain rate drops, and objects pile up.

Large Object Heap (LOH) Finalizers Without Compaction

Objects on the Large Object Heap aren’t compacted by default in most GC modes. When a finalizable object lives on the LOH, the delay in finalization can add to fragmentation. The backlog holds memory hostage and can make the heap’s fragmentation profile worse, leading to out-of-memory exceptions even when there’s plenty of total memory available.

Spotting a Finalizer Queue Backlog

You won’t catch a backlog by glancing at surface-level memory counters. Total managed heap size can look fine while a big chunk of memory is tied up in objects waiting for cleanup.

Performance Counters

Windows Performance Monitor has a counter under .NET CLR Memory: Finalization Survivors. This tracks objects that survived a collection only because they’re waiting for finalization. If that number keeps climbing or stays persistently high, you’ve got a backlog. Pair it with Promoted Finalization-Memory from Gen 0 to see how fast new finalizable objects are entering the queue.

Memory Dump Analysis

In a memory dump, reach for the SOS debugging extension command !finalizequeue. It dumps all objects in the freachable queue and those registered for finalization. Look for a large total count or types that dominate the queue. For example:

0:000> !finalizequeue
SyncBlocks to be cleaned up: 0
Free-Threaded Interfaces to be released: 0
MTA Interfaces to be released: 0
STA Interfaces to be released: 0
----------------------------------
generation 0 has 18 finalizable objects (000001e0b8101a30->000001e0b8101ac0)
generation 1 has 17 finalizable objects (000001e0b81019a8->000001e0b8101a30)
generation 2 has 214 finalizable objects (000001e0b8101300->000001e0b81019a8)
Ready for finalization 0 objects (000001e0b8101ac0->000001e0b8101ac0)
      MT    Count    TotalSize Class Name
00007ff8e6a1c6a0        1           24 System.WeakReference
...

If “Ready for finalization” shows a zero but hundreds of objects are still registered in generation 2, the finalizer thread is processing them but can’t keep up with the influx. Use !threads to check the finalizer thread’s state—see if it’s actually running or sitting there blocked.

ETW Tracing

Event Tracing for Windows providers like Microsoft-Windows-DotNETRuntime emit events for finalization activity. The GarbageCollection/FinalizeObject event fires for each finalizer execution. Analyzing timestamps and object counts over time lets you measure queue depth and processing latency directly.

Software engineer analyzing performance data on multiple monitors

Performance Fallout from a Backlog

A finalizer queue backlog isn’t just a memory leak. It kicks off a cascade of trouble that can make your whole application wobble.

Increased Memory Pressure and GC Frequency

As finalizable objects pile up, the managed heap grows. The GC responds by triggering more frequent collections—especially expensive generation 2 sweeps that scan the entire heap. You end up in a nasty cycle: high allocation, delayed finalization, aggressive GC, and a lot of CPU cycles burned for very little reclaimed memory. Some engineers call this a “mid-life crisis” for objects on the heap.

Gen 2 Heap Fragmentation

Objects that survive into generation 2 because of pending finalization can leave holes when they finally get released. In workloads without compaction, freed memory turns into gaps that are too small for new allocations. The heap size creeps up, and you might face an OutOfMemoryException.

Resource Leaks and Handle Exhaustion

Finalizers often release native resources—file handles, sockets, database connections. If the backlog delays finalization, those handles stay open far longer than you’d expect. In server applications, this can blow through the process’s handle limit, causing failures when opening new connections or files. You’ll see “Too many open files” errors or socket exceptions, even though your higher-level wrappers are being disposed correctly.

Application Pauses and Latency Spikes

When the GC does a blocking generation 2 collection to claw back memory, application threads are suspended. A backlog-driven spike in Gen 2 collections means more frequent—and longer—pauses. For latency-sensitive applications, that’s a fast way to miss your SLAs and annoy users.

How to Mitigate Finalizer Queue Backlogs

Preventing backlogs takes a mix of careful design choices and operational monitoring. There isn’t one magic switch.

Reduce Finalizable Object Count

The bluntest fix: don’t use finalizers unless you truly have to. Modern .NET gives you patterns like SafeHandle and IAsyncDisposable that sidestep finalization. When you wrap native resources, reach for SafeHandle instead of a raw IntPtr with a finalizer. The runtime handles SafeHandle finalization more efficiently and ties into the critical finalizer infrastructure.

Dispose Promptly and Deterministically

Use IDisposable and using statements with discipline. When you explicitly dispose an object, you can suppress its finalizer with GC.SuppressFinalize(this). That removes the object from the finalizer queue entirely—no two-phase delay. Audit your code paths to make sure every IDisposable gets disposed, even when exceptions fly.

Keep Finalizers Fast and Non-Blocking

If you absolutely can’t avoid a finalizer, strip its work down to the bare minimum. No I/O, no lock acquisition, no calls to external services. A common pattern: set a flag inside the finalizer that says “I’m done,” then hand off the real cleanup to a background thread or a dedicated resource-reclaim pool. The finalizer itself should finish in microseconds.

Monitor and Alert on Backlog Indicators

Instrument your application to track the Finalization Survivors performance counter. Set thresholds that trigger alerts when the count drifts above a baseline that makes sense for your workload. In production, periodic memory dumps can be analyzed offline to confirm whether the queue is growing.

Consider GC Mode Tuning

For server applications, switching to sustained low latency mode or workstation GC with background finalization might help the finalizer thread keep pace. But tread carefully—tuning GC modes changes memory management behavior across the board, so validate under realistic load before you commit.

FAQ

What’s the difference between the finalizer queue and the freachable queue?

The finalizer queue holds every object that registered a finalizer, reachable or not. The freachable queue is a subset—only objects the GC has marked unreachable and that are now waiting for their finalizers to run. After finalization, they leave the freachable queue and become eligible for normal collection.

Can a finalizer queue backlog cause an OutOfMemoryException even if the heap isn’t full?

Absolutely. The backlog can fragment the heap—especially the Large Object Heap or Gen 2—so free memory exists but no single block is big enough for an allocation. Also, if finalizers are holding up the release of native handles, the process can hit handle limits long before managed memory runs out.

How can I tell if a specific type is causing the backlog?

In a memory dump, !finalizequeue groups objects by method table (MT) and shows count and total size per type. Look for types with unusually high counts. Cross-reference with !dumpheap -stat to see the full heap picture. ETW events give you timestamps, so you can tie finalization delays to specific type names.

Is it safe to call GC.Collect() to force finalization?

Calling GC.Collect() followed by GC.WaitForPendingFinalizers() can drain the queue temporarily, but it’s not a production fix. It introduces blocking pauses and doesn’t touch the root cause. Keep it for testing or diagnostics, and focus on cutting finalizable allocations or tightening disposal patterns for a permanent solution.

Understanding Finalizer Queue Backlogs and Their Impact on .NET Application Performance

The Mechanics of Finalization in .NET

Inside the .NET runtime, any object that wraps unmanaged resources—file handles, network sockets, raw memory allocations—depends on a finalizer to release those resources when the garbage collector (GC) reclaims the managed wrapper. You declare a finalizer with the ~ClassName() syntax. At allocation time the runtime places the object onto a dedicated structure called the finalizer queue. That queue acts as a root set, keeping the object and the whole graph it references alive well beyond their normal lifetime.

When the GC marks an object as unreachable and sees it has a finalizer, it moves the reference from the finalizer queue to the freachable queue. A separate, dedicated thread then walks that queue and calls each finalizer one after another. This design adds real latency: the object survives at least one full GC cycle before its finalizer even runs, and the managed memory it occupies cannot be reclaimed until finalization finishes.

Abstract representation of a .NET garbage collection process with memory blocks and queues

How the Finalizer Queue Drains

The finalizer thread pulls entries from the freachable queue sequentially. If a finalizer blocks—maybe it hits a synchronous I/O call, waits on a lock, or spins in an infinite loop—the whole queue grinds to a halt. No other finalizers execute until that call returns or the thread gets aborted. By default you get one finalizer thread per process. In high-stress situations the runtime might inject an extra thread, but the behaviour is not guaranteed and you should never count on it. The serial execution model means any backlog in the freachable queue turns directly into a growing pile of dead objects that still hold native resources and cannot let them go.

Tools such as PerfView and dotnet-counters expose Finalization Survivors and Promoted Finalization-Memory. When those counters trend upward, objects are being bumped into higher GC generations purely because their finalizers have not run yet. That promotion drives up the frequency of expensive Gen2 collections and fragments the managed heap over time.

How Backlogs Develop and Degrade Performance

A finalizer queue backlog almost always starts quietly. Imagine a burst of network connections that all time out at once. Each socket’s SafeHandle-derived wrapper queues up a finalizer, but the finalizer thread burns 50 milliseconds per object releasing the underlying handle because of a blocking closesocket call. If 200 sockets become unreachable inside a second, the freachable queue jumps to 200 entries. The finalizer thread needs 10 seconds just to drain them. During that window the GC promotes these dead objects into Gen2, where they inflate the working set and force premature Gen2 collections. Application throughput drops as GC thread suspensions get more frequent.

Memory Pressure and Gen2 Heap Growth

An object sitting in the finalization pipeline hangs onto its own memory plus the whole transitive closure of objects it references. A DbConnection that never got closed might hold a 4 KB internal buffer, a command object, and a transaction reference. Let 500 such connections pile up in the freachable queue and the retained memory easily crosses several megabytes. That memory stays resident; it pushes up the process’s private bytes. When the GC eventually promotes these objects to Gen2, the Gen2 heap expands. A larger Gen2 heap means longer collection pauses because the GC has to scan more memory. In server-side apps running under sustained load, this feedback loop can shove a process straight toward an OutOfMemoryException even when the live object count looks modest.

Monitoring dashboard showing .NET memory counters and finalization backlog alerts

Thread Pool Starvation and Finalizer Deadlocks

The finalizer thread doesn’t work in a bubble. If a finalizer tries to grab a lock that a user thread is holding while that user thread waits for a GC to finish, you get a deadlock. A more common mess: a finalizer blocks on an async operation—misusing Task.Wait() or Task.Result—and starves the thread pool when the synchronization context dispatches back to a pool thread. The runtime’s finalizer thread isn’t a thread-pool thread, but the chain of blocked dependencies can stall progress across the whole application. I have walked into production hangs where a single finalizer calling FileStream.Flush() on a network share blocked for 30 seconds, the freachable queue ballooned past 10,000 entries, and the process turned completely unresponsive.

Diagnosing a Finalizer Queue Backlog

Spotting a backlog means correlating several diagnostic signals. Grab a memory dump during high memory usage first. Use !finalizequeue in SOS (Son of Strike) to inspect the freachable and finalizer queues directly. The output gives you object counts in each queue, broken down by type. A fat count of objects from a single type—say, System.Net.Sockets.Socket or Microsoft.Win32.SafeHandles.SafeFileHandle—points straight at a specific resource leak.

0:000> !finalizequeue
SyncBlocks to be cleaned up: 0
Free-Threaded Interfaces to be released: 0
----------------------------------
Generation 0 has 12 finalizable objects (000001e0b5c01040->000001e0b5c010a0)
Generation 1 has 5 finalizable objects (000001e0b5c01018->000001e0b5c01040)
Generation 2 has 1483 finalizable objects (000001e0b5c00018->000001e0b5c01018)
Ready for finalization 1847 objects (000001e0b5c010a0->000001e0b5c01448)
Statistics for all finalizable objects:
              MT    Count    TotalSize Class Name
00007ff8e5a8c7a0      912       145920 System.Net.Sockets.Socket
...

A high number under Ready for finalization confirms the backlog. Next, grab the finalizer thread’s call stack with ~* kb or !threads and check whether it is blocked. The ThreadState often reads WaitSleepJoin. The managed stack shows which finalizer method is executing or stuck. I once tracked down a SafeHandle.ReleaseHandle() override that called a blocking DeviceIoControl API; the native stack showed the thread sitting in ntdll!NtWaitForSingleObject for minutes.

Using ETW and PerfView for Real-Time Analysis

ETW (Event Tracing for Windows) lets you monitor without touching the process. The Microsoft-Windows-DotNETRuntime provider emits FinalizeObject and IncreaseMemoryPressure events. Inside PerfView the GCStats view graphs the Finalization Survivors rate. If that rate stays above zero during steady-state operation, finalizers can’t keep up with allocation. The FinalizeObject event includes the object’s type name and how long the finalizer call took. Sort by duration and the worst offenders jump right out. I have seen finalizer durations over 2 seconds for objects that should finish in microseconds, which isolates the root cause immediately.

PerfView trace analysis showing finalizer durations and GC survival metrics

Mitigation Strategies and Best Practices

Keeping backlogs from forming starts with ripping blocking work out of finalizers. The only safe code inside a finalizer releases native resources without waiting. That usually means calling CloseHandle, ReleaseSemaphore, or writing a flag to a memory-mapped file. Anything that might block—file I/O, network calls, lock acquisition—has to be deferred. The Dispose pattern gives you the primary path for deterministic cleanup. When a consumer forgets to call Dispose, the finalizer serves as a safety net, but it must be designed to finish fast and reliably.

Reach for SafeHandle instead of writing custom finalizers. The runtime treats SafeHandle-derived objects specially during finalization, which lowers the risk of premature release and async-dispose race conditions. The SafeHandle.ReleaseHandle() method runs in a constrained execution region (CER), yet it still executes on the finalizer thread. Keep that method non-blocking. For resources that need a complicated shutdown sequence, consider hooking the AppDomain.ProcessExit event or spinning up a dedicated background thread that drains a concurrent queue of disposal actions. That separates the finalizer’s tiny notification from the potentially slow resource release.

Monitoring and Alerting in Production

Set up performance counters that track finalization activity. The .NET CLR Memory / Finalization Survivors counter and the # Bytes in all Heaps counter together tell you whether a backlog is driving memory retention. A threshold of 100 finalization survivors sustained for more than 60 seconds should trigger an alert. In containerized environments, pair these with memory limit alerts. A process that stabilises near its memory limit because of promoted finalizable objects will eventually hit an OOM kill, so the diagnostic data has to be captured before that point.

For applications with high object churn, call GC.SuppressFinalize() right after deterministic cleanup. That removes the object from the finalizer queue entirely, bypasses the freachable queue, and cuts down on promotion. In hot paths, the difference between an object that gets Gen0-collected and one that survives into Gen2 can add up to hundreds of milliseconds of pause time across the application’s lifetime.

FAQ

What is the difference between the finalizer queue and the freachable queue?

The finalizer queue holds live objects that have been allocated and registered for finalization. When the GC decides an object is unreachable, it moves the reference to the freachable queue. The finalizer thread reads from the freachable queue and runs each object’s finalizer. The separation guarantees objects are not finalized while the application might still be using them.

How can I tell if my application has a finalizer backlog without a memory dump?

Use dotnet-counters to watch dotnet.gc.finalization_survivors and dotnet.gc.time_in_gc. If finalization survivors keep climbing and the time spent in GC tracks that climb, a backlog is likely. Also, when the process’s Gen2 heap size grows monotonically without a matching increase in live objects, finalizable objects are being promoted.

Why does the finalizer thread sometimes appear stuck even when finalizers are fast?

A single slow finalizer blocks everything queued behind it. The thread can also get stuck in a native call that never returns, deadlock with a user thread, or end up waiting on a task that can’t complete because of thread-pool exhaustion. Looking at the finalizer thread’s native and managed stacks during the hang is the most direct way to find the blocking call.