Inside PerfView: Diagnosing Real-World .NET Performance Problems Like a Seasoned Engineer

When a production incident hits and the CPU graph jumps past 95%, junior engineers open Task Manager, guess at the problem, and restart the service. Senior engineers reach for PerfView. Not the GUI-first, click-through version — the one you drive from the console, with custom collection flags, symbol servers configured ahead of time, and a working mental model of the CLR’s runtime internals. This is that PerfView.

Close-up of server rack LED indicators blinking in a dark data center
Production diagnostics demand precision — PerfView delivers it when configured correctly.

Why the Default Collection Is a Trap

Double-clicking PerfView.exe and hitting “Collect” with whatever providers are ticked by default — that’s the first mistake I spot even in mid-level engineers. The default kernel and CLR providers pull in so much noise that the .etl file balloons to gigabytes, and the trace processing time outlasts the incident window. You have to be specific about what you ask for.

Start from the command line. Always. The GUI is fine for an initial poke around a trace, not for collection. Use the /DataFile and /Providers flags to narrow the scope. For a high-CPU investigation, I typically disable kernel events beyond Profile and ContextSwitch, and I specify CLR providers at the Informational level only for the modules I actually suspect.

PerfView.exe /DataFile:HighCPU.etl /BufferSizeMB:256 /Providers:"Microsoft-Windows-DotNETRuntime:0x1:5,Microsoft-Windows-DotNETRuntimeRundown:0x1:5" collect

That provider GUID specifies only GC, JIT, and Exception events at the verbose level. No loader callbacks, no interop stubs, no assembly resolve noise. The trace stays small, focused, and parseable in under 30 seconds.

Symbol Resolution: The Divide Between Guesswork and Certainty

If your call stacks show [Unknown] or generic mscorlib.ni.dll entries, you’re flying blind. Senior engineers set up a symbol server path before collection, not after. The /SymbolPath parameter tells PerfView where to find PDBs, and it must include both your build server’s symbol store and the Microsoft public symbol server.

I keep a small script that maps a network drive to the latest release symbols and appends the Microsoft URL. The syntax is semicolon-delimited:

/SymbolPath:\"\\build-server\symbols;https://msdl.microsoft.com/download/symbols"

When the trace loads, every frame resolves. You see the exact method name, the source file, and the line number. No guesswork. This alone separates a useful trace from a waste of disk space.

Filtering CPU Stacks to the Thread That Matters

Open the trace in the GUI, navigate to the CPU Stacks view, and immediately group by process and thread. The default flame graph aggregates all threads, which hides the one thread spinning at 100%. Right-click the thread column, select Filter to Selection, and now the flame graph shows only that thread’s activity.

Look for flat tops — methods with no children. Those are the CPU consumers. Often you will see a while loop inside an IDisposable implementation that never yields. The call stack tells the story: MyService.Dispose() → SpinWait.SpinOnce() → kernel wait. That is a bug, not a feature.

Software engineer analyzing code on multiple monitors with debugging tools open
Interpreting a PerfView flame graph requires understanding the runtime’s execution model.

GC Pauses: The Silent Latency Killer

High memory traffic doesn’t always show up as CPU spikes. It shows up as GC pause time. In PerfView, the GCStats view gives you the total pause duration per generation. If gen-2 collections are taking 200ms on a server that should respond in under 50ms, you have a real problem.

Drill into GC Heap Alloc Ignore Free to see which types allocate the most. Sort by Alloc MB. I once found a logging wrapper that allocated a new StringBuilder on every call, totaling 80MB per minute under load. The fix was a pooled instance. No profiler needed — just PerfView and a careful read of the allocation graph.

Cross-Referencing with ETW Events

The Events view is underused. Filter to Microsoft-Windows-DotNETRuntime/GC/Start and /Stop events, and export the timestamps. Plot them in Excel alongside your latency metrics. The correlation is often striking: every spike in request duration aligns precisely with a gen-2 collection. Then you know the fix isn’t in the code path itself but in the object lifetime management.

Contention and Blocking: The Lock That Broke Production

PerfView captures Monitor.Enter contention by default if you include the right CLR keywords. In the .NET ThreadPool view, look for threads in the Waiting state with a Monitor wait reason. The stack trace shows the lock location.

I diagnosed a four-thread deadlock in under ten minutes using this technique. The classic pattern: Thread A holds lock X and waits for lock Y; Thread B holds lock Y and waits for lock X. PerfView’s Thread Time view, sorted by wait time, gave me the exact four stacks. No magic, just methodical filtering.

Memory Dumps vs. Traces: When to Switch Tools

PerfView is not a memory dump analyzer. If the issue is a slow leak over hours, you need a series of dumps compared in dotMemory or WinDbg+SOS. PerfView excels at active analysis — what is happening right now, in the running process. Use it to capture the burst, then switch to a dump for the snapshot.

I combine them: start a PerfView collection, wait for the spike, stop it, then take a dump with procdump immediately after. The trace shows the dynamic behavior; the dump shows the state. Together, they form a complete picture.

Server hardware with glowing indicators and detailed circuit boards
A well-configured PerfView trace captures the exact moment of failure without disrupting the production service.

Automating PerfView in CI/CD for Regression Detection

Senior engineers don’t wait for production. They integrate PerfView into the performance test pipeline. A short five-second trace during a load test can capture GC pause times and allocation rates automatically. Write a script that invokes PerfView.exe with the /NoGui and /AcceptEula flags, parses the XML output, and fails the build if gen-2 pauses exceed a threshold.

PerfView.exe /NoGui /AcceptEula /DataFile:LoadTest.etl /Providers:... collect
PerfView.exe /NoGui /AcceptEula /DataFile:LoadTest.etl GcStats > GcStats.txt

Parse GcStats.txt for “Gen2 Pause Time Mean” and assert. This catches allocation regressions before they reach staging, let alone production.

FAQ

Why do my call stacks show [Broken] even with symbols configured?

[Broken] typically indicates a mismatch between the trace’s recorded module version and the PDB version. Make sure the symbol path points to the exact build artifacts deployed to the profiled machine. Also verify that the _NT_SYMBOL_PATH environment variable isn’t overriding your PerfView argument.

How do I reduce the trace file size without losing critical events?

Use provider filter keywords aggressively. For GC-only analysis, specify 0x1 for GC events instead of the default 0x1F. Increase the circular buffer size (/BufferSizeMB) to 1024 and set /CircularMB to a reasonable limit. This drops old events when the buffer fills, keeping the file size constant while preserving the most recent activity.

Can PerfView collect on a production server with minimal overhead?

Yes, if configured carefully. Use the /KernelEvents=Default flag to limit kernel providers, set CLR providers to 0x1:4 (Informational level for GC only), and reduce the sampling interval with /SampleIntervalMS=100. This lowers CPU overhead to under 2% in typical scenarios, making it safe for short production collections.

What is the difference between CPU Samples and CPU Stacks?

CPU Samples shows a flat list of methods with sample counts, useful for ranking hot methods. CPU Stacks shows a hierarchical call tree, allowing you to trace the exact path to the hot method. Always start with CPU Stacks grouped by thread to understand the context, then switch to CPU Samples for a quantitative summary.