Why Most Engineers Under-Use PerfView
PerfView is not simply a performance profiler. It is a forensic toolkit for .NET production incidents, designed by the .NET runtime team to expose what is happening at the memory, thread, and CPU level. Many developers, even those with several years of .NET experience, reach for PerfView only when they need a quick CPU trace or a managed memory dump analysis. The tool is capable of far more than that. When I mentor engineers, I do not teach them just the commands. I teach them a systematic approach: how to capture data without destabilizing a live server, how to interpret the output, and how to present findings that leave no room for speculation.
This guide is not a beginner tutorial. I assume you already understand basic performance concepts, ETW events, and the difference between managed and native allocations. What I will show you is the workflow I use when I get a call at 3 a.m. because a production service is consuming 100% CPU or the GC pause time has jumped from 200 ms to 3 seconds. I will show you how to use PerfView the way a senior engineer uses it: with precise commands, a critical eye on overhead, and an understanding of the underlying runtime mechanics.

Building a Reliable Collection Strategy
Before you look at a single stack trace, you must plan the collection. The biggest mistake I see is engineers running high-overhead traces on a production server during peak load and then being surprised when the service tips over. PerfView can collect kernel events, .NET runtime events, and sample call stacks with minimal impact—if you configure it correctly.
Choosing the Right Provider Set
The default collection is often too broad or too narrow. For CPU investigations, I start with CPU Samples and enable only the Microsoft-Windows-DotNETRuntime provider with the GC and Contention keywords if I suspect lock issues. Avoid enabling the JIT keyword unless you specifically need to see method compilation overhead. The JIT events are verbose and can add measurable pressure on the ETW session. The exact command I use looks like this:
PerfView.exe /DataFile:HighCpuTrace.etl /BufferSizeMB:256 /CircularMB:2000 /Providers:Microsoft-Windows-DotNETRuntime:0x1:5 collect
The keywords 0x1 give you GC events only. The 5 is the verbosity level, which filters out unnecessary payload. If I need thread time or context switch data, I also add the kernel provider with Loader and ThreadTime flags, but I always measure the impact in a staging environment first. A senior engineer knows that a trace that is too large becomes unreadable anyway.
Managing Trace Size and Duration
Circular buffers are essential for capturing intermittent problems. I set /CircularMB to at least 2000 for a service handling hundreds of requests per second. The buffer wraps; when you stop the trace, you get the last N megabytes of events. This lets you leave PerfView running for hours with negligible overhead, then trigger a stop when the issue occurs. For high-memory services, I pair this with a low /BufferSizeMB to avoid ETW session allocation failures. The balance is always between event loss and system impact.
For heap snapshots, the rules change. A full heap dump with GCCollectOnly or a forced GC via PerfView’s HeapSnapshot provider will block managed threads. I never run this on a production node without first draining traffic. The trace itself must be short—30 seconds is usually sufficient—and I warn the on-call team that a pause will occur. There is no magic: a heap snapshot triggers a full blocking GC, and that is a design decision, not a bug.

Analyzing CPU Traces with Surgical Precision
Once the ETL file is on your machine, the real work begins. I open the trace in PerfView and immediately go to the CPU Stacks view. The default grouping by module is nearly useless for managed code. I switch to the By Name grouping and set the time range to exclude the startup and shutdown periods. What I want is a flat cost table that shows me inclusive and exclusive CPU time for each method, filtered to the time window where the problem was observed.
Interpreting the Flame Graph
PerfView’s flame graph is not a toy. I expand nodes from the bottom up, looking for broad plateaus where one method dominates the sample count. If I see System.String.Concat taking 40% of samples inside a request handler, I do not immediately blame string concatenation. I look at the calling context. Often, the real problem is a loop that builds a large string for logging or serialization, and the fix is to use a StringBuilder or restructure the log pipeline. The flame graph shows you the relationship; the flat list shows you the cost. You need both.
A senior engineer also knows when the trace itself is lying. ETW sampling is statistical. If you have many short-lived threads, a method that runs for 1 ms but is called millions of times might be under-sampled. I always cross-reference PerfView CPU data with dotnet-counters or Windows Performance Monitor metrics collected during the same window. The CPU percentage reported by the OS should roughly align with the trace’s sampled busy time. If they diverge significantly, your sampling interval is too coarse, or kernel overhead is masking the signal.
Spotting Hidden Bottlenecks: GC and JIT
CPU traces often hide GC pauses because the GC runs on dedicated threads that may not be sampled in the same way. I always open the GCStats view immediately after a CPU analysis. A sudden increase in % Time in GC from 5% to 45% tells you more than any stack trace. If Gen 2 collections are frequent, I then take a heap snapshot to find the pinning or large object heap fragmentation causing them. Do not tune code if the GC is the bottleneck. Tune allocations.
Similarly, if I see high CPU in clr!ThePreStub or JIT-related stacks, it means methods are being compiled at runtime under load. The fix is not always to enable background JIT; sometimes you need to run a warm-up script or use ReadyToRun images. PerfView’s JITStats view gives you the exact method names and compilation times. This is data you can take to the build team.

Memory Diagnostics Without Guessing
Memory leaks are the most misdiagnosed problems in .NET. A senior engineer does not guess. We collect a heap snapshot with PerfView using the /GCCollectOnly flag, which forces a full GC before the snapshot so we see only live objects. Then we open the Heap Snapshot view and sort by Inclusive Size.
Rooting Paths, Not Just Object Counts
The object list is distracting. What matters is the reference graph. I select a suspicious type—say, System.Byte[] taking 800 MB—and click Open in GC Heap Explorer. From there, I pick a random large instance and trace its roots back to a static field, an event handler, or a pinned Gen 2 segment. The tool shows me the full chain: StaticVar -> List -> Byte[]. That is the leak. Fixing it means breaking that chain, usually by nulling out the static reference after use or unsubscribing from an event.
For managed memory, I ignore shallow size. A 24-byte object that holds a reference to a 2 MB array is the real problem. PerfView’s Inclusive Size column accounts for this. I also filter by generation. Objects in Gen 2 that are not supposed to be there—configuration data, cached responses—indicate a leak that survives collections. The Gen 2 Object Count trend over multiple snapshots is a dead giveaway.
Native Memory and VirtualAlloc
Managed memory is only part of the story. When the private working set grows but managed heap size is stable, I switch to the Native Memory view. PerfView can show you allocations from VirtualAlloc by call stack, provided you collected kernel events. I look for stacks ending in System.Net.Http or Socket buffers, which often point to pinned object arrays that the GC cannot move. The fix might be to pool buffers or reduce the send/receive buffer sizes at the socket level. PerfView gives you the exact allocation size and count, so you can calculate the potential savings.
Advanced Diffing and Baseline Comparisons
One trace is a story; two traces are evidence. When I suspect a regression, I capture a baseline trace from a known-good build and a comparison trace from the suspect build, using identical collection settings and load patterns. PerfView’s Diff feature lets me subtract one trace from another, showing which methods increased in CPU time or which types grew in memory.
Automating Comparisons with the Command Line
The GUI is fine for exploration, but for repeatable analysis, I script PerfView. The command PerfView.exe /Diff baseline.etl regression.etl /Out:diff.xml produces an XML report I can parse in a CI pipeline. I focus on methods with a delta greater than 5% and an absolute inclusive time over 100 ms. These thresholds filter out noise. The report becomes a checklist for the developer who introduced the change. No opinions, just data.
For memory, the diff shows you which types increased in count and total size. I once caught a leak of SemaphoreSlim objects this way. The diff showed 50,000 more instances in the regression trace than in the baseline. The root cause was a missing Dispose in a fire-and-forget task. Without the diff, the object count looked normal because the process had been running for days and the absolute numbers were large. The baseline made the leak obvious.
Production-Safe Practices and Team Integration
Collecting data is the technical part. Acting on it without causing an outage is the engineering part. I never run PerfView on a production server without a pre-agreed rollback plan. The collection commands are stored in a runbook, reviewed by the ops team, and tested on a canary node. The trace files are pulled immediately after collection and analyzed offline. Leaving an ETW session running for days because someone forgot to stop it is a common and dangerous mistake.
I also insist that every PerfView analysis ends with a concise report: the trace duration, the collection flags, the key findings, and the recommended change. Screenshots of the flame graph are not enough. The report must link to the exact stack frames and object types, so any engineer can reproduce the analysis. When I train teams, I make them write this report before they touch any code. It forces clarity.
FAQ
What is the difference between CPU Samples and CPU Stacks in PerfView?
CPU Samples collects ETW sampling events at a configurable interval (default 1 ms) and shows you a flat list of methods with inclusive and exclusive time. CPU Stacks is the flame-graph visualization of the same data, showing call hierarchies. Use the flat list for cost ranking and the flame graph for understanding the call path that led to the cost. They are two views of the same underlying trace data.
How can I collect a trace without blocking the finalizer thread?
When you enable the GCCollectOnly provider for a heap snapshot, PerfView triggers a full blocking garbage collection across all generations. This suspends managed threads, including the finalizer, until the collection completes. To avoid blocking, do not use GCCollectOnly in production. Instead, use the HeapSnapshot provider without the forced GC flag, which takes a snapshot of the current heap state without inducing a collection. The trade-off is that you will see dead objects that have not yet been collected, which can complicate the analysis.
Why does my PerfView trace show high CPU in ntoskrnl.exe and not in my managed code?
High kernel CPU (ntoskrnl.exe or ntdll.dll) typically indicates that your application is spending time in Windows system calls rather than managed code. Common causes include excessive context switching due to thread contention, heavy file I/O, or socket operations that are not async. Use the Thread Time and Context Switch views in PerfView to correlate the kernel stacks with your managed threads. You may find that a seemingly innocent File.ReadAllBytes call is triggering synchronous I/O that burns kernel CPU.
Can PerfView analyze dumps from Linux containers?
PerfView can open managed memory dumps (`.dmp` files) collected on Linux via `dotnet-dump` or `createdump`, provided the dump is in the Windows minidump format. You must have the correct version of `mscordaccore.dll` and `mscordbi.dll` from the target runtime accessible to PerfView. The tool will prompt you for the DAC path if it cannot locate them automatically. Symbol resolution may require manual configuration of `_NT_SYMBOL_PATH`. CPU traces (ETL files) are Windows-only; for Linux, use `perf` or `dotnet-trace` and convert the format if needed.