How to Use PerfView Like a Senior Engineer

How to Use PerfView Like a Senior Engineer

PerfView is the lens. When a production service starts burning CPU or leaking memory, I don’t open a memory profiler first. I open PerfView. The raw ETW (Event Tracing for Windows) data it captures separates what you think is happening from what the runtime actually did. This isn’t a getting-started guide. I’m writing for engineers who have clicked around the tool but still feel like they’re fumbling when an incident wakes them at 3 a.m. By the end, the goal is simple: stop guessing, start proving.

Close-up of performance profiling charts on a monitor
PerfView captures high-resolution ETW traces that reveal thread-level timing and allocation patterns.

Setting Up for Production-Grade Collection

Nobody gets excited about trace collection setup. But senior engineers treat it like a biopsy, not a hammer. You want exactly the right data, minimal overhead, and a clean detach. If you start the collector without a plan, you’ll either miss the problem or drown in noise.

Choosing the Right Collection Mode

PerfView gives you three main modes: Collect, Heap Snapshot, and Thread Time. For a CPU-bound issue, Thread Time with the “/CpuSample” flag is the boring, correct choice. For memory investigations, a Heap Snapshot with /ForceGC gives you a steady baseline to compare against. There’s a subtler option tucked into the Collect menu: CPU Sample with GC Only Events. The overhead it adds is tiny, but it still exposes GC pauses in the trace. If you’re hunting a managed memory leak, don’t skip /GCOnly. It shrinks the trace an order of magnitude while keeping allocation ticks intact.

Preparing the Target Process

Before you launch the collector, set up your symbol paths. Use _NT_SYMBOL_PATH to point at your private symbol server or a local cache. I’ve watched a single missing PDB stretch a 10-minute investigation into two hours of dead ends. Also, make sure no other profiler is latched onto the process. ETW sessions conflict, and PerfView will fail silently sometimes. Run logman query -ets and kill any leftover sessions from earlier debugging attempts.

Collecting the Trace with Precision

Clicking through the GUI is fine when you’re poking at a dev box. But production incidents demand a scripted response. Command-line collection is the only way to guarantee you capture the same thing every time.

Command-Line Invocation for CPU Sampling

PerfView.exe /DataFile:HighCPU.etl /AcceptEula /CpuSample /MaxCollectSec:60 /NoView /Process:1234 /StopOnMaxCollectSec /Zip:true collect

That one-liner attaches to process 1234, samples CPU for 60 seconds, stops on its own, and packages the result into a compressed ETL.ZIP file. /NoView keeps the GUI from popping up—irrelevant on a headless server, but you’d be surprised how many people forget it. /Zip:true often squeezes a 600 MB trace down to 40 MB, which matters when you’re pulling it over a throttled VPN.

Adjusting Provider Keywords for Managed Code

The default .NET provider keyword set is moderate. For deeper IL-level work, add /Providers:Microsoft-Windows-DotNETRuntime:0x1F000080018:5. That keyword mask enables GC, JIT, loader, and exception events at informational level 5. It gives you thread time attribution and exception callbacks without filling the trace with verbose ETW chatter. It’s the minimal set I reach for in almost every managed-code investigation.

Engineer analyzing a detailed performance trace on a laptop
Analyzing a collected ETL trace requires methodical filtering to isolate hot paths.

Analyzing CPU Spikes Like a Surgeon

This is where most engineers lose an hour. They open the CPU stacks view and start scrolling. Thousands of methods scroll past, and nothing stands out. Senior engineers filter aggressively from the first click.

Filtering by Time Range and Thread

Go to the Events view. You’ll see the CPU utilization graph. Drag-select the spike—the exact region where CPU was pegged—right-click, and choose Filter to Selected Time Range. You’ve just thrown out the quiet periods that dilute the signal. If the spike is tied to a specific thread pool thread, filter further by ThreadID. Now the call tree shows what the CPU was executing during the incident window, not an aggregate of the entire trace.

Reading the CallTree with the “By Name” Tab

Switch to the CallTree tab and pick By Name. This collapses methods by assembly and namespace. Immediately you’ll see whether the hot path sits in your code or a third-party library. Focus on the Exc % column—exclusive CPU percentage. That’s time spent directly inside the method, not in its children. A method with high inclusive time and low exclusive time is just a dispatcher; the real work is further down. Triage from the highest exclusive percentage downward.

Spotting Async Continuations

Async code makes a mess of call stacks. A Task.Delay continuation might show up under a completely unrelated thread. Use the Thread Time (with Start Stop Activities) view. It reconstructs the logical activity chain using System.Threading.Tasks.TplEventSource events. Without it, you can misattribute 30% of CPU to a thread pool dispatch loop and never see the actual business logic sitting behind an await.

Diagnosing Managed Memory Leaks

Heap size alone is a distraction. A leak isn’t just about how much memory an object uses; it’s about rootedness. An object is a leak only if it’s reachable from a static root or an active thread and shouldn’t be.

Taking and Comparing Snapshots

Take a baseline snapshot with /HeapSnapshot /ForceGC. Exercise the application. Take a second snapshot. Open both in PerfView and use the Diff feature from the Heap Stacks view. It highlights object types whose instance count grew between the two points. Ignore transient types—byte arrays from I/O buffers come and go. You’re looking for types with a steady, monotonic increase.

Tracing Roots with the “Ref Tree”

Found a suspicious type? Open the References view and expand the Ref Tree. Start from the Static vars root or the Thread root and walk the reference chain. What you’re looking for is the unexpected holder—maybe a static ConcurrentDictionary that accumulates entries with no eviction policy. The reference graph is directed; PerfView renders it as a tree by picking one path per object, so verify a few times so you’re not staring at a coincidental reference.

Server room with blinking lights symbolizing memory pressure
Heap snapshot analysis reveals root causes that heap size alone cannot explain.

Correlating GC Pauses with Application Latency

GC pauses are the silent latency killers in managed applications. PerfView’s GCStats view logs every garbage collection during the trace, including pause duration and generation.

Interpreting GCStats Output

Open GCStats and sort by Pause MSec. You’re hunting for full blocking Gen 2 collections with pauses above 100 ms. These often point to Large Object Heap pressure. If you see frequent Gen 2 collections with high pause times, drill into the Heap Snapshot you took during the trace and examine the LOH for arrays larger than 85,000 bytes that are pinned or have long lifetimes.

Using the “GC Reason” Column

The Reason column tells you why the GC fired. AllocSmall means a small-object allocation hit the budget. Induced means someone called GC.Collect explicitly. An excessive number of induced GCs in production is a code smell—usually a well-intentioned attempt to manage memory that backfires by promoting objects to higher generations too early. Trace the call stack of the induced GC through the Any Stacks view to find the method responsible.

Advanced Analysis with Custom Views

PerfView’s built-in views handle 90% of what you’ll encounter. The remaining 10% forces you to write custom queries over the ETW data. That’s the moment the tool stops being a profiler and becomes a data platform.

Using the PerfViewExcel Extensions

Export trace data to XML and pull it into Excel with the PerfViewData.xlsx template. It lets you build pivot tables that aggregate CPU time by namespace, module, and time window simultaneously—something the native views won’t do. For example, pivot on Module!Namespace over 10-second intervals to catch a background service whose CPU consumption climbs gradually. A static call tree would average that rise into invisibility.

Querying with LINQPad and TraceEvent

For programmatic access, reference Microsoft.Diagnostics.Tracing.TraceEvent in a LINQPad script. Open an ETL file and write LINQ queries against the events:

var source = new ETWTraceEventSource("trace.etl");
var clr = new TraceEventParser(source);
clr.GCStart += (data) => {
    if (data.Depth >= 2) Console.WriteLine($"Gen2 GC at {data.TimeStamp}");
};
source.Process();

That script prints a timestamp for every Gen 2 GC. Extend it with HTTP request events from Microsoft-Windows-HttpService and you can build a latency impact report the GUI can’t generate. I keep a library of these scripts because a one-off investigation today becomes a repeatable diagnostic tomorrow.

Automating Trace Collection for CI/CD Pipelines

PerfView isn’t just a desktop tool. It can sit inside a performance regression suite. A nightly build that runs a benchmark and collects a PerfView trace catches CPU regressions before they ever hit production.

Headless Trace Collection in a Test Runner

Wrap your test in a batch script that starts PerfView collection beforehand and stops it afterward. The /StopOnPerfCounter flag lets you stop collection automatically when a performance counter crosses a threshold:

start /B PerfView.exe /DataFile:Benchmark.etl /CpuSample /StopOnPerfCounter:Process("%ProcessorTime","MyApp")>50 /NoView collect

That command watches the “%ProcessorTime” counter for “MyApp” and stops collection the moment it drops below 50%—meaning the benchmark ended. Pair it with a script that triggers PerfView’s HeapSnapshotFromProcess action to capture final memory state. The resulting ETL and GC dump files become build artifacts you can diff against the previous build.

Common Pitfalls and How to Avoid Them

Even people who’ve used PerfView for years fall into traps that waste time or produce garbage data. Here are three I see in post-mortems over and over.

Collecting Too Much Data

Enable every ETW provider at verbose level and you’ll get a trace so big PerfView can’t open it—or it takes 20 minutes to load. The temptation to “capture everything just in case” is strong, but a trace that crashes your analysis tool has zero value. Start with the minimal provider set, look at the data, and add providers only if gaps appear.

Ignoring Symbol Loading Failures

When PerfView shows ?!? instead of method names, stop. Don’t proceed. Fix the symbol path, reload the trace, and verify your own assemblies resolve. If third-party assemblies won’t resolve, try /NGenSymbols to include precompiled native images. A call stack full of question marks isn’t a basis for any conclusion.

Misreading Inclusive vs. Exclusive Time

Inclusive time lies. A method that calls Thread.Sleep can show high inclusive CPU if the sampling interval catches the sleep, but exclusive time will hover near zero. When you’re dealing with I/O-bound or waiting operations, cross-reference the Wall Clock Time view. CPU time and wall clock time diverge sharply in those scenarios.

FAQ

Why does PerfView show “Broken Stacks” in my trace?

Broken stacks happen when ETW can’t unwind the call stack. The usual suspects are missing symbols, optimized native code without frame pointers, or a corrupted stack. Fix it by ensuring your symbol path includes Microsoft public symbols (srv*C:\Symbols*https://msdl.microsoft.com/download/symbols) and enable /NGenSymbols if you’re using NGEN’d assemblies. In some cases, disabling “Suppress JIT Optimization” in your build and collecting with /ClrStackWalk improves stack resolution.

Can PerfView collect traces on Linux?

PerfView itself is a Windows application and relies on ETW, which is Windows-specific. But the trace analysis engine (TraceEvent) can open traces collected on Linux with perfcollect or dotnet-trace, as long as they’re converted to ETL format. For .NET Core applications on Linux, collect with dotnet-trace, then analyze the .nettrace file using PerfView on a Windows machine, or run dotnet-trace convert to change the format.

How do I reduce the size of a PerfView trace without losing critical data?

The most effective method is to limit ETW providers and keywords to exactly what you need. For CPU analysis, the default /CpuSample provider set is already minimal. For memory, use /GCOnly instead of a full heap snapshot unless you need object references. Also, use /MaxCollectSec to limit duration and /BufferSizeMB to set a smaller buffer (e.g., 64 MB) if the system is memory-constrained. Post-collection, /Zip:true compresses the trace significantly without losing fidelity.