Debugging Large Object Heap Fragmentation Step by Step

You have plenty of free virtual memory, yet the process throws OutOfMemoryException and dies. I have seen this exact scenario on production servers more times than I care to count. The usual suspect? Large Object Heap fragmentation. I have spent too many late nights staring at WinDbg dumps, and the story is almost always the same: short-lived big arrays, pinned buffers that stick around, and allocation patterns nobody noticed during code review. Bit by bit, the LOH turns into a mess of gaps that the runtime cannot clean up. This article is the debugging workflow I use to find the real cause—no guesswork, just a sequence of steps that lead straight to the offending code.

Fragmented memory blocks visualized as scattered puzzle pieces

Understanding the Large Object Heap Layout

The LOH stores objects that are 85,000 bytes or larger. Right away, that is the first thing to remember: the threshold is strict, and anything that crosses it gets a special treatment. Unlike the Small Object Heap, the LOH does not compact during garbage collection. The runtime sweeps dead objects and sticks their freed space into a free list. Then it walks that list when a new big allocation comes in. If no single free block fits, you get OutOfMemoryException—even though the total free memory on the LOH might be several times the requested size. That is classic external fragmentation, and it is maddening the first time you see it.

Generation 2 collections include the LOH sweep. Adjacent dead objects get coalesced into bigger free blocks. But there is a catch. A pinned object sitting right in the middle acts like a concrete pillar. Coalescing cannot happen around it, so you end up with two smaller free zones instead of one large one. When the application pins LOH objects regularly—say ArrayPool<byte> buffers lent to unmanaged code or network I/O—the heap slowly becomes a checkerboard of pinned chunks and free pockets. I have watched a single pinned byte[] from a SocketAsyncEventArgs pool cut the usable free space in half for hours.

How Allocations Land on the LOH

Any object that needs more than 85,000 bytes goes straight to the LOH. That covers byte[] buffers, big strings, and collections like Dictionary<,> with thousands of entries. The runtime aligns every allocation to 8 bytes, so a 100,000-byte array plus header fields still qualifies. The trouble is that many of these allocations are fleeting. An image processing pipeline grabs a large buffer, does its work, and lets it go. The region turns into a free block. If the next request is for a buffer that is even a few bytes larger, the free block is ignored, and the heap grows. Over hours of uptime the heap size balloons while the free spaces get smaller and more scattered. It is a slow-burning problem that eventually catches fire.

Debugger interface analyzing memory segments

Step-by-Step Diagnostic Workflow

I have settled on a reliable sequence for diagnosing LOH fragmentation. The aim is simple: find the objects that block coalescing—pinned handles and long-lived temporary arrays—and trace them back to the code that created them. My toolbox is WinDbg with SOS, PerfView for heap snapshots, and ETW traces when I need to see allocations in real time.

Step 1: Capture a Memory Dump at the Failure Point

You need a full memory dump captured exactly when the OutOfMemoryException fires. For IIS-hosted apps I use <legacyCorruptedStateExceptionsPolicy> plus a small debugger script. For standalone processes, ProcDump with -e 1 does the job. Timing matters. A dump taken even a couple of minutes later can be worthless. A subsequent GC might have compacted the heap, or freed space might have been reclaimed, hiding the fragmentation that caused the crash. I learned this the hard way after chasing ghosts in a dump that was too late.

Step 2: Examine the LOH with !heapstat and !dumpheap

Open the dump in WinDbg. Load SOS with .loadby sos clr. Start with !eeheap -gc to see the GC heap sizes. Look at the LOH segment: total size and, more importantly, how much of it is free. A healthy LOH has only a sliver of free space. If the free percentage is above 30% and you are debugging an OOM crash, fragmentation is basically confirmed. I have seen numbers above 50% on servers that ran for weeks without a restart.

Now run !dumpheap -stat -type Free. This lists every free object on the LOH with its size. You are looking for a pattern: lots of small free blocks, none large enough for the allocation that failed. To see the actual addresses, use !dumpheap -type Free -min 85000. The gaps between free blocks are your suspects—pinned objects or survivors that are holding the heap open. I often jot the addresses down and cross-reference them later.

Step 3: Identify Pinned Objects Using !gchandles and !gcroot

Pinned objects are the top offender. !gchandles dumps all pinned handles. Each handle points to an object the GC cannot move. Take the address of a handle that looks suspicious—maybe there are dozens of them—and run !gcroot <address>. Follow the reference chain. It frequently ends at a SocketAsyncEventArgs buffer, a WCF internal buffer, or a custom ArrayPool that forgot to return its buffers. If you see no pinned handles, don’t relax yet. The heap can still suffer from temporary arrays that survive multiple GCs simply because a static collection or a cache holds a reference. I once found a ConcurrentDictionary that cached 200 MB of byte arrays indefinitely.

Step 4: Use !maddress to Check Free Block Coalescing

!maddress -summary gives you the state of every memory region. For the LOH, focus on regions marked Free and how they sit next to each other. When two free regions have just one allocated object between them, that object is the blocker. Use !maddress <address> on it to get its type and size. This technique is a scalpel. It saves you from scanning the whole heap and quickly points at the exact allocation that is breaking coalescing. I have cut hours-long investigations down to minutes with this.

Step 5: Trace Allocations with ETW and PerfView

Linking fragmentation to your source code requires an ETW trace. Fire up PerfView with the GCAllocationTick event turned on. Let the application run until the OOM condition is close, then stop the trace. Open the GCStats view and filter for LOH allocations. The call stack column shows the method that created the large objects. I look for methods that allocate arrays inside loops, or that hold references to temporary buffers longer than necessary. The usual trouble spots: Stream.CopyTo with a buffer size that gets promoted to LOH, XmlSerializer generating huge temporary strings, and hand-rolled BinaryWriter code that lets a MemoryStream buffer grow unchecked.

Code editor highlighting a memory allocation in C#

Practical Mitigation Techniques

Once you know the source, the fix depends on the allocation pattern. If the application honestly needs big temporary buffers, switch to ArrayPool<byte>.Shared and make absolutely sure every Rent has a matching Return. The pool reuses buffers and keeps them on the SOH when the size allows. For buffers that must go to unmanaged code, pin them only for the duration of the call: use fixed or GCHandle.Alloc with GCHandleType.Pinned and free the handle immediately. Caching pinned handles is asking for trouble.

For long-lived cache objects, I lean on WeakReference or a custom eviction policy that kicks items out before the LOH gets too chopped up. If your code uses MemoryStream heavily, set a capacity that matches your typical data size to avoid repeated resizing and reallocation. In a real pinch, you can compact the LOH manually. Set GCSettings.LargeObjectHeapCompactionMode to GCLargeObjectHeapCompactionMode.CompactOnce and then trigger a full GC. This is a sledgehammer: it halts all threads and moves large objects, which can be painfully slow. I reserve it for maintenance windows or as a one-shot recovery when fragmentation has already taken the service down.

Monitoring Fragmentation in Production

Proactive monitoring stops fragmentation from becoming a 3 a.m. pager storm. Expose a performance counter that reports the LOH free space ratio via GC.GetGCMemoryInfo(). Set an alert when the ratio climbs above 25% and the LOH size is still growing—that combination means free blocks are not getting reused. Pair this with a lightweight ETW session that logs LOH allocation rates. A sudden spike in large allocations per second often signals that fragmentation is about to bite. I have built dashboards around these exact metrics, and they have saved our team from more than one outage.

FAQ

What is the difference between LOH fragmentation and a managed memory leak?

A managed memory leak happens when references keep objects alive forever, so the heap just grows and grows. LOH fragmentation is sneakier. Objects get freed properly, but they leave holes that are too small to use. In a leak, the heap size climbs endlessly. With fragmentation, the heap size might level off, but the free blocks become useless for new allocations. I have seen cases where the LOH was half empty and still throwing OOM.

Can I force the LOH to compact like the SOH?

Not by default. The LOH normally sweeps and reuses free space without moving objects. You can opt into compaction by setting GCSettings.LargeObjectHeapCompactionMode = GCLargeObjectHeapCompactionMode.CompactOnce and then calling GC.Collect(). This forces the next full blocking GC to compact the LOH. It works from .NET Framework 4.5.1 and .NET Core 2.0 onward. But it is a stop-the-world operation—every thread pauses. Use it sparingly, and never in the middle of peak traffic.

How do I determine the exact size of the failing allocation?

When the OutOfMemoryException fires, the stack trace often shows the call—something like new byte[length]. If length is a variable, grab its value from the dump. !clrstack -a lists arguments and locals for the faulting frame. The size that could not be satisfied is the key piece of the puzzle. I have found that the failing allocation is sometimes only a few kilobytes larger than the biggest free block. That single fact explains the whole crash.

Why does my application work fine for days and then suddenly crash?

Because fragmentation builds silently. Early on, the LOH has big open spaces. Over time, pinned objects and surviving temporary arrays chop it into smaller and smaller pieces. The crash happens when a new allocation asks for a size that no single free block can provide. Often the trigger is a slightly larger request than usual—maybe a file upload that exceeds the typical buffer size—and it exposes all the fragmentation that has been accumulating for days. The application did not suddenly break; it just finally ran out of usable gaps.