Heap corruption in a .NET process is one of the most difficult classes of bugs to diagnose in production. The CLR’s managed heap provides a layer of safety, but that layer is not impenetrable. When corruption occurs, the symptoms are often delayed, misleading, and inconsistent across runs. This article examines the specific patterns that cause managed heap corruption and the diagnostic techniques that expose them.

What Heap Corruption Looks Like
Heap corruption rarely announces itself at the point of origin. Instead, you observe secondary effects: an AccessViolationException in unrelated code, a corrupted method table pointer during garbage collection, or an OutOfMemoryException despite adequate memory. The GC assumes the heap is coherent. When it is not, the collector’s traversal of object graphs produces undefined behavior.
Common observable symptoms include:
- Random
NullReferenceExceptioninstances in code paths that cannot produce null references under normal conditions. - GC crashes (SegFault in
coreclr!WKS::gc_heap::mark_throughor similar frames). - Object header corruption visible via
!DumpObjshowing inconsistent method tables. - Heap verification failures when running
!VerifyHeapin SOS.
Primary Corruption Patterns
Pattern 1: Unsafe Code Overwriting Managed Objects
The most straightforward corruption pattern involves unsafe blocks that use pointer arithmetic to write past the bounds of a managed object. When a fixed statement pins a byte array and code writes beyond its allocation, the overwritten memory belongs to an adjacent object on the heap.
Consider this common mistake:
fixed (byte* ptr = buffer)
{
// Writing beyond the buffer's actual size
for (int i = 0; i <= buffer.Length; i++) // Off-by-one
{
ptr[i] = ComputeValue(i);
}
}
The off-by-one error writes one byte past the allocation. On the small object heap, this corrupts the object header of the next heap object. The corruption may not surface until the GC compacts the heap and attempts to relocate the damaged object, which could be minutes or hours after the write occurred.
Pattern 2: P/Invoke Marshaling Mismatches
When native interop declarations specify incorrect buffer sizes or layout attributes, the runtime copies more data into a managed buffer than the buffer can hold. This is especially common with struct marshaling and StringBuilder capacity parameters.
A frequent scenario involves [Out] parameters where the native side writes more data than the managed side allocated:
[DllImport("legacy.dll")]
public static extern int GetData(
[Out] byte[] buffer,
int bufferSize // Declared correctly but native code ignores it
);
If the native implementation unconditionally writes 4096 bytes regardless of bufferSize, any managed buffer smaller than that will be overwritten. The overflow corrupts adjacent heap objects. Diagnosing this requires comparing the declared interop signature against the actual behavior of the native function, often requiring access to the native source or binary analysis.

Pattern 3: Pinning-Induced Fragmentation Leading to False Corruption
Long-lived pins prevent the GC from relocating objects. Over time, this fragments the managed heap and can produce conditions where the GC’s internal bookkeeping appears inconsistent. While not corruption in the strict sense, pinning-induced fragmentation causes !VerifyHeap to report anomalies and produces allocation failures that resemble corruption.
The pattern typically involves:
- A
GCHandleof typePinnedthat is never freed. - Long-lived
fixedpointers in ausingscope that spans a large code block. - Repeated pinning of the same object in a high-throughput loop, preventing compaction.
Excessive pinning is documented in Microsoft’s .NET GC fundamentals as a known source of performance degradation, but its role in producing corruption-like symptoms is less widely understood.
Pattern 4: COM Interop and RCW/CCW Lifetime Violations
Runtime Callable Wrappers (RCWs) and COM Callable Wrappers (CCWs) manage object identity across the managed-native boundary. When a COM object is released prematurely (for example, via explicit Marshal.ReleaseComObject called too many times), the managed RCW still holds a reference. Subsequent calls through the RCW operate on freed native memory, which can corrupt the native heap and, indirectly, the managed heap if the native code writes back into shared buffers.
This pattern is particularly insidious because the crash occurs in native code, far from the managed call site that triggered it. The stack trace shown in a crash dump points to the native DLL, obscuring the managed origin.
Diagnostic Techniques
Using SOS and SOSEX
The !VerifyHeap command in SOS is the primary tool for detecting managed heap corruption. It walks every object on the heap, checking method table pointers, object sizes, and sync block consistency. When corruption is present, !VerifyHeap reports the specific object and the nature of the inconsistency.
!VerifyHeap — Full heap validation!DumpObj <address> — Inspect a suspect object!GCWhere <address> — Determine which generation an address belongs to!DumpHeap -stat — Statistical overview of heap contents!ObjSize <address> — Check an object’s true size including references
When !VerifyHeap identifies a corrupted object, note the method table it reports. If the method table pointer points to something that is not a valid method table, !DumpObj on the previous heap object often reveals the source: it may be a buffer that was overwritten, with the overflow spilling into the next object’s header.
Enabling GC Stress and Heap Verification
In pre-production environments, the COMPLUS_GCStress environment variable forces the GC to collect more aggressively, surfacing corruption sooner. Setting COMPLUS_GCStress=3 causes the GC to run a full collection on every allocation, which transforms delayed corruption into immediate failure.
The COMPLUS_HeapVerify variable enables additional runtime heap checks. When set to 1, the CLR validates heap consistency at each GC, providing an earlier signal than a crash dump taken after the fact.

Memory Dump Analysis Workflow
When a production process crashes with an access violation, capture a full dump immediately. Use procdump -ma (Windows) or createdump (.NET 5+) with full memory. The dump must include all heap memory; minidumps without full heap data are not useful for corruption analysis.
Once loaded in a debugger:
- Run
!analyze -vto identify the crash context. - Run
!VerifyHeapto locate all corrupted objects. - For each corrupted object, examine the preceding object with
!DumpObj. - Check the method table of the corrupted object:
!DumpMT -md <method_table_address>. - If the method table is invalid, search the dump for byte patterns that match the corrupted region to find the code that wrote there.
This workflow is methodical but time-consuming. The key discipline is to resist the temptation of fixating on the crashing thread. The crash is a symptom; the corruption is the disease, and it originated elsewhere.
Prevention Strategies
Preventing heap corruption requires addressing the root causes rather than working around symptoms:
- Eliminate unsafe code where possible. Replace pointer-based buffer manipulation with
Span<T>andMemory<T>, which provide bounds checking by default. When unsafe code is unavoidable, add explicit bounds assertions before every write operation. - Validate P/Invoke signatures against native headers. Automated tools like the dotnet/pinvoke library provide verified signatures for common Win32 functions. For custom interop, cross-reference every parameter size and direction against the native declaration.
- Minimize pinning duration. Pin objects for the shortest possible scope. Never store pinned GC handles in long-lived collections. If you must pin across an async boundary, consider copying the data instead.
- Avoid explicit COM reference management. Let the GC manage RCW lifetimes rather than calling
Marshal.ReleaseComObject. When explicit release is required for correctness, audit the reference count carefully.
FAQ
Can heap corruption occur without any unsafe code?
Yes. P/Invoke signature mismatches, COM interop lifetime violations, and bugs in the CLR itself can all corrupt the managed heap without any unsafe blocks in your C# code. A native dependency writing past a buffer allocated by the runtime's marshaling layer is the most common cause in purely safe managed projects.
Why does !VerifyHeap sometimes miss corruption?
!VerifyHeap checks structural consistency: method table pointers, object sizes, and sync blocks. It cannot detect semantic corruption where the bytes within an object are wrong but the object's header is intact. For example, if unsafe code overwrites the contents of a string without corrupting its header, !VerifyHeap reports no errors, yet the string's contents are garbage.
Is heap corruption more common on x86 or x64?
Both architectures are susceptible, but the manifestation differs. On x86, the smaller address space makes heap denser, so buffer overruns are more likely to corrupt an adjacent object's header. On x64, the same overrun may write into padding or alignment slack, sometimes going unnoticed. The underlying bug exists on both platforms; only the probability of immediate detection varies.
Conclusion
Diagnosing .NET heap corruption in production demands a methodical approach that looks past the crash symptom to find the originating write. The four patterns covered here—unsafe pointer overruns, P/Invoke mismatches, pinning fragmentation, and COM lifetime violations—account for the majority of corruption cases I have investigated. Each has a distinct signature in a crash dump, and each is preventable with disciplined interop practices and bounds verification. When corruption does occur, the SOS verification commands combined with a full memory dump provide the evidence needed to locate and fix the root cause.