When Your String Intern Pool Becomes a Memory Bomb: A WinDbg Autopsy

The alert hit at 03:14 UTC: System.OutOfMemoryException across three nodes in West Europe. The service—a content enrichment API that had hummed along for eighteen months—was cycling every four minutes. The on-call engineer grabbed a full memory dump from the last surviving instance before it collapsed. The dump weighed 2.7 GB. The managed heap, per process counters, sat at 1.9 GB. That was the first crack in the assumption: the service’s steady-state working set had never topped 400 MB. Something had anchored a massive object graph, and it had done so recently enough that the GC hadn’t reclaimed it—or couldn’t.

I loaded the dump in WinDbg Preview 1.2402.24001.0, SOS pulled from the matching runtime (Microsoft .NET 8.0.3). First command, always: !dumpheap -stat. It answers one question—which types own the heap, by instance count and total size. The output scrolled for a few seconds and stopped on a line that made the diagnosis trivial:

              MT    Count    TotalSize Class Name
00007ffc3e4c7d10   142837   342808800 System.String

142,837 string instances chewing 327 MB. That’s not a normal string population for a service that processes JSON payloads and returns small result sets. The second-largest type, System.Char[], clocked in at 18 MB. The ratio was off by a factor of twenty. These strings weren’t transient request buffers; they were long-lived, promoted objects squatting in Generation 2.

What to check first: When !dumpheap -stat shows a single type dominating the heap by an order of magnitude, don’t fire !gcroot at random instances. Sample the values first. A high count of identical strings points to interning or caching. A high count of unique strings points to unbounded generation.

I sampled twenty random addresses from the string MT list with !do. Every string was a book title. Not JSON payloads, not error messages, not log lines. Titles like The Last Ember of Dawn, Whispers Through the Static, A Crown of Rust and Bone. Syntactically plausible, emotionally evocative, entirely synthetic. The service wasn’t a publishing platform. It had no business holding a corpus of generated book titles. But somewhere in its code, a utility was producing them—and storing every one.

I needed the root. !gcroot on a sampled title traced through a ConcurrentDictionary<string, CachedTitle> held by a static field in TitleCacheManager. The dictionary’s key was the generated title string; the value, a small metadata object. The dictionary held 142,837 entries. No eviction policy. No size limit. A write-only cache fed by a generator that could produce an effectively infinite stream of unique strings. The generator was a custom utility that called an external service to produce book title ideas for internal testing of a content classification model. The irony isn’t lost: naming things is one of the two hard problems in computer science, and here the act of generating names had become the failure mode.

What would mislead you: A memory profiler that groups by allocation call stack would show the string allocations inside the HTTP client response deserialization. You’d chase the external service’s response size, the JSON parser, the string decoder. You’d miss the cache. The heap statistics tell the truth: the strings are alive, not transient. The call stack tells you where they were born, not why they survived.

Distinguishing Pathological Interning from Legitimate Caching

String interning is a legitimate optimization when the set of possible values is small and bounded. Configuration keys, enum names, HTTP header names—canonical candidates. The CLR’s internal interning table is limited and garbage-collected under certain conditions, but a custom ConcurrentDictionary used as an intern pool has no such guardrails. The diagnostic signature of a pathological intern pool shows up in three heap statistics:

  1. String count dominates total object count. In a healthy service, strings are numerous but short-lived. They live in Gen 0 and Gen 1 and get collected fast. When strings become the plurality of the heap by instance count, they’ve been promoted.
  2. String total size dwarfs other types. The 327 MB of strings in this dump was 17× the next-largest type. That ratio is a red flag before you even inspect the values.
  3. Gen 2 heap size is disproportionately large. Run !eeheap -gc. In this dump, Gen 2 was 1.4 GB. Gen 0 and Gen 1 combined were under 50 MB. The promotion pressure came from the cache, not from request allocation.

Legitimate caches have bounded size, eviction policies, or time-to-live constraints. A legitimate cache of generated titles might hold the last 1,000 results for a few minutes. It wouldn’t hold every title ever generated. The absence of any eviction logic is the forensic fingerprint of a memory bomb.

How the Cache Fragmented the Heap

The strings themselves weren’t the only cost. Each CachedTitle value object contained a DateTime and a List<string> of tags. The dictionary’s internal buckets array and the linked-list nodes for collision chains added another 40 MB. Total retained graph: roughly 380 MB. But the fragmentation damage was worse.

Run !dumpheap -type System.String -min 85000 to find strings on the Large Object Heap. I found 1,200 strings over 85,000 bytes. These weren’t individual titles; they were the internal char[] arrays of the dictionary’s buckets after resizing. The dictionary had resized its internal storage dozens of times as it grew from 0 to 142,837 entries. Each resize allocated a new, larger array on the LOH. The old arrays became unreachable, but the LOH doesn’t compact by default in .NET 8 workstation GC. The free blocks left behind fragmented the LOH, inflating the process’s committed memory even after collections.

!heapstat -inclUnrooted confirmed 92 MB of free LOH blocks interleaved with live arrays. The GC couldn’t satisfy a subsequent large allocation request—likely another dictionary resize or a large request buffer—because no single free block was large enough. That was the proximate cause of the OutOfMemoryException: not total memory exhaustion, but LOH fragmentation blocking a contiguous allocation.

Reconstructing the Code Path from the Dump

With the root identified, I needed to confirm the code path that fed the cache. The TitleCacheManager static constructor had registered a Timer callback that called an internal GenerateTitlesAsync method every 30 seconds. The method fetched 10 titles from an external service, called GetOrAdd on the dictionary for each, and logged the count. No TryRemove, no size check, no MemoryCache with a sliding expiration. The timer had fired 14,284 times since the last process start, producing 142,840 titles (three duplicates were rejected by GetOrAdd).

The timer callback was visible in the thread pool queue snapshot from !threadpool. One thread was executing the callback at the moment of the dump. Its call stack, retrieved with !clrstack, showed:

OS Thread Id: 0x4a38 (42)
        Child SP               IP Call Site
000000C1B7F3E8B8 00007ffc3e1a4b44 [GCFrame: 000000c1b7f3e8b8] 
000000C1B7F3E9A0 00007ffc3e1a4b44 [HelperMethodFrame_1OBJ: 000000c1b7f3e9a0] System.Threading.TimerQueueTimer.Fire()
000000C1B7F3EAC8 00007ffc1a2c3e90 System.Threading.TimerQueueTimer.FireNextTimers()
...

The stack confirmed the timer was active and had just fired. The dictionary’s count field, read from the object with !do, was 142,837. The math matched: 14,284 timer fires × 10 titles per batch = 142,840, minus three duplicates. The evidence was self-consistent.

Why the GC Did Not Save You

A common assumption: the GC will eventually collect unused objects. The cache’s entries were all rooted by the static dictionary. They were Gen 2 objects. Gen 2 collections are infrequent and expensive. The workstation GC in .NET 8 triggers a Gen 2 collection only when Gen 2 itself is full or when GC.Collect is called explicitly. The dictionary’s steady growth pushed Gen 2 to 1.4 GB, but the GC did collect Gen 2 several times during the process lifetime. Each collection freed the unreachable bucket arrays from previous resizes, but the live entries remained. The LOH fragmentation accumulated because the LOH wasn’t compacting. The GC was doing its job; the code was defeating it.

What to check first: If you suspect a cache is causing GC pressure, run !dumpheap -stat and note the Gen 2 heap size from !eeheap -gc. Then run !gcroot on a few large objects. If the root chain ends in a static field, you’ve found the anchor. The fix isn’t to tune GC settings; it’s to change the code.

Implementing a Bounded, Evictable Cache

The remediation was straightforward once the root cause was proven. The TitleCacheManager was replaced with an IMemoryCache instance from Microsoft.Extensions.Caching.Memory, configured with a size limit of 10,000 entries and a sliding expiration of 5 minutes. The size limit uses the cache’s built-in compaction logic, which evicts least-recently-used entries when the limit is exceeded. The sliding expiration ensures that entries unused for 5 minutes are removed even if the limit isn’t reached.

The key configuration:

var cache = new MemoryCache(new MemoryCacheOptions
{
    SizeLimit = 10000,
    CompactionPercentage = 0.25,
    ExpirationScanFrequency = TimeSpan.FromMinutes(1)
});

var entryOptions = new MemoryCacheEntryOptions
{
    SlidingExpiration = TimeSpan.FromMinutes(5),
    Size = 1 // Each entry counts as 1 unit toward SizeLimit
};

The SizeLimit and Size properties are critical. Without them, MemoryCache won’t enforce a count-based limit. The CompactionPercentage controls how many entries are evicted when the limit is exceeded; 0.25 means 25% of entries are removed in a single compaction pass. This prevents thrashing when the cache is near the limit.

For services that can’t take a dependency on Microsoft.Extensions.Caching, a custom ConcurrentDictionary wrapper with a periodic cleanup timer is an alternative. The cleanup must run on a background thread, iterate the dictionary, and remove entries older than a threshold. The threshold must be enforced by a timestamp stored in the value, not by relying on external wall-clock comparisons that can drift.

Validating the Fix in a Second Dump

After deploying the bounded cache, I requested a follow-up dump from the same environment 24 hours later. String count: 9,847. Total string size: 2.1 MB. Gen 2: 12 MB. The LOH had zero free blocks over 85,000 bytes. Process working set: 180 MB. The fix was confirmed not by absence of errors—the service hadn’t thrown OOM before the fix either, until it did—but by the heap statistics matching the expected steady-state profile.

What would mislead you: A performance counter dashboard showing “Gen 2 collections per second” would have looked normal throughout the incident. The collections were happening, but they weren’t reducing the live set. Counters tell you that work is being done; dumps tell you whether the work is effective.

Generalizing the Diagnostic Pattern

This incident is a specific instance of a general failure class: unbounded in-memory accumulation rooted in a static collection. The diagnostic pattern is repeatable:

  1. Identify the dominant type with !dumpheap -stat. Look for a single type whose total size is an order of magnitude larger than the next type.
  2. Sample instances with !do to understand whether values are unique, duplicated, or patterned. This distinguishes caches from genuine business data.
  3. Trace the root with !gcroot on a representative instance. If the root chain terminates in a static field, the object graph is effectively immortal.
  4. Inspect the collection that holds the objects. Check its count, its resizing history (via bucket array sizes on LOH), and any eviction logic in the source code.
  5. Measure fragmentation with !heapstat -inclUnrooted and !eeheap -gc. Distinguish true memory pressure from LOH fragmentation.
  6. Fix the code, not the GC settings. Add a size limit, a time-to-live, or an eviction policy. Validate with a follow-up dump.

This pattern applies to any unbounded collection: List<T> accumulating log events, ConcurrentBag<T> holding orphaned tasks, ConditionalWeakTable<TKey, TValue> with long-lived keys. The forensic signature is always the same: a dominant type in !dumpheap -stat, a static root in !gcroot, and a Gen 2 heap that grows monotonically.

The Broader Context: AI-Generated Content in Production Pipelines

The utility that triggered this incident was a custom wrapper around an external title generation service. The team had built it to produce training data for a content classification model. They hadn’t anticipated that the generator would be called on a timer, that the results would be cached indefinitely, or that the cache would grow without bound. The use of AI-assisted tools in content pipelines is increasingly common. The Authors Guild, in its AI Best Practices for Authors, notes that writers are experimenting with AI for drafting and research, and that ethical boundaries are still being negotiated. In production engineering, the boundary is simpler: any data you generate and store must have a defined lifetime. If you can’t state the maximum size of a collection, you have a memory leak by design.

The Reedsy Book Title Generator is an example of a tool that produces multiple title options per session, calibrated to genre and tone. Such generators are designed for human authors iterating on a manuscript, not for automated timer-driven ingestion into a server process. The mismatch between the tool’s intended use and the engineering team’s integration pattern is the root of the incident. The tool worked perfectly; the integration was the failure.

Conclusion

When a production service dies with OutOfMemoryException, the heap tells a story. In this case, the story was 142,837 book titles, each a small string that collectively formed a 380 MB memory bomb. The root was a static dictionary with no eviction policy. The fragmentation was LOH blocks left by repeated dictionary resizes. The fix was a bounded cache with sliding expiration. The diagnostic method was !dumpheap -stat → !do → !gcroot → !eeheap -gc → code change → follow-up dump. Every .NET engineer who supports production systems should be able to execute this sequence from memory. The tools are free. The method is repeatable. The only prerequisite is the willingness to read the dump before you read the code.