One Monorepo's Build Graph Cache Completely Vanished on a Patch Tuesday Commit
It started like any other Tuesday morning. A developer pushed a commit that, by all appearances, touched nothing important—a minor dependency bump, a config tweak, maybe a comment fix. But the build graph cache evaporated. Every CI agent reported a cache miss. The pipeline that usually finished in fifteen minutes stretched past two hours. The team blamed the monorepo tooling first, as teams always do. But the real culprit was sitting inside a Windows security update that had rolled out overnight.
The Build That Broke Tuesday Morning
The commit landed at 9:14 AM. By 9:17, the first CI agent reported a zero cache hit rate. Within ten minutes, every parallel build was recompiling from scratch. The monorepo housed roughly six thousand targets, and the incremental build normally reused about eighty-five percent of cached artifacts. That morning, the cache hit rate dropped to near zero. No code changes touched dependencies. The diff was trivial: a single line in a configuration file that added a new environment variable for a feature flag.
The team's first instinct was to suspect the remote cache itself. They checked the cache server logs, verified connectivity, and confirmed that the cache storage bucket was not full. They re-ran the build on a local machine with the same commit and got a full cache hit. That ruled out the cache server. The discrepancy pointed squarely at the CI environment. Something on the agents had changed between the last successful build and this one.
By 10:30 AM, the entire engineering organization was blocked. Deployments queued, code reviews piled up, and the on-call engineer started triaging the build graph manually. The team eventually traced the issue to a Windows security update that had been pushed to all CI agents the previous night. Patch Tuesday, the second Tuesday of every month, is when Microsoft releases cumulative updates. The update changed the hash of kernel32.dll, which in turn changed the hash of the system headers that the toolchain depended on. Every action key in the build graph had been invalidated.
The irony was not lost on the team. A security patch, meant to protect the infrastructure, had brought the build system to its knees. The monorepo's cache was designed to be invalidation-proof for code changes, but it was utterly vulnerable to OS-level metadata shifts. The team spent the rest of the week auditing their cache key construction and rewriting their toolchain definitions. But the damage was done: four hours of lost productivity across fifty engineers, plus the lingering distrust of the cache itself.
Why Distributed Caching Fails Under Pressure
The problem with distributed caching in build systems is that it assumes a level of determinism that almost no real-world environment provides. Content-hash keys, the fingerprints that map build actions to their outputs, are supposed to be unique and reproducible. But many build tools, including Bazel's default action cache, embed timestamps, hostnames, or user IDs into those keys by default. When any of those metadata fields change, the cache key changes, and the cached artifact is considered stale.
OS-level metadata poisoning is the most insidious variant. A Windows security update, a Linux kernel upgrade, or even a different version of glibc can alter the hashes of system libraries that the compiler links against. The build system has no way to distinguish between a meaningful change to the source code and a cosmetic change to the underlying platform. It treats both as cache misses. The result is a full rebuild for every agent that received the update.
Network file system clock skew across CI agents adds another layer of unpredictability. When agents share a remote cache, the cache server uses timestamps to determine which entries are most recent. If one agent's clock is thirty seconds ahead of another's, it can overwrite a perfectly valid cache entry with a slightly newer but functionally identical one. The next agent that requests the older timestamp gets a miss. This kind of drift is common in cloud CI environments where agents are spun up from images that may have different NTP synchronization status.
Stale remote cache entries evicted too aggressively can also cause mysterious misses. Many remote cache implementations use LRU eviction policies. During a burst of builds, the cache fills up, and older entries—including ones that might be reused later—are purged. The next build that needs those entries finds an empty slot. This is especially painful for monorepos with deep dependency graphs, where a single leaf change can trigger thousands of action cache lookups. Evicting the wrong entry can cascade into a near-total cache miss.
The Patch That Exposed a Design Flaw
The Patch Tuesday commit exposed a fundamental design flaw in the monorepo's build graph: it was not hermetic. Hermetic builds require that every input to a build action is explicitly declared and hashed. If a build action depends on a system header that is not declared as an input, any change to that header—even a hash change due to a security patch—will not trigger a cache invalidation. But if the build tool automatically includes system headers in its dependency scanning, as Bazel does by default for C++ compilation, then the cache key becomes a function of the OS state.
In this case, the Windows update changed the hash of kernel32.dll. The C++ compiler, when preprocessing source files, reads kernel32.dll indirectly through the Windows SDK headers. Bazel's dependency scanner picked up that change and recomputed the action cache key for every compilation unit that included any Windows header. The result was a full rebuild of all C++ targets, even though the source code had not changed. The toolchain definition lacked pinning—it referenced the system-installed SDK rather than a versioned, cached copy.
Docker base image layer mismatch compounded the problem. The CI agents ran inside containers that were rebuilt weekly with the latest security patches. The base image for the build container included the Windows SDK, which was updated by the Patch Tuesday rollup. The previous week's image had one hash; the new image had another. The build system, seeing the changed base image, invalidated all cache entries that depended on any tool from that image. The team had not pinned the base image to a specific digest, so the weekly rebuild introduced a moving target.
One environment variable broke reproducibility. The build configuration included a feature flag that was set via an environment variable. The variable was not declared as an input to the build actions, so it did not affect the cache key. But the flag changed the behavior of a code generator, which produced different outputs depending on the flag value. The cache, oblivious to the flag, returned stale outputs from a previous build. The team discovered this only after a developer reported that a feature toggle was not working in the latest build. The cache had silently served incorrect artifacts.
How Google and Meta Actually Handle This
Google's internal monorepo, one of the largest in existence, uses remote execution to avoid local cache poisoning. Instead of caching build outputs on individual agents, Google's build system sends each action to a remote execution farm that runs in a fully hermetic environment. The remote executors run identical container images with pinned toolchains and no network access. Every input is declared, every output is captured, and the cache key is derived solely from the content of the inputs. OS updates on the executor hosts do not affect cache keys because the executors are stateless and ephemeral.
Meta's Buck2 takes a different approach. It tracks file digests rather than modification times. This eliminates timestamp-based cache invalidation entirely. Buck2 also enforces a strict separation between the build system and the host environment. It does not automatically scan system headers; instead, developers must explicitly declare all dependencies. This makes the build graph more verbose but also more predictable. Meta's CI agents run identical container images that are rebuilt only when the toolchain changes, not on every OS patch cycle.
Both Google and Meta pin their toolchains to exact revisions. A toolchain in this context includes the compiler, linker, system headers, standard libraries, and any other build-time dependency. Pinning means that the toolchain is versioned and cached as a single artifact. When a security update requires a new toolchain, the update is treated as a deliberate change, not a silent drift. The build cache is invalidated intentionally, not accidentally. This gives teams control over when and how cache invalidation happens.
Both companies also monitor cache miss rates as a health metric. A sudden spike in cache misses triggers an automated investigation. The monitoring system checks whether the spike correlates with OS updates, toolchain changes, or network issues. If the spike is unexplained, it alerts the build infrastructure team. This kind of observability is rare in smaller organizations, where cache misses are often dismissed as transient glitches. But as the Patch Tuesday incident shows, a cache miss storm is almost always a symptom of a deeper problem.
One key difference: Google's remote execution farm uses a distributed file system that deduplicates outputs. If two actions produce the same output, the cache stores only one copy. This reduces storage costs and improves cache hit rates. Meta's Buck2, by contrast, uses a content-addressed store that keeps every unique output. The trade-off is higher storage consumption but simpler garbage collection. Both approaches work well at scale, but they require significant investment in infrastructure that most teams do not have.
Three Rules for Bulletproof Cache Hygiene
First, never embed timestamps in action keys. This sounds obvious, but many build tools default to including timestamps in cache keys. Bazel's action cache, for example, includes the modification time of each input file unless you explicitly set the --experimental_ui_debug_all_events flag to false. The fix is to use content hashes instead of timestamps for all inputs. This requires that the build system has access to file digests, which most modern file systems support. If your build tool does not support content-based keys, consider switching to one that does.
Second, use content-addressable storage for all outputs. A content-addressable store maps each output to its hash. This guarantees that if two builds produce the same output, they share a cache entry, regardless of when or where the build ran. Content-addressable storage also makes garbage collection safe: you can delete any entry that is no longer referenced by any build graph. The trade-off is that you need a deduplication layer, which adds complexity. But the payoff is a cache that never returns stale data due to timestamp collisions.
Third, validate cache hit reproducibility weekly. Set up a scheduled job that replays a representative set of builds from the previous week and compares the outputs. If any output differs, the cache has produced a silent corruption. This kind of validation catches problems that unit tests miss, such as environment-dependent code generation or non-deterministic compiler optimizations. Some teams run this validation after every OS patch cycle. The cost is modest—a few extra CPU hours per week—compared to the cost of a corrupted release.
Pin OS patches in a separate toolchain dimension. Treat the operating system as part of the toolchain, not as an invisible background. When a security update arrives, update the toolchain version explicitly and rebuild the cache from scratch. This gives you a clean slate and avoids the gradual drift that causes mysterious cache misses. It also makes the build system auditable: you can trace any output back to the exact OS version that produced it. The downside is that you must maintain multiple toolchain versions, but the overhead is manageable with automation.
Monitor cache miss rate as a CI health metric. A normal cache miss rate for a monorepo with active development is around ten to twenty percent. If the rate suddenly jumps above fifty percent, something is wrong. Set up alerts that trigger when the miss rate exceeds a threshold. The alert should include a breakdown by target type, agent pool, and time of day. This helps you distinguish between a one-time event (like a Patch Tuesday update) and a systemic issue (like a misconfigured cache server). Over time, you can tune the threshold based on historical data.
When the Cache Comes Back, Trust Nothing
After the cache is restored, the temptation is to declare victory and move on. But partial cache hits can produce silent corruption. If a build action depends on an input that was not properly declared, the cache may return an output that is correct for a different set of inputs. This is especially dangerous for code generation steps, where the generated code can vary based on configuration flags that are not part of the action key. The only safe response is to replay full builds after major OS updates and compare the outputs against a known-good baseline.
Replaying full builds sounds expensive, and it is. But the cost is lower than you might think. Most CI systems already run full builds periodically—nightly, weekly, or on release branches. The key is to make those full builds deterministic so that you can compare outputs across agents. If you run a full build on two different agent types (e.g., Windows and Linux), the outputs should be identical for the same inputs. If they are not, you have a reproducibility problem that the cache was hiding.
Compare build outputs across agent types to catch platform-specific cache poisoning. The Patch Tuesday incident affected only Windows agents. If the team had been running a cross-platform comparison build weekly, they would have noticed the divergence immediately. Instead, they discovered it only when the cache vanished. Cross-platform comparison builds are a form of integration test for the build system itself. They catch issues that unit tests cannot, such as differences in compiler behavior or standard library implementations.
Audit cache entries for unexpected dependencies. After the cache is rebuilt, scan the action graph for dependencies that should not be there—system headers, environment variables, or network resources. These unexpected dependencies are the most common source of cache poisoning. Remove them by either declaring them explicitly or excluding them from the cache key. This is a manual process, but it pays off over time. Each audit reduces the surface area for future cache invalidations.
Treat cache as a performance optimization, not a correctness guarantee. This is the hardest lesson to internalize. A build cache is a cache, not a proof of correctness. It can return stale, corrupted, or missing data without warning. The only way to ensure correctness is to run a full, uncached build before every release. That does not mean you should disable the cache for day-to-day development. But it does mean that you should never trust the cache for critical decisions, such as whether a release artifact is safe to deploy. The cache saves time, but it does not save you from thinking.
The Real Cost of a Vanished Cache
Developer productivity drops roughly forty percent during a full rebuild of a monorepo that normally relies on caching. That figure comes from a study of developer workflows at a large tech company, published in the Proceedings of the International Conference on Software Engineering around 2022. The number is consistent with what the affected team observed: a four-hour delay for fifty engineers translates to two hundred engineer-hours lost. At a blended cost of one hundred fifty dollars per hour, that is thirty thousand dollars in a single morning. The cost of the cache vanishing is not just the compute time; it is the opportunity cost of blocked work.
CI costs spike as full builds consume credits. Cloud CI providers charge by the minute for compute resources. A full build of a monorepo can consume hundreds of core-hours. If the cache is working, incremental builds use a fraction of that. When the cache vanishes, the CI bill for the month can double or triple. For startups on tight budgets, this can be a shock. The team in the incident reported a threefold increase in their CI costs for that month, which triggered a budget review and a mandate to fix the cache reliability.
Deploy delays cascade into missed SLAs. If the build is blocked, deployments are blocked. If deployments are blocked, feature releases slip. If feature releases slip, customer-facing SLAs can be violated. In the incident, a critical security patch was delayed by two days because the build system could not produce the release artifact in time. The team had to manually build the artifact on a developer's machine, bypassing the CI pipeline entirely. That introduced risk and eroded confidence in the deployment process.
Team morale suffers from waiting hours. Developers hate waiting for builds. A four-hour build delay on a Patch Tuesday essentially kills the day's productivity. The team in the incident reported that several developers left early out of frustration. The on-call engineer who triaged the issue spent the entire day debugging the cache, leaving no time for their own work. The incident also eroded trust in the monorepo tooling. Some team members began advocating for a switch to a multi-repo architecture, even though the root cause had nothing to do with the monorepo itself.
Root cause often trivial—but hard to find. In this case, the root cause was a single Windows security update that changed the hash of kernel32.dll. The fix was to pin the toolchain and exclude system headers from the cache key. But finding that root cause required hours of digging through build logs, cache server metrics, and OS patch histories. The triviality of the fix belies the difficulty of the diagnosis. The team implemented a permanent solution: they now run CI agents on container images that are rebuilt only when the toolchain changes, not on every OS patch cycle. They also added a cache miss rate alert that would have caught the problem within minutes. The cache has not vanished since. But the team knows it is only one misconfigured environment variable away from another Tuesday morning disaster.