Claude Code Leaked 500K Lines of Source Code — 5 AI Agent Architecture Lessons I Took Away
AI Engineering·10 min

Claude Code Leaked 500K Lines of Source Code — 5 AI Agent Architecture Lessons I Took Away

Claude Code v2.1.88 accidentally shipped its full TypeScript source via npm. Memory systems, KV Cache Fork-Join, tool design, permission layers, ULTRAPLAN — cross-referencing the leak with my own multi-agent setup.

Y
Young

On March 31st, someone noticed Claude Code v2.1.88 on npm had an extra 57–60MB source map file.

The full TypeScript source code just... leaked out.

Boris Cherny (Claude Code's creator) responded: "Human error, no one was fired."

Anthropic's first move was a misfired DMCA that took down thousands of researchers' repos. They walked it back, limiting it to 1 repo and 96 forks.

The AI world went wild, but what I was watching was something else: several deep technical analyses pulled the architecture apart and explained it.

I use Claude Code daily to manage multiple parallel projects. I've built my own memory system and agent architecture. After reading these analyses, a few things jumped out.


1. Memory System: Three Layers, Not One Big File

Claude Code's internal memory system, based on cross-referencing multiple leak analyses, uses 5 context compaction strategies with 14 cache-break vectors that track when the cache needs to be invalidated.

The user-facing memory structure has three layers:

  • MEMORY.md: A small index file, capped at 200 lines / 25KB
  • Topic files: Split by subject, loaded on demand
  • Session transcripts: Searchable full conversation records

The most interesting part is autoDream — confirmed by name in the leaked code. It triggers when: more than 24 hours since last consolidation, more than 5 sessions completed, and no other consolidation running. It runs in 4 phases: Orient → Gather Recent Signal → Consolidate → Prune and Index.

Boris described this as "REM sleep" in his YC Lightcone interview. The /dream command is one of his 42 Claude Code techniques.

My own setup is almost identical — a main index pointing to topic files, with periodic cleanup runs.

One line summary: Memory shouldn't be an append-only junk pile. Someone (or some agent) needs to manage its quality.


2. KV Cache Fork-Join: Parallel Agents Are Almost Free

This was the most surprising discovery in the leak analyses.

Two concepts first:

KV Cache: When an LLM processes your prompt, it stores the computation results for each token (KV = Key-Value). Next time it sees the same prefix, it reuses the cache instead of recomputing. Simply put: "Store what you've already calculated, save money next time."

Fork-Join: Fork splits the main agent into multiple sub-agents. Join merges results back.

Combined, this is Claude Code's killer design: sub-agents inherit the main agent's KV Cache.

The cost impact:

  • 1 agent's context = 1 cache cost
  • 3 parallel agents each inherit the main agent's cache, only computing their unique delta
  • Total cost ≈ 1 shared cache + 3 individual deltas

Swyx (Latent Space founder) put it simply in his analysis: "Parallelism basically free."

This explains why my parallel agent bills weren't as high as expected — sub-agents share the bulk of the prefix context. Only their individual outputs are incremental cost.


3. Tool Design: 40+ Registered, But the Point Isn't Turning Them All On

The leak revealed 40+ registered tools with the base tool definition spanning 29,000 lines of code.

Core tools look like this:

  • AgentTool: spawn sub-agents
  • BashTool: run shell commands
  • FileReadTool, FileEditTool: read/write files
  • WebFetchTool: fetch web pages
  • SkillTool: invoke skill templates

The point is loading tools on demand, not enabling everything.

I made this mistake before — figured more tools meant a more capable agent. The opposite happened. Too many tools made the agent's choices confused, producing bizarre decisions.

Less is more. Default to few, add when needed.


4. Four Permission Levels: Not All-or-Nothing

I'd been stuck on this — caught between "give the agent more autonomy" and "the agent is messing with my repo."

The leaked code confirms 4 permission levels: Plan (read-only), Standard (requires confirmation), Auto (wildcard approval), Bypass (full access). Every bash command also passes through 23 security checks in bashSecurity.ts.

Boris himself uses auto mode — a built-in safety classifier decides what needs human approval and what can proceed autonomously.

The design philosophy is clear: Autonomy and safety aren't opposites. They're managed separately.

My approach: batch queries (reading git log, checking issues) get high autonomy. Mutations (editing code, committing) require explicit confirmation.


5. ULTRAPLAN + Three Sub-Agent Execution Models

The leaked code contains a mode called ULTRAPLAN: runs on Opus 4.6 in a cloud container with a 30-minute planning window, complete with a browser interface for real-time monitoring and plan approval/rejection.

Sub-agents aren't one-size-fits-all either. The leak confirms 3 execution models:

  • Fork: Inherits parent context + KV cache. Cheapest.
  • Teammate: Independent session, communicates via file-based mailbox (this is Agent Teams).
  • Worktree: Gets its own git worktree, one isolated branch per agent.

Boris mentioned Two-Claude planning in his YC Lightcone interview: one agent writes the plan, another acts as a staff engineer reviewing it. My approach is similar — write a spec before major architecture decisions, then open a separate session to challenge the assumptions.


My Cost Control Practices

After managing a dozen parallel projects, I mapped my agent usage to four tiers — which directly correspond to the three sub-agent models revealed in the leak:

Direct (1–2 tasks, 1x cost)

Read a file, answer a question, single edit — just do it in the main conversation. Most daily work falls here.

Task() sub-agents = Fork mode (4+ independent parallel tasks)

This is the Fork model from the leak. I use Task() daily to check status across a dozen projects in parallel — each sub-agent inherits the parent's KV cache and only computes its own delta. The bill feels like roughly 1.3–1.5x rather than scaling linearly, but that's my estimate based on principles, not official Anthropic numbers.

Worktree isolation = Worktree mode (needs independent git environment)

The leak confirms an isolation: "worktree" parameter. I use claude -w to spin up parallel dev sessions, each getting its own git worktree and merging back when done. For mass cross-file changes (migrations, renames), /batch fans out to multiple worktree agents running in parallel. Cross-repo work uses --add-dir to mount additional directories.

Agent Teams = Teammate mode (haven't used yet)

The Teammate model from the leak — independent sessions communicating via file-based mailbox. Cost is 3–7x (some report up to 15x), KV cache isn't shared.

Writing this section, I actually paused — all my projects are FastAPI + Next.js full-stack. Shouldn't simultaneous frontend-backend changes be the perfect Agent Teams use case?

Then I realized: my workflow is sequential. I change the API first, verify the response, then update the frontend to consume it. I'm not changing both sides simultaneously with real-time schema coordination. Agent Teams would earn its cost when you're under a tight deadline and need to ship both frontend and backend in one afternoon, with two agents coordinating API formats in real time — that's when paying 3–7x for speed makes sense.

But my bottleneck isn't "writing code too slowly." Across a dozen client projects, what slows me down is waiting for client replies, contracts, and meeting schedules. Agent Teams can't fix that. So Task() for parallel status checks + worktree for isolated development is the right combination for my current work pattern.

Model tiering saves even more

I hardcode model routing rules in my CLAUDE.md: status checks and file ops specify model="haiku" (cheapest), research and analysis use Sonnet, architecture decisions get Opus. Sub-agents that don't specify a model inherit the parent session's model — running Haiku-level tasks on Opus is just burning money.

For automation scripts, claude -p --bare skips CLAUDE.md and MCP loading entirely — 10x faster startup.


The Biggest Takeaway From Comparing the Leak to My Own System

Boris Cherny said something on Lenny's Podcast (February 2026) that stuck with me: "Coding is solved."

He didn't mean AI is perfect. He meant the bottleneck for writing code is no longer the technology itself — it's ideas and prioritization.

Looking at the leaked architecture, Anthropic spent enormous effort on "making agents safely autonomous," "making memory manageable," and "making parallelism cheap." These aren't AI capability problems. They're systems design problems.

Get the systems design right, and the AI capability ceiling actually opens up.

Running dozens of agents across multiple projects daily, looking back at the leaked architecture, my deepest feeling isn't "wow, look what they built." It's "we're solving the same problems."

Memory rots. Clean it up. Parallelism seems expensive. It's not as bad as you think. Autonomy needs layers. Not all-or-nothing.

These aren't Anthropic-specific insights. They're what anyone running multi-agent systems at any scale runs into.

The leak just lets us compare notes on how close our solutions are.


References

Note: Architecture details are from third-party analyses of the leaked code, cross-referenced across multiple sources. Anthropic has not officially confirmed these designs. KV cache cost estimates (~1.3–1.5x) are based on principles, not official figures.

Claude CodeAI Agentarchitecturecost-optimizationmulti-agent