After Reading Karpathy's llm-wiki (2): What My System Actually Looks Like
AI Development Practice·12 min

After Reading Karpathy's llm-wiki (2): What My System Actually Looks Like

A year and a half, a year and a half of daily use, 38 knowledge files. Not a tutorial — a collection of screw-ups. How the six-layer architecture came to be, four tools, the story behind every rule. And what I'd do differently if starting over.

Y
Young Tsai

Part 1 covered Karpathy's llm-wiki concept, and how the same architecture got used in opposite ways — Lex Fridman deletes his, I keep mine and verify it

This one opens up the system. Not an ideal architecture diagram — real file structures, design decisions, and the screw-up behind each one


Numbers

  • Used every day for a year and a half
  • 38 knowledge files — covering client status, behavior rules, tech notes, strategy
  • AI starts every session already knowing: which clients I manage, what mistakes to avoid, which contracts are expiring
  • 6 automated pipelines running daily in the background (LINE / Email / Recordings / Calendar / RSS)

Six-layer architecture

Karpathy has three layers: raw → wiki → schema

Mine has six, because managing knowledge isn't enough — AI behavior drifts too

CLAUDE.md — Always loaded. Tells AI "what you tend toget wrong"Rules — Path-triggered. Different directories, differentrulesPlugins — Official marketplace. 17 installedSkills — On-demand SOPs. 111 totalHooks — Code-enforced gates. Where AI judgment can't betrustedAgents — Isolated workers. Only when context isolationis needed

Why three extra layers? Example:

My CLAUDE.md has this rule: "When Young says redo, start over, or try a different way — STOP. Don't patch the old direction. Start fresh"

This came from a session with 48 corrections. Every time I said "redo," AI kept patching the old approach, making it worse each round

Karpathy's llm-wiki focuses on knowledge management — this kind of behavior issue isn't in its scope. But in my day-to-day, this comes up constantly


Put the right things in the right place

Karpathy's architecture keeps everything in one directory: raw/ + wiki/ + schema/ together

I manage 15+ projects, each in its own GitHub repo. Can't put everything in one place

So my architecture is distributed:

~/.claude/Global config (skills, agents, hooks shared across allprojects)young_job_maanger/Command center (reads all project status, never modifiescode)~/project/healthcare-app/Its own CLAUDE.md, its own rules, its own knowledge~/project/education-app/Same — completely independent

The command center does three things: sync all project status, dispatch tasks (via GitHub Issues), track progress. It doesn't write code, doesn't commit, doesn't push — read-only

Each project has its own CLAUDE.md with project-specific rules. When AI enters a project, it loads that project's config

Global skills and agents live in ~/.claude/ — available everywhere. Project-specific ones live in the project's own .claude/

The benefit: a lesson learned in the healthcare project becomes a global rule, and the education project automatically benefits. But healthcare's GCP config doesn't pollute education's Vercel config


From knowledge base to engineering system

Karpathy's llm-wiki solves knowledge management — and he defines that problem beautifully. My situation is different — I need more than organized knowledge, I need the whole workflow to run. So the knowledge base became one layer of a larger system

The full system also includes:

  • 6 automated pipelines — LINE messages, email, voice recordings, calendar, RSS feeds, all running daily without manual collection
  • GitHub Actions CI/CD — code gets pushed, tests run, deployment happens automatically
  • Crawler system — auto-scraping 40+ AI/tech sources, producing digests
  • Cross-platform integration — GCP Cloud Run, Vercel, local demos, three deployment environments managed with one system

Each of these is independent and can be swapped or upgraded individually, but they're integrated through the command center

Karpathy's llm-wiki solves "how to organize knowledge." My system solves "how to make one person's entire workflow run automatically" — knowledge management is part of it, but knowledge management alone wouldn't let me handle 15+ projects

llm-wiki defines knowledge management beautifully. If organizing knowledge is what you need, it's already enough. My situation involves multiple projects, multiple clients, and external system integrations — so I grew these additional layers on top of the knowledge base foundation


Four knowledge types

Every file is tagged:

feedback (behavior rules) — never delete

Once, AI read a knowledge file saying "email system awaiting OAuth2 setup" and told me the system wasn't running. One ls showed 30 days of data — system had been running fine, the knowledge file was just stale

After that I added: "Must run actual commands to check system state. Can't just read the knowledge base"

project (status tracking) — goes stale fastest

Revenue, progress, TODOs. Changes weekly. Quality checks focus here

reference (technical notes) — links break

Deployment gotchas, architecture notes. Rarely change, but URLs go dead

user (identity) — never delete

Who I am, tech stack, communication style. Loaded at every session start


Four tools

Karpathy mentions Ingest, Query, Lint

After reading his framework, I checked it against my own situation: are these three enough? Managing multiple projects, I found one more need he didn't run into — proactively discovering cross-file patterns. So I built a fourth tool

That's what distillation looks like: read someone's architecture, don't copy it, compare it to your own context, find the gap, fill it. In the process, your system grows a new capability

Quality check (Lint)

Scans all files, graded by severity:

  • 🔴 Overdue TODOs, index contradicts content
  • 🟡 Files missing from index, "in progress" but untouched 30+ days
  • 🟢 Incomplete formatting, index nearing capacity

Conversation save (File-back)

When a conversation produces deep analysis (research report, comparison table, action items), prompts to save it

Key: distill, don't copy. Only core conclusions and key data

Before saving, searches existing files first. I learned this one the hard way: same research saved in three files, updated one, forgot the other two

Index sync

Keeps the index consistent with actual files. Catches ghost entries, missing entries, and stale descriptions

Proactive growth (Evolve)

This one came from my own use case

Reads all knowledge files, finds cross-file patterns:

  • Three education projects scattered across files, never viewed together
  • Strategy says something is important but no action matches
  • "Waiting for reply" items older than 14 days

Only suggests, never auto-modifies. I don't want to wake up and find AI merged three files overnight


Hooks: every rule traces to an incident

Quality checks are after-the-fact. Hooks intercept in real-time

HookWhat it doesWhat incident created it
Completion verifierAI claims done without evidence → blockedAI said "deployed" but URL returned 404
Cleanup reminder7+ days since cleanup → reminderIndex hit 193/200 lines before anyone noticed
Plan mode enforcer3+ files changing → forced planning first29,799-line session with no plan, 8.9% correction rate
Commit guardNo explicit instruction = no commitAI committed and pushed code it shouldn't have

a year and a half of daily use, Plan Mode used exactly once. That's why a hook has to force it. Relying on AI to decide when to plan doesn't work


/dream: deep cleaning

Knowledge files bloat. 38 is the cleaned-up number

I have an operation called /dream — merge duplicates, delete stale ones, resolve contradictions

Different from "proactive growth" above: growth is the analyst (makes suggestions), dream is the cleaner (actually does it)

Last run:

  • 3 groups merged (7 files → 3)
  • 4 outdated deleted
  • 6 got missing formatting added
  • Index dropped from 193 to 187 lines

Without periodic cleaning, the knowledge base becomes a junkyard. Hook enforces it: 7+ days without cleaning = notification


92 sessions of behavior analysis

This is where I think my approach differs most from Karpathy's

His quality check scans knowledge accuracy. Mine also scans AI behavior patterns

92 session analysis:

  • 46 with zero corrections (50% clean rate)
  • Clean sessions: open-ended discussion, single clear task
  • Problem sessions: involve code commits, multiple projects, or full document output

Extracted 7 behavior rules from this data, wrote them all into system config

Every new rule or hook improves the clean rate. That's the feedback loop

One question I'm still working on: how do you test the components that have LLM inside them?

Hooks are easy — they're code, deterministic, run once and you know if it works. But a skill is a markdown prompt — how do you know its trigger conditions are still accurate? An agent is an autonomous worker — how do you know its output quality hasn't degraded?

Traditional software has unit tests. ML has loss functions. But what does "health" mean for a markdown prompt?

I don't have the answer yet. The 92-session behavior analysis was a first step, but that's looking backward, not real-time monitoring. What I'm thinking about is whether you can give each skill and agent a clear reward signal — correct trigger = positive, output gets corrected = negative — and track that trend over time

That's the problem I'm solving next


Cost

I chose Claude MAX 20x from day one — about USD $200/month. I burn through the quota every single week

This isn't a money-saving move. But do the math: roughly USD $3,600 over a year and a half, and in return one person can manage 15+ projects and 3 platforms after hours. Once the system is built, most tasks run near-automatically. No Notion, no Trello, no assistant

Two other costs:

  • Discipline — quality checks and cleanup must run regularly, or things rot
  • Learning curve — hooks are scripts, skills are document templates. This isn't an app you install


Back to Karpathy

Karpathy's llm-wiki made me look at my own system with fresh eyes. Not because he told me what to do — because he gave me a framework to understand what I was already doing

I distilled the quality check concept from his architecture, and confirmed from Lex's approach why I chose to keep mine

Your system won't look like mine, and it shouldn't look like Karpathy's either. But if you're using AI for work, one question is worth asking: is your knowledge alive between sessions, or does it start over every time?

The unsolved problem

This system has been running for a year and a half. Hooks are code — when they break, I know immediately. But skills are prompts, agents are autonomous workers — these components have AI inside them, and I still don't know how to test whether they're working correctly

That's what I'm solving next. If you're thinking about the same thing, let's talk


This is the final part of the Karpathy llm-wiki series. Part 1 here

claude-codememory-systemknowledge-managementhookspostmortemai-architecture