Part 1 covered Karpathy's llm-wiki concept, and how the same architecture got used in opposite ways — Lex Fridman deletes his, I keep mine and verify it
This one opens up the system. Not an ideal architecture diagram — real file structures, design decisions, and the screw-up behind each one
Numbers
- Used every day for a year and a half
- 38 knowledge files — covering client status, behavior rules, tech notes, strategy
- AI starts every session already knowing: which clients I manage, what mistakes to avoid, which contracts are expiring
- 6 automated pipelines running daily in the background (LINE / Email / Recordings / Calendar / RSS)
Six-layer architecture
Karpathy has three layers: raw → wiki → schema
Mine has six, because managing knowledge isn't enough — AI behavior drifts too
Why three extra layers? Example:
My CLAUDE.md has this rule: "When Young says redo, start over, or try a different way — STOP. Don't patch the old direction. Start fresh"
This came from a session with 48 corrections. Every time I said "redo," AI kept patching the old approach, making it worse each round
Karpathy's llm-wiki focuses on knowledge management — this kind of behavior issue isn't in its scope. But in my day-to-day, this comes up constantly
Put the right things in the right place
Karpathy's architecture keeps everything in one directory: raw/ + wiki/ + schema/ together
I manage 15+ projects, each in its own GitHub repo. Can't put everything in one place
So my architecture is distributed:
The command center does three things: sync all project status, dispatch tasks (via GitHub Issues), track progress. It doesn't write code, doesn't commit, doesn't push — read-only
Each project has its own CLAUDE.md with project-specific rules. When AI enters a project, it loads that project's config
Global skills and agents live in ~/.claude/ — available everywhere. Project-specific ones live in the project's own .claude/
The benefit: a lesson learned in the healthcare project becomes a global rule, and the education project automatically benefits. But healthcare's GCP config doesn't pollute education's Vercel config
From knowledge base to engineering system
Karpathy's llm-wiki solves knowledge management — and he defines that problem beautifully. My situation is different — I need more than organized knowledge, I need the whole workflow to run. So the knowledge base became one layer of a larger system
The full system also includes:
- 6 automated pipelines — LINE messages, email, voice recordings, calendar, RSS feeds, all running daily without manual collection
- GitHub Actions CI/CD — code gets pushed, tests run, deployment happens automatically
- Crawler system — auto-scraping 40+ AI/tech sources, producing digests
- Cross-platform integration — GCP Cloud Run, Vercel, local demos, three deployment environments managed with one system
Each of these is independent and can be swapped or upgraded individually, but they're integrated through the command center
Karpathy's llm-wiki solves "how to organize knowledge." My system solves "how to make one person's entire workflow run automatically" — knowledge management is part of it, but knowledge management alone wouldn't let me handle 15+ projects
llm-wiki defines knowledge management beautifully. If organizing knowledge is what you need, it's already enough. My situation involves multiple projects, multiple clients, and external system integrations — so I grew these additional layers on top of the knowledge base foundation
Four knowledge types
Every file is tagged:
feedback (behavior rules) — never delete
Once, AI read a knowledge file saying "email system awaiting OAuth2 setup" and told me the system wasn't running. One ls showed 30 days of data — system had been running fine, the knowledge file was just stale
After that I added: "Must run actual commands to check system state. Can't just read the knowledge base"
project (status tracking) — goes stale fastest
Revenue, progress, TODOs. Changes weekly. Quality checks focus here
reference (technical notes) — links break
Deployment gotchas, architecture notes. Rarely change, but URLs go dead
user (identity) — never delete
Who I am, tech stack, communication style. Loaded at every session start
Four tools
Karpathy mentions Ingest, Query, Lint
After reading his framework, I checked it against my own situation: are these three enough? Managing multiple projects, I found one more need he didn't run into — proactively discovering cross-file patterns. So I built a fourth tool
That's what distillation looks like: read someone's architecture, don't copy it, compare it to your own context, find the gap, fill it. In the process, your system grows a new capability
Quality check (Lint)
Scans all files, graded by severity:
- 🔴 Overdue TODOs, index contradicts content
- 🟡 Files missing from index, "in progress" but untouched 30+ days
- 🟢 Incomplete formatting, index nearing capacity
Conversation save (File-back)
When a conversation produces deep analysis (research report, comparison table, action items), prompts to save it
Key: distill, don't copy. Only core conclusions and key data
Before saving, searches existing files first. I learned this one the hard way: same research saved in three files, updated one, forgot the other two
Index sync
Keeps the index consistent with actual files. Catches ghost entries, missing entries, and stale descriptions
Proactive growth (Evolve)
This one came from my own use case
Reads all knowledge files, finds cross-file patterns:
- Three education projects scattered across files, never viewed together
- Strategy says something is important but no action matches
- "Waiting for reply" items older than 14 days
Only suggests, never auto-modifies. I don't want to wake up and find AI merged three files overnight
Hooks: every rule traces to an incident
Quality checks are after-the-fact. Hooks intercept in real-time
| Hook | What it does | What incident created it |
|---|---|---|
| Completion verifier | AI claims done without evidence → blocked | AI said "deployed" but URL returned 404 |
| Cleanup reminder | 7+ days since cleanup → reminder | Index hit 193/200 lines before anyone noticed |
| Plan mode enforcer | 3+ files changing → forced planning first | 29,799-line session with no plan, 8.9% correction rate |
| Commit guard | No explicit instruction = no commit | AI committed and pushed code it shouldn't have |
a year and a half of daily use, Plan Mode used exactly once. That's why a hook has to force it. Relying on AI to decide when to plan doesn't work
/dream: deep cleaning
Knowledge files bloat. 38 is the cleaned-up number
I have an operation called /dream — merge duplicates, delete stale ones, resolve contradictions
Different from "proactive growth" above: growth is the analyst (makes suggestions), dream is the cleaner (actually does it)
Last run:
- 3 groups merged (7 files → 3)
- 4 outdated deleted
- 6 got missing formatting added
- Index dropped from 193 to 187 lines
Without periodic cleaning, the knowledge base becomes a junkyard. Hook enforces it: 7+ days without cleaning = notification
92 sessions of behavior analysis
This is where I think my approach differs most from Karpathy's
His quality check scans knowledge accuracy. Mine also scans AI behavior patterns
92 session analysis:
- 46 with zero corrections (50% clean rate)
- Clean sessions: open-ended discussion, single clear task
- Problem sessions: involve code commits, multiple projects, or full document output
Extracted 7 behavior rules from this data, wrote them all into system config
Every new rule or hook improves the clean rate. That's the feedback loop
One question I'm still working on: how do you test the components that have LLM inside them?
Hooks are easy — they're code, deterministic, run once and you know if it works. But a skill is a markdown prompt — how do you know its trigger conditions are still accurate? An agent is an autonomous worker — how do you know its output quality hasn't degraded?
Traditional software has unit tests. ML has loss functions. But what does "health" mean for a markdown prompt?
I don't have the answer yet. The 92-session behavior analysis was a first step, but that's looking backward, not real-time monitoring. What I'm thinking about is whether you can give each skill and agent a clear reward signal — correct trigger = positive, output gets corrected = negative — and track that trend over time
That's the problem I'm solving next
Cost
I chose Claude MAX 20x from day one — about USD $200/month. I burn through the quota every single week
This isn't a money-saving move. But do the math: roughly USD $3,600 over a year and a half, and in return one person can manage 15+ projects and 3 platforms after hours. Once the system is built, most tasks run near-automatically. No Notion, no Trello, no assistant
Two other costs:
- Discipline — quality checks and cleanup must run regularly, or things rot
- Learning curve — hooks are scripts, skills are document templates. This isn't an app you install
Back to Karpathy
Karpathy's llm-wiki made me look at my own system with fresh eyes. Not because he told me what to do — because he gave me a framework to understand what I was already doing
I distilled the quality check concept from his architecture, and confirmed from Lex's approach why I chose to keep mine
Your system won't look like mine, and it shouldn't look like Karpathy's either. But if you're using AI for work, one question is worth asking: is your knowledge alive between sessions, or does it start over every time?
The unsolved problem
This system has been running for a year and a half. Hooks are code — when they break, I know immediately. But skills are prompts, agents are autonomous workers — these components have AI inside them, and I still don't know how to test whether they're working correctly
That's what I'm solving next. If you're thinking about the same thing, let's talk
This is the final part of the Karpathy llm-wiki series. Part 1 here
