From Prompt to Graph: What Two Years of AI Engineering Actually Looks Like
AI Engineering·12 min

From Prompt to Graph: What Two Years of AI Engineering Actually Looks Like

In 2022 everyone was studying how to write prompts. By July 2026, everyone's arguing about Graph Engineering. Five layers emerged in between — each because the previous one wasn't enough. Here's what I learned walking through all five.

Y
Young

Have you noticed that the AI world spawns a new "X Engineering" every few months?

2022 was Prompt Engineering. 2025 became Context Engineering. February 2026 introduced Harness Engineering. June was Loop Engineering. July? Graph Engineering.

Looks like buzzword churn. But after spending two years walking from layer one to layer five, I can tell you: each layer emerged because the previous one wasn't enough. Not new words — new walls.


The Big Picture: Five Layers

Prompt2022Context2025Harness2026/2Loop2026/6Graph2026/7
LayerWhat You ControlOne-liner
PromptWording of a single instructionWhat you tell the AI
ContextAll information the AI sees during inferenceWhat the AI knows
HarnessThe complete execution environment around the agent (context + constraints + tools + verification + state + feedback)Agent = Model + Harness
LoopThe autonomous cycle design of the agentHow the AI runs itself
GraphThe topology between multiple loopsHow multiple AIs collaborate

Each layer doesn't replace the previous — it stacks on top. You don't stop writing prompts when you learn Context; your prompts become one small piece inside the Context system.

Note: this table looks like five parallel layers, but the actual relationship is containment. Context ⊂ Harness (Context is the Knowledge Base component of Harness). Loop runs inside Harness. Graph connects multiple Loops together. Russian dolls, not Lego bricks.


Layer 1: Prompt Engineering (2022)

What you're doing: Figuring out how to get better answers with a single instruction.

This is everyone's starting point. Chain-of-thought, few-shot, role prompting, system prompts — the core belief was "if the wording is right, the AI will be right."

My mistake: I thought if I made the prompt precise enough, AI could do everything for me. Turns out, the same prompt breaks when you change the context. It wasn't ignoring instructions — it simply didn't know the background.

Why it's not enough: There's a ceiling on how much information one instruction can carry. You can't stuff "who I am, my style, my project state, my past decisions" into a single prompt. A prompt is one email, but your AI needs an entire office of documents.


Layer 2: Context Engineering (2025)

What you're doing: Managing what the AI "sees" during inference.

Karpathy dropped a metaphor in June 2025 that woke everyone up:

"The LLM is the CPU, the context window is RAM, and you're the OS — responsible for loading the right information."

The implication: AI output quality doesn't depend on how well you write the prompt. It depends on whether the data you feed in is correct. RAG, memory, tool output, session history — all of it is context.

My approach: I started building CLAUDE.md (system config) + rules/ (constraints) + skills/ (capabilities) + memory (persistent knowledge). Every time the AI boots, it auto-loads relevant context. No more relying on a single prompt — instead, I design "what the AI sees at startup."

Why it's not enough: Context solves "what does the AI know" but doesn't solve "what if the AI goes rogue." You can show it all the right information and it'll still commit to main, delete the wrong file, or print your secrets. Knowing ≠ behaving.


Layer 3: Harness Engineering (2026/2)

What you're doing: Designing the complete execution environment surrounding the agent.

The core concept came from Mitchell Hashimoto (creator of Terraform) in February 2026. His formula: Agent = Model + Harness

Harness isn't just guardrails. It's the entire "bridle" — context (knowledge), constraints (limits), tools (capabilities), verification (testing), state (persistence), and feedback loops (error correction). Hashimoto's core discipline:

"Every time an agent makes a mistake, you engineer a solution so that mistake becomes structurally impossible. The fix is never 'try harder next time' — it's making failure impossible by design." — Paraphrased from Mitchell Hashimoto, source

This is fundamentally different from "write a better prompt." Prompts rely on wording. Harness relies on structure.

My mistake: I initially thought Harness meant writing a few guard rails. Then I understood — a complete Harness has six components:

  1. Knowledge Base: Documents, architecture decisions, all context (Context Engineering lives here)
  2. Architectural Constraints: Linters, hooks, physical blockers (what the agent cannot do)
  3. Tools & Integrations: Tools given to the agent (MCP, CLI, browser control)
  4. Verification Gates: Tests the agent must pass before claiming "done"
  5. State Management: Cross-session memory and progress persistence
  6. Feedback Loops: Error detection, doom loop interruption, auto-correction

Key point: Context Engineering is a subset of Harness, not a parallel layer. Harness contains context + everything else that makes an agent reliable.

Why it's not enough: Harness solves "what environment does the agent run in," but the agent still needs you to push it step by step. You say one thing, it does one thing. You stop, it stops. If you want AI to complete an entire task autonomously (start to finish), you need the next layer.


Layer 4: Loop Engineering (2026/6)

What you're doing: Designing the agent's autonomous cycle.

Loop Engineering has no single creator — it's a consensus that emerged across the community in early 2026. IBM, Augment Code, and ML Mastery started using the term almost simultaneously. The core definition: design the system that auto-prompts, verifies, and reruns the agent — instead of you pushing it one message at a time.

The core shift: humans are no longer inside the loop.

Before: Human says → AI does → Human checks → Human says → AI does (human in loop) Now: Human sets conditions → AI runs autonomously (observe → act → verify → correct → repeat) → Human reviews results

My approach: I use Codex (OpenAI) for unattended long tasks — give it a spec, set up eval (acceptance criteria), and it runs TDD red-green cycles on its own. When done, I only review the results and diff. If eval fails, it rolls back automatically.

Another example: a cron job runs sync every 12 hours. The AI decides what to do (merge PRs, close issues, update dashboards) without me clicking anything.

Why it's not enough: One loop is powerful, but real tasks aren't linear. You need: one agent does research → feeds results to another agent for implementation → implementation goes to a third agent for review → review rejects back to implementation. That's not a loop — it's a graph.


Layer 5: Graph Engineering (2026/7, This Month)

What you're doing: Composing multiple Loops into a directed acyclic graph (DAG).

On 2026/7/18, Peter Steinberger asked on X: "Are we still talking loops or did we shift to graphs yet?" Hours later, Hamel Husain posted: "Loop Engineering Is Dead. Enter Graph Engineering." By the weekend, the entire timeline had courses, roadmaps, and toolkits.

The concept isn't new (workflow engines have drawn DAGs for a decade). What's new is treating agent loops as nodes in a graph:

  • nodes = individual agents (each running their own loop), plus deterministic functions, routers, human checkpoints
  • edges = data flow (whose output is whose input)
  • shared state = state shared across nodes
ResearchImplementReviewreject

My approach: Today I analyzed NVIDIA stock using a Graph:

  • Node 1: Technical analyst A (chart patterns and support/resistance)
  • Node 2: Quantitative trader B (RSI and Fibonacci)
  • Node 3: Synthesis (combines both outputs)

All three complete in < 2 minutes, each judging independently, with the final node synthesizing.

Core challenges (pits I'm currently falling into):

  1. State sharing: How do you format Node A's output so Node B can directly use it? I'm currently passing natural language — not structured enough.
  2. Error propagation: If an earlier node is wrong, downstream follows. How do you set checkpoints to stop errors from spreading?
  3. Topology design: When to parallelize (independent judgments) vs. serialize (need prior results)? Wrong design means either too slow (all serial) or poor quality (force-parallelize what should be serial).

Honestly: I'm currently in an orchestrator-worker architecture (main Claude as foreman dispatching work). Not yet multiple autonomous agents communicating through shared state in a full DAG. It's early-stage Graph, heading toward the real thing.

Why it's blowing up now: Because agents are stable enough. In 2024, letting AI run autonomously meant it crashed in three steps. By mid-2026, agents can reliably complete entire loops without breaking (most of the time). So "connecting multiple stable loops together" finally becomes practical. LangGraph has existed for three years, but only now is it truly needed.


Two Concepts That Get Confused (Not on the Main Axis)

Vibe Coding (2025/2, named by Karpathy): Don't look at code, just vibe with AI to write programs. Essentially "giving up precise control." Karpathy himself said it was outdated by 2026/2 — because real engineering can't run on vibes.

Agentic Engineering (2026/2, named by Karpathy): Vibe Coding's successor. But it's more of an umbrella term covering Harness + Loop + Graph.


Which Layer Am I Using Now?

All of them. Not five separate puzzle pieces — Russian dolls nested inside each other.

The most important correction: Harness isn't just "guardrails" — it's the complete execution environment surrounding the agent. Hashimoto's original definition is "Agent = Model + Harness" — Harness contains context, constraints, tools, verification, state, and feedback loops.

So my system looks like this:

Prompt (innermost): Every message I send to the AI. Most basic, always present.

Harness (the entire surrounding layer): The complete execution environment I've built over two years.

  • Knowledge Base: CLAUDE.md + 296 skills + 15 rules + 1,465 memory files → controls what the AI knows (= Context Engineering lives here)
  • Architectural Constraints: 84 hooks physically blocking dangerous actions + rules with NEVER/STOP soft limits
  • Tools & Integrations: gstack/OpenClaw (browser control) + 5 executor wrappers (codex/cursor/kiro/grok/gemini) + MCP
  • Verification Gates: stop guards + eval-driven acceptance + QA checklists
  • State Management: job repo (logs/events/crm/projects) for cross-session persistence + session handoff
  • Feedback Loops: error-reflection + session-correction-capture + /reflect auto-behavior modification

Loop (autonomous cycles): Codex unattended + cron sync every 12 hours + eval-driven termination. Loop runs inside Harness — Harness determines what it can do, Loop determines how it runs itself.

Graph (multi-agent collaboration): subagent pipelines + multi-model dispatch. Honestly, I'm currently orchestrator-worker (main Claude as foreman). Not yet fully autonomous agents communicating through shared state. Early-stage Graph, heading toward the real thing.

This isn't a "which layer should you learn" question — it's a "which layer are you stuck at" question.

If your AI output quality is unstable → You're stuck at Context (wrong information fed in) If your AI takes dangerous actions → You're stuck at Harness (no guardrails) If you have to manually push every step → You're stuck at Loop (no autonomous cycle) If a single AI can't complete complex tasks → You're stuck at Graph (need multi-agent collaboration)


Advice for Different Levels

Just starting with AI: Get Prompt and Context right. Write clearly what you want, give AI enough background. These two layers done well already puts you ahead of 90%.

Already using AI for work: Start building Harness. Your AI will eventually do something wrong (commit secrets, edit wrong files, give wrong answers). Hooks and rules are your insurance.

Want AI to run autonomously: Learn Loop. Design eval (how to tell if it's right) + rollback (what happens when it's wrong) + termination conditions (when to stop). Nail these three, and you can let go.

Building AI systems: Graph is your next step. Not because it's trendy — because one loop genuinely isn't enough. Real workflows have branches, merges, and rejection loops. That's a graph.


None of these layers is a silver bullet. But each one is the natural solution that emerged when the previous layer hit its ceiling.

I spent two years hitting walls to arrive at today's five-layer architecture. Hopefully this saves you some of those bruises.

AIClaude CodeAgentPrompt EngineeringContext EngineeringGraph Engineering