The Case One Detective Couldn't Crack — Multi-Agent AI Workflow Guide
AI Dev Practice·7 min

The Case One Detective Couldn't Crack — Multi-Agent AI Workflow Guide

One AI window doing everything is like one detective collecting evidence, interrogating, and running forensics alone. Context explosion, blind spots, quality collapse. The fix mirrors real investigations: scout, execute, review — a complete Multi-Agent AI workflow guide.

Y
Young Tsai

A bug spanning three microservices took me three hours to diagnose — not because the technical challenge was extreme, but because one AI conversation window couldn't hold all the context. It was like a detective pushing through the door into a chaotic crime scene.

Log entries splattered across the scene like blood spatter — error messages contradicting each other, three microservices failing simultaneously, every lead pointing to a different suspect. This isn't petty theft. This is a carefully orchestrated serial case.

The detective is alone. He needs to collect evidence, interrogate suspects, run forensic analysis, and write the final report. All by himself.

That's exactly what it felt like when I ran everything through a single Claude window last year.


The Lone Detective's Dilemma

At first, it worked. Requirements, code, tests, code review, client follow-ups — all handled in one conversation window. Efficient.

Then the case got complicated.

Imagine a detective piling all evidence on a single desk. Crime scene photos, fingerprint reports, witness statements, surveillance footage, phone records. The desk keeps growing. When he digs for a critical piece of early evidence, it's buried at the bottom, and he forgets it exists.

AI conversation windows work the same way. There's a context limit. Pack requirements, code, error logs, and review notes into one session, and the AI starts losing track of earlier context. Response quality degrades gradually — like a detective hallucinating after seventy-two sleepless hours on the case.

But memory isn't the fatal flaw. Blind spots are.


You Can't Spot Holes in Your Own Confession

There's an iron rule in criminal investigation: the person who collects evidence should never be the one drawing conclusions. Because people confirm their own hypotheses. Once you decide suspect A did it, every ambiguous piece of evidence starts looking like it points to A.

AI does the same thing. When the same AI writes code and then reviews that code, it tends to confirm its own reasoning. It can't see its own blind spots.

This isn't an AI defect. It's a fundamental property of any thinking system. Self-review is always weaker than external review.

You can't spot the bug in your own code. Your colleague finds it in thirty seconds. The detective thinks his report is airtight. The prosecutor flips through it and finds three holes.

Cases get solved not because one person is a genius, but because different roles have competing objectives. Evidence collectors just collect — they don't decide guilt. Forensic analysts just analyze — they don't pick suspects. The prosecutor and defense attorney have opposing goals. That tension is what surfaces the truth.

Software companies run on the same logic. PM, engineer, QA, designer — each role defines "success" differently. That built-in friction produces better output than any single person optimizing alone.

Multi-agent AI workflows bring that investigation logic into the AI world.


Assembling the Squad

I categorize AI agents into three roles, mapping to three phases of an investigation:

The Scout Unit (Scout Agent)

First on the scene. Their job: read the codebase, analyze git history, summarize issue states, flag risk areas. They survey the scene but don't touch evidence and don't theorize. Like a forensic team pulling up the perimeter tape, photographing everything, dusting for prints — observe only, never act.

The Action Unit (Executor Agent)

They move based on the scout's report. Write code, fix bugs, build tests. The key detail: the action unit receives stripped-down intel — only what's relevant to this specific task. They don't need the full case history. Smaller context means sharper execution. Like a bomb squad that only needs to know the device's location and type, not the suspect's childhood trauma.

The Internal Affairs Unit (Reviewer Agent)

Their goal is to find problems. They don't care how the code was produced. They don't know what decisions the action unit made. They only look at the final result: hunting for gaps, logic errors, security issues. The reviewer's objectives are deliberately opposed to the executor's — one ships, the other tears apart. One closes the case, the other picks it apart with a magnifying glass.

That opposition is by design. Like internal affairs in a police department — their entire purpose is institutionalized distrust.


Three Modes of Investigation

With the squad assembled, the question becomes: how do you run the operation?

Direct Action (1-3 step cases)

A parking ticket doesn't need a task force. Changing a CSS color, updating a config value, answering a question — just do it. Splitting it across three agents is like sending ten detectives to investigate a stolen bicycle. Pure waste.

Parallel Deployment (independent tasks)

Ten projects need status checks. These ten tasks don't affect each other. Dispatch ten scout units simultaneously, aggregate the intel when they report back. This isn't about making the detectives smarter — it's about having them work at the same time.

Joint Operation (interdependent tasks)

A backend database change affects the frontend. You can't investigate these in isolation — what the backend scout discovers must reach the frontend action unit before it makes decisions. This is a cross-jurisdiction joint investigation: not just division of labor, but genuine intelligence sharing. What Team A finds changes Team B's course of action.


Case File: A Bug in Production

A concrete case. A production bug.

The old way (lone detective):

I describe the bug in one conversation window, paste the logs, ask for analysis. The AI offers a theory. I follow it — wrong lead. Paste new logs. The AI revises. Back and forth, the conversation grows longer, early critical evidence gets buried at the bottom. The detective gets lost in his own notebook.

The new way (detective squad):

The scout unit reads the logs and relevant code, files an investigation report: root cause, blast radius, recommended direction. The action unit writes the fix based on that report — its context contains only the report and affected files. Clean and focused. The internal affairs unit reviews the fix independently. They don't know what the first two units said. They evaluate the code on its own merits.

Three layers of review. Each layer works with a clean scene, undistorted by noise from the other stages. From intake to resolution, every step has an independent perspective.


The Trap: Ten Detectives Fighting Over One Clue

I've seen the temptation: if three agents are better than one, ten must be even better.

No. Ten detectives crammed into the same crime scene step on each other's footprints, contaminate evidence, and fight over who gets to interrogate the suspect. Cost explodes. Speed actually drops. And the coordination between agents becomes its own source of complexity. Every additional agent is another mind to manage, another point of failure.

Don't send ten detectives after a parking ticket.

The working principle: use the minimum number of agents to accomplish the maximum amount of work.

Before assembling a squad, ask two questions. Does this task actually need roles with different objectives? Do these tasks actually need parallelism or information sharing? If the answer to both is no, handle it directly. Don't assemble.


2026: The Detectives Learn a Common Language

Something is quietly happening: communication protocols between AI agents are being standardized.

Until recently, different AI tools were isolated precincts. Each had its own evidence format, report templates, communication channels. Making them collaborate required mountains of translation code.

Now standards are emerging — such as Anthropic's MCP (Model Context Protocol) — that let agents call tools, read resources, and hand off tasks in a consistent way. Like Interpol establishing a unified intelligence exchange format — cross-border cases used to require manual translation; now there's a shared language.

This is an infrastructure-level shift. Just as HTTP let different web servers speak the same protocol, MCP and similar agent communication standards are turning multi-agent systems from "advanced technique for experts" into "a tool anyone can reach for."


Closing Statement

Multi-agent AI workflows aren't a flashy technical demo. They solve a specific problem: one detective working until judgment fails, forever blind to their own blind spots.

The logic of role separation is nothing new. Police departments have always split investigation, forensics, and oversight. Software teams have always balanced PM, engineer, and QA. Bringing that logic into the AI world gives you a more reliable, more consistent process for working cases.

You don't need to memorize any configuration. Just hold onto this:

One detective writes the report. Another detective reviews it. A third detective verifies. Three layers, each doing their job.

That's the standard operating procedure for AI-assisted development in 2026.


Young Tsai is a freelance AI developer managing 15 client projects simultaneously. These posts are notes from what actually worked, and what didn't.

aimulti-agentsoftware-developmentautomationworkflow