A bit of background first.
I used to work as a data engineer at a large tech company, managing over 400 e-commerce crawlers simultaneously. Every crawler had checkpoints, health checks, and a unified pipeline framework. I ran Airflow for scheduling, dbt for data transformation, Scrapy for crawler management. I wasn't just "familiar" with the design philosophy behind these tools — I was "getting paged at 3am to fix a production incident" familiar.
Then I left to freelance solo, and started using Claude Code to build tools for myself.
"Write me a script that pulls my emails every day and saves them as Markdown."
Ten minutes. It ran. Then I asked for calendar sync, then messaging apps, then audio transcripts, then family shared files.
Five data pipelines. Five shell scripts. Each one generated by Claude Code in under ten minutes.
I even wired up CI/CD — GitHub Actions on a cron, fully automated.
You might ask: why didn't I use Airflow? dbt? Python?
Because I thought, "this isn't production, it's just a personal tool, why bother going all out."
The guy who managed 400 crawlers vibe-coded his way into 5 frameworkless scripts and thought nothing of it.
They all ran, so what was there to think about?
Three months later I came back and noticed one thing.
I was afraid to touch any of them.
The Sweet Trap of Vibe Coding
Vibe coding has a particular quality: it makes you feel like the problem is solved.
Claude Code generates a working script, CI/CD schedules it, GitHub Actions shows green checkmarks every day, and you feel like you've shipped. Done. Moving on.
But "runs" and "maintainable" are two different things.
Month one: all five pipelines running, everything's great.
Month two: I added a new field, updated the email script, forgot to touch the calendar script. No visible breakage, still running.
Month three: the messaging pipeline's tracker broke. Malformed JSON — but the script had || true everywhere, so errors were silently swallowed. It kept running, kept showing green checkmarks, just stopped syncing anything.
I didn't notice it was broken for over a month.
Five scripts, five separate tracking mechanisms. Some used JSON, some plain text, one had nothing at all. Each script was generated independently, each one inventing its own format, with no awareness of the others.
It got worse. Inside Claude Code I also had a set of skills doing similar things. When I updated the shell script's output format, the skill didn't know. When the skill added new logic, the shell didn't follow. Two systems doing the same job, permanently out of sync.
If this had been production, I would have had Airflow for unified scheduling and dbt for data quality checks from day one. But because it was "just a personal tool," I used nothing.
I finally had to admit: I had grown a system through vibe coding, with no design underneath.
Refactor One: Extract Shared Functions (Fixed!)
Alright, time to clean this up.
I told Claude Code: "Extract the duplicated logic across the five scripts into a shared lib.sh."
Fifteen minutes later, a beautiful shared function library: health checks, checkpoint read/write, error reporting — all centralized.
Ran it. No problems. Done!
Then I changed the checkpoint format in the email pipeline.
The calendar pipeline exploded.
The reason: shell has no classes. All state lives in global variables. EMAIL_LAST_RUN, CALENDAR_LAST_RUN — five pipelines' worth of variables crammed into one shared namespace. Change one name, and the other four might silently read the wrong value.
DRY problem solved. Architecture problem: untouched.
Refactor Two: Simulate a Framework (This Time for Sure!)
If shell functions aren't enough, let me build a framework.
Each pipeline declares a config, a generic runner executes it:
PIPELINE_NAME="email"
COLLECT_FN="collect_email"
The runner uses bash indirect expansion to call dynamically:
collect_fn="${PIPELINE_NAME}_collect"
${!collect_fn}
Do you know what ${!collect_fn} does? It's a bash magic syntax that takes a variable's value and uses it as another variable name to dereference. Sounds clever. Debugging it makes you want to quit.
No IDE support, no type checking. When something breaks you reach for set -x, print every line, then squint through 100 lines of trace trying to find where it exploded.
I sat in front of the screen for three seconds and it hit me:
I was simulating Python in shell.
So why not just use Python?
This was not the life I wanted.
The Coffee Machine Epiphany
After the second failed refactor, I went to make coffee.
Standing there, I suddenly thought: wait, what are these five pipelines actually doing?
Stripped to the core, four things:
- Fetch something from out there
- Parse what came back
- Remember where you left off, so next time you continue from there
- Report when done: still alive, fetched N items
If you've built web crawlers, this sounds familiar. Because that's exactly how Google crawls pages — fetch, parse, track progress, report status. It's also exactly how Scrapy's Spider is designed.
My five pipelines and web crawlers are the same problem. I'm just not fetching HTML — I'm fetching emails, calendar events, message history.
I managed crawlers for two years. When I built my own tools, I forgot all of it.
Not because I didn't know. Vibe coding let me skip the "thinking" step. Claude Code would run at a single sentence, so I kept talking, kept running, and never stopped to ask: "what's the best practice here?"
The answer was in my head the whole time. The crawler architecture pattern. The Python community has solved this hundreds of times. Scrapy's Spider pattern is a ready-made answer.
I just forgot to use it.
Refactor Three: Back to Basics
No more forcing shell to do things it wasn't built for. I wrote a BasePipeline class in Python.
Imagine you're onboarding a new engineer to manage crawlers. You don't have them start from scratch — you give them a skeleton: "You define what to fetch, the framework handles progress tracking and status reporting."
That's exactly this:
class BasePipeline:
def fetch(self):
"""Fetch from external source — you define this"""
raise NotImplementedError
def parse(self, raw):
"""Parse what came back — you define this"""
raise NotImplementedError
def run(self):
"""The skeleton: fetch → parse → checkpoint → report"""
checkpoint = self.load_checkpoint()
raw_items = self.fetch()
new_items = [i for i in raw_items if i["id"] > checkpoint]
for item in new_items:
try:
self.save(self.parse(item))
except Exception as e:
self.report_error(e, item)
self.update_checkpoint(new_items[-1]["id"])
self.report_health(len(new_items))
Adding a new pipeline means inheriting the skeleton and filling in two methods:
class CalendarPipeline(BasePipeline):
def fetch(self):
return pull_ics_events(self.config["url"])
def parse(self, raw):
return {"title": raw["summary"], "start": raw["dtstart"]}
Health checks, checkpoints, error reporting — all inherited. Never written twice.
The Result
The morning after the refactor, I ran pytest. 0.04 seconds. 15 tests, all green. I stared at the screen for five seconds.
Deleted 1,490 lines (583 shell + 907 duplicated skill code). Added 1,042 lines (668 Python + 374 tests).
Net: -155 lines. Deleted more than I added. 29 files changed.
Adding a new pipeline now takes 15 minutes and doesn't touch anything else.
AI Convenience vs Systematization: Not Either/Or
This whole episode clarified something.
LLMs are fast — one sentence generates code, ten minutes and it runs. But for some things, structured Python modules and rule-based frameworks are simply more reliable.
It's not that AI is bad. They solve different layers of problems:
LLMs are good at: understanding vague requirements, handling unstructured data, generating prototypes quickly
Systematic frameworks are good at: state tracking, error recovery, consistency guarantees, testability
My pipelines now use LLM for some fetch and parse steps — like turning a pile of messy message history into a structured to-do list. But checkpointing and health checks are deterministic Python code, period.
You don't want an LLM tracking your bookmarks. It'll hallucinate one.
AI handles understanding. The system handles memory. You need both — but don't mix up which one goes where.
This balance doesn't reveal itself in one shot. It's an iterative process. My first version was all AI-generated (too dependent). My second was all hand-designed (too slow). The third finally hit the sweet spot: LLM for prototypes, human brain for architecture, Python for the system.
Shell Scripts: Knowing Their Place
To be clear: shell isn't bad.
It's fast, it's light, it runs everywhere. No virtualenv, no pip install, just chmod +x and go.
But once your prototype has lived for more than three months, other parts of the system depend on it, and it needs to track state — it's no longer a prototype.
Shell script is a prototype, not an architecture.
The signal: when your shell script starts containing python3 -c "import json...", it's time to switch languages. Four of my five scripts had that line. It took me six months to see it.
Freelancing solo makes this trap particularly easy to fall into, because no one pumps the brakes. No code review, no one asking "are you confident this design holds up long-term?" You're rushing, so you just merge. You're busy, so you tell yourself "I'll refactor this later."
Later never comes — until the pipeline quietly breaks while you're sleeping.
In the AI Era, Crawler Thinking Matters More Than Ever
One last thing I think is genuinely important.
Traditional crawlers scrape HTML. Structured, DOM-based, parseable with CSS selectors. But my five pipelines fetch emails, calendar events, message logs, audio transcripts.
No consistent format. Some of it isn't even text.
Before, you'd write a custom parser for every source — tedious. But now with multimodal AI, you can drop a voice recording in and get back a transcript plus summary. Dump a thread of emails and pull out three action items.
The crawler skeleton — fetch → parse → checkpoint → report — hasn't changed. But the "parse" step has been completely rewritten by AI.
Before, parsing meant regex and custom parsers. Now parsing can be "hand the raw data to an LLM and let it understand it."
So any information source can become a pipeline. Social media mentions, client voice messages, photos of handwritten notes — the framework is the same. Just inherit BasePipeline.
Crawler thinking matters more in the AI era, not less. Not because the technology changed, but because there's so much more you can now crawl.
Back to That Cup of Coffee
This post isn't a Python crawler tutorial.
The point is that moment at the coffee machine. When you suddenly remember you already knew the answer — you were just moving so fast, carried along by the tool's speed, that you forgot to stop and think.
If you're using AI to write code — Claude Code, Cursor, Copilot, whatever — here's my one piece of advice:
After it runs, make a cup of coffee, and ask yourself: how would I design this if AI didn't exist?
The answer in your head is probably worth more than the code AI gave you.
I'm Young — data engineering background, now freelancing solo with AI, managing 10 projects at once. If you're building AI-assisted tools and finding your system getting harder to understand over time, let's talk.
