They found a bug I missed in week two
Sean joined the project in his second week and reported an issue while testing the reading feature: "Some characters show the wrong pronunciation."
I thought it was a minor display bug. He kept digging. Turned out it wasn't a display problem — it was a fundamental flaw in how the system handled polyphonic characters. This bug had been sitting in the core module since day one. Three months of development, and I'd never caught it.
Because I always tested with the same set of articles. He used different ones.
This isn't because he's brilliant (though he is pretty good). It's PBL doing what PBL does: when you give students real problems, they approach them from angles you never considered.
Why bring on interns at all?
Some context.
I'm building LingoLeap, a Chinese reading platform that uses AI to help struggling readers practice reading aloud and comprehension. Tech stack: React + FastAPI + GCP, with Vertex AI Gemini powering Socratic-style tutoring dialogues.
The development team is just me. Using Claude Code for AI-assisted development, one person can deliver an entire product — during Demo 4, I merged 44 PRs in a single session.
So the question is: If one person + AI can ship everything, why bring on interns?
Two reasons.
First, this platform exists to serve education. If the person building an education product isn't willing to teach, what does that say about the product's credibility?
Second, I wanted to test something: What does PBL look like in the AI era?
PBL isn't a buzzword — mapping to 7 design elements
PBL (Project-Based Learning) gets talked about a lot. But most "PBL" looks like this: teacher designs a fake prompt, students produce a disposable report, then receive a grade.
PBLWorks (formerly Buck Institute) defined 7 Essential Design Elements for Gold Standard PBL. Here's how our intern program maps:
| Gold Standard Element | How we did it |
|---|---|
| Challenging Problem | "How do we help struggling readers learn to read aloud?" — a real social problem |
| Sustained Inquiry | 6 months of continuous development, not a one-time assignment |
| Authenticity | Real GitHub Issues, real users (teachers + students), real Production deployment |
| Student Voice & Choice | Pick your own Issues, decide your own approach, write your own PRs |
| Reflection | Weekly meetings for review + Skill Tree for self-tracking |
| Critique & Revision | Code Review = real mentor feedback loops |
| Public Product | GitHub contribution graph — goes straight into college applications |
All seven. I didn't specifically reference this framework when designing the curriculum — the practices came from years of experience in EdTech and mentoring. When I later searched for academic references for a talk, I found a high degree of alignment between the frameworks and what we'd been doing. The full literature review is in a separate post.
What does the research say?
Two important 2025 papers directly support this approach.
MDPI Education Sciences (2025) surveyed 300 teachers and found that, based on teacher assessments, AI-enhanced PBL significantly outperforms traditional PBL, with an effect size of Cohen's d = 1.30 — classified as "large." AI's greatest value isn't replacing teachers but providing personalized learning paths and continuous feedback.
Frontiers in Education (2025) tested AI + PBL in programming education and found:
- Student engagement: η² = 0.694
- Intrinsic motivation: η² = 0.690
- Academic achievement: η² = 0.519
η² above 0.5 means over half the variance is attributable to the AI-PBL intervention. This isn't a marginal difference. It's massive.
Curriculum design: four ascending tiers
I designed a four-tier progressive curriculum for our two interns — Ryan and Sean. Every tier uses real GitHub Issues, not exercises.
Tier 1: Bug Fix Challenge
The safest entry point. 5 real bugs, each with clear reproduction steps and expected behavior. Skills learned: reading someone else's code, Git operations, submitting PRs, receiving Code Review.
Ryan merged 2 PRs in his first week — both approved on first review.
Tier 2: UX Improvement
Starting to touch UI. 5 improvement tasks including responsive design and WCAG accessibility. Skills learned: reading design specs, understanding user needs, writing cross-browser CSS.
Tier 3: Feature Development
Independent feature work. 4 tasks including writing their own Acceptance Criteria in Given/When/Then format. Skills learned: the complete journey from requirements to implementation.
Tier 4: Technical Deep Dive
Tracing data flows, writing technical documentation, conducting Code Reviews, teaching others. Skills learned: systems thinking and knowledge transfer.
No grades — tracking growth with Skill Trees
Traditional education uses grades. I use Skill Trees.
20 skills, 600 total XP, distributed across four tiers. Each skill has clear unlock conditions (usually completing a specific PR or passing a Code Review).
Progress as of March 13:
- Ryan: 6/20 skills, 85 XP — all Tier 1 complete, starting Tier 2
- Sean: 10/20 skills, 215 XP — Tier 1 + most of Tier 2, starting Tier 3
I also built an interactive Skill Tree webpage where you can switch between their profiles. They update their own progress every Friday.
Why no grades?
Because grades are endpoints. Skill Trees are maps. Grades tell you "what score you got." Skill Trees tell you "where you are and where to go next." And they never max out — you can always keep climbing.
What role does AI play here?
There are two layers of AI at work, intertwined.
Layer 1: The platform itself uses AI to teach reading. Vertex AI Gemini drives Socratic dialogues with a "warm but firm" tone. AI handles real-time reading error detection (LCS diff algorithm) and fluency analysis (CPM calculation).
Layer 2: The development process uses AI. I use Claude Code as my primary development tool. One person can ship the entire product. This frees me to be a mentor instead of constantly chasing code deadlines.
Together, these create a pattern that barely exists in academic literature:
Platform uses AI to teach reading × Development process uses AI to teach programming = Dual-layer AI-PBL
The critical point: this project isn't simulated. It's not a practice exercise. It's a real education product running in Production. The code students write actually gets used by real teachers and students.
Subtraction-based development — AI opened 400+ Issues, humans judge necessity
During our March 13 weekly meeting, we discussed an interesting phenomenon.
Through AI-assisted development, the project accumulated over 400 GitHub Issues. Bugs, improvement suggestions, architecture refactors — everything.
If we did them all, we'd be working until next year.
So one thing I taught the interns: Not every Issue needs to be done. Judging "what not to do" is more important than "what to do."
This is "subtraction-based development." AI excels at generating output but struggles with judging value. The human role is filtering, prioritizing, and setting direction.
For high schoolers just starting to code, this mindset might be more valuable than any technical skill.
What do they take away?
Not a certificate. Not a grade.
GitHub contribution graph — real green dots representing real contributions.
Skill Tree screenshot — progress across 20 skills, each linked to specific PRs and Code Reviews.
A portfolio for college applications — "I contributed to a real AI reading platform used in schools" is more compelling than any class project.
But the most important takeaway is a shift in perspective:
Programming isn't an exam subject. Programming is a tool for solving real problems. And AI is an accelerator that lets you solve them faster.
What Ryan and Sean learned in these 6 months isn't "how to write React" or "how to use Git." What they learned is: when you face a problem you've never seen before, how to break it down, find resources, ask for help, and deliver.
That's what PBL actually teaches.
And AI makes it all possible — not because AI replaces learning, but because AI gives one person the capacity to be both architect and mentor simultaneously.
Postscript: for those who want to try this
If you're thinking about using this model with students or interns, a few suggestions:
Issues must be real. Fake prompts produce fake learning. Let them touch real codebases, even if they're scared at first.
Code Reviews must be thorough. Not rubber stamps. Specifically state what's good, what could improve, and why. This is the strongest implementation of PBL's Critique & Revision element.
Skill Trees > Grades. Letting students see where they are and where they can go is far more meaningful than assigning a number.
AI is your leverage, not their shortcut. You use AI to accelerate development; spend the saved time reviewing their PRs and discussing design decisions. Don't let students use AI to cut corners — that defeats the purpose of PBL.
Lower the barrier, not the standard. Tier 1 starts with simple bug fixes, but Code Review standards stay high. Low barriers get them in the door. High standards make sure they actually learn.
References
- MDPI Education Sciences (2025) — 10.3390/educsci15020150
- Frontiers in Education (2025) — 10.3389/feduc.2025.1674320
- PBLWorks Gold Standard PBL — pblworks.org
