Speaking a Language Educators Understand
I'm giving a talk at an AI literacy education forum hosted by CommonWealth Parenting — one of the largest education media platforms in Taiwan with 22,000+ subscribers. The topic is my experience mentoring high school students building real products with AI.
After finishing the script, I started thinking: the audience is full of teachers, principals, and education researchers. They have their own professional vocabulary and frameworks. If I just talk about "GitHub Issues" and "Code Review," they'll understand the words but might not connect them to their own practice.
I wanted to try describing what I do using their language. To find common ground and make the conversation click.
So I started searching for PBL and AI education literature. I found 6 papers closely related to what I'd been doing, and reading them gave me another angle to look at my own work from the past six months.
Background: A Decade in EdTech, Not Education Research
I've spent ten years in education technology. At Junyi Academy (Taiwan's equivalent of Khan Academy), I led the architecture for assessment systems and AI learning features. I've helped experimental schools build curricula and teaching systems from scratch. I've mentored engineering teams and non-technical people alike.
When I designed this intern program, I drew on years of accumulated experience and judgment. But I'm not an education researcher — I don't have time to design control groups, run statistics, or publish papers. My validation method is more practical: can students ship working software, and can they work independently after I leave?
This literature search wasn't about proving I was right. It was about learning how education researchers describe similar practices — picking up their frameworks and vocabulary.
What I Did
I brought two high school students with zero coding background into a real, running Chinese reading learning platform (LingoLeap). Not a simulation — real GitHub Issues, real Code Review, real Production deployment.
The design:
- Four-tier curriculum: Bug Fix → UX Improvement → Feature Development → Technical Deep Dive. Every tier uses real Issues, not exercises
- Skill Tree: 20 skills, 600 total XP. No grades — growth tracked through a skill tree. Unlock conditions are completing PRs or passing Code Reviews
- Weekly rhythm: Pick Issues Monday → develop during the week → PR review + Skill Tree update Friday
- Code Review as teaching: No rubber-stamping. Specific feedback every time — what's good, what to change, and why
- Dual-layer AI: The platform uses AI to teach reading; the development process uses AI to assist coding. Time I save with AI goes toward mentoring
These practices came from years of mentoring experience. I didn't specifically reference education papers when designing them.
Then I Went Looking for Papers
Here are the 6 papers I found, and which of my practices they correspond to.
1. Gold Standard PBL — PBLWorks Framework
The most cited framework in PBL, defining 7 design elements. I mapped them:
| Gold Standard Element | What I Did |
|---|---|
| Challenging Problem | "How do we help struggling readers learn to read aloud?" — a real social problem |
| Sustained Inquiry | 6 months of continuous development, not a one-off assignment |
| Authenticity | Real GitHub Issues, real users, real Production |
| Student Voice & Choice | Students choose their own Issues and solutions |
| Reflection | Weekly meetings + Skill Tree self-tracking |
| Critique & Revision | Code Review = real feedback loops |
| Public Product | GitHub contribution graph — goes straight into college applications |
All seven aligned. This framework helped me organize what had been a vague sense of "I think this works" into a clearer structure.
Source: PBLWorks Gold Standard PBL
2. AI + PBL Effect Size Cohen's d = 1.30
MDPI Education Sciences (2025) surveyed 300 teachers and found AI-enhanced PBL significantly outperformed traditional PBL, with an effect size of Cohen's d = 1.30 — a "large effect" in statistics.
The paper says AI's greatest value is personalization + continuous feedback.
What I built: Skill Tree is personalized tracking, Code Review is continuous feedback. Anyone who's mentored people knows one-size-fits-all teaching rarely works. Seeing data back this up resonated with me.
3. AI × PBL × Programming Education: Engagement η² = 0.694
Frontiers in Education (2025) tested AI + PBL in programming education:
| Metric | Effect Size η² |
|---|---|
| Engagement | 0.694 |
| Intrinsic Motivation | 0.690 |
| Academic Achievement | 0.519 |
η² above 0.5 means more than half the variance is attributable to the AI-PBL intervention.
I didn't run statistics, but I observed the same phenomena: Ryan merged 2 PRs in his first week, Sean proactively traced a root cause bug in the phonetic system — nobody forced them to. They wanted to.
DOI: 10.3389/feduc.2025.1674320
4. AI in PBL Co-Design — Student Agency
CHI 2024 (ACM's top HCI conference) explored AI's role in PBL, emphasizing the importance of student agency: scaffolding, feedback, and personalization must not sacrifice students' decision-making power.
My approach: students choose their own Issues, decide their own solutions, and write their own PRs. I review but don't write for them.
5. PBL + Agile Are Natural Partners
IEEE Transactions on Education (2023) found Scrum + PBL works well in programming education — Agile's iterative development and PBL's iterative inquiry are essentially the same thing.
Our weekly rhythm (Monday pick → weekday build → Friday review) is a simplified Sprint. Anyone who's led engineering teams probably arrives at this rhythm naturally. Seeing academic research connect the two was interesting.
6. AI-PBL Improves Both Programming Skills and Critical Thinking
Springer Nature (2025) found AI-assisted PBL doesn't just teach students to code — it also improves critical thinking. AI's role is adaptive support, adjusting to each student's level.
One thing I teach interns: AI generated 400+ Issues, but not all of them should be done. Judging "what not to do" matters more than "what to do." That's critical thinking in practice.
DOI: 10.1007/s44322-025-00041-0
Seeing My Practices Through the Papers' Lens
After finishing, I pulled together a comparison table:
| What Papers Say | What I Do |
|---|---|
| AI-PBL effect size d = 1.30 | AI-assisted development + PBL curriculum |
| Personalization + continuous feedback matter most | Skill Tree + Code Review |
| Student engagement η² = 0.694 | Interns proactively chase bugs, choose Issues |
| Student agency must be preserved | Choose Issues, decide solutions, write own PRs |
| Agile + PBL are natural partners | Monday pick → weekday build → Friday review |
| AI-PBL improves critical thinking | Teaching them to judge "what not to do" |
Before, I could only say "I think this way of mentoring works better." Now I can say "this practice corresponds to PBL's Sustained Inquiry and Critique & Revision." Same meaning, but in words educators recognize.
That's the real takeaway for me — not proving I was right, but learning a more precise language to describe what I do. It'll make future collaborations with educators much smoother.
What We Did Differently
Most AI-PBL research focuses on "students using AI tools."
Our model is different: the mentor also uses AI to develop, freeing up time for mentoring. The platform uses AI to teach reading; the development process uses AI to teach programming — dual-layer AI-PBL.
And the project isn't simulated. It's a real education product running in Production. Code students write actually gets used by real teachers and students.
I haven't found this combination in the current literature.
What This Preparation Taught Me
Searching for papers was originally just about adding academic credibility to a talk. But after completing this review, I got more than I expected.
First, I finally have a shared vocabulary. When I used to talk with educators, I'd say "Code Review" and they'd think of grading exams. I'd say "Sprint" and they'd think of rushing. Now I know: Code Review is PBL's Critique & Revision. Sprint is the rhythm of Sustained Inquiry. Same thing, different communities, different words. Knowing their words makes dialogue possible.
Second, I saw my own blind spots. The papers mention Equity Levers — understanding students, cognitive demand, literacy, shared power. I partially addressed this in my design (letting students choose Issues, lowering the threshold without lowering the bar), but I never systematically thought about equity. My two interns had different backgrounds and learning speeds; I adjusted by instinct, without a framework to check against. That's something I can do better next time.
Third, I want to start documenting more. Not publishing papers, but at least keeping more complete records. Interns' growth trajectories, weekly changes, which designs worked and which didn't — if documented, these could help not just me but anyone trying a similar approach.
Next time I mentor interns, I'll try using PBL's 7 elements as a design checklist — not following it by the book, but having an extra lens to check whether I'm missing something.
References
- PBLWorks Gold Standard PBL — pblworks.org
- MDPI Education Sciences (2025) — 10.3390/educsci15020150
- Frontiers in Education (2025) — 10.3389/feduc.2025.1674320
- CHI 2024 (ACM) — 10.1145/3613904.3642807
- IEEE Transactions on Education (2023) — 10.1109/TE.2023.3293519
- Springer Nature (2025) — 10.1007/s44322-025-00041-0
