Rediscovering What I Do Through Educators' Language — 6 Papers vs. One Intern Project
Education·8 min

Rediscovering What I Do Through Educators' Language — 6 Papers vs. One Intern Project

Preparing for a talk at a major education forum, I searched for PBL and AI education research to describe my intern program in language educators would recognize. Six papers later, I understood my own work differently.

Y
Young Tsai

Speaking a Language Educators Understand

I'm giving a talk at an AI literacy education forum hosted by CommonWealth Parenting — one of the largest education media platforms in Taiwan with 22,000+ subscribers. The topic is my experience mentoring high school students building real products with AI.

After finishing the script, I started thinking: the audience is full of teachers, principals, and education researchers. They have their own professional vocabulary and frameworks. If I just talk about "GitHub Issues" and "Code Review," they'll understand the words but might not connect them to their own practice.

I wanted to try describing what I do using their language. To find common ground and make the conversation click.

So I started searching for PBL and AI education literature. I found 6 papers closely related to what I'd been doing, and reading them gave me another angle to look at my own work from the past six months.


Background: A Decade in EdTech, Not Education Research

I've spent ten years in education technology. At Junyi Academy (Taiwan's equivalent of Khan Academy), I led the architecture for assessment systems and AI learning features. I've helped experimental schools build curricula and teaching systems from scratch. I've mentored engineering teams and non-technical people alike.

When I designed this intern program, I drew on years of accumulated experience and judgment. But I'm not an education researcher — I don't have time to design control groups, run statistics, or publish papers. My validation method is more practical: can students ship working software, and can they work independently after I leave?

This literature search wasn't about proving I was right. It was about learning how education researchers describe similar practices — picking up their frameworks and vocabulary.


What I Did

I brought two high school students with zero coding background into a real, running Chinese reading learning platform (LingoLeap). Not a simulation — real GitHub Issues, real Code Review, real Production deployment.

The design:

  • Four-tier curriculum: Bug Fix → UX Improvement → Feature Development → Technical Deep Dive. Every tier uses real Issues, not exercises
  • Skill Tree: 20 skills, 600 total XP. No grades — growth tracked through a skill tree. Unlock conditions are completing PRs or passing Code Reviews
  • Weekly rhythm: Pick Issues Monday → develop during the week → PR review + Skill Tree update Friday
  • Code Review as teaching: No rubber-stamping. Specific feedback every time — what's good, what to change, and why
  • Dual-layer AI: The platform uses AI to teach reading; the development process uses AI to assist coding. Time I save with AI goes toward mentoring

These practices came from years of mentoring experience. I didn't specifically reference education papers when designing them.


Then I Went Looking for Papers

Here are the 6 papers I found, and which of my practices they correspond to.

1. Gold Standard PBL — PBLWorks Framework

The most cited framework in PBL, defining 7 design elements. I mapped them:

Gold Standard ElementWhat I Did
Challenging Problem"How do we help struggling readers learn to read aloud?" — a real social problem
Sustained Inquiry6 months of continuous development, not a one-off assignment
AuthenticityReal GitHub Issues, real users, real Production
Student Voice & ChoiceStudents choose their own Issues and solutions
ReflectionWeekly meetings + Skill Tree self-tracking
Critique & RevisionCode Review = real feedback loops
Public ProductGitHub contribution graph — goes straight into college applications

All seven aligned. This framework helped me organize what had been a vague sense of "I think this works" into a clearer structure.

Source: PBLWorks Gold Standard PBL


2. AI + PBL Effect Size Cohen's d = 1.30

MDPI Education Sciences (2025) surveyed 300 teachers and found AI-enhanced PBL significantly outperformed traditional PBL, with an effect size of Cohen's d = 1.30 — a "large effect" in statistics.

The paper says AI's greatest value is personalization + continuous feedback.

What I built: Skill Tree is personalized tracking, Code Review is continuous feedback. Anyone who's mentored people knows one-size-fits-all teaching rarely works. Seeing data back this up resonated with me.

DOI: 10.3390/educsci15020150


3. AI × PBL × Programming Education: Engagement η² = 0.694

Frontiers in Education (2025) tested AI + PBL in programming education:

MetricEffect Size η²
Engagement0.694
Intrinsic Motivation0.690
Academic Achievement0.519

η² above 0.5 means more than half the variance is attributable to the AI-PBL intervention.

I didn't run statistics, but I observed the same phenomena: Ryan merged 2 PRs in his first week, Sean proactively traced a root cause bug in the phonetic system — nobody forced them to. They wanted to.

DOI: 10.3389/feduc.2025.1674320


4. AI in PBL Co-Design — Student Agency

CHI 2024 (ACM's top HCI conference) explored AI's role in PBL, emphasizing the importance of student agency: scaffolding, feedback, and personalization must not sacrifice students' decision-making power.

My approach: students choose their own Issues, decide their own solutions, and write their own PRs. I review but don't write for them.

DOI: 10.1145/3613904.3642807


5. PBL + Agile Are Natural Partners

IEEE Transactions on Education (2023) found Scrum + PBL works well in programming education — Agile's iterative development and PBL's iterative inquiry are essentially the same thing.

Our weekly rhythm (Monday pick → weekday build → Friday review) is a simplified Sprint. Anyone who's led engineering teams probably arrives at this rhythm naturally. Seeing academic research connect the two was interesting.

DOI: 10.1109/TE.2023.3293519


6. AI-PBL Improves Both Programming Skills and Critical Thinking

Springer Nature (2025) found AI-assisted PBL doesn't just teach students to code — it also improves critical thinking. AI's role is adaptive support, adjusting to each student's level.

One thing I teach interns: AI generated 400+ Issues, but not all of them should be done. Judging "what not to do" matters more than "what to do." That's critical thinking in practice.

DOI: 10.1007/s44322-025-00041-0


Seeing My Practices Through the Papers' Lens

After finishing, I pulled together a comparison table:

What Papers SayWhat I Do
AI-PBL effect size d = 1.30AI-assisted development + PBL curriculum
Personalization + continuous feedback matter mostSkill Tree + Code Review
Student engagement η² = 0.694Interns proactively chase bugs, choose Issues
Student agency must be preservedChoose Issues, decide solutions, write own PRs
Agile + PBL are natural partnersMonday pick → weekday build → Friday review
AI-PBL improves critical thinkingTeaching them to judge "what not to do"

Before, I could only say "I think this way of mentoring works better." Now I can say "this practice corresponds to PBL's Sustained Inquiry and Critique & Revision." Same meaning, but in words educators recognize.

That's the real takeaway for me — not proving I was right, but learning a more precise language to describe what I do. It'll make future collaborations with educators much smoother.


What We Did Differently

Most AI-PBL research focuses on "students using AI tools."

Our model is different: the mentor also uses AI to develop, freeing up time for mentoring. The platform uses AI to teach reading; the development process uses AI to teach programming — dual-layer AI-PBL.

And the project isn't simulated. It's a real education product running in Production. Code students write actually gets used by real teachers and students.

I haven't found this combination in the current literature.


What This Preparation Taught Me

Searching for papers was originally just about adding academic credibility to a talk. But after completing this review, I got more than I expected.

First, I finally have a shared vocabulary. When I used to talk with educators, I'd say "Code Review" and they'd think of grading exams. I'd say "Sprint" and they'd think of rushing. Now I know: Code Review is PBL's Critique & Revision. Sprint is the rhythm of Sustained Inquiry. Same thing, different communities, different words. Knowing their words makes dialogue possible.

Second, I saw my own blind spots. The papers mention Equity Levers — understanding students, cognitive demand, literacy, shared power. I partially addressed this in my design (letting students choose Issues, lowering the threshold without lowering the bar), but I never systematically thought about equity. My two interns had different backgrounds and learning speeds; I adjusted by instinct, without a framework to check against. That's something I can do better next time.

Third, I want to start documenting more. Not publishing papers, but at least keeping more complete records. Interns' growth trajectories, weekly changes, which designs worked and which didn't — if documented, these could help not just me but anyone trying a similar approach.

Next time I mentor interns, I'll try using PBL's 7 elements as a design checklist — not following it by the book, but having an extra lens to check whether I'm missing something.


References

pbleducationairesearchedtech