A Day Writing Chinese TTS Rules, Deleted by One Prompt
AI Implementation in Practice·8 min

A Day Writing Chinese TTS Rules, Deleted by One Prompt

AI mispronounced 'garbage', so I built a 37-rule fix table. One afternoon prompt beat it completely — a dev log of Chinese TTS moving from rules to LLM.

Y
Young Tsai

Have you ever had that moment where you finish something, then realize you never needed to build it

That was me that day — serious work in the morning writing 37 Chinese pronunciation rules, all dead code by afternoon

If you are an engineer, this is a real case of LLMs flipping old rule-based systems If you are a PM or founder building Taiwan-Chinese products, this tells you something practical: your team might cut half of TTS maintenance cost and still improve audio quality

How it started

I have a Chinese reading product where students listen to AI read textbook content aloud

One day PM messaged me: "It pronounced garbage as lā jī, that's mainland pronunciation. We need Taiwan pronunciation lè sè"

Wrong pronunciation in a Chinese learning product is an education incident, not a "let's observe a bit more" bug

I opened TTS and listened once, yes, mainland accent. That is where this story starts


Four PM requirements

Before coding, align requirements. PM's list was clean:

  • Taiwan accent (for example, "attack" should be gōng jí, not gōng jī)
  • Natural, not robotic
  • Polyphonic characters must be correct (in "cheer", the character should be pronounced hè, not hē)
  • Names and places must be correct (Tai Tzu-ying, Chen Yen-po, week 214, year 2021)

There are more than three usable TTS providers, but we tested these first:

ProviderTaiwan AccentCostQuality
Azure Speech (zh-TW HsiaoChen)✅ NativeExpensiveGood, slightly rigid
Google Chirp3-HD❌ Mainland accentMediumVery good
Gemini Flash TTS (preview)🟡 Prompt-controllableCheapGood

We tried all three, and each had tradeoffs


Version 1: Azure + SSML <phoneme>

Our earliest production version used Azure plus SSML <phoneme> tags to manually fix polyphonic characters:

<phoneme alphabet="x-microsoft-zhuyin" ph="ㄏㄜˋ">he</phoneme>cai

One look at this XML tells the story: each polyphonic character needs manual zhuyin tagging

Azure's upside is native Taiwan pronunciation. Downsides:

  1. Expensive, bill keeps climbing every month
  2. Every new polyphonic case needs manual engineering rules
  3. Slightly rigid voice quality, less human

We ran this for a while, and I kept thinking there had to be a cheaper and more natural way


Version 2: Gemini + 37-rule homophone replacement table

When Google released Gemini Flash TTS preview, it caught my attention

Ultra-low cost (over two thousand sentences for only about $0.30), great quality, except

Default pronunciation leans mainland

So what did I do. I came up with a method that now feels pretty wild in hindsight

TTS cares about pronunciation more than literal meaning, so what if I swap characters with Taiwan-pronounced homophones

Take the word "attack" as an example:

  • Mainland accent: final character read as jī (tone 1) -> whole word becomes gōng jī
  • Taiwan accent: final character should be jí (tone 2) -> whole word should be gōng jí
  • My hack: find another character pronounced jí in Taiwan speech
  • Rewrite input and feed TTS, betting it reads by phonetics so output sounds like correct gōng jí

And this is not an edge case. Taiwan vs mainland has many same-character different-pronunciation words:

WordTaiwanMainlandDifference
garbagelè sèlā jīboth characters differ
attackgōng jígōng jīfinal syllable tone 2 vs tone 1
enterpriseqì yèqǐ yèfirst syllable tone 4 vs tone 3
researchyán jiùyán jiūfinal syllable tone 4 vs tone 1
dangerwéi xiǎnwēi xiǎnfirst syllable tone 2 vs tone 1
smilewéi xiàowēi xiàofirst syllable tone 2 vs tone 1
expectationqí dàiqī dàifirst syllable tone 2 vs tone 1
qualityzhí liàngzhì liàngfirst syllable tone 2 vs tone 4
Francefà guófǎ guófirst syllable tone 4 vs tone 3
cheerhè cǎihē cǎifirst syllable tone 4 (shout) vs tone 1 (drink)
recognizerèn shìrèn shifinal syllable full tone vs neutral tone

Same traditional Chinese surface form, noticeably different pronunciation. This has always been a pain point for Chinese TTS in Taiwan products

The references are public and official:

The problem is: we had the lists, but could not modify the model

Rule-style TTS like Azure stays rigid; high-quality Chirp3-HD stays mainland-accented; training a Taiwan-accent model ourselves was beyond budget and data

That left two paths:

  1. Character-swap workaround: trick TTS with homophone replacement — this was Version B
  2. Prompt-level control: use a model that understands natural-language instructions — later we found Gemini 3.1 preview fits this

The rest of this post is about how these two paths played out, and why path #1 looked smart but still broke other things

The hidden assumption behind my swap strategy: TTS reads text as phonetic symbols and does not care whether the word is real

_TAIWAN_TTS_REPLACEMENTS = [
    ("garbage", "music-color"),      # TW le se vs CN la ji
    ("research", "study-old"),      # final syllable: TW jiu vs CN jiu(alt tone)
    ("danger", "surround-risk"),    # initial syllable: TW wei vs CN wei(alt tone)
    ("attack", "attack-urgent"),    # final syllable tone fix requested by PM
    # ... 37 rules
]

Every line followed the same pattern: find a Taiwan-pronounced homophone and force it in, using TTS like a phonetic machine

I also added Arabic-number-to-Chinese conversion: 214 -> two hundred fourteen, so Gemini would not read it in English

Built it, ran batch, regenerated 2417 audio files, pushed to staging for PM acceptance

Then I hit a pitfall


Pitfall: one function was never called

PM came back: "'attack' still sounds mainland"

What

I checked logs and found the issue: during batch execution, the code only removed punctuation and never applied my 37 replacement rules

Gemini received original text, so the whole table did nothing and the entire batch run was wasted

One-line fix, rerun batch, another $0.30 burned, then finally correct

That round alone cost half a day


To trust the output, I built an audit system

At that point PM raised a requirement that later felt very insightful:

"I need to validate every single audio sentence"

Reasonable concern — rules might misfire, Gemini might misread, I might push an unchecked version

So I built:

  1. Append-only JSONL log (29 fields per record): original sentence sent to Gemini, replaced text, triggered rules, audio SHA, storage path, generation time, generator identity, and which old version got replaced
  2. Back-office audit page: PM can browse 2000+ lines in a browser, filter by "replacement applied", "contains numbers", "Taiwan-specific terms", inspect full metadata, and play audio directly

Later I realized the biggest value is not debugging, it is trust

Engineers can live with rough logs for debugging, but product teams accepting AI-generated output need a sense of accountability and traceability

I had underestimated that layer


Version 3: one prompt was enough

At noon I sat down and stared at that 37-rule table

Coincidentally, another engineer in the same industry wrote a post (the Gemini 3.1 TTS hands-on from evanlin.com) showing decent results by directly prompting Taiwan accent and Taiwan wording

I listened again and noticed something

Gemini default was not pure mainland accent. It was a mixed accent with some Taiwan friendliness

Only certain words (garbage, attack, research) leaned mainland. Overall it felt "mostly right, locally wrong"

So what if I only said "please read in Taiwan accent". Would that also fix those local wrong spots

I did not know, but a batch run was only 20 minutes + $0.30. Cost was too low not to test

So I built a variant system:

  • Variant A: no replacements, only prompt prefix
  • Variant B: replacement-table version
prompt prefix = "Please read the following in traditional Chinese used in Taiwan, with a warm and natural tone:"

Ran Variant A batch: 2417 sentences / ~20 minutes / ~$0.30


A/B blind listening: 44 samples, 6 categories, PM made the call

To turn "feels better" into a verifiable conclusion, I built an A/B blind-listening page

44 samples covering:

  • Taiwan pronunciation high-frequency words (garbage / research / danger / attack / enterprise / score)
  • Other Taiwan pronunciation cases (smile / as much as possible / rest / quality / recognize / knowledge)
  • Personal names with hard characters (Chen Yen-po / Yang Chun-han / Tai Tzu-ying / Chen Yu-fei)
  • Places (Tokyo / Taiwan)
  • Numbers (214 / 2021 / 100)
  • Chinese-English mixed reading (NASA / Hemsworth / PEACE / Frankenstein)

Left in green was A (prompt-only), right in orange was B (replacement-table), each pair shown with side-by-side audio elements

I should also explain who listened

Our team has two high-school engineering interns. Both can independently write React, use git, and handle PR review. Their coding ability is honestly on par with many fresh graduates entering software jobs. They are also very willing to use AI as an amplifier, not trapped by old "handcrafted code only" beliefs

For this project, they were unexpectedly well matched — high-schoolers just came out of years of short-video and YouTube consumption, and their sensitivity to mainland vs Taiwan accent differences is better than mine. For me, checking whether "attack" was tone 1 or tone 2 took multiple replays. For them, one play was enough. I have been in a tech echo chamber too long, and my accent perception got dulled

They also helped me tune prompt wording — phrases like "warm tone" and "natural narration" were iterated together. The final prompt line for A included their contributions

So this blind test was not only me. It was me + two interns + PM, and the result was consistent:

  • A sounded more natural: smoother sentence flow, warmer tone
  • B was slightly stiff: replacement characters disturbed semantic rhythm and pause patterns
  • Hard names: A surprisingly pronounced uncommon names correctly via prompt alone
  • Numbers: A made Gemini read week 214 as "week two hundred fourteen" directly, so my number-conversion rules became unnecessary

Decision: use A from now on


One detail that stayed in my head: why did "attack-urgent" sound weird

After blind listening, one intern told me: "In version B, the rewritten word sounds uncomfortable, not sure why"

It took me a while to understand

Recall why we did this replacement — mainland accent reads one syllable as tone 1, so we swapped with a Taiwan tone-2 homophone to force pronunciation

That strategy assumes: TTS only reads phonetics and ignores lexical meaning

But Gemini did not play by that rule

The rewritten token is not a real word

Old-era TTS (like Azure) roughly follows text -> phoneme -> acoustic features -> waveform pipelines (Tacotron / FastSpeech style, explained in Microsoft neural TTS survey). So for it, two strings can be mostly just two phonetic sequences — homophone substitution can work

But new LLM-based TTS (VALL-E, NaturalSpeech 3, Gemini TTS) reframes TTS as a language-model task. In Google's official blog, Gemini TTS "not only knows what to say, but how to say it," deciding delivery based on transcript context

So the point that semantic context affects prosody is backed by official docs and papers, not just my guess

As for the more specific claim "fake words damage sentence-level prosody" — that is my reasonable architecture-level inference, but I have not seen a direct benchmark paper yet. Happy to be corrected

In practice, though, the observation was clear: feed that fake rewritten word into Gemini, and prosody gets strange. The syllable is right, the breath is off. If this inference is right, the implication is simple — we fixed pronunciation by interfering with what the model is actually good at

Old TTS (Tacotron / FastSpeech family)LLM TTS (VALL-E / Gemini family)
String -> phoneme -> acoustic features -> waveformString -> context-aware interpretation -> prosody -> waveform
OOV often misread with G2P fallbackFake words do not crash, but context prosody drifts
Better with richer rules/dictionariesBetter with more precise prompts

We also noticed one subtle detail — mainland-style "mo" can carry a slight curled sound, while Taiwan style is cleaner and flatter

Variant A read it in the Taiwan way. Variant B lost that subtle friendliness after replacement

This is exactly the kind of "mouthfeel" engineers rarely discuss, but users feel immediately


Deleting one day's work

After merging the decision PR, I sat there looking at code I had written that morning

  • _TAIWAN_TTS_REPLACEMENTS table -> dead code
  • _apply_taiwan_pronunciation -> dead code
  • _numbers_to_chinese_tw -> dead code
  • _clean_for_gemini -> dead code
  • Polyphonic-audit docs built for Variant B -> PR closed, not merged
  • 5 research issues (polyphonic chars / mixed language / pausing / names / places) -> all closed with "Variant A OK"

One day earlier I thought these were core architecture. One day later they became museum pieces kept only for rollback

I did not delete them, because deletion cost was higher than keeping them — if Variant A gets complaints later, reverting one PR + switching storage path brings B back immediately


Three lessons I will carry forward

1. Try prompt first, build rules second

I spent half a day on 37 rules, and ten minutes on the prompt

Honestly, I only understood this after doing it, not before. My old habit was "rules first, prompt as fallback" — veteran engineer inertia

Now I reverse it: prompt first, rules only if prompt is not enough

LLMs push experiment cost so low that this should change development order

2. The value of audit trail is trust, not debugging

I thought provenance logs were for my own debugging. After shipping, I realized

The real user is PM, who needs the right to validate line by line

Engineers getting bitten by bugs is routine, but product teams have real anxiety signing off AI-generated content they did not verify with their own eyes

In AI products, audit trail is less a debug tool and more a sleep-at-night mechanism for non-engineers

3. Long-term value of the variant system

tts-variants.yaml + --variant A|B was built for one A/B experiment, but after using it I saw it can stay long-term

If Gemini ships new voice options, we can add Variant C If we need lower-grade-specific tone, add Variant D Storage split by folder avoids cache contamination

A well-designed one-time experiment can become infrastructure by itself

This is a principle I often underuse — spending 30 extra minutes on a light abstraction now can remove a full architecture refactor later

4. High-school interns who can code are more useful than you think

Across this whole story, several key judgments on pronunciation quality were made by two high-school engineering interns — accent differences, prompt wording, blind-test conclusions

My old mental model of "intern mentoring" was to assign low-risk chores, teach git, and call it good if the pipeline runs

Reality was the opposite — their coding ability was on par with many junior professionals, and age was an advantage:

  • Their ears were not dulled by the tech echo chamber, so accent intuition was sharper
  • Their LLM mindset was more native than ours — they were not stuck in "rules are engineering, prompts are cheating"
  • They tried wild experiments I would not try, and several became final product decisions

If you are in industry and hesitating on high-school interns, my take is: worth it. Condition: treat them like engineers, not helpers


What happened in the next few weeks (bonus)

After deleting 37 rules, where did that bandwidth go

I thought things would calm down, but more hidden problems surfaced

1. Sentence splitting also moved from regex to LLM

We originally split sentences by regex over punctuation (period, exclamation, question mark, ellipsis), maintaining many edge cases

Since prompt-style thinking won on pronunciation, I copied it to splitting — moved to semantic splitting with Opus 4.7

Same pattern as TTS: rule era needs hundreds of corner cases, LLM reads semantics and outputs 2301 clean chunks

2. Audit trail evolved into a merge gate

PM's "line-by-line acceptance" request turned into stricter engineering discipline

I wrote a verifier with --strict. After each batch regeneration, JSONL records must match cloud audio objects 1:1 (no extra, no missing, exact keys), or PR cannot merge

It evolved from "PM acceptance dashboard" into "engineering merge gate"

3. Elegant fallback for safety filter

Out of 2000+ lines, Gemini safety filter refused two lines (unclear why, content looked normal)

Fix was elegant: keep the same cache key and call Chirp3-HD (the mainland-accent provider) for those two lines

Runtime does not care about provider source, it only plays cache hits. Because variant/provider abstraction was already clean, this fallback added only ~30 lines

4. Previously hidden UX problems surfaced

Once pronunciation was fixed, deeper issues became visible:

  • Loudness jumps between paragraphs -> added loudnorm to standardize at -16 LUFS
  • Subtitle highlights drifted from audio timing -> switched to character-weighted progress + prefetch
  • TTS button had only idle/playing states, loading and error were confusing -> redesigned into 4-state UX

Those issues were always there, but hidden by the big fire of wrong accent

After putting out the big fire, small fires become visible — this is normal engineering reality


Cost comparison (final bill)

SolutionHuman MaintenanceMonthly API CostAudio QualityTaiwan Accent
Azure + SSML phonemeHigh (manual per character)ExpensiveGood, slightly rigid✅ Native
Chirp3-HD (mainland accent)0MediumVery good❌
Gemini Variant B (replacement table)Medium (37 rules + upkeep)~$0.30 per 2k linesSlightly rigid✅ via rules
Gemini Variant A (prompt-only)Zero~$0.30 per 2k linesBest✅ via prompt

A wins across the board: zero maintenance, best quality, lowest cost


My judgment

At least in this project, it felt like my old work — phoneme tagging, sentence rules, polyphonic dictionaries, homophone replacement, SSML tags — now has much lower ROI on LLM TTS

The higher-output time now goes to prompt design, audit trail, variant system, cost monitoring, and trust mechanisms

I am not saying old skills are useless. If LLM prosody becomes unstable, old phoneme-level control still works as fallback. But the main path clearly moved

Have you also built something and watched it become obsolete the same day? These moments are some of the most valuable learning artifacts in this era


If you are building Chinese education or edtech products and are stuck on AI implementation, feel free to book 30 minutes — there is a good chance the pitfall I already hit is exactly the one you are hitting now

aittschinese-ttsllmprompt-engineeringedtech