cd /news/artificial-intelligence/ai-in-practice-shadowing-assessment-… · home › topics › artificial-intelligence › article
[ARTICLE · art-141586] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

[AI in Practice] Shadowing Assessment with Gemini 3.5 Transcribe: Let Song Lingo Listen to You Read Lyrics

A developer added a shadowing-scoring feature to Song Lingo, a Japanese-learning web app, by integrating Google's newly released Gemini 3.5 Transcribe speech-to-text model to compare learner pronunciation against reference lyrics. The developer chose the batch `gemini-3.5-transcribe` model over the live streaming variant and multimodal Flash models because it returns stable word-level timestamps, costs roughly 40 tokens per call, and does not consume the app's 100-call daily TTS quota. The post also details the Files API upload and `interactions` call pattern, and notes the model's cited 2.6% average word error rate for full transcription.

by read11 min views3 publishedSep 29, 2026

In the first post, I used Gemini 3.8 Flash TTS to create Song Lingo: paste a YouTube MV URL, Gemini transcribes the lyrics, adds furigana, translation, and grammar, and then a teacher designed with voice design demonstrates pronunciation sentence by sentence. In the second post, I deployed it to Cloud Run and locked it with IAP so only I can use it.

After deployment, I opened a roadmap on GitHub with three items: mobile layout, shadowing score, and flashcards. The mobile layout is finished; this post is about the second item.

The teacher demonstrates, but I never know if I'm pronouncing it correctly. Listening to the teacher ten times is not as good as saying it once and being told where you're wrong. This is a key step in Song Lingo's evolution from "listening to songs and reading lyrics" to "practicing pronunciation."

Google released Gemini 3.5 Transcribe on 8/26, which consists of two models specialized for speech-to-text:

Model Usage
gemini-3.5-transcribe Recorded audio files, transcribed in full
gemini-3.5-transcribe-live Real-time streaming via Live API WebSocket, transcribing while speaking

Key features:

smart handles self-corrections, removes filler words, and auto-formats; verbatim transcribes word-for-word. The official blog cites Artificial Analysis: average Word Error Rate (WER) of 2.6% for full transcription and 4.0% for streaming; compared to the previous generation Chirp 3, the time to get final results is 70% faster. Pricing has not been announced.

The actual call method is the same as TTS via the interactions API, but audio files must first be uploaded using the Files API:

uploaded = client.files.upload(file=audio_path, config={"mime_type": "audio/m4a"})
interaction = client.interactions.create(
    model="gemini-3.5-transcribe",
    input=[{"type": "audio", "uri": uploaded.uri, "mime_type": uploaded.mime_type}],
    generation_config={
        "transcription_config": {
            "language_codes": ["ja-JP"],
            "mode": {"type": "verbatim", "timestamp_granularities": ["word"]},
        }
    },
)

Many formats are supported; audio/webm happens to be the format for Chrome recordings, and audio/m4a corresponds to AAC recorded by iPhone Safari.

For shadowing scores, there were three candidates:

Model Suitable for Use in Shadowing
gemini-3.5-transcribe Transcribing after recording the whole segment ✅ Has word-level timestamps, takes about 5s in testing
gemini-3.5-transcribe-live Real-time streaming A line of lyrics is only a few seconds; transcribing after finishing is enough. Maintaining WebSockets behind Cloud Run and IAP is much more troublesome.
gemini-3.8-flash and other multimodal models Listening to audio directly to provide text feedback Results vary every time, no guaranteed word-level timestamps, not suitable as the core for scoring.

I finally chose gemini-3.5-transcribe. As a specialized speech-to-text model, the results are stable, it provides word-level info, and it only uses about 40 tokens per call, not consuming the 100-call daily quota for TTS.

The most expensive lesson from the first post was: the web version of TTS followed the Python SDK's way of reading output_audio, but that field didn't exist in the REST response, burning a whole day's quota in an hour.

So this time, I wrote a rule directly in the roadmap issue: Verify the parsing logic with real API responses before going live.

Test audio was generated using macOS's built-in say command with sentences I wrote myself: one saying "今日は晴れです" (Today is sunny) and another intentionally saying "今日は雨です" (Today is rainy):

say -v Kyoko -o ok.aiff "今日は晴れです"
afconvert -f m4af -d aac ok.aiff ok.m4a # Same AAC format as iPhone recordings

Then I intercepted the raw HTTP content sent and received by the SDK and printed the response structure. I discovered two things not clearly explained in the documentation that would have caused errors if followed blindly.

output_text is assembled by the SDK again Documentation says transcription results are in interaction.output_text. The actual REST response looks like this:

steps[] → { type: "model_output",
            content[] → { type: "text", text: "今日は晴れです。",
                          annotations[] → { type: "word_info", text: "今日",
                                            start_index: 0, end_index: 6,
                                            start_offset: "0.100s", end_offset: "0.400s" } } }

output_text, like output_audio last time, is a convenience field the Python SDK assembles from steps. Knowing this beforehand, I read directly from steps[].content[].

The start_index for the word "今日" is 0 and end_index is 6, not 0 to 2. Positions are calculated using UTF-8 bytes, where one CJK character takes 3 bytes. If sliced according to JavaScript string indices, Japanese text would be completely misaligned.

These two things can only be known by looking at real responses. This time, spending two API calls meant no more guessing after going live.

The official smart mode handles self-corrections and removes filler words. While a plus for meeting minutes, it's a minus for shadowing: if you mispronounce and then correct yourself, smart might only keep the corrected version; if you miss a particle, it might fill it in smoothly. Shadowing needs "what you actually said," so use verbatim.

Setting the original text as custom vocabulary seems like it would improve accuracy, but it makes the model biased toward hearing the correct answer, making it harder to catch mistakes. Also, documentation states custom vocabulary and word-level timestamps cannot be used together.

Transcription results might be in Kanji "今日" or Hiragana "きょう"; both are correct. Comparing only text would mark correct pronunciations as wrong.

So for Japanese, I compare both ways simultaneously; it counts as correct if either matches:

The comparison itself uses Python's built-in difflib.SequenceMatcher for character alignment, mapping results back to each word. English and Korean are compared word-by-word.

The first version only checked "if the word's characters matched." Consequently, "晴れ" in "今日は雨です" was marked as missed, when it was actually replaced by another word.

After changing to record alignment results for each character (match, replaced, deleted), they can be distinguished:

Scenario Transcription Result Judgment
Correct 今日は晴れです。 All correct, 100 points
Replaced a word 今日は雨です。 "晴れ" mispronounced, 75 points
Skipped middle word 今日はです "晴れ" missed, 75 points
Stopped halfway 今日は "晴れ" "です" missed, 50 points
Transcribed as Hiragana きょうははれです All correct
Katakana with full-width symbols キョウハ ハレデス! All correct
Added filler words before/after えっと今日は晴れですね All correct

These are unit tests that don't call the API, allowing the comparison logic to be tested instantly after any modification.

Speech-to-text checks " whether others can understand which word you are saying," not fine-grained pronunciation scoring:

Therefore, the card description explicitly states: "Only checks if the word is recognizable; pitch and vowel length are not within the scoring scope." If more detailed feedback is needed later, the recording and the teacher's demo can be sent to gemini-3.8-flash for a textual explanation, but that's a separate feature.

The first version's flow was: Click "Shadow" -> Button immediately changes to "■ Stop and Score" -> Request microphone.

The problem is that on the first use, the browser pops up a microphone permission prompt. At this point, the button already says "Stop and Score," but the mic hasn't started recording. If clicked again, the program thinks it hasn't started and requests the microphone again instead of stopping.

Cause & Solution: Added a "Preparing Microphone" state. After clicking, the button shows "Preparing Microphone..." and is temporarily disabled. Only when MediaRecorder actually starts recording does it switch to "Stop and Score."

By the way, the issue originally specified "Hold to record, release to send," but I changed it to "Click to start, click to stop" during implementation because of this permission prompt: if holding, the finger would definitely release when clicking the permission button, making the first attempt fail. Plus, long-pressing on iPhone easily triggers text selection, and saying a whole lyric takes several seconds; clicking twice is more user-friendly than holding.

For end-to-end testing, I've been using Puppeteer with headless Chrome. Chrome has a set of test flags to "use an audio file to fake a microphone":

--use-fake-ui-for-media-stream
--use-fake-device-for-media-stream
--use-file-for-fake-audio-capture=ok.wav

However, getUserMedia just hung without responding. I tried 48kHz wav, not specifying a file, pre-granting mic permissions—nothing worked. It couldn't even get a basic fake device; it seems macOS's microphone permission mechanism blocks headless Chrome.

Cause & Solution: Instead of getting stuck on the test environment, I changed the approach. Before the page loads, I replace getUserMedia to decode a test audio file and play it via Web Audio, creating a real audio stream:

navigator.mediaDevices.getUserMedia = async () => {
  const ctx = new AudioContext({ sampleRate: 48000 });
  const buffer = await ctx.decodeAudioData(testWavBytes.buffer);
  const src = ctx.createBufferSource();
  src.buffer = buffer;
  const dest = ctx.createMediaStreamDestination();
  src.connect(dest);
  src.start();
  return dest.stream;
};

Except for the "system microphone" part, the entire flow—recording (MediaRecorder to webm/opus), up, transcribing, comparing, and displaying results—is real. The test result was 100 points, all four words green, only one scoring request sent, and no residual files in the server's temp folder.

The bug from Pitfall 1 was also caught in the first version of this test: the button showed "Stop and Score," but MediaRecorder was never created.

Shadowing involves two sensitive things: my recordings and API quota.

Design Implementation
Recordings not saved Only written to the server's temp folder during scoring, then deleted immediately along with the folder (in a finally block); not stored in buckets or logged.
Don't trust lyrics from the browser The server reads the original text and pronunciation itself based on song ID and line number.
No double billing Only one scoring process at a time; SDK auto-retry is disabled.
Stop when quota is exhausted After receiving a 429, requests are rejected during the Retry-After period without calling Gemini.
Inaccessible to unauthenticated users The shadowing API is also behind IAP; unauthenticated requests are blocked at the IAP layer before reaching the app, so no quota is consumed.

Sending recordings to Gemini for transcription is necessary for the feature itself, but on the service side, the recording only exists for the few seconds of that call.

The first item on the roadmap was the mobile layout: MV fixed at the top, cards swipeable left/right, thumb-accessible bottom control bar, plus a PWA that can be added to the home screen. The layout itself isn't special, but there's a pitfall with PWA behind IAP worth noting.

When browsers fetch the PWA manifest, they don't include cookies by default. Behind IAP, requests without cookies are redirected to the Google login page, so the manifest can't be read, and the PWA can't be installed.

The HTML solution is to add crossorigin="use-credentials" to the <link rel="manifest">. However, looking at the Next.js 16 source code, I found its built-in manifest link only adds this attribute in Vercel preview environments:

crossOrigin: !manifestOrigin && process.env.VERCEL_ENV === 'preview' ? 'use-credentials' : undefined

Cause & Solution: Instead of using Next's built-in app/manifest.ts, I used a regular route to provide the manifest and manually placed the link with use-credentials in the <head>. Icons in the manifest were also embedded as data URLs to prevent the browser from making another cookie-less request.

Figures
Commits this time 2 (Mobile layout & PWA, Shadowing score)
Total repo commits 18
Closed issues 2 / 3 (roadmap: flashcards remaining)
One shadowing score Approx. 9–10s (upload ~2s, transcription ~4s, rest is Python startup)
Usage per score Approx. 40 tokens, doesn't consume TTS quota
Calls spent verifying API 5–6 transcribe calls

The workflow for learning a song on mobile is now: View card → Listen to teacher → Press shadow and speak → See which words were wrong → Listen to your recording vs. teacher → Swipe to next line.

"See real responses first" must be a rule, not a memory. I explicitly wrote this in the issue and made two API calls before starting. As a result, I knew about output_text and UTF-8 bytes before writing a single line of code.

The officially promoted mode isn't always best for your use case. smart mode is a plus for most scenarios, but shadowing needs exactly what it removes. Think about whether you want "organized results" or "what actually happened" before choosing a mode.

Features that improve accuracy might actually be counterproductive. Custom vocabulary biases the model toward the correct answer, while scoring needs to catch errors.

UI state should reflect real state, not assumed state. When the button says "Recording," the mic isn't actually on yet. This gap is most obvious during the first authorization, which is exactly when users are most likely to click randomly.

When the test environment hangs, change the approach. Instead of fighting macOS mic permissions, replace the "system microphone" segment and let the rest run for real.

Clarify capabilities upfront. Speech-to-text checks "understandability," not pronunciation. Writing this limitation on the screen is more honest than letting users think a 100-point score means perfect pronunciation.

Code is at kkdai/song-lingo, shadowing implementation is in shadow.py and web/app/api/songs/[id]/lines/[index]/shadow/. Official resources: Gemini 3.5 Transcribe Launch Post, Transcribe Documentation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-in-practice-shado…] indexed:0 read:11min 2026-09-29 · —