[AI in Practice] Shadowing Assessment with Gemini 3.5 Transcribe: Let Song Lingo Listen to You Read Lyrics A developer added a shadowing-scoring feature to Song Lingo, a Japanese-learning web app, by integrating Google's newly released Gemini 3.5 Transcribe speech-to-text model to compare learner pronunciation against reference lyrics. The developer chose the batch `gemini-3.5-transcribe` model over the live streaming variant and multimodal Flash models because it returns stable word-level timestamps, costs roughly 40 tokens per call, and does not consume the app's 100-call daily TTS quota. The post also details the Files API upload and `interactions` call pattern, and notes the model's cited 2.6% average word error rate for full transcription. In the first post https://dev.to/gde/ai-in-practice-gemini-38-flash-tts-launch-i-built-a-learn-japanese-with-mvs-web-app-and-4o79 , I used Gemini 3.8 Flash TTS to create Song Lingo: paste a YouTube MV URL, Gemini transcribes the lyrics, adds furigana, translation, and grammar, and then a teacher designed with voice design demonstrates pronunciation sentence by sentence. In the second post https://dev.to/gde/ai-in-practice-deploying-song-lingo-to-cloud-run-making-a-private-lyrics-website-just-for-me-mb3 , I deployed it to Cloud Run and locked it with IAP so only I can use it. After deployment, I opened a roadmap on GitHub with three items: mobile layout, shadowing score, and flashcards. The mobile layout is finished; this post is about the second item. The teacher demonstrates, but I never know if I'm pronouncing it correctly. Listening to the teacher ten times is not as good as saying it once and being told where you're wrong. This is a key step in Song Lingo's evolution from "listening to songs and reading lyrics" to "practicing pronunciation." Google released Gemini 3.5 Transcribe https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/ on 8/26, which consists of two models specialized for speech-to-text: | Model | Usage | |---|---| | gemini-3.5-transcribe | Recorded audio files, transcribed in full | | gemini-3.5-transcribe-live | Real-time streaming via Live API WebSocket, transcribing while speaking | Key features: smart handles self-corrections, removes filler words, and auto-formats; verbatim transcribes word-for-word. The official blog cites Artificial Analysis: average Word Error Rate WER of 2.6% for full transcription and 4.0% for streaming; compared to the previous generation Chirp 3, the time to get final results is 70% faster. Pricing has not been announced. The actual call method is the same as TTS via the interactions API, but audio files must first be uploaded using the Files API: uploaded = client.files.upload file=audio path, config={"mime type": "audio/m4a"} interaction = client.interactions.create model="gemini-3.5-transcribe", input= {"type": "audio", "uri": uploaded.uri, "mime type": uploaded.mime type} , generation config={ "transcription config": { "language codes": "ja-JP" , "mode": {"type": "verbatim", "timestamp granularities": "word" }, } }, Many formats are supported; audio/webm happens to be the format for Chrome recordings, and audio/m4a corresponds to AAC recorded by iPhone Safari. For shadowing scores, there were three candidates: | Model | Suitable for | Use in Shadowing | |---|---|---| | gemini-3.5-transcribe | Transcribing after recording the whole segment | ✅ Has word-level timestamps, takes about 5s in testing | | gemini-3.5-transcribe-live | Real-time streaming | A line of lyrics is only a few seconds; transcribing after finishing is enough. Maintaining WebSockets behind Cloud Run and IAP is much more troublesome. | | gemini-3.8-flash and other multimodal models | Listening to audio directly to provide text feedback | Results vary every time, no guaranteed word-level timestamps, not suitable as the core for scoring. | I finally chose gemini-3.5-transcribe . As a specialized speech-to-text model, the results are stable, it provides word-level info, and it only uses about 40 tokens per call, not consuming the 100-call daily quota for TTS . The most expensive lesson from the first post was: the web version of TTS followed the Python SDK's way of reading output audio , but that field didn't exist in the REST response, burning a whole day's quota in an hour. So this time, I wrote a rule directly in the roadmap issue: Verify the parsing logic with real API responses before going live. Test audio was generated using macOS's built-in say command with sentences I wrote myself: one saying "今日は晴れです" Today is sunny and another intentionally saying "今日は雨です" Today is rainy : say -v Kyoko -o ok.aiff "今日は晴れです" afconvert -f m4af -d aac ok.aiff ok.m4a Same AAC format as iPhone recordings Then I intercepted the raw HTTP content sent and received by the SDK and printed the response structure. I discovered two things not clearly explained in the documentation that would have caused errors if followed blindly. output text is assembled by the SDK again Documentation says transcription results are in interaction.output text . The actual REST response looks like this: steps → { type: "model output", content → { type: "text", text: "今日は晴れです。", annotations → { type: "word info", text: "今日", start index: 0, end index: 6, start offset: "0.100s", end offset: "0.400s" } } } output text , like output audio last time, is a convenience field the Python SDK assembles from steps . Knowing this beforehand, I read directly from steps .content . The start index for the word "今日" is 0 and end index is 6, not 0 to 2. Positions are calculated using UTF-8 bytes , where one CJK character takes 3 bytes. If sliced according to JavaScript string indices, Japanese text would be completely misaligned. These two things can only be known by looking at real responses. This time, spending two API calls meant no more guessing after going live. The official smart mode handles self-corrections and removes filler words. While a plus for meeting minutes, it's a minus for shadowing: if you mispronounce and then correct yourself, smart might only keep the corrected version; if you miss a particle, it might fill it in smoothly. Shadowing needs "what you actually said," so use verbatim . Setting the original text as custom vocabulary seems like it would improve accuracy, but it makes the model biased toward hearing the correct answer , making it harder to catch mistakes. Also, documentation states custom vocabulary and word-level timestamps cannot be used together. Transcription results might be in Kanji "今日" or Hiragana "きょう"; both are correct. Comparing only text would mark correct pronunciations as wrong. So for Japanese, I compare both ways simultaneously; it counts as correct if either matches : The comparison itself uses Python's built-in difflib.SequenceMatcher for character alignment, mapping results back to each word. English and Korean are compared word-by-word. The first version only checked "if the word's characters matched." Consequently, "晴れ" in "今日は雨です" was marked as missed , when it was actually replaced by another word . After changing to record alignment results for each character match, replaced, deleted , they can be distinguished: | Scenario | Transcription Result | Judgment | |---|---|---| | Correct | 今日は晴れです。 | All correct, 100 points | | Replaced a word | 今日は雨です。 | "晴れ" mispronounced, 75 points | | Skipped middle word | 今日はです | "晴れ" missed, 75 points | | Stopped halfway | 今日は | "晴れ" "です" missed, 50 points | | Transcribed as Hiragana | きょうははれです | All correct | | Katakana with full-width symbols | キョウハ ハレデス! | All correct | | Added filler words before/after | えっと今日は晴れですね | All correct | These are unit tests that don't call the API, allowing the comparison logic to be tested instantly after any modification. Speech-to-text checks " whether others can understand which word you are saying ," not fine-grained pronunciation scoring: Therefore, the card description explicitly states: "Only checks if the word is recognizable; pitch and vowel length are not within the scoring scope." If more detailed feedback is needed later, the recording and the teacher's demo can be sent to gemini-3.8-flash for a textual explanation, but that's a separate feature. The first version's flow was: Click "Shadow" - Button immediately changes to "■ Stop and Score" - Request microphone. The problem is that on the first use, the browser pops up a microphone permission prompt . At this point, the button already says "Stop and Score," but the mic hasn't started recording. If clicked again, the program thinks it hasn't started and requests the microphone again instead of stopping. Cause & Solution : Added a "Preparing Microphone" state. After clicking, the button shows "Preparing Microphone..." and is temporarily disabled. Only when MediaRecorder actually starts recording does it switch to "Stop and Score." By the way, the issue originally specified "Hold to record, release to send," but I changed it to "Click to start, click to stop" during implementation because of this permission prompt: if holding, the finger would definitely release when clicking the permission button, making the first attempt fail. Plus, long-pressing on iPhone easily triggers text selection, and saying a whole lyric takes several seconds; clicking twice is more user-friendly than holding. For end-to-end testing, I've been using Puppeteer with headless Chrome. Chrome has a set of test flags to "use an audio file to fake a microphone": --use-fake-ui-for-media-stream --use-fake-device-for-media-stream --use-file-for-fake-audio-capture=ok.wav However, getUserMedia just hung without responding. I tried 48kHz wav, not specifying a file, pre-granting mic permissions—nothing worked. It couldn't even get a basic fake device; it seems macOS's microphone permission mechanism blocks headless Chrome. Cause & Solution : Instead of getting stuck on the test environment, I changed the approach. Before the page loads, I replace getUserMedia to decode a test audio file and play it via Web Audio, creating a real audio stream: js navigator.mediaDevices.getUserMedia = async = { const ctx = new AudioContext { sampleRate: 48000 } ; const buffer = await ctx.decodeAudioData testWavBytes.buffer ; const src = ctx.createBufferSource ; src.buffer = buffer; const dest = ctx.createMediaStreamDestination ; src.connect dest ; src.start ; return dest.stream; }; Except for the "system microphone" part, the entire flow—recording MediaRecorder to webm/opus , uploading, transcribing, comparing, and displaying results—is real. The test result was 100 points, all four words green, only one scoring request sent, and no residual files in the server's temp folder. The bug from Pitfall 1 was also caught in the first version of this test: the button showed "Stop and Score," but MediaRecorder was never created. Shadowing involves two sensitive things: my recordings and API quota . | Design | Implementation | |---|---| | Recordings not saved | Only written to the server's temp folder during scoring, then deleted immediately along with the folder in a finally block ; not stored in buckets or logged. | | Don't trust lyrics from the browser | The server reads the original text and pronunciation itself based on song ID and line number. | | No double billing | Only one scoring process at a time; SDK auto-retry is disabled. | | Stop when quota is exhausted | After receiving a 429, requests are rejected during the Retry-After period without calling Gemini. | | Inaccessible to unauthenticated users | The shadowing API is also behind IAP; unauthenticated requests are blocked at the IAP layer before reaching the app, so no quota is consumed. | Sending recordings to Gemini for transcription is necessary for the feature itself, but on the service side, the recording only exists for the few seconds of that call. The first item on the roadmap was the mobile layout: MV fixed at the top, cards swipeable left/right, thumb-accessible bottom control bar, plus a PWA that can be added to the home screen. The layout itself isn't special, but there's a pitfall with PWA behind IAP worth noting. When browsers fetch the PWA manifest, they don't include cookies by default . Behind IAP, requests without cookies are redirected to the Google login page, so the manifest can't be read, and the PWA can't be installed. The HTML solution is to add crossorigin="use-credentials" to the