# [AI in Practice] Shadowing Assessment with Gemini 3.5 Transcribe: Let Song Lingo Listen to You Read Lyrics

> Source: <https://dev.to/evanlin/ai-in-practice-shadowing-assessment-with-gemini-35-transcribe-let-song-lingo-listen-to-you-read-55ih>
> Published: 2026-09-29 08:38:38+00:00

In the [first post](https://dev.to/gde/ai-in-practice-gemini-38-flash-tts-launch-i-built-a-learn-japanese-with-mvs-web-app-and-4o79), I used Gemini 3.8 Flash TTS to create Song Lingo: paste a YouTube MV URL, Gemini transcribes the lyrics, adds furigana, translation, and grammar, and then a teacher designed with voice design demonstrates pronunciation sentence by sentence. In the [second post](https://dev.to/gde/ai-in-practice-deploying-song-lingo-to-cloud-run-making-a-private-lyrics-website-just-for-me-mb3), I deployed it to Cloud Run and locked it with IAP so only I can use it.

After deployment, I opened a roadmap on GitHub with three items: mobile layout, shadowing score, and flashcards. The mobile layout is finished; this post is about the second item.

The teacher demonstrates, but I never know if I'm pronouncing it correctly. **Listening to the teacher ten times is not as good as saying it once and being told where you're wrong.** This is a key step in Song Lingo's evolution from "listening to songs and reading lyrics" to "practicing pronunciation."

Google released [Gemini 3.5 Transcribe](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/) on 8/26, which consists of two models specialized for speech-to-text:

| Model | Usage | 
|---|---|
| `gemini-3.5-transcribe` | Recorded audio files, transcribed in full | 
| `gemini-3.5-transcribe-live` | Real-time streaming via Live API WebSocket, transcribing while speaking | 

Key features:

`smart` handles self-corrections, removes filler words, and auto-formats; `verbatim` transcribes word-for-word.
The official blog cites Artificial Analysis: average Word Error Rate (WER) of 2.6% for full transcription and 4.0% for streaming; compared to the previous generation Chirp 3, the time to get final results is 70% faster. Pricing has not been announced.

The actual call method is the same as TTS via the `interactions` API, but audio files must first be uploaded using the Files API:

```
uploaded = client.files.upload(file=audio_path, config={"mime_type": "audio/m4a"})
interaction = client.interactions.create(
    model="gemini-3.5-transcribe",
    input=[{"type": "audio", "uri": uploaded.uri, "mime_type": uploaded.mime_type}],
    generation_config={
        "transcription_config": {
            "language_codes": ["ja-JP"],
            "mode": {"type": "verbatim", "timestamp_granularities": ["word"]},
        }
    },
)
```

Many formats are supported; `audio/webm` happens to be the format for Chrome recordings, and `audio/m4a` corresponds to AAC recorded by iPhone Safari.

For shadowing scores, there were three candidates:

| Model | Suitable for | Use in Shadowing | 
|---|---|---|
| **`gemini-3.5-transcribe`** | Transcribing after recording the whole segment | ✅ Has word-level timestamps, takes about 5s in testing | 
| `gemini-3.5-transcribe-live` | Real-time streaming | A line of lyrics is only a few seconds; transcribing after finishing is enough. Maintaining WebSockets behind Cloud Run and IAP is much more troublesome. | 
| `gemini-3.8-flash` and other multimodal models | Listening to audio directly to provide text feedback | Results vary every time, no guaranteed word-level timestamps, not suitable as the core for scoring. | 

I finally chose `gemini-3.5-transcribe`. As a specialized speech-to-text model, the results are stable, it provides word-level info, and it only uses about 40 tokens per call, **not consuming the 100-call daily quota for TTS**.

The most expensive lesson from the first post was: the web version of TTS followed the Python SDK's way of reading `output_audio`, but that field didn't exist in the REST response, burning a whole day's quota in an hour.

So this time, I wrote a rule directly in the roadmap issue: **Verify the parsing logic with real API responses before going live.**

Test audio was generated using macOS's built-in `say` command with sentences I wrote myself: one saying "今日は晴れです" (Today is sunny) and another intentionally saying "今日は雨です" (Today is rainy):

```
say -v Kyoko -o ok.aiff "今日は晴れです"
afconvert -f m4af -d aac ok.aiff ok.m4a # Same AAC format as iPhone recordings
```

Then I intercepted the raw HTTP content sent and received by the SDK and printed the response structure. I discovered two things not clearly explained in the documentation that would have caused errors if followed blindly.

`output_text` is assembled by the SDK again
Documentation says transcription results are in `interaction.output_text`. The actual REST response looks like this:

```
steps[] → { type: "model_output",
            content[] → { type: "text", text: "今日は晴れです。",
                          annotations[] → { type: "word_info", text: "今日",
                                            start_index: 0, end_index: 6,
                                            start_offset: "0.100s", end_offset: "0.400s" } } }
```

`output_text`, like `output_audio` last time, is a convenience field the Python SDK assembles from `steps`. Knowing this beforehand, I read directly from `steps[].content[]`.

The `start_index` for the word "今日" is 0 and `end_index` is 6, not 0 to 2. **Positions are calculated using UTF-8 bytes**, where one CJK character takes 3 bytes. If sliced according to JavaScript string indices, Japanese text would be completely misaligned.

These two things can only be known by looking at real responses. This time, spending two API calls meant no more guessing after going live.

The official `smart` mode handles self-corrections and removes filler words. While a plus for meeting minutes, it's a minus for shadowing: **if you mispronounce and then correct yourself, smart might only keep the corrected version; if you miss a particle, it might fill it in smoothly.** Shadowing needs "what you actually said," so use `verbatim`.

Setting the original text as custom vocabulary seems like it would improve accuracy, but it makes the model **biased toward hearing the correct answer**, making it harder to catch mistakes. Also, documentation states custom vocabulary and word-level timestamps cannot be used together.

Transcription results might be in Kanji "今日" or Hiragana "きょう"; both are correct. Comparing only text would mark correct pronunciations as wrong.

So for Japanese, I compare both ways simultaneously; **it counts as correct if either matches**:

The comparison itself uses Python's built-in `difflib.SequenceMatcher` for character alignment, mapping results back to each word. English and Korean are compared word-by-word.

The first version only checked "if the word's characters matched." Consequently, "晴れ" in "今日は雨です" was marked as **missed**, when it was actually **replaced by another word**.

After changing to record alignment results for each character (match, replaced, deleted), they can be distinguished:

| Scenario | Transcription Result | Judgment | 
|---|---|---|
| Correct | 今日は晴れです。 | All correct, 100 points | 
| Replaced a word | 今日は雨です。 | "晴れ" mispronounced, 75 points | 
| Skipped middle word | 今日はです | "晴れ" missed, 75 points | 
| Stopped halfway | 今日は | "晴れ" "です" missed, 50 points | 
| Transcribed as Hiragana | きょうははれです | All correct | 
| Katakana with full-width symbols | キョウハ ハレデス！ | All correct | 
| Added filler words before/after | えっと今日は晴れですね | All correct | 

These are unit tests that don't call the API, allowing the comparison logic to be tested instantly after any modification.

Speech-to-text checks " **whether others can understand which word you are saying**," not fine-grained pronunciation scoring:

Therefore, the card description explicitly states: "Only checks if the word is recognizable; pitch and vowel length are not within the scoring scope." If more detailed feedback is needed later, the recording and the teacher's demo can be sent to `gemini-3.8-flash` for a textual explanation, but that's a separate feature.

The first version's flow was: Click "Shadow" -> Button immediately changes to "■ Stop and Score" -> Request microphone.

The problem is that **on the first use, the browser pops up a microphone permission prompt**. At this point, the button already says "Stop and Score," but the mic hasn't started recording. If clicked again, the program thinks it hasn't started and **requests the microphone again** instead of stopping.

**Cause & Solution**: Added a "Preparing Microphone" state. After clicking, the button shows "Preparing Microphone..." and is temporarily disabled. **Only when `MediaRecorder` actually starts recording does it switch to "Stop and Score."**

By the way, the issue originally specified "Hold to record, release to send," but I changed it to "Click to start, click to stop" during implementation because of this permission prompt: if holding, the finger would definitely release when clicking the permission button, making the first attempt fail. Plus, long-pressing on iPhone easily triggers text selection, and saying a whole lyric takes several seconds; clicking twice is more user-friendly than holding.

For end-to-end testing, I've been using Puppeteer with headless Chrome. Chrome has a set of test flags to "use an audio file to fake a microphone":

```
--use-fake-ui-for-media-stream
--use-fake-device-for-media-stream
--use-file-for-fake-audio-capture=ok.wav
```

However, `getUserMedia` just hung without responding. I tried 48kHz wav, not specifying a file, pre-granting mic permissions—nothing worked. It couldn't even get a basic fake device; it seems macOS's microphone permission mechanism blocks headless Chrome.

**Cause & Solution**: Instead of getting stuck on the test environment, I changed the approach. Before the page loads, I replace `getUserMedia` to decode a test audio file and play it via Web Audio, creating a real audio stream:

``` js
navigator.mediaDevices.getUserMedia = async () => {
  const ctx = new AudioContext({ sampleRate: 48000 });
  const buffer = await ctx.decodeAudioData(testWavBytes.buffer);
  const src = ctx.createBufferSource();
  src.buffer = buffer;
  const dest = ctx.createMediaStreamDestination();
  src.connect(dest);
  src.start();
  return dest.stream;
};
```

Except for the "system microphone" part, the entire flow—recording (`MediaRecorder` to webm/opus), uploading, transcribing, comparing, and displaying results—is real. The test result was 100 points, all four words green, only one scoring request sent, and no residual files in the server's temp folder.

The bug from Pitfall 1 was also caught in the first version of this test: the button showed "Stop and Score," but `MediaRecorder` was never created.

Shadowing involves two sensitive things: **my recordings** and **API quota**.

| Design | Implementation | 
|---|---|
| Recordings not saved | Only written to the server's temp folder during scoring, then deleted immediately along with the folder (in a `finally` block); not stored in buckets or logged. | 
| Don't trust lyrics from the browser | The server reads the original text and pronunciation itself based on song ID and line number. | 
| No double billing | Only one scoring process at a time; SDK auto-retry is disabled. | 
| Stop when quota is exhausted | After receiving a 429, requests are rejected during the `Retry-After` period without calling Gemini. | 
| Inaccessible to unauthenticated users | The shadowing API is also behind IAP; unauthenticated requests are blocked at the IAP layer before reaching the app, so no quota is consumed. | 

Sending recordings to Gemini for transcription is necessary for the feature itself, but on the service side, the recording only exists for the few seconds of that call.

The first item on the roadmap was the mobile layout: MV fixed at the top, cards swipeable left/right, thumb-accessible bottom control bar, plus a PWA that can be added to the home screen. The layout itself isn't special, but there's a pitfall with PWA behind IAP worth noting.

When browsers fetch the PWA manifest, **they don't include cookies by default**. Behind IAP, requests without cookies are redirected to the Google login page, so the manifest can't be read, and the PWA can't be installed.

The HTML solution is to add `crossorigin="use-credentials"` to the `<link rel="manifest">`. However, looking at the Next.js 16 source code, I found its built-in manifest link **only adds this attribute in Vercel preview environments**:

```
crossOrigin: !manifestOrigin && process.env.VERCEL_ENV === 'preview' ? 'use-credentials' : undefined
```

**Cause & Solution**: Instead of using Next's built-in `app/manifest.ts`, I used a regular route to provide the manifest and manually placed the link with `use-credentials` in the `<head>`. Icons in the manifest were also embedded as data URLs to prevent the browser from making another cookie-less request.

|  | Figures | 
|---|---|
| Commits this time | 2 (Mobile layout & PWA, Shadowing score) | 
| Total repo commits | 18 | 
| Closed issues | 2 / 3 (roadmap: flashcards remaining) | 
| One shadowing score | Approx. 9–10s (upload ~2s, transcription ~4s, rest is Python startup) | 
| Usage per score | Approx. 40 tokens, doesn't consume TTS quota | 
| Calls spent verifying API | 5–6 transcribe calls | 

The workflow for learning a song on mobile is now: View card → Listen to teacher → Press shadow and speak → See which words were wrong → Listen to your recording vs. teacher → Swipe to next line.

**"See real responses first" must be a rule, not a memory.** I explicitly wrote this in the issue and made two API calls before starting. As a result, I knew about `output_text` and UTF-8 bytes before writing a single line of code.

**The officially promoted mode isn't always best for your use case.** `smart` mode is a plus for most scenarios, but shadowing needs exactly what it removes. Think about whether you want "organized results" or "what actually happened" before choosing a mode.

**Features that improve accuracy might actually be counterproductive.** Custom vocabulary biases the model toward the correct answer, while scoring needs to catch errors.

**UI state should reflect real state, not assumed state.** When the button says "Recording," the mic isn't actually on yet. This gap is most obvious during the first authorization, which is exactly when users are most likely to click randomly.

**When the test environment hangs, change the approach.** Instead of fighting macOS mic permissions, replace the "system microphone" segment and let the rest run for real.

**Clarify capabilities upfront.** Speech-to-text checks "understandability," not pronunciation. Writing this limitation on the screen is more honest than letting users think a 100-point score means perfect pronunciation.

Code is at [kkdai/song-lingo](https://github.com/kkdai/song-lingo), shadowing implementation is in `shadow.py` and `web/app/api/songs/[id]/lines/[index]/shadow/`. Official resources: [Gemini 3.5 Transcribe Launch Post](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/), [Transcribe Documentation](https://ai.google.dev/gemini-api/docs/transcribe).
