{"slug": "i-ported-meta-s-mms-forced-aligner-to-typescript-word-level-timestamps-in-node", "title": "I ported Meta's MMS forced aligner to TypeScript: word-level timestamps in Node.js, 42% faster than Python", "summary": "A developer published mms-forced-align, an npm package that runs Meta's MMS forced aligner in Node.js via ONNX Runtime with a from-scratch TypeScript CTC Viterbi decoder of roughly 170 lines. The package reports being about 42% faster than Python's torchaudio on the same fp32 model with roughly 32% less peak memory, matching torchaudio timings to 0.0 ms on 37 of 37 reference words and 110 of 116 words on a real clip. It supports Hindi, Tamil, Bengali and 17 other Indic languages through a TypeScript port of uroman's romanizer.", "body_md": "Picture a Hinglish food-review reel. The voiceover says *\"aaj maine kaafi kuchh try kiya\"*, and each word lights up on screen the moment it's spoken. To build that, you need to know exactly when every word starts and ends, down to a few hundredths of a second.\n\nThe best open tool for that job is Meta's MMS forced aligner. The usual way to run it is Python: `torchaudio`, PyTorch, and about 2.6 GB of RAM. I'm building my caption tool in Node.js and TypeScript, and I didn't want a Python process sitting next to it. When I looked for a Node.js version, I couldn't find one. So I built it.\n\n`mms-forced-align`\n**TL;DR**\n\n`mms-forced-align` runs Meta's MMS forced-aligner model in Node.js through ONNX Runtime, with a CTC Viterbi decoder I wrote from scratch in TypeScript (~170 lines).\n**~42% faster** than Python's `torchaudio` on the same fp32 model, with **~32% less peak memory**.\n**37 of 37** reference words match `torchaudio` to **0.0 ms**. On a real 116-word clip, 110 words are identical.- Works with Hindi, Tamil, Bengali and 17 more Indic languages through a TypeScript port of uroman's romanizer.\n\n```\nnpm install mms-forced-align\n```\n\nYou already have the transcript: the exact words someone said. Forced alignment answers the remaining question: **when** was each word said?\n\nIt is **not** speech-to-text. Speech-to-text is a guessing game (\"what did they say?\"). Forced alignment is a matching game (\"I know what they said, so where is each word?\"). Because the words are already known, the problem is much more constrained, and that's why the timings come out so precise.\n\nThink of it like a lyrics sheet and a song. You have the lyrics. Alignment puts a timestamp on every word, like a karaoke file.\n\nThat's the first three seconds of the clip I benchmark with, and those boxes are real output from the package. You need these timings for karaoke-style highlighting, word-by-word burned-in captions, subtitle sync and read-along audiobooks.\n\nYou can. Python's `torchaudio.pipelines.MMS_FA` and Mahmoud Ashraf's `ctc-forced-aligner` both work well. But for a Node.js service, a Python sidecar costs a lot:\n\n`import torch, torchaudio` alone took 3.91 s, before any model loads. `import(\"mms-forced-align\")` took 0.15 s.\nI wanted `npm install` and a function call.\n\nBefore I wrote this post, I searched npm, GitHub and Hugging Face properly. As far as I could find, `mms-forced-align` (first published on 9 August 2026) was the **first npm package to do MMS forced alignment with a proper CTC Viterbi decoder in Node.js**. Plenty of good work sits close to it, and it deserves credit:\n\n`@storyteller-platform/align` added its own MMS CTC Viterbi aligner for audiobooks. I take that as a good sign the approach is useful.\nThis is the heart of the post. Here's the whole pipeline first, then each piece in plain English.\n\nOnly one box, the acoustic model, runs inside ONNX Runtime. Everything else is TypeScript in the package. The model files are an ONNX export (a portable, framework-neutral model format) of Meta's MMS aligner, `onnx-community/mms-300m-1130-forced-aligner-ONNX`, and they download from Hugging Face on first use.\n\nThe input is mono audio at exactly **16,000 samples per second**, as a `Float32Array`. If you pass any other sample rate, the package throws instead of quietly resampling, because silently \"fixing\" audio is how you get timings that are wrong in ways nobody notices. I resample with `ffmpeg` beforehand.\n\nThe audio is normalized (zero mean, unit variance, the same formula as Hugging Face's `Wav2Vec2FeatureExtractor`) and sent to the model. The model doesn't hear the audio as one stream. It hears it in **frames** of 320 samples, which is **20 ms** each. One second of audio is about 50 frames.\n\nThe acoustic model (a 300-million-parameter wav2vec2-style network) returns **logits**: raw, unnormalized scores, one for each of 31 tokens, for every frame. The 31 tokens are a special **blank**, a few housekeeping tokens, the letters `a`–` z`, and the apostrophe. So a 10-second clip gives you a grid of about 500 frames × 31 scores.\n\nAlignment works in log-probabilities, not raw scores, so the first thing my code does is run a numerically stable `log_softmax` on each frame. (`log_softmax` turns a row of raw scores into log-probabilities that add up properly.) It's an easy step to miss, because the ONNX export outputs raw logits, not log-probs. **Read what the export actually returns; don't assume.**\n\nThat picture is the key intuition. Most of the time, the model's strongest opinion is **blank**. Letters spike briefly when they're heard.\n\nThink of a stenographer who mostly writes \"–\" (nothing new yet) and only writes a letter at the moment they hear it. The blank is the model's way of saying \"still the same sound\" or \"nothing new\". This scheme is called **CTC** (Connectionist Temporal Classification), and it's how the model can be trained without anyone hand-labelling where each letter starts.\n\nNext, the transcript. Each word becomes a list of letter IDs. Anything outside `a–z'` throws a `VocabError` instead of being guessed at.\n\nThen I build the CTC **extended sequence**: a blank around every letter. `hello` becomes:\n\n```\n_ h _ e _ l _ l _ o _\n```\n\nN letters give 2N + 1 states. Those blanks matter, especially the one between the two `l` s. Without it, there would be no way to tell \"one long `l`\" from \"two `l` s\".\n\nNow picture a grid. Time runs left to right, one column per 20 ms frame. The extended sequence runs top to bottom, one row per state. Every cell has a cost from the model: how unlikely that state is at that moment.\n\nThe task is to find the cheapest route from the top-left to the bottom-right, with three rules for each step to the next frame:\n\n`l` straight to The **Viterbi algorithm** is dynamic programming. Like a sat-nav, it remembers the best route to *every* cell, so it never re-explores a path. For each cell it stores the best score so far and a **backpointer** saying which move got it there. When it reaches the last frame, it follows the backpointers backwards to recover the single best path. That grid is called a **trellis**.\n\nHere's the inner loop from the package, close to verbatim:\n\n``` js\nfor (let t = 1; t < numFrames; t++) {\n  const alphaCur = new Float64Array(L).fill(-Infinity);\n  const bp = new Int8Array(L);\n  for (let s = 0; s < L; s++) {\n    let best = alphaPrev[s];          // stay\n    let bestFrom = 0;\n    if (s >= 1 && alphaPrev[s - 1] > best) {\n      best = alphaPrev[s - 1];        // advance\n      bestFrom = 1;\n    }\n    const canSkip = s >= 2 && ext[s] !== blankId && ext[s] !== ext[s - 2];\n    if (canSkip && alphaPrev[s - 2] > best) {\n      best = alphaPrev[s - 2];        // skip a blank\n      bestFrom = 2;\n    }\n    alphaCur[s] = best === -Infinity ? -Infinity : best + at(t, ext[s]);\n    bp[s] = bestFrom;\n  }\n  backpointers.push(bp);\n  alphaPrev = alphaCur;\n}\n```\n\nThe `canSkip` line is the \"no skipping from `l` to `l`\" rule. The whole decoder, including backtracking and the error cases, is about 170 lines of TypeScript.\n\nThe best path says which state was active in every frame. Odd-numbered states are letters and even-numbered states are blanks, so runs of letter frames become per-letter spans, and letters group back into words using each word's letter count.\n\nThen frames become seconds: `seconds = frame × samplesPerFrame / 16000`.\n\n**The bug I hit here:** I first hard-coded `samplesPerFrame = 320`. The timings drifted, and the drift grew the further you got into the clip. wav2vec2's convolutional front-end doesn't downsample at exactly 320 for every input length. The fix, which is also what torchaudio does, is to measure the ratio for each clip:\n\n``` js\nconst samplesPerFrame = waveform.length / numFrames;\n```\n\nOne innocent-looking constant became a bug that only showed up on long audio.\n\nEach word also gets a `score`: the average probability of the winning token over the word's frames. Blank frames count toward it, so scores skew high. It's useful for comparing words, not as a calibrated probability.\n\n``` js\nimport { createAligner } from \"mms-forced-align\";\n\n// Load once, reuse for many files. Loading the ONNX session is the expensive part.\nconst aligner = await createAligner({ quantized: false }); // fp32: faster and more accurate on CPU (see below)\n\nconst timings = await aligner.align(\n  waveform,   // Float32Array, mono PCM in [-1, 1]\n  16000,      // sample rate, must be exactly 16000\n  [\"aaj\", \"albeck\", \"mein\", \"maine\", \"kaafi\", \"kuchh\", \"try\", \"kiya\"] // transcript words, in order, a–z and ' only\n);\n// → [{ word: \"aaj\", start: 0.04, end: 0.12, score: ... }, { word: \"albeck\", start: 0.18, end: 0.52, score: ... }, ...]\n\nawait aligner.dispose();\n```\n\nReal transcripts have punctuation and numbers, so clean them first:\n\n``` js\nconst words = transcript\n  .split(/\\s+/)\n  .map((w) => w.replace(/[^a-zA-Z']/g, \"\"))\n  .filter(Boolean);\n```\n\nThis drops tokens with no letters (like `\"2024\"`), so they get no timing. If you need numbers timed, spell them out first.\n\nA port is only worth something if it gives the same answers as the original. So I didn't eyeball the output. I tested against the reference.\n\nI generated **golden fixtures** with Python's `torchaudio.pipelines.MMS_FA`: the reference word timings for a known clip. My test suite feeds the TypeScript aligner the same audio and compares every word.\n\nThen a harder test: a full 36-second, 116-word Hinglish clip, compared with Python's fp32 output.\n\n**110 of 116 words have identical start and end times.** The other 6 differ by exactly one 20 ms frame (one word by three frames, 60 ms), all at pauses. The average difference is 0.54 ms for starts and 0.89 ms for ends. The likely cause is my benchmark harness, not the decoder: I resampled the 44.1 kHz audio to 16 kHz with a simple linear resampler, while torchaudio uses a windowed-sinc one, so the two models heard very slightly different audio. When both sides get identical audio, as in the golden tests, they match exactly.\n\nHere are the first 10 words. JS fp32 and Python fp32 agree on every one:\n\n| word | start (s) | end (s) | \n|---|---|---|\n| aaj | 0.040 | 0.120 | \n| albeck | 0.180 | 0.520 | \n| mein | 0.580 | 0.660 | \n| maine | 0.700 | 0.840 | \n| kaafi | 0.920 | 1.160 | \n| kuchh | 1.180 | 1.341 | \n| try | 1.381 | 1.521 | \n| kiya | 1.601 | 1.721 | \n| kya | 2.481 | 2.581 | \n| aur | 2.721 | 2.841 | \n\n**Setup:** Windows, 4 CPU cores, no GPU on either side. A 36.45 s clip resampled to 16 kHz mono, with a 116-word Hinglish transcript. Python used `torch` 2.13.0+cpu and `torchaudio` 2.11.0+cpu (fp32; this pipeline has no quantized variant). JS used `mms-forced-align` with the int8 model (317 MB on disk) and the fp32 model (1.2 GB). Models were already on disk, and each config ran 3 times.\n\n| Config | Run 1 | Run 2 | Run 3 | Avg | × realtime | \n|---|---|---|---|---|---|\n| Python fp32 | 37.17 s | 37.11 s | 36.92 s | **37.07 s** | 1.017× | \n| JS int8 | 27.01 s | 26.87 s | 27.03 s | **26.97 s** | 0.740× | \n| JS fp32 | 22.36 s | 20.93 s | 20.69 s | **21.33 s** | 0.585× | \n\nJS fp32 aligned **~42% faster** than Python at the same precision. Below 1.0× realtime means it finishes faster than the audio plays.\n\n**RSS** (resident set size) is how much RAM the process actually holds.\n\n| Config | Before load | After load | Peak (after align) | \n|---|---|---|---|\n| Python fp32 | 193 MB | 2,608 MB | **2,657 MB** | \n| JS int8 | 55 MB | 379 MB | **875 MB** | \n| JS fp32 | 87 MB | 1,300 MB | **1,797 MB** | \n\nPython peaked at ~2.6 GB, about 1.5× JS fp32 (~32% less memory for JS) and ~3× JS int8.\n\n| Config | Import | Model load | First align | **Total** | \n|---|---|---|---|---|\n| Python fp32 | 3.91 s | 4.01 s | 37.21 s | **~45.1 s** | \n| JS int8 | 0.15 s | 2.10 s | 26.71 s | **~29.0 s** | \n| JS fp32 | 0.15 s | 3.76 s | 21.36 s | **~25.3 s** | \n\nEnd to end, JS fp32 was ~44% faster. Python's numbers also depend heavily on multiple cores: with 1 thread instead of 4, its alignment took 102.93 s instead of 37.21 s (2.77× slower).\n\n**Caveats.** One machine, one 36-second clip, 3 runs: treat this as a strong signal, not a universal law. CPU only on both sides; I have no GPU numbers. Model-load times swing a lot with the OS disk cache (Python's first load was 8.74 s cold vs 4.01 s warm, and JS fp32's was 7.36 s vs 3.76 s), so compare alignment times, which were stable. And I couldn't pin `onnxruntime-node`'s thread count through its public API (` OMP_NUM_THREADS=1` had no effect), so the JS side ran with ONNX Runtime's defaults.\n\nI expected the int8 quantized model to be the fast option. That's the usual story: smaller numbers, less work. On this CPU it was **slower** than fp32 (26.97 s vs 21.33 s), **and** less accurate (8 of 37 golden words outside tolerance, vs 37 of 37 exact).\n\nQuantization isn't free speed. It depends on whether your hardware and runtime have fast int8 kernels for these operations. Where int8 does win is size: a ~4× smaller download (317 MB vs 1.2 GB), lower memory, and faster loading. That matters for a serverless function that aligns one short clip per cold start. For a long-running worker that loads the model once and aligns many files, fp32 wins.\n\nThat's why every example in this post passes `quantized: false`. (The package default is still `true` in 0.2.x, so existing users don't suddenly download 1.2 GB.) The general lesson: **benchmark on your own hardware before you trust \"quantized = faster\".**\n\nThe model's vocabulary is Latin letters only. Hindi in Devanagari, Tamil, Bengali and the rest can't go in directly. Meta's own approach for MMS is to **romanize** text first with **uroman**, a romanizer from USC's Information Sciences Institute.\n\nSo I ported the relevant part of uroman's algorithm to TypeScript, using character data extracted directly from uroman, and validated it against fixtures generated by the real uroman. `alignNative()` romanizes your words, aligns the romanized form, and hands back timings keyed to **your original words**, matched by index:\n\n``` js\nimport { createAligner, alignNative } from \"mms-forced-align\";\n\nconst aligner = await createAligner({ quantized: false });\nconst timings = await alignNative(\n  aligner, waveform, 16000,\n  [\"आज\", \"मैंने\", \"काफी\", \"कुछ\", \"किया\"],\n  \"hi-IN\"\n);\n// → timings keyed to the original Devanagari words\n```\n\nNotice the romanized forms: `maimne` and `kaaphii`, not the \"maine\" and \"kaafi\" a Hinglish speaker would type. uroman is mechanical on purpose. The romanization only has to be good enough for the acoustic model, not pretty for humans, and since you get your original words back, you never see it.\n\nIt covers **21 language codes**: 20 Indic languages plus `en-IN` passthrough. The scripts are Devanagari (Hindi, Marathi, Nepali, Sanskrit, Maithili, Dogri, Konkani, Bodo), Bengali (Bengali, Assamese), Gurmukhi (Punjabi), Gujarati, Odia, Tamil, Telugu, Kannada, Malayalam and Perso-Arabic (Urdu, Kashmiri, Sindhi). English words code-switched into an Indic sentence, which is constant in Hinglish, pass through unchanged.\n\nTwo scripts are deliberately left out: Santali (Ol Chiki) and Manipuri (Meitei Mayek), because uroman's own data for them is unreliable. A few rare Perso-Arabic characters that romanize to non-letters (Sindhi `۽`, for example) throw a clear error that names the character.\n\nThat's a design rule across the whole package: **fail loudly, never guess.** Wrong sample rate, unknown character, word-count mismatch: all of them throw. For captions, a wrong timestamp is worse than an error, because nobody notices it until a viewer does.\n\nThe package code is MIT. **The model weights are CC-BY-NC-4.0 (non-commercial)**, per both the ONNX export and the source model on Hugging Face. They aren't bundled in the npm package; they download on first use. If you want to use this commercially, check that license yourself first.\n\n`onnxruntime-node`), no GPU yet.`<star>` token for words that are in the transcript but missing from the audio.`320` looked like a fact about wav2vec2. It was an approximation, and it caused drift that grew with clip length.\n\n```\nnpm install mms-forced-align\n```\n\nWord-level timestamps in Node.js/TypeScript with Meta's MMS forced-aligner model and a TypeScript CTC Viterbi decoder. No Python needed.\n\nBuilt by [Arham Sayyed](https://github.com/arham-sayyed).\n\nGive it audio and the transcript you already know was spoken in it, and it\nreturns the start and end time of every word. This is **not** speech-to-text\nyou already know the words, and this package only works out *when* each one\nwas said. That's what you need for karaoke-style highlighting, word-by-word\ncaptions, subtitle sync and read-along audio.\n\nThe neural forward pass runs Meta's MMS (Massively Multilingual Speech)\nforced-aligner model through an ONNX export and `onnxruntime-node`. The\nalignment itself is a from-scratch TypeScript CTC Viterbi decoder. As far as I\ncould find, this is the first npm package to do MMS forced alignment with a\nproper CTC Viterbi decoder in Node.js. See [Prior art](https://github.com/arham-sayyed/mms-forced-align#prior-art) for related\nprojects.\n\n```\nnpm install mms-forced-align\njs\nimport {\n```\n\n…\nI'm building this into **Awaaz**, a caption tool for Hinglish and Indian-language video: upload a video, get a transcript from Sarvam AI's speech API, get word timings from `mms-forced-align`, edit the captions, pick a style, and render with Remotion. Indian creators speak Hinglish, and most caption tools mangle it. Awaaz is still in development; I'll write about it when it's live.\n\nIf you try the package, I'd love to hear what you build. Issues, benchmarks on your hardware (especially ARM and GPU machines) and pull requests are all welcome on GitHub.\n\n`MahmoudAshraf/mms-300m-1130-forced-aligner`` onnx-community/mms-300m-1130-forced-aligner-ONNX``ctc-forced-aligner`\n*I'm **Arham Sayyed**, a full-stack developer from Mumbai who works mostly in Node.js and TypeScript. I built `mms-forced-align` and I'm building Awaaz. Find me on [GitHub](https://github.com/arham-sayyed), [LinkedIn](https://www.linkedin.com/in/arham-sayyed/) and here on [Dev.to](https://dev.to/arhamsayyed).*", "url": "https://wpnews.pro/news/i-ported-meta-s-mms-forced-aligner-to-typescript-word-level-timestamps-in-node", "canonical_source": "https://dev.to/arhamsayyed/i-ported-metas-mms-forced-aligner-to-typescript-word-level-timestamps-in-nodejs-42-faster-than-34c5", "published_at": "2026-10-09 23:18:23+00:00", "updated_at": "2026-10-09 23:29:22.424615+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "developer-tools", "ai-infrastructure"], "entities": ["Meta", "mms-forced-align", "ONNX Runtime", "torchaudio", "Hugging Face", "uroman", "Mahmoud Ashraf", "ctc-forced-aligner"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-ported-meta-s-mms-forced-aligner-to-typescript-word-level-timestamps-in-node", "markdown": "https://wpnews.pro/news/i-ported-meta-s-mms-forced-aligner-to-typescript-word-level-timestamps-in-node.md", "text": "https://wpnews.pro/news/i-ported-meta-s-mms-forced-aligner-to-typescript-word-level-timestamps-in-node.txt", "jsonld": "https://wpnews.pro/news/i-ported-meta-s-mms-forced-aligner-to-typescript-word-level-timestamps-in-node.jsonld"}}