cd /news/ai-tools/i-ported-meta-s-mms-forced-aligner-t… · home › topics › ai-tools › article
[ARTICLE · art-148531] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

I ported Meta's MMS forced aligner to TypeScript: word-level timestamps in Node.js, 42% faster than Python

A developer published mms-forced-align, an npm package that runs Meta's MMS forced aligner in Node.js via ONNX Runtime with a from-scratch TypeScript CTC Viterbi decoder of roughly 170 lines. The package reports being about 42% faster than Python's torchaudio on the same fp32 model with roughly 32% less peak memory, matching torchaudio timings to 0.0 ms on 37 of 37 reference words and 110 of 116 words on a real clip. It supports Hindi, Tamil, Bengali and 17 other Indic languages through a TypeScript port of uroman's romanizer.

by read15 min views4 publishedOct 9, 2026

Picture a Hinglish food-review reel. The voiceover says "aaj maine kaafi kuchh try kiya", and each word lights up on screen the moment it's spoken. To build that, you need to know exactly when every word starts and ends, down to a few hundredths of a second.

The best open tool for that job is Meta's MMS forced aligner. The usual way to run it is Python: torchaudio, PyTorch, and about 2.6 GB of RAM. I'm building my caption tool in Node.js and TypeScript, and I didn't want a Python process sitting next to it. When I looked for a Node.js version, I couldn't find one. So I built it.

mms-forced-align TL;DR

mms-forced-align runs Meta's MMS forced-aligner model in Node.js through ONNX Runtime, with a CTC Viterbi decoder I wrote from scratch in TypeScript (~170 lines). ~42% faster than Python's torchaudio on the same fp32 model, with ~32% less peak memory. 37 of 37 reference words match torchaudio to 0.0 ms. On a real 116-word clip, 110 words are identical.- Works with Hindi, Tamil, Bengali and 17 more Indic languages through a TypeScript port of uroman's romanizer.

npm install mms-forced-align

You already have the transcript: the exact words someone said. Forced alignment answers the remaining question: when was each word said?

It is not speech-to-text. Speech-to-text is a guessing game ("what did they say?"). Forced alignment is a matching game ("I know what they said, so where is each word?"). Because the words are already known, the problem is much more constrained, and that's why the timings come out so precise.

Think of it like a lyrics sheet and a song. You have the lyrics. Alignment puts a timestamp on every word, like a karaoke file.

That's the first three seconds of the clip I benchmark with, and those boxes are real output from the package. You need these timings for karaoke-style highlighting, word-by-word burned-in captions, subtitle sync and read-along audiobooks.

You can. Python's torchaudio.pipelines.MMS_FA and Mahmoud Ashraf's ctc-forced-aligner both work well. But for a Node.js service, a Python sidecar costs a lot:

import torch, torchaudio alone took 3.91 s, before any model loads. import("mms-forced-align") took 0.15 s. I wanted npm install and a function call.

Before I wrote this post, I searched npm, GitHub and Hugging Face properly. As far as I could find, mms-forced-align (first published on 9 August 2026) was the first npm package to do MMS forced alignment with a proper CTC Viterbi decoder in Node.js. Plenty of good work sits close to it, and it deserves credit:

@storyteller-platform/align added its own MMS CTC Viterbi aligner for audiobooks. I take that as a good sign the approach is useful. This is the heart of the post. Here's the whole pipeline first, then each piece in plain English.

Only one box, the acoustic model, runs inside ONNX Runtime. Everything else is TypeScript in the package. The model files are an ONNX export (a portable, framework-neutral model format) of Meta's MMS aligner, onnx-community/mms-300m-1130-forced-aligner-ONNX, and they download from Hugging Face on first use.

The input is mono audio at exactly 16,000 samples per second, as a Float32Array. If you pass any other sample rate, the package throws instead of quietly resampling, because silently "fixing" audio is how you get timings that are wrong in ways nobody notices. I resample with ffmpeg beforehand.

The audio is normalized (zero mean, unit variance, the same formula as Hugging Face's Wav2Vec2FeatureExtractor) and sent to the model. The model doesn't hear the audio as one stream. It hears it in frames of 320 samples, which is 20 ms each. One second of audio is about 50 frames.

The acoustic model (a 300-million-parameter wav2vec2-style network) returns logits: raw, unnormalized scores, one for each of 31 tokens, for every frame. The 31 tokens are a special blank, a few housekeeping tokens, the letters a– z, and the apostrophe. So a 10-second clip gives you a grid of about 500 frames × 31 scores.

Alignment works in log-probabilities, not raw scores, so the first thing my code does is run a numerically stable log_softmax on each frame. (log_softmax turns a row of raw scores into log-probabilities that add up properly.) It's an easy step to miss, because the ONNX export outputs raw logits, not log-probs. Read what the export actually returns; don't assume.

That picture is the key intuition. Most of the time, the model's strongest opinion is blank. Letters spike briefly when they're heard.

Think of a stenographer who mostly writes "–" (nothing new yet) and only writes a letter at the moment they hear it. The blank is the model's way of saying "still the same sound" or "nothing new". This scheme is called CTC (Connectionist Temporal Classification), and it's how the model can be trained without anyone hand-labelling where each letter starts.

Next, the transcript. Each word becomes a list of letter IDs. Anything outside a–z' throws a VocabError instead of being guessed at.

Then I build the CTC extended sequence: a blank around every letter. hello becomes:

_ h _ e _ l _ l _ o _

N letters give 2N + 1 states. Those blanks matter, especially the one between the two l s. Without it, there would be no way to tell "one long l" from "two l s".

Now picture a grid. Time runs left to right, one column per 20 ms frame. The extended sequence runs top to bottom, one row per state. Every cell has a cost from the model: how unlikely that state is at that moment.

The task is to find the cheapest route from the top-left to the bottom-right, with three rules for each step to the next frame:

l straight to The Viterbi algorithm is dynamic programming. Like a sat-nav, it remembers the best route to every cell, so it never re-explores a path. For each cell it stores the best score so far and a backpointer saying which move got it there. When it reaches the last frame, it follows the backpointers backwards to recover the single best path. That grid is called a trellis.

Here's the inner loop from the package, close to verbatim:

for (let t = 1; t < numFrames; t++) {
  const alphaCur = new Float64Array(L).fill(-Infinity);
  const bp = new Int8Array(L);
  for (let s = 0; s < L; s++) {
    let best = alphaPrev[s];          // stay
    let bestFrom = 0;
    if (s >= 1 && alphaPrev[s - 1] > best) {
      best = alphaPrev[s - 1];        // advance
      bestFrom = 1;
    }
    const canSkip = s >= 2 && ext[s] !== blankId && ext[s] !== ext[s - 2];
    if (canSkip && alphaPrev[s - 2] > best) {
      best = alphaPrev[s - 2];        // skip a blank
      bestFrom = 2;
    }
    alphaCur[s] = best === -Infinity ? -Infinity : best + at(t, ext[s]);
    bp[s] = bestFrom;
  }
  backpointers.push(bp);
  alphaPrev = alphaCur;
}

The canSkip line is the "no skipping from l to l" rule. The whole decoder, including backtracking and the error cases, is about 170 lines of TypeScript.

The best path says which state was active in every frame. Odd-numbered states are letters and even-numbered states are blanks, so runs of letter frames become per-letter spans, and letters group back into words using each word's letter count.

Then frames become seconds: seconds = frame × samplesPerFrame / 16000.

The bug I hit here: I first hard-coded samplesPerFrame = 320. The timings drifted, and the drift grew the further you got into the clip. wav2vec2's convolutional front-end doesn't downsample at exactly 320 for every input length. The fix, which is also what torchaudio does, is to measure the ratio for each clip:

const samplesPerFrame = waveform.length / numFrames;

One innocent-looking constant became a bug that only showed up on long audio.

Each word also gets a score: the average probability of the winning token over the word's frames. Blank frames count toward it, so scores skew high. It's useful for comparing words, not as a calibrated probability.

import { createAligner } from "mms-forced-align";

// Load once, reuse for many files.  the ONNX session is the expensive part.
const aligner = await createAligner({ quantized: false }); // fp32: faster and more accurate on CPU (see below)

const timings = await aligner.align(
  waveform,   // Float32Array, mono PCM in [-1, 1]
  16000,      // sample rate, must be exactly 16000
  ["aaj", "albeck", "mein", "maine", "kaafi", "kuchh", "try", "kiya"] // transcript words, in order, a–z and ' only
);
// → [{ word: "aaj", start: 0.04, end: 0.12, score: ... }, { word: "albeck", start: 0.18, end: 0.52, score: ... }, ...]

await aligner.dispose();

Real transcripts have punctuation and numbers, so clean them first:

const words = transcript
  .split(/\s+/)
  .map((w) => w.replace(/[^a-zA-Z']/g, ""))
  .filter(Boolean);

This drops tokens with no letters (like "2024"), so they get no timing. If you need numbers timed, spell them out first.

A port is only worth something if it gives the same answers as the original. So I didn't eyeball the output. I tested against the reference.

I generated golden fixtures with Python's torchaudio.pipelines.MMS_FA: the reference word timings for a known clip. My test suite feeds the TypeScript aligner the same audio and compares every word.

Then a harder test: a full 36-second, 116-word Hinglish clip, compared with Python's fp32 output.

110 of 116 words have identical start and end times. The other 6 differ by exactly one 20 ms frame (one word by three frames, 60 ms), all at s. The average difference is 0.54 ms for starts and 0.89 ms for ends. The likely cause is my benchmark harness, not the decoder: I resampled the 44.1 kHz audio to 16 kHz with a simple linear resampler, while torchaudio uses a windowed-sinc one, so the two models heard very slightly different audio. When both sides get identical audio, as in the golden tests, they match exactly.

Here are the first 10 words. JS fp32 and Python fp32 agree on every one:

word start (s) end (s)
aaj 0.040 0.120
albeck 0.180 0.520
mein 0.580 0.660
maine 0.700 0.840
kaafi 0.920 1.160
kuchh 1.180 1.341
try 1.381 1.521
kiya 1.601 1.721
kya 2.481 2.581
aur 2.721 2.841

Setup: Windows, 4 CPU cores, no GPU on either side. A 36.45 s clip resampled to 16 kHz mono, with a 116-word Hinglish transcript. Python used torch 2.13.0+cpu and torchaudio 2.11.0+cpu (fp32; this pipeline has no quantized variant). JS used mms-forced-align with the int8 model (317 MB on disk) and the fp32 model (1.2 GB). Models were already on disk, and each config ran 3 times.

Config Run 1 Run 2 Run 3 Avg × realtime
Python fp32 37.17 s 37.11 s 36.92 s 37.07 s 1.017×
JS int8 27.01 s 26.87 s 27.03 s 26.97 s 0.740×
JS fp32 22.36 s 20.93 s 20.69 s 21.33 s 0.585×

JS fp32 aligned ~42% faster than Python at the same precision. Below 1.0× realtime means it finishes faster than the audio plays.

RSS (resident set size) is how much RAM the process actually holds.

Config Before load After load Peak (after align)
Python fp32 193 MB 2,608 MB 2,657 MB
JS int8 55 MB 379 MB 875 MB
JS fp32 87 MB 1,300 MB 1,797 MB

Python peaked at ~2.6 GB, about 1.5× JS fp32 (~32% less memory for JS) and ~3× JS int8.

Config Import Model load First align Total
Python fp32 3.91 s 4.01 s 37.21 s ~45.1 s
JS int8 0.15 s 2.10 s 26.71 s ~29.0 s
JS fp32 0.15 s 3.76 s 21.36 s ~25.3 s

End to end, JS fp32 was ~44% faster. Python's numbers also depend heavily on multiple cores: with 1 thread instead of 4, its alignment took 102.93 s instead of 37.21 s (2.77× slower).

Caveats. One machine, one 36-second clip, 3 runs: treat this as a strong signal, not a universal law. CPU only on both sides; I have no GPU numbers. Model-load times swing a lot with the OS disk cache (Python's first load was 8.74 s cold vs 4.01 s warm, and JS fp32's was 7.36 s vs 3.76 s), so compare alignment times, which were stable. And I couldn't pin onnxruntime-node's thread count through its public API ( OMP_NUM_THREADS=1 had no effect), so the JS side ran with ONNX Runtime's defaults.

I expected the int8 quantized model to be the fast option. That's the usual story: smaller numbers, less work. On this CPU it was slower than fp32 (26.97 s vs 21.33 s), and less accurate (8 of 37 golden words outside tolerance, vs 37 of 37 exact).

Quantization isn't free speed. It depends on whether your hardware and runtime have fast int8 kernels for these operations. Where int8 does win is size: a ~4× smaller download (317 MB vs 1.2 GB), lower memory, and faster . That matters for a serverless function that aligns one short clip per cold start. For a long-running worker that loads the model once and aligns many files, fp32 wins.

That's why every example in this post passes quantized: false. (The package default is still true in 0.2.x, so existing users don't suddenly download 1.2 GB.) The general lesson: benchmark on your own hardware before you trust "quantized = faster".

The model's vocabulary is Latin letters only. Hindi in Devanagari, Tamil, Bengali and the rest can't go in directly. Meta's own approach for MMS is to romanize text first with uroman, a romanizer from USC's Information Sciences Institute.

So I ported the relevant part of uroman's algorithm to TypeScript, using character data extracted directly from uroman, and validated it against fixtures generated by the real uroman. alignNative() romanizes your words, aligns the romanized form, and hands back timings keyed to your original words, matched by index:

import { createAligner, alignNative } from "mms-forced-align";

const aligner = await createAligner({ quantized: false });
const timings = await alignNative(
  aligner, waveform, 16000,
  ["आज", "मैंने", "काफी", "कुछ", "किया"],
  "hi-IN"
);
// → timings keyed to the original Devanagari words

Notice the romanized forms: maimne and kaaphii, not the "maine" and "kaafi" a Hinglish speaker would type. uroman is mechanical on purpose. The romanization only has to be good enough for the acoustic model, not pretty for humans, and since you get your original words back, you never see it.

It covers 21 language codes: 20 Indic languages plus en-IN passthrough. The scripts are Devanagari (Hindi, Marathi, Nepali, Sanskrit, Maithili, Dogri, Konkani, Bodo), Bengali (Bengali, Assamese), Gurmukhi (Punjabi), Gujarati, Odia, Tamil, Telugu, Kannada, Malayalam and Perso-Arabic (Urdu, Kashmiri, Sindhi). English words code-switched into an Indic sentence, which is constant in Hinglish, pass through unchanged.

Two scripts are deliberately left out: Santali (Ol Chiki) and Manipuri (Meitei Mayek), because uroman's own data for them is unreliable. A few rare Perso-Arabic characters that romanize to non-letters (Sindhi ۽, for example) throw a clear error that names the character.

That's a design rule across the whole package: fail loudly, never guess. Wrong sample rate, unknown character, word-count mismatch: all of them throw. For captions, a wrong timestamp is worse than an error, because nobody notices it until a viewer does.

The package code is MIT. The model weights are CC-BY-NC-4.0 (non-commercial), per both the ONNX export and the source model on Hugging Face. They aren't bundled in the npm package; they download on first use. If you want to use this commercially, check that license yourself first.

onnxruntime-node), no GPU yet.<star> token for words that are in the transcript but missing from the audio.320 looked like a fact about wav2vec2. It was an approximation, and it caused drift that grew with clip length.

npm install mms-forced-align

Word-level timestamps in Node.js/TypeScript with Meta's MMS forced-aligner model and a TypeScript CTC Viterbi decoder. No Python needed.

Built by Arham Sayyed.

Give it audio and the transcript you already know was spoken in it, and it returns the start and end time of every word. This is not speech-to-text you already know the words, and this package only works out when each one was said. That's what you need for karaoke-style highlighting, word-by-word captions, subtitle sync and read-along audio.

The neural forward pass runs Meta's MMS (Massively Multilingual Speech) forced-aligner model through an ONNX export and onnxruntime-node. The alignment itself is a from-scratch TypeScript CTC Viterbi decoder. As far as I could find, this is the first npm package to do MMS forced alignment with a proper CTC Viterbi decoder in Node.js. See Prior art for related projects.

npm install mms-forced-align
js
import {

… I'm building this into Awaaz, a caption tool for Hinglish and Indian-language video: upload a video, get a transcript from Sarvam AI's speech API, get word timings from mms-forced-align, edit the captions, pick a style, and render with Remotion. Indian creators speak Hinglish, and most caption tools mangle it. Awaaz is still in development; I'll write about it when it's live.

If you try the package, I'd love to hear what you build. Issues, benchmarks on your hardware (especially ARM and GPU machines) and pull requests are all welcome on GitHub.

MahmoudAshraf/mms-300m-1130-forced-aligner`` onnx-community/mms-300m-1130-forced-aligner-ONNX``ctc-forced-aligner I'm Arham Sayyed, a full-stack developer from Mumbai who works mostly in Node.js and TypeScript. I built mms-forced-align and I'm building Awaaz. Find me on GitHub, LinkedIn and here on Dev.to.

── more in #ai-tools 4 stories · sorted by recency
── more on @meta 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-ported-meta-s-mms-…] indexed:0 read:15min 2026-10-09 · —