# Tighter on-device Cleanup with Epilude Model 4.1

> Source: <https://www.epilude.com/news/introducing-model-4-1>
> Published: 2026-07-28 17:36:15+00:00

Today we are shipping Epilude Model 4.1, a refinement of the on-device model that powers [Local Mode](/help/dictation/local-mode). It is built for the longer dictations where Model 4 could occasionally ramble or repeat itself. The result is tighter writing, fewer critical errors, and the same private, offline experience on your Mac.

## What Model 4.1 is

Local Mode runs two models directly on your Mac. A speech model turns your voice into text. A cleanup model turns that raw transcript into writing you would actually send: it removes filler, fixes punctuation, resolves false starts and mid-sentence corrections, and follows the tone you asked for.

Model 4.1 is a new version of the cleanup model. It is our fine-tune of an open-weight model in the Qwen family. The speech model is unchanged. Together they keep the same on-device footprint and hardware requirements as Model 4.

| Spec | Model 4.1 |
|---|---|
| Download size | 1.5 GB, one time |
| Hardware | Apple Silicon Macs |
| Full AI Cleanup | Macs with 16 GB of memory or more |
| Decoding | Greedy, fully deterministic |
| Connectivity | Works entirely offline |

## How we evaluate

We maintain an internal benchmark of 90 dictation scenarios spanning punctuation, formatting, self-corrections, tone control, multiple languages, mixed-language speech, and long-input structure. Every candidate runs the full suite repeatedly. Frontier-model judges cast multiple independent votes on every output, and a scenario counts as passed only when it passes in every judged run. We also re-grade identical outputs so grader noise does not look like a model change.

We do not publish this benchmark, and its results are not comparable to anything external. It is how we decide whether a cleanup model is safe enough to ship.

## Results

On that internal benchmark, Model 4.1 clears one more scenario than Model 4 and reduces the critical-error set without introducing a new one:

| Internal benchmark | Model 4 | Model 4.1 |
|---|---|---|
| Scenarios passed, of 90 | 78 | 79 |
| Critical errors | 4 | 3 |

A scenario counts as passed only when it passes in every repeated judged run. The remaining gains concentrate in the work that becomes visible only after you have been talking for a while: keeping a long answer on track, avoiding repetition, and staying faithful to the parts that are easy to over-clean.

## What we learned building it

This release reinforced a lesson that has shaped each model generation: evaluation quality has to match the failures you are trying to prevent. A broad score can hide the one sentence a person would never send. We made the quality bar more specific around long-input faithfulness, then held the new model to it across repeated runs.

It also reinforced the value of model averaging, a known technique in the research literature. We selected a stable combined model to avoid rare failure modes and deliver more dependable cleanup.

## Limitations

Model 4.1 is not perfect. Long, highly structured dictations remain harder than short messages, especially when a single thought contains several corrections, quoted material, and a change of tone. Cloud mode is still ahead on our internal suite overall. Local Mode remains the choice when keeping your audio and text on your Mac matters most.

## What's next

We will keep working on the difficult cases that matter most in real writing: longer inputs, multilingual dictation, and models that preserve a speaker's intent under cleanup. If Model 4.1 changes how dictation feels on your Mac, or if you catch it making a mistake, we want to hear about it through the [Help Center](/help).
