{"slug": "tighter-on-device-cleanup-with-epilude-model-4-1", "title": "Tighter on-device Cleanup with Epilude Model 4.1", "summary": "Epilude released Model 4.1, a refined on-device cleanup model for Local Mode on Mac, that reduces critical errors from 4 to 3 and passes 79 of 90 internal benchmark scenarios, up from 78. The model is a fine-tune of an open-weight Qwen model, maintains the same 1.5 GB download and Apple Silicon hardware requirements, and runs entirely offline. Epilude said the release focuses on improving long-input faithfulness and avoiding repetition.", "body_md": "Today we are shipping Epilude Model 4.1, a refinement of the on-device model that powers [Local Mode](/help/dictation/local-mode). It is built for the longer dictations where Model 4 could occasionally ramble or repeat itself. The result is tighter writing, fewer critical errors, and the same private, offline experience on your Mac.\n\n## What Model 4.1 is\n\nLocal Mode runs two models directly on your Mac. A speech model turns your voice into text. A cleanup model turns that raw transcript into writing you would actually send: it removes filler, fixes punctuation, resolves false starts and mid-sentence corrections, and follows the tone you asked for.\n\nModel 4.1 is a new version of the cleanup model. It is our fine-tune of an open-weight model in the Qwen family. The speech model is unchanged. Together they keep the same on-device footprint and hardware requirements as Model 4.\n\n| Spec | Model 4.1 |\n|---|---|\n| Download size | 1.5 GB, one time |\n| Hardware | Apple Silicon Macs |\n| Full AI Cleanup | Macs with 16 GB of memory or more |\n| Decoding | Greedy, fully deterministic |\n| Connectivity | Works entirely offline |\n\n## How we evaluate\n\nWe maintain an internal benchmark of 90 dictation scenarios spanning punctuation, formatting, self-corrections, tone control, multiple languages, mixed-language speech, and long-input structure. Every candidate runs the full suite repeatedly. Frontier-model judges cast multiple independent votes on every output, and a scenario counts as passed only when it passes in every judged run. We also re-grade identical outputs so grader noise does not look like a model change.\n\nWe do not publish this benchmark, and its results are not comparable to anything external. It is how we decide whether a cleanup model is safe enough to ship.\n\n## Results\n\nOn that internal benchmark, Model 4.1 clears one more scenario than Model 4 and reduces the critical-error set without introducing a new one:\n\n| Internal benchmark | Model 4 | Model 4.1 |\n|---|---|---|\n| Scenarios passed, of 90 | 78 | 79 |\n| Critical errors | 4 | 3 |\n\nA scenario counts as passed only when it passes in every repeated judged run. The remaining gains concentrate in the work that becomes visible only after you have been talking for a while: keeping a long answer on track, avoiding repetition, and staying faithful to the parts that are easy to over-clean.\n\n## What we learned building it\n\nThis release reinforced a lesson that has shaped each model generation: evaluation quality has to match the failures you are trying to prevent. A broad score can hide the one sentence a person would never send. We made the quality bar more specific around long-input faithfulness, then held the new model to it across repeated runs.\n\nIt also reinforced the value of model averaging, a known technique in the research literature. We selected a stable combined model to avoid rare failure modes and deliver more dependable cleanup.\n\n## Limitations\n\nModel 4.1 is not perfect. Long, highly structured dictations remain harder than short messages, especially when a single thought contains several corrections, quoted material, and a change of tone. Cloud mode is still ahead on our internal suite overall. Local Mode remains the choice when keeping your audio and text on your Mac matters most.\n\n## What's next\n\nWe will keep working on the difficult cases that matter most in real writing: longer inputs, multilingual dictation, and models that preserve a speaker's intent under cleanup. If Model 4.1 changes how dictation feels on your Mac, or if you catch it making a mistake, we want to hear about it through the [Help Center](/help).", "url": "https://wpnews.pro/news/tighter-on-device-cleanup-with-epilude-model-4-1", "canonical_source": "https://www.epilude.com/news/introducing-model-4-1", "published_at": "2026-07-28 17:36:15+00:00", "updated_at": "2026-07-28 17:52:22.667126+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-products", "ai-tools"], "entities": ["Epilude", "Epilude Model 4.1", "Qwen", "Mac", "Apple Silicon"], "alternates": {"html": "https://wpnews.pro/news/tighter-on-device-cleanup-with-epilude-model-4-1", "markdown": "https://wpnews.pro/news/tighter-on-device-cleanup-with-epilude-model-4-1.md", "text": "https://wpnews.pro/news/tighter-on-device-cleanup-with-epilude-model-4-1.txt", "jsonld": "https://wpnews.pro/news/tighter-on-device-cleanup-with-epilude-model-4-1.jsonld"}}