cd /news/artificial-intelligence/we-analyzed-1532-audio-cleanup-jobs-… · home topics artificial-intelligence article
[ARTICLE · art-92332] src=vidclean.net ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

We Analyzed 1,532 Audio Cleanup Jobs. Two Thirds Weren't Audio Files

VidClean analyzed 1,533 audio cleanup jobs between July 20 and August 11, 2026, finding that 67% (1,034) of files were videos, not audio files. The company's three tools—Remove Background Noise, Repair Audio, and Enhance Speech—all use the same DeepFilterNet3 neural noise model, but Enhance Speech's share of jobs doubled to 19% in the last three weeks, while Remove Background Noise still dominates at half of all jobs. Nearly half of all cleanup jobs were under one minute of media, and 86% were under five minutes.

read7 min views1 publishedAug 11, 2026
We Analyzed 1,532 Audio Cleanup Jobs. Two Thirds Weren't Audio Files
Image: source

Blog

Melvin Bucio6 min read VidClean runs three tools that clean up sound: Remove Background Noise, Repair Audio, and Enhance Speech. All three are built on the same neural noise model (DeepFilterNet3), but they package it differently. Remove Background Noise strips noise and mains hum and touches nothing else. Repair Audio does the same and also normalizes loudness to broadcast level. Enhance Speech runs the full repair chain, then lets you dial how strongly the noise removal is applied with a strength slider.

Between July 20 and August 11, 2026, people ran 1,533 files through those three tools. This post is about what those jobs looked like: which tool people actually reach for, what kind of files they upload, and how long those files run.

One difference from our silence study, up front: that dataset counted seconds, so it could report hours of dead air. The audio counters store job counts and coarse buckets, not total durations, so this piece is about counts. No "hours of noise removed" number appears below because no honest one exists.

WHICH TOOL PEOPLE ACTUALLY REACH FOR #

Of the 1,533 jobs in the window:

Half the demand is the bluntest promise: make the noise go away. The two tools that do more than that, normalizing loudness or offering a blend control, split the rest.

The mix is moving, though. These three tools have lifetime counters going back to their launches in May, and the before and after look different. Before July 20, Remove Background Noise was 70% of all cleanup jobs, Repair Audio 21%, and Enhance Speech 9%. In the last three weeks Enhance Speech's share has doubled to 19%.

Part of that is a launch artifact: Enhance Speech shipped about two and a half weeks after the other two, so its lifetime share starts diluted. But correcting for days live does not close the gap. Adjusted to per-day rates, all three tools grew between the two periods, because the site's traffic grew, but they did not grow equally: Remove Background Noise runs at 3.3 times its earlier daily rate, Repair Audio at 5.5 times, and Enhance Speech at 6.5 times. The simplest tool still dominates, but the growth is concentrated in the tools that do more than delete noise.

"AUDIO CLEANUP" IS MOSTLY A VIDEO PROBLEM #

Here is the number that surprised me. Of the 1,533 files uploaded for audio cleanup, 1,034 were videos. That is 67%: two out of three.

I assumed a noise removal tool would mostly see audio files: podcast WAVs, voice memos, narration tracks. Those exist, 499 of them, but they are the minority. The typical job is a video whose audio needs rescuing: a talking head clip, a screen recording, a phone video with a fan or an air conditioner humming behind it.

That reframes what these tools are for. Nobody records a video and then wants a clean MP3 back; they want the same video with the noise gone. Accordingly, all three tools hand back the same video with the cleaned audio swapped in, leaving the picture untouched wherever the format allows. The finding is that this path is not the edge case. It is the main road.

It also says something about where bad audio comes from. Dedicated audio recordings are usually made on purpose, with at least a passable microphone. Video audio is very often an afterthought: whatever the camera or phone picked up from across the room. The uploads reflect that.

THE FILES ARE CLIPS, NOT EPISODES #

How long is a typical cleanup job? Short. Very short.

Nearly half of all cleanup jobs are under one minute of media. 86% are under five minutes.

One disclosure before reading too much into the right-hand tail: the free tier caps these three tools at 10 minutes per file (Pro raises it to 3 hours), so this distribution is truncated. The absence of half-hour files in the chart is partly the product's doing, not purely the users'.

The left side of the chart needs no such asterisk. Everyone in this dataset was free to upload a 10 minute file, and half of them uploaded less than one minute instead. Noise cleanup, as actually practiced, is dominated by clips: the short-form video, the single take, the one soundbite that has to be usable.

The same counters cover VidClean's other tools in the same window, which makes the shape easier to see by contrast. Files under five minutes make up 86% of audio cleanup jobs, 68% of silence removal jobs, and 55% of transcription jobs. Every one of these tools lets a free user upload a file well past the five minute mark, so the differences here are behavior, not limits. People transcribe episodes. They remove silence from recordings of every length. But when audio is broken, what they fix is a clip.

I can offer an interpretation, though the data cannot confirm it: cleanup is triage, not process. Transcription and silence removal are things you do to everything as part of a workflow. Noise removal is something you do when a specific short thing went wrong.

METHODOLOGY #

VidClean deletes uploaded files within an hour of processing and keeps no per-file records, so this analysis is built on aggregate counters. When a cleanup job completes, the backend increments a handful of running totals: one for the tool used, one coarse duration bucket, and one flag for whether the input was an audio file or a video. No filenames, no user identifiers, no timestamps, and no per-file rows exist anywhere. No individual upload can be reconstructed from the data, and the analysis is limited to exactly the cuts shown above.

The window is July 20 to August 11, 2026. July 20 is simply when these counters shipped. The lifetime numbers in the trend section come from separate per-tool completion counters that have existed since each tool launched in May 2026.

Two disclosures. First, VidClean's own automated production tests run through the same pipeline, roughly one short synthetic file per tool per test run, on files a few seconds long. They inflate the under-one-minute bucket somewhat: excluding a deliberately generous estimate of test volume still leaves the sub-minute share around 45%, so "about half" survives, but the exact 48% should be read with that grain of salt. The video versus audio split barely moves under the same exclusion. Second, all three tools cap file length as described above, and the duration chart should be read as a distribution over files that fit the caps.

This data is licensed CC BY 4.0. Feel free to reuse the numbers with credit to VidClean.

LIMITATIONS #

This is VidClean's user base, not a random sample of anyone's audio. People who upload to a noise removal tool believe they have a noise problem, and people who found a free web tool skew toward exactly the short-clip, phone-video use case the data shows. I make no claim that these proportions generalize to all recorded audio.

The counters pool some things this post would love to split. Duration and the audio-versus-video flag are recorded for the tool family as a whole, not per tool, so I cannot tell you whether Repair Audio jobs run longer than Enhance Speech jobs. The per-tool numbers here are exactly the ones the counters can support: job counts and their trend.

And the three-week window is what it is: long enough for 1,533 jobs, short enough that the trend figures, especially, deserve a re-check in a few months. The counters keep running. When they say something new, we will write it down.

WHICH TOOL FITS YOUR PROBLEM #

If the data suggests anything practical, it is that most people reach for the broadest tool first, and the split between the three is worth knowing before you upload. Use Remove Background Noise when the recording level is fine and something in the room is not. Use Repair Audio when it is also too quiet or uneven, because that one normalizes loudness as well. Use Enhance Speech when full noise removal sounds too aggressive and you want to dial it back. All three are free in the browser, no account needed, and all three accept a video file and give you the video back.

VidClean is a solo project. Questions about the data or the methodology are welcome at hello@vidclean.net.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @vidclean 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/we-analyzed-1532-aud…] indexed:0 read:7min 2026-08-11 ·