An ordinary video records color and motion, but it does not explicitly tell an editor which pixels are near the camera and which belong to the background.
I added browser-local depth estimation to Timeline Studio, my open-source video editor, to turn that missing information into an editable resource. The current implementation supports depth-aware blur, depth-based relighting, and exporting an independent grayscale depth video.
This post covers the implementation and what happened when I used the resulting depth videos as references for AI video generation.
The implementation uses Depth Anything V2 Small to estimate relative depth from RGB images. In our default visualization, brighter pixels generally represent nearer surfaces, and darker pixels represent farther ones. The display can also be inverted.
Relative depth is not a distance measurement. A grayscale value does not tell us that an object is exactly two meters from the camera. It is useful for visual layering and effects, but it is not a calibrated 3D reconstruction.
The processing flow is:
Video segment
-> map timeline time to source time
-> decode and resize a frame
-> run depth inference in a Web Worker
-> normalize and refine depth pixels
-> preview an effect or encode a depth video
The worker loads the model through Transformers.js with WebGPU and a Q4F16 configuration. The essential setup looks like this:
const estimator = await pipeline(
"depth-estimation",
modelPath,
{
device: "webgpu",
dtype: "q4f16",
progress_callback: reportProgress,
},
);
Production setup also handles pinned model revisions, mirrored downloads, progress reporting, optional caching, and failures.
A worker keeps inference work off the main JavaScript thread. It does not eliminate resource contention: decoding, inference, and editor rendering still share the device. That is why bounded work matters.
Video frames stay local during this depth-analysis path. Sending the exported depth video to a remote generation service is a separate operation that uploads reference material.
This was one of the most important integration details.
An editor segment may be trimmed, sped up, slowed down, or assigned a speed curve. Sampling the original file at the same numerical timestamp as the timeline would produce incorrect depth frames.
The analysis uses the editor's shared source-time mapping:
const sourceTime = getVisualSourceTime(segment, localTime);
Preview and depth-video output use the same mapping. Analysis cache signatures include the source, trim, playback rate, speed curve, quality setting, and model revision. Relevant changes invalidate the old analysis.
A good depth estimate attached to the wrong video frame is still a broken effect.
The current analysis presets are:
| Preset | Target sample rate | Maximum input dimension |
|---|---|---|
| Fast | 8 fps | 392 px |
| Balanced | 16 fps | 504 px |
| Quality | 24 fps | 518 px |
These are sampling settings, not measured real-time inference speeds. Actual throughput depends on the browser, GPU, and workload.
Each analysis has a 720-sample budget. Longer clips reduce their effective sampling density to cover the segment within that budget.
The implementation overlaps inference for the current frame with preparation of the next frame. It keeps only a limited number of tasks in flight instead of decoding a whole video into memory.
Independent depth-video output runs at 24 fps. Fast and Balanced modes use RIFE interpolation between depth samples. Interpolation can make transitions smoother, but it does not add genuine depth observations and can still fail around rapid motion or occlusion.
For depth-aware blur, the renderer prepares a sharp layer, a blurred layer, and a depth-derived mask.
Pixels close to the selected focus depth stay sharp. Pixels outside the focus band progressively mix in the blurred layer, with a smooth transition at the boundary.
This is a compositing approximation rather than a physically exact lens simulation. Hair, transparent materials, and complex occlusion remain challenging.
The depth output also receives a small edge-guided smoothing pass. Neighboring pixels contribute less when their depth or source-color differences are large. This reduces unwanted mixing across foreground/background boundaries; image texture guides smoothing rather than becoming new geometry.
The editor can encode depth results as a new asset in My assets. The current browser path uses WebM with VP9 video, preserves the source aspect ratio, and caps the longest output dimension at 1080 pixels. Where source audio exists, the export prepares audio matched to the segment timing.
Output size and inference resolution are different. Encoding a larger image does not recover geometry the model never predicted.
I also tested a workflow combining three inputs:
There is an important qualification: these experiments submitted the depth video as an ordinary reference video, not through a dedicated depth-conditioning API. The generation service had to interpret the grayscale reference itself.
Some broad movements survived, including turning toward the camera and moving into a closer shot. But the results were not frame-exact. In one experiment, a foreground limb shape was interpreted as a different body action, and head/shoulder proportions drifted from the reference.
A separate two-character sleeve-transition experiment followed the main action structure more closely when I supplied the original RGB video instead. Sampled-frame inspection showed the sleeve raise, full-frame occlusion, and subsequent close-up, while hand details still differed.
The practical lesson is to choose reference inputs according to the task. Depth emphasizes spatial structure, while RGB retains semantic information about gestures, fabric, and objects.
A relative depth map also does not directly provide a calibrated camera FOV. Framing fidelity needs to be checked through screen-space proportions, crop boundaries, and camera changes. Up a depth reference alone is not proof of accurate reconstruction.
The browser implementation connects depth analysis to editable effects and independent video output. The engineering work is as much about timing, cancellation, progress, cache validity, and bounded resources as it is about running a model.
The generation experiments are useful evidence of possibilities and limitations, not a promise of precise motion transfer. Better temporal stability, difficult boundaries, and dedicated depth-conditioning integrations remain areas to improve.
If you are building browser AI tools or experimenting with depth-guided editing, I would be interested in how you handle these issues.
Timeline Studio is free and open source. Feedback, reproducible issues, and contributions are welcome.