# VideoDB leads a 100-hour action retrieval benchmark vs. AWS Nova and Twelve Labs

> Source: <https://videodb.io/blog/a-faster-path-to-the-next-policy>
> Published: 2026-10-09 15:14:59+00:00

# VideoDB leads a 100-hour action retrieval benchmark against AWS Nova and Twelve Labs

A head-to-head retrieval benchmark, and what it means for robotics teams choosing what their next policy should learn from.

On a 100-hour benchmark of 2,835 egocentric videos and 11,369 text queries, **VideoDB Action Segmentation + Search** returned a relevant moment in its top 10 results for **40.39%** of queries, about **5 points ahead of AWS Nova (35.40%)** and **7.5 points ahead of Twelve Labs Marengo 3.5 (32.87%)**. The gap widens when the result must match the action’s actual start and end: at temporal IoU ≥ 0.5, VideoDB reaches **20.49%**, about **2.05×** the next-best pipeline, and at 0.75 it reaches **12.83%**, about **3.81×**.

We also gave five open embedding models VideoDB’s action boundaries in place of fixed four-second windows. All five improved on temporally matched retrieval, by 1.71–3.57× at tIoU ≥ 0.5 (top 10), which suggests that the choice of temporal intervals is an important part of retrieval design alongside the embedding model.

Why this matters: before training the next robot policy, researchers have to find specific moments in hours of recordings, like a slip, a clean transfer or a recovery, and decide what the model should learn from. Below we explain how we measured retrieval, what we tested, and how this fits into that workflow.

Throughout this article, we use an illustrative cup-transfer task to follow that workflow: a robot loses its grip while moving a cup toward a tray, and a researcher locates the attempt and compares it with successful transfers and recoveries.

## Retrieve the action, not just the video

For this kind of investigation, a search result needs to capture enough of an action to explain what happened. A clip showing a cup in a gripper may be visually relevant, yet reveal little about whether the grasp was stable or why the transfer failed. The researcher needs a temporal unit that preserves the relationship between the attempt and its outcome, with a link back to the recording when more context is needed.

Action segmentation provides a way to organize video into those units by assigning descriptions and boundaries to individual actions. Our [revised WGO-Bench study](https://videodb.io/blog/wgo-bench-action-annotations) evaluated how accurately VideoDB could identify action boundaries and label actions across 100 human and robot videos. VideoDB reached **28.93% semantic F1 at temporal IoU 0.5**, compared with **19.92%** for the public MacroData Refiner baseline. That study concerns the quality of the annotations; the next question is whether a retrieval pipeline can use action records to find relevant moments across a larger collection.

The boundaries of each result are important to that evaluation. In the cup example, a four-second window might capture the grasp without the slip, while another captures the drop without the approach. Either could return a relevant fragment, but the researcher would still have to reconstruct the sequence before deciding what it shows.

We measure this aspect of retrieval with **temporal intersection over union**, or tIoU: the shared time between a returned interval and a reference action, divided by the total time covered by either. A four-second fragment inside a ten-second action can reach at most 0.4 tIoU, even if every returned frame belongs to the action. Returning a much longer interval creates the opposite problem, preserving the event while adding unrelated activity that also lowers temporal agreement. The aim is to retrieve a well-bounded action while keeping the surrounding recording available for review.

## What 100 hours of retrieval tells us

To evaluate retrieval at a larger scale, we compared eight pipelines on **a 100-hour benchmark curated from [Lightwheel’s EgoStandard](https://huggingface.co/datasets/LightwheelAI/EgoStandard)**, comprising 2,835 videos, 11,369 text queries, and 17,967 reference intervals. The corpus captures human activity from an egocentric perspective, with reference intervals identifying individual actions within the recordings.

### How we compare the pipelines

Each pipeline searches the same videos with the same text queries and returns up to 50 timestamped candidates. What differs is how it divides the video, represents each interval, and ranks the results.

**VideoDB Action Segmentation + Search** generates action records with descriptions and temporal boundaries, then retrieves and ranks those records against the query. The approach connects semantic search to source-linked intervals, as described in our [search architecture report](https://labs.videodb.io/papers/search-over-the-visual-world.pdf).

**AWS Nova** uses Nova 2 Multimodal Embeddings through Amazon Bedrock Knowledge Bases, returning provider-managed, timestamped video chunks. **Twelve Labs** uses [Marengo 3.5](https://docs.twelvelabs.io/v1.3/docs/concepts/models/marengo/marengo-3-5) visual clip embeddings and the start and end times returned by its embedding service. Twelve Labs supports its own [temporal segmentation](https://docs.twelvelabs.io/v1.3/docs/guides/create-embeddings/at-scale/video). Marengo’s 512-dimensional embeddings are normalized and searched locally with FAISS.

For the local comparison, we selected five prominent open embedding models: [SigLIP2](https://huggingface.co/google/siglip2-so400m-patch14-384), [Qwen3-VL Embedding](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B), [Cosmos-Embed](https://huggingface.co/nvidia/Cosmos-Embed1-336p), [InternVideo2](https://github.com/OpenGVLab/InternVideo/tree/main/InternVideo2/multi_modality), and [PE-AV](https://huggingface.co/facebook/pe-av-large). We encode **four-second windows sampled at 2 FPS**, giving eight frame samples per full window, with each model’s native preprocessing. We use **FAISS as the local vector index** to store normalized embeddings.

### What counts as a successful result

We score the first K results in two ways, both requiring a match in the **same source video**:

- **Hit@K** uses a minimum-overlap threshold. At least one returned interval must overlap a reference action by**min(0.5 seconds, 50% of the reference duration)** . For an action lasting ten seconds, just half a second of overlap is enough; for a 0.4-second action, the threshold is 0.2 seconds. This is a permissive overlap check, so it can count a fragment without capturing the full action.
- **Hit@K at temporal IoU 0.5 or 0.75** uses a stricter temporal-overlap threshold. At least one result must meet the specified ratio of shared time to total covered time. Both where the interval starts and where it ends affect this score: missing part of the action or including extra footage outside it lowers temporal IoU. A higher threshold demands closer temporal agreement.

Both metrics report the percentage of queries with at least one qualifying result. The difference is how much temporal agreement is needed for a result to count as a match.

VideoDB leads Hit@1, Hit@5, and Hit@10, with **40.39% at Hit@10**, compared with **35.40%** for AWS Nova and **32.87%** for Twelve Labs. Its lead is larger on temporally matched retrieval: at Hit@10 (tIoU ≥ 0.5), VideoDB reaches **20.49%**, about **2.05×** the next-best pipeline’s 10.01%. At the stricter 0.75 threshold, it reaches **12.83%**, about **3.81×** the next-best pipeline’s 3.37%, and leads every reported IoU-qualified column.

AWS Nova and Twelve Labs achieve higher Hit@50 under the permissive overlap rule. VideoDB leads at every reported retrieval depth on Hit@K at both tIoU thresholds, which require the retrieved interval to match the reference action more closely.

For researchers reviewing search results, early ranking and temporal agreement are useful properties because they determine what appears first and how much of the action it contains. The comparison evaluates each complete pipeline, including its representation, temporal units, and ranking.

## A search result is not a training decision

The distinction between relevance and suitability becomes clear when we return to the cup transfer. A successful attempt could help a researcher compare approaches and grasps, a failed attempt could support diagnosis, and a recovery could be useful for a different learning objective. All three may belong in the search results, but deciding which should enter training requires inspecting the sequence and understanding the supervision available in its source.

Each selected interval should carry enough context to explain where it came from, why it was chosen, and how it is intended to be used. That includes the source episode, timestamps, the origin of its labels, and the reviewer’s decision. Where compatible robot observations, actions, and state are available, the selection should preserve their relationship to the video and the control schema needed to interpret them. This makes it possible to revisit a selection when an experiment succeeds, fails, or reveals an unexpected regression.

The source of the experience also constrains how it can be used. Human video can support visual representations, task structure, or semantic supervision, but it does not automatically provide robot control targets. Similarly, a failed action may be informative for diagnosis without being an appropriate behavior-cloning target. Retrieval should help researchers make these distinctions rather than treat every relevant clip as interchangeable training data.

Existing research illustrates several ways action structure can contribute to learning. [π0.5](https://arxiv.org/abs/2504.16054) combines semantic subtask supervision with low-level action learning, while [SARM](https://arxiv.org/abs/2509.25358v4) uses subtask annotations for stage and progress supervision, then learned rewards to filter and reweight demonstrations. [COLLAGE](https://arxiv.org/abs/2508.01131) approaches demonstration selection through multiple retrieval cues and task-specific weighting. These methods use different learning objectives, but each makes the role of the selected or annotated experience explicit.

Search could also help investigate changes between policy versions. If a new checkpoint appears to drop cups more often, researchers could retrieve attempts linked to each version and compare grasps, slips, and recoveries under similar conditions. This helps distinguish a possible policy regression from changes in objects or hardware, while consistent outcome labels and attempt counts provide the basis for measuring it.

That comparison could then guide the next data decision: select successful grasps under the affected conditions, review recovery examples, or collect new demonstrations if those cases are missing. More generally, when the available recordings do not cover the behavior a researcher needs to study, the search can help turn a broad request for more data into a specific collection target.

## Do VideoDB’s action boundaries help other embedding models?

On the same 100-hour benchmark, we also compared five embedding models using the action boundaries from VideoDB, which leads the temporal-IoU retrieval metrics above. Rather than dividing the videos into four-second windows, we gave Cosmos-Embed, SigLIP2, PE-AV, Qwen3-VL Embedding, and InternVideo2 the same action intervals and let each model encode the original frames within them. Only the segmentation is shared; each model computes its own embeddings. Comparing these with each model’s four-second-window scores shows how much the choice of interval affects retrieval.

The comparison uses the same corpus, queries, and reference intervals. Results for each model are below.

### Four-second windows vs. VideoDB action intervals, by model

Values are query success percentages.

| Each cell shows four-second windows followed by VideoDB action intervals. |  |  |  | 
|---|---|---|---|
| Encoder | Hit@10 | Hit@10 tIoU ≥ 0.5 | Hit@10 tIoU ≥ 0.75 | 
|---|---|---|---|
| Cosmos-Embed | 23.78 → **24.94** | 6.08 → **12.06** | 2.08 → **7.50** | 
| SigLIP2 | 24.96 → **26.91** | 6.22 → **12.25** | 1.90 → **6.93** | 
| PE-AV | 12.13 → **16.47** | 2.44 → **8.70** | 0.82 → **5.19** | 
| Qwen3-VL Embedding | 24.20 → **26.58** | 5.30 → **12.75** | 1.79 → **7.88** | 
| InternVideo2 | 19.50 → **20.02** | 4.98 → **8.51** | 1.61 → **5.07** | 

Showing top 10 results for five embedding models.

Across all five tested embedding models, replacing four-second windows with VideoDB action intervals improves temporally matched retrieval at every reported depth. At top 10, success rates increase by **1.71–3.57× at tIoU ≥ 0.5** and **3.15–6.33× at tIoU ≥ 0.75**. The improvement extends across encoders, supporting the hypothesis that the choice of temporal intervals is an important part of retrieval design alongside the embedding model.

The improvement is most consistent when retrieval is evaluated against the action’s temporal boundaries. Hit@K counts a result once it meets a minimum-overlap threshold, even if it captures only a small part of the action. Temporal IoU demands closer agreement with the reference interval. Under this stricter measure, VideoDB action intervals improve retrieval across all five encoders at every reported depth, showing a consistent benefit in retrieving intervals that better match the actions being searched for.

## Measure progress at the next model

Our broader goal is to make the work between model versions faster by reducing the effort it takes to turn recorded robot experience into useful training selections. VideoDB Action Segmentation + Search helps researchers locate relevant actions across recordings, inspect their context, and decide which examples belong in the next experiment. We want researchers to spend less time finding evidence and more time testing what the next policy should learn, so the experience they have already collected becomes a better foundation for the next model.
