# TwelveLabs debuts Pegasus 1.6 to turn first-person video into robot training data

> Source: <https://cryptobriefing.com/twelvelabs-pegasus-1-6-robotics-training-data/>
> Published: 2026-10-06 18:20:55+00:00

Photo: Steve A Johnson / Pexels

# TwelveLabs debuts Pegasus 1.6 to turn first-person video into robot training data

The video intelligence model adds native egocentric understanding, aiming to automate the slow manual labeling that robotics teams rely on

Teaching a robot to fold laundry starts with someone watching hours of footage of humans folding laundry. Then that person writes down every move.

TwelveLabs wants to take that second job off the table. On October 6, 2026, the company released **Pegasus 1.6**, a video intelligence model with native support for egocentric video. That means footage shot from a first-person point of view, built into the model’s core design.

The target audience is clear. Robotics and physical AI teams have plenty of raw footage, and they increasingly need ways to turn it into usable training data faster.

## What Pegasus 1.6 actually does

Egocentric video is the kind captured by wearable cameras, teleoperated systems and similar sources. Picture a GoPro strapped to a worker’s head, or the camera feed from a human remotely steering a robotic arm.

Pegasus 1.6 is built to handle it. The model automates several workflows that robotics teams typically grind through by hand:

**Action segmentation and labeling.** The model breaks a video into discrete steps and tags what is happening in each one.

**Dense captioning of spatial interactions.** Rather than a one-line summary, the model generates detailed descriptions of how hands and objects move relative to each other.

### AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

**Quality control.** The model can also help flag issues in footage before it ever reaches a training pipeline.

Beyond video, Pegasus 1.6 supports image analysis of 1 to 20 images per request. TwelveLabs also points to improved entity recognition and richer metadata extraction.

## The specs, and the manual labor they target

Pegasus 1.6 ships with a **261,120-token context window**, large enough to process long videos of up to 2 hours.

Labeling egocentric footage by hand can take **70 to 155 hours for every hour of video**. Run the math on a single two-hour clip and a human annotator could be looking at weeks of work for one file.

TwelveLabs is positioning the model as a way for automakers and robotics firms to move from small, carefully curated datasets toward a scalable pipeline built on real human experience captured on video.

Access runs through the TwelveLabs API. One caveat for teams planning large jobs: batch analysis is still handled by Pegasus 1.5, so high-volume bulk processing has not fully moved to the new model yet.

## Six months after Pegasus 1.5

The release comes roughly half a year after its predecessor. Pegasus 1.5 launched on April 20, 2026, and introduced Time-Based Metadata Extraction for 2-hour videos without requiring indexing.

In practical terms, 1.5 let users pull structured, timestamped information out of long videos without first running them through a separate indexing step. Version 1.6 builds on that foundation and points it at a narrower, more demanding use case: footage where the camera is effectively the actor’s eyes.

**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

[Editorial Policy](https://cryptobriefing.com/editorial-policy/).
