TwelveLabs said its latest video-understanding model, Pegasus 1.6, can now analyze first-person footage, more accurately identify people and objects, and process still images through the same API used for video, broadening its use for robotics, media and security customers.

The model builds on Pegasus’s core function: converting video into structured, searchable text such as summaries, tags and timestamps that applications can use.
The company said the biggest addition is support for egocentric video — footage shot from the perspective of a person or machine, including body cameras and robot-mounted cameras. Such data is increasingly used by robotics and embodied-AI teams for labeling and training, a category TwelveLabs previously did not support well.
Pegasus 1.6 also improves entity recognition, making it better at identifying people, characters and objects and getting their names right. TwelveLabs said this is critical for workflows such as media cataloging and security review, where a confident but incorrect label can create errors that require human correction.
In addition, the upgrade allows still images to be analyzed natively through the same API, model and prompts used for video. That removes the need for a separate image-analysis pipeline, the company said, benefiting teams that work with both video and images, such as product videos and photos or screen recordings and screenshots.
TwelveLabs said Pegasus differs from general-purpose multimodal models that are built mainly for text and adapted to video, often treating video as a series of still frames. Pegasus was designed for video from the outset, allowing it to reason over an entire two-hour video in a single call and return structured, timestamped output natively, the company said.
The company listed media, sports and broadcasting as beneficiaries, with more complete summaries that can include on-screen graphics and jersey numbers. Security teams can use entity recognition across long footage, while robotics and physical AI teams can use first-person video for training and labeling. Developers get one API for video and images, with output shaped by their application’s schema.








the 2026 Bund Summit