Modeling the world. Remodeling video.
TwelveLabs models can see and reason about video like no other AI – and they set the standard for a new era of video data interaction.
Marengo 3.5

Our breakthrough multimodal foundation model that analyzes video, audio, images, text, and documents—powering faster, more precise any-to-any retrieval with 512‑dimensional embeddings and an uncertainty signal on every match.
Pegasus 1.5
Our powerful video-first language model integrates visual, audio, and speech information — and employs this deep video understanding to reach new heights in text generation.
Marengo 3.5 is a multimodal embedding model for video, audio, images, text, and documents.
Marengo
3.5

Powered Features
Where power meets potential.
Embed
Introducing ‘rich embeddings’
Compose video, audio, image, text, and documents into a single query (16K-token sync endpoint) and get back a 512-dimensional embedding with an uncertainty signal on every match.

Marengo’s latest breakthroughs.
Pegasus understands video and generates accurate descriptions and analysis.
Pegasus
1.5

Powered Features
Where words and moving image unite.
Analyze
Generate understanding with Pegasus
With video-to-text generation, Pegasus redefines how humans interact with video data. Intuitive, versatile, and powerful – this is human-level reasoning, at AI scale.










