Modeling the world. Remodeling video.
TwelveLabs models can see and reason about video like no other AI – and they set the standard for a new era of video data interaction.
Marengo 3.0

Our breakthrough video foundation model analyzes frames and their temporal relationships, along with speech and sound — a huge leap forward for search and any-to-any retrieval tasks.
Pegasus 1.5
Our powerful video-first language model integrates visual, audio, and speech information — and employs this deep video understanding to reach new heights in text generation.
Marengo transforms text, audio, image, and video into numerical representations called embeddings.
Marengo
3.0
Powered Features
Where power meets potential.
Search
Introducing ‘any-to-any’ search
Marengo’s state-of-the-art ‘any-to-any’ search helps you pinpoint exact moments in vast video libraries, or allow customers to find any video moment within your platform.
Embed
Introducing ‘rich embeddings’
With Marengo, it’s easy to build complex features like semantic search, hybrid search, anomaly detection, and more.
Marengo’s latest breakthroughs.
Pegasus understands video and generates accurate descriptions and analysis.
Pegasus
1.5

Powered Features
Where words and moving image unite.
Analyze
Generate understanding with Pegasus
With video-to-text generation, Pegasus redefines how humans interact with video data. Intuitive, versatile, and powerful – this is human-level reasoning, at AI scale.













