Modeling the world. 

Remodeling video.

TwelveLabs models can see and reason about video like no other AI – and they set the standard for a new era of video data interaction.

Main hero accent video
Cover
Marengo 3.5
Graphic

Our breakthrough multimodal foundation model that analyzes video, audio, images, text, and documents—powering faster, more precise any-to-any retrieval with 512‑dimensional embeddings and an uncertainty signal on every match.

Cover
Pegasus 1.5

Our powerful video-first language model integrates visual, audio, and speech information — and employs this deep video understanding to reach new heights in text generation.

Marengo 3.5 is a multimodal embedding model for video, audio, images, text, and documents.

Embeddings enable unsurpassed information retrieval. Now, you can perform powerful cross-modal searches – across text, audio, image, video and documents.

This any-to-any retrieval can transform applications. Content discovery, recommendation systems, description and analysis will all change for good.

At TwelveLabs, we’re developing video-native AI systems that can solve problems with human-level reasoning. Helping machines learn about the world — and enabling humans to retrieve, capture, and tell their visual stories better.

Marengo

3.5
Graphic

Powered Features

Where power meets potential.

Embed

Introducing ‘rich embeddings’

Compose video, audio, image, text, and documents into a single query (16K-token sync endpoint) and get back a 512-dimensional embedding with an uncertainty signal on every match.

Graphic

Marengo is moving at lightning speed.

Marengo is moving at lightning speed.

Pegasus understands video and generates accurate descriptions and analysis.

A powerful intelligence interface, Pegasus can answer questions, generate creative outputs, and provide detailed analysis of any video.




Simply describe what you need in natural language. Get marketing suggestions, high-impact captions, or even a child-friendly summary of a video instantly.

At TwelveLabs, we’re developing video-native AI systems that can solve problems with human-level reasoning. Helping machines learn about the world — and enabling humans to retrieve, capture, and tell their visual stories better.

Pegasus

1.5
Graphic

Powered Features

Where words and moving image unite.

Analyze

Generate understanding with Pegasus

With video-to-text generation, Pegasus redefines how humans interact with video data. Intuitive, versatile, and powerful – this is human-level reasoning, at AI scale.

Graphic
Graphic

Pegasus is reaching ever-higher peaks.

Image

Pegasus is reaching ever-higher peaks.

Image

We’re fundamentally transforming how people see, experience and use video through AI.

Grounded in research and knowledge sharing, we’re learning and growing alongside our partners, with consideration for the future.

At TwelveLabs, we’re developing video-native AI systems that can solve problems with human-level reasoning. Helping machines learn about the world — and enabling humans to retrieve, capture, and tell their visual stories better.

Cover CTA

No one sees your video like TwelveLabs.

See for yourself what our AI can do - try one of your videos in our Playground.

Cover CTA

No one sees your video like TwelveLabs.

See for yourself what our AI can do - try one of your videos in our Playground.

Cover thread

No one sees your video like TwelveLabs.

See for yourself what our AI can do - try one of your videos in our Playground.