Product

Jockey: Your Entire Video and Photo Library Organized, Searchable, and Understood the Way You Think About It

James Le, Travis Couture

Jockey has since been rebuilt and relaunched as a first-party TwelveLabs product, not an open-source LangGraph project, now in research preview.

Jockey has since been rebuilt and relaunched as a first-party TwelveLabs product, not an open-source LangGraph project, now in research preview.

In this article

No headings found on page

Join our newsletter

Receive the latest advancements, tutorials, and industry insights in video understanding

Search, analyze, and explore your videos with AI.

Jul 3, 2024

5 Min

Copy link to article

TL;DR

Jockey is now in research preview. It's TwelveLabs' video AI agent layer built on top of Marengo, Pegasus, and Search that lets developers, and consumers work with a video and photo library the way they actually think about it: by person, by moment, by insight; not by filename or timestamp. Jockey reaches people through two surfaces: the API for builders, and MCP on Claude. This article walks through what's new: the full-stack flywheel, global entity understanding, and how to get started.

1 - Overview of Twelve Labs APIs

TwelveLabs has shipped Pegasus, Marengo, and Search as separate model capabilities - each best-in-class on its own. Jockey is the first product to integrate all three into a single agent surface. Pegasus handles understanding and description; Marengo handles multimodal indexing and retrieval across video and photos; Search adds multi-step, iterative retrieval on top. Because Jockey sits on top of the full stack, it isn't limited to a single clip or a single query; it reasons across your entire library, and it gets better automatically every time one of those models improves.


2 - The Stack Compounds

Every model upgrade is an automatic Jockey upgrade. TwelveLabs has shipped Pegasus, Marengo, and Search as separate model capabilities, released on their own timelines. Jockey is the first product to integrate all of them into one agent, which means every foundation-model upgrade automatically makes Jockey better. When Pegasus improves, Jockey's metadata extraction improves. When Marengo improves, Jockey's search improves. When Search lands new operators, Jockey's cognition engine gets new tools to plan over. No API migration required; the same Jockey session just gets smarter underneath you.


3 - One Agent, Reachable Everywhere

Jockey reaches every audience through a single underlying engine; you don't get a different AI depending on how you connect, you get the same one:

  • API: for builders who want to integrate video and photo intelligence directly into their own applications, with structured citations and asset grouping in every response.

  • MCP on Claude: for teams and consumers who want to connect their own photo and video library and just ask. No new interface to learn: Jockey meets people inside tools they already use.

Under the hood, Jockey combines your library's index (built on Marengo + Pegasus) with an agentic search loop (Search) that can iteratively refine a query until it finds the right person, moment, or insight; and returns a structured, citable answer, not just a wall of text.


4 - The first AI that understands who's in your library

Most video AI tools process a file. Jockey processes an entire video and photo library. The headline capability of this release is global entity understanding: recognizing the same person across thousands of videos and photos in your library, and automatically grouping everything they appear in. This isn't a filter you apply after the fact; it's a new primitive Jockey reasons over:

  • Digital Marketing Teams: “Find ads that overperformed on Instagram relative to YouTube. Have Jockey look at their content and explain what about them suits the IG feed vs YouTube."

  • Creators: "I'm launching a new series. Based on my existing content, find the topics my audience engages with most and draft a 4-episode arc with hooks for each."

Every other Jockey capability (content assembly, content organization, performance analysis, even search itself) gets dramatically better once "the same person across many videos" is a stable building block.

5 - Build with the Jockey API

Jockey isn't something you fork and modify, it's something you build on. Builders integrate through the Jockey API to add video and photo intelligence to their product without standing up their own ML infrastructure: query a knowledge store, get back structured, timestamped, citable results you can render directly in your app. In research preview, that includes the Filtered Response API for scoped, grouped results, and full parity between video and photo search, so an app built on Jockey can reason over mixed-media libraries out of the box. Early research preview builders report a 70%+ "could not have built this" rate and a median time-to-working-prototype under 4 hours.

6 - Try Jockey

Jockey is now live at twelvelabs.io/jockey with limited spots during our research preview. Sign up through Playground, where you'll find guidance to:

  1. Connect via MCP: add Jockey as a connector for Claude and upload your photo and video library.

  2. Build with the API: get an API key and start querying a knowledge store directly.

Pricing:

  • Free: 5GB storage, 15 responses calls/day

  • Plus ($20/mo): 100GB storage, 40 responses calls/day

  • Pro ($100/mo): 500GB storage, 200 responses calls/day

TL;DR

Jockey is now in research preview. It's TwelveLabs' video AI agent layer built on top of Marengo, Pegasus, and Search that lets developers, and consumers work with a video and photo library the way they actually think about it: by person, by moment, by insight; not by filename or timestamp. Jockey reaches people through two surfaces: the API for builders, and MCP on Claude. This article walks through what's new: the full-stack flywheel, global entity understanding, and how to get started.

1 - Overview of Twelve Labs APIs

TwelveLabs has shipped Pegasus, Marengo, and Search as separate model capabilities - each best-in-class on its own. Jockey is the first product to integrate all three into a single agent surface. Pegasus handles understanding and description; Marengo handles multimodal indexing and retrieval across video and photos; Search adds multi-step, iterative retrieval on top. Because Jockey sits on top of the full stack, it isn't limited to a single clip or a single query; it reasons across your entire library, and it gets better automatically every time one of those models improves.


2 - The Stack Compounds

Every model upgrade is an automatic Jockey upgrade. TwelveLabs has shipped Pegasus, Marengo, and Search as separate model capabilities, released on their own timelines. Jockey is the first product to integrate all of them into one agent, which means every foundation-model upgrade automatically makes Jockey better. When Pegasus improves, Jockey's metadata extraction improves. When Marengo improves, Jockey's search improves. When Search lands new operators, Jockey's cognition engine gets new tools to plan over. No API migration required; the same Jockey session just gets smarter underneath you.


3 - One Agent, Reachable Everywhere

Jockey reaches every audience through a single underlying engine; you don't get a different AI depending on how you connect, you get the same one:

  • API: for builders who want to integrate video and photo intelligence directly into their own applications, with structured citations and asset grouping in every response.

  • MCP on Claude: for teams and consumers who want to connect their own photo and video library and just ask. No new interface to learn: Jockey meets people inside tools they already use.

Under the hood, Jockey combines your library's index (built on Marengo + Pegasus) with an agentic search loop (Search) that can iteratively refine a query until it finds the right person, moment, or insight; and returns a structured, citable answer, not just a wall of text.


4 - The first AI that understands who's in your library

Most video AI tools process a file. Jockey processes an entire video and photo library. The headline capability of this release is global entity understanding: recognizing the same person across thousands of videos and photos in your library, and automatically grouping everything they appear in. This isn't a filter you apply after the fact; it's a new primitive Jockey reasons over:

  • Digital Marketing Teams: “Find ads that overperformed on Instagram relative to YouTube. Have Jockey look at their content and explain what about them suits the IG feed vs YouTube."

  • Creators: "I'm launching a new series. Based on my existing content, find the topics my audience engages with most and draft a 4-episode arc with hooks for each."

Every other Jockey capability (content assembly, content organization, performance analysis, even search itself) gets dramatically better once "the same person across many videos" is a stable building block.

5 - Build with the Jockey API

Jockey isn't something you fork and modify, it's something you build on. Builders integrate through the Jockey API to add video and photo intelligence to their product without standing up their own ML infrastructure: query a knowledge store, get back structured, timestamped, citable results you can render directly in your app. In research preview, that includes the Filtered Response API for scoped, grouped results, and full parity between video and photo search, so an app built on Jockey can reason over mixed-media libraries out of the box. Early research preview builders report a 70%+ "could not have built this" rate and a median time-to-working-prototype under 4 hours.

6 - Try Jockey

Jockey is now live at twelvelabs.io/jockey with limited spots during our research preview. Sign up through Playground, where you'll find guidance to:

  1. Connect via MCP: add Jockey as a connector for Claude and upload your photo and video library.

  2. Build with the API: get an API key and start querying a knowledge store directly.

Pricing:

  • Free: 5GB storage, 15 responses calls/day

  • Plus ($20/mo): 100GB storage, 40 responses calls/day

  • Pro ($100/mo): 500GB storage, 200 responses calls/day