Skip to content

Pegasus 1.5 by TwelveLabs Review

TwelveLabs delivers enterprise video AI powered by multimodal intelligence, enabling search, analysis, and understanding of video across vision, audio, and language.

shipped Apr 21, 2026videofreemium
Domain rating69Monthly visits3.9K/mo
videocoderesearch
Pegasus 1.5 by TwelveLabs — product screenshot

Why it matters

1Processes videos up to two hours long in a single request.
2Achieved SOC2 compliance, ensuring data security and privacy.
3Outperforms Gemini 3.1 Pro by 13.1% in aggregate segmentation quality.
4Supports batch analysis for up to 1,000 video requests in a single call.

Stork’s verdict on Pegasus 1.5 by TwelveLabs

Pegasus 1.5 offers precise Time Based Metadata Extraction with custom JSON schemas, but its API-driven nature demands developer resources.

Pegasus 1.5 by TwelveLabs reviewed by Stork AI · stork.ai/en/pegasus-1-5-by-twelvelabs

Stork Quadrant

Becomes the API· 27/100

Replaceable as a UI, but kept alive as the API the agents call.

TwelveLabs built a capable multimodal video understanding API before the frontier labs caught up. That window is closing. GPT-4o, Gemini 1.5 Pro, and Claude already handle video natively, and they're getting faster and cheaper. There's no proprietary data, no network, no regulatory gate — just a specialized model that bigger players will commoditize.

Claude Sonnet 4.6, scored 2026-05-30

Defensibility · 0/100

  • Physical-world coupling
  • Regulatory moat
  • Network liquidity
  • Proprietary refreshing data
  • High-trust catastrophic workflows
  • Multi-party coordination
  • Brand / community / taste

An LLM alone could replace

  • Summarize what happens in a video by describing its content
  • Transcribe audio and extract key topics or themes from spoken content
  • Answer questions about a video's subject matter given a transcript or description
  • Generate metadata tags or chapter markers for video content

Agent-Readiness · 60/100

  • Verified MCPStork MCP listing: io-twelvelabs-twelvelabs-mcp-server (untested)
  • Listed on agent surfacesStork:io-twelvelabs-twelvelabs-mcp-server
  • Usage-based pricingpricing page heuristic match: https://www.twelvelabs.io/pricing
  • Headless agent auth
  • Public OpenAPIhttps://docs.twelvelabs.io/v1.3/docs/resources/platform-overview
  • Active changeloghttps://www.twelvelabs.io/blog/introducing-pegasus-1-5 (2026-04-19)
  • llms.txthttps://www.twelvelabs.io/llms.txt

Score history · +5 pts over 3 re-scores

How to defend

Go vertical and own the liability: pick one industry where wrong video analysis has real consequences — insurance claims, legal evidence, broadcast compliance — and become the vendor that signs the contract and bears the risk. That's the only move that creates a moat here.

  • Ship an MCP server and list it on Stork — biggest single point gain (+25).
  • Expose API-key auth with a self-serve sandbox tier; remove sales-call gates (+15).

About Pegasus 1.5 by TwelveLabs

Headquarters
San Francisco, USA
Founded
2020
Team Size
51-100
Funding
Series A

Specs

API Available

Yes, public API

overview

What is Pegasus 1.5 by TwelveLabs?

Pegasus 1.5 by TwelveLabs is a video language model developed by Twelve Labs that enables developers and enterprises to analyze, search, and understand video content using multimodal AI. It transforms raw video into structured, queryable data, specializing in video-to-text generation and segmentation. Released on April 20, 2026, Pegasus 1.5 represents a significant advancement in video reasoning models, offering capabilities such as Time Based Metadata Extraction (TBM) which allows users to define custom JSON schemas for timestamped, structured metadata from videos up to two hours long. The model processes multiple modalities within videos—visual, audio, and textual information—to produce contextually relevant text and structured data. It supports direct video analysis from URLs, assets, or base64 strings, and features multimodal prompting, allowing the inclusion of reference images for enhanced context. Pegasus 1.5 utilizes a shared context window of 261,120 tokens for input and output, supporting responses up to 98,304 tokens. As of May 28, 2026, it supports synchronous analysis, and batch analysis for up to 1,000 requests was introduced on June 18, 2026.

features

Key Features of Pegasus 1.5 by TwelveLabs

Pegasus 1.5 by TwelveLabs offers a comprehensive suite of features designed for deep video understanding and analysis, leveraging multimodal AI to process visual, audio, and language data within video assets.

  • Generates concise summaries and detailed textual descriptions of video content.
  • Identifies and extracts timestamped segments, including editorial narratives, sports plays, and speaker changes.
  • Enables Time Based Metadata Extraction (TBM) using custom JSON schemas for structured data output.
  • Supports multimodal prompting, allowing users to include reference images for specific object or scene identification.
  • Processes videos directly from URLs, assets, or base64 strings without pre-indexing.
  • Utilizes a shared context window of 261,120 tokens for extended input and output, supporting responses up to 98,304 tokens.
  • Offers synchronous analysis for immediate results and batch analysis for processing up to 1,000 videos concurrently.
  • Integrates the Marengo Multimodal Embedding Model for spatiotemporal embeddings, making every video moment findable.
  • Provides content safety checks to identify sensitive or rule-violating material.
  • Facilitates content analysis for marketing and advertising effectiveness, identifying areas for improvement.

use cases

Who Should Use Pegasus 1.5 by TwelveLabs?

Pegasus 1.5 by TwelveLabs is designed for organizations and professionals who manage and derive insights from video content at scale, transforming raw footage into a strategic, queryable asset.

  • Developers and Enterprises: For building custom video intelligence applications, integrating advanced video understanding into existing workflows, and turning large video archives into strategic assets.
  • Media Companies and Creatives: For content summarization, generating timestamped clips for highlights, and efficiently searching vast video libraries for specific actions, scenes, or dialogue.
  • Sports Organizations: For detailed play analysis, creating highlight reels, coaching reviews, and extracting specific events with precise temporal boundaries.
  • AdTech and Marketing: For contextual targeting, ensuring brand-safe ad placements, and evaluating video content for persuasive effectiveness and engagement.
  • Public Sector and Security Operators: For evidence management, anomaly detection in surveillance footage, and generating after-incident reports with timestamped events.

how to use

How to Use Pegasus 1.5 by TwelveLabs

To utilize Pegasus 1.5 by TwelveLabs, users typically interact with its API to submit video content for analysis and receive structured data or textual outputs.

  • 1Access the TwelveLabs API documentation (https://docs.twelvelabs.io/v1.3/docs/resources/platform-overview) to understand available endpoints and parameters.
  • 2Ingest video content by providing a URL, uploading an asset, or sending a base64 encoded string to the API.
  • 3Formulate natural language queries or prompts, optionally including reference images for multimodal context.
  • 4Specify a custom JSON schema for Time Based Metadata Extraction (TBM) to receive structured, timestamped data.
  • 5Choose between synchronous analysis for immediate results or batch analysis for processing multiple videos efficiently.
  • 6Receive and integrate the generated summaries, detailed descriptions, segmented data, or extracted insights into downstream applications.

pricing

Pegasus 1.5 by TwelveLabs Pricing & Plans

TwelveLabs offers a flexible pricing model for Pegasus 1.5, including a freemium tier and scalable plans designed for developers and enterprises, with usage-based components for API calls and token consumption.

  • Free: Provides basic limits at no cost, suitable for initial exploration and small-scale projects.
  • Developer: Offers three tiers with increasing usage limits, priced based on monthly spending, catering to growing development needs.
  • Enterprise: Custom plans with tailored limits and dedicated support, designed for large-scale deployments and specific organizational requirements.
  • Input text (Pegasus model): Priced at $0.001 per 1,000 tokens.
  • Output text (Pegasus Analyze API): Priced at $0.007 per 1,000 tokens.
  • API rate limits vary by plan and usage type, measured across dimensions such as Duration per day (DPD), Duration per hour (DPH), Requests per day (RPD), Requests per minute (RPM), Tokens per day (TPD), and Tokens per minute (TPM).

Pros

  • +Achieves high accuracy in video segmentation and time-based metadata extraction.
  • +Provides reliable, schema-compliant JSON outputs for structured data.
  • +Supports processing of long-form videos up to two hours in a single request.
  • +Offers flexible analysis options including synchronous and efficient batch processing.
  • +Multimodal prompting with image references enhances the precision of video queries.
  • +Demonstrates strong competitive performance against leading general-purpose AI models in video understanding tasks.

Cons

  • Not aligned with HIPAA compliance standards, limiting use in certain healthcare contexts.
  • Detailed pricing for all usage dimensions (e.g., video duration processing) requires deeper inquiry beyond publicly listed token costs.
  • Primarily API-driven, which may necessitate developer resources for full implementation and integration.
  • The platform's focus on enterprise and developer use cases might mean a less intuitive out-of-the-box user interface for non-technical users.

Similar Tools

Pegasus 1.5 by TwelveLabs vs Competitors

Pegasus 1.5 by TwelveLabs is positioned as a purpose-built solution for production video workflows, distinguishing itself through specialized multimodal video understanding, robust segmentation, and reliable structured data output.

1

Mixpeek is designed for teams building video intelligence applications, offering a comprehensive platform that handles ingestion, extraction, indexing, and retrieval of video data.

While both offer video AI, Mixpeek provides an end-to-end solution for building custom video intelligence applications with deep content analysis. TwelveLabs focuses on quick cloud-based video understanding with natural language queries and generative text outputs from its foundation models like Pegasus.

2
Google Video Intelligence API

This API provides robust video annotation and content categorization, deeply integrated with the Google Cloud ecosystem for large-scale analytics and developer-focused applications.

Google Video Intelligence API is a developer-centric service for annotating and categorizing video content within the Google Cloud environment. TwelveLabs offers a more integrated platform for natural language video search and generative text outputs, built on its own multimodal foundation models.

3
Clarifai Video

Clarifai Video is a visual AI platform that provides dedicated video analysis models and a visual workflow builder, enabling non-ML engineers to train and chain custom concept detection models.

Clarifai Video emphasizes customizable concept detection and a user-friendly workflow builder for tailored AI solutions. TwelveLabs, with Pegasus 1.5, focuses on advanced multimodal video understanding, summarization, and the generation of structured, time-based metadata through natural language queries.

4
Memories.ai

Memories.ai is an AI video intelligence platform focused on large-scale search, summarization, and multimodal understanding, with an emphasis on contextual memory and timeline insights for streamlined workflows.

Memories.ai provides a platform for broad video intelligence tasks including search and summarization, leveraging contextual memory. TwelveLabs' Pegasus 1.5 specifically advances video understanding by generating structured, time-based metadata across entire videos, moving beyond clip-based answers to enable schema-first interaction for precise temporal boundaries.

More on Stork

Related AI Tools

Other tools in this category, matched by shared tags