Skip to content
research

Meta's Secret AI Video Models Just Leaked

Two unidentified AI video models have mysteriously appeared on public leaderboards, outperforming established players. This unexpected arrival could signal a massive shift in the generative video landscape, and the prime suspect is a tech giant you already know.

Aki Tanaka
Meta's Secret AI Video Models Just Leaked

Two New Contenders Enter the Arena

Two mysterious new AI video models, code-named Polaris and Vega, have suddenly appeared on the Artificial Analysis text-to-video leaderboards, generating significant industry buzz. Their unannounced emergence, first noted by Brent Lynch, immediately ignited intense speculation about their origins and capabilities, signalling a potent, unexpected force in the rapidly evolving generative AI landscape.

Initial public demonstrations reveal robust performance for both models. Outputs are currently limited to 10-second clips at a sharp 1080p resolution, crucially including integrated audio. This high fidelity, combined with observed multi-shot consistency and prompt adherence in early previews, suggests these models are already remarkably well-developed and potentially production-ready for diverse short-form content applications.

Their abrupt arrival fundamentally disrupts the established hierarchy of AI video generation. Frontrunners like OpenAI’s Sora and Google’s Veo now face unexpected, formidable challengers that seemingly materialized overnight. The models also preemptively challenge the much-anticipated Kling 4, forcing a rapid re-evaluation of the competitive field. Though official confirmation is pending, strong industry conjecture points to Meta as the developer of at least one model, aligning with recent AI advancements from Meta’s Superintelligence Labs, including Muse Image and Muse Glimmer.

Connecting the Dots to Zuck's Lab

Leading theory points directly to Meta. Its Superintelligence Labs, under Alexandr Wang, has recently engaged in a rapid release cycle, introducing significant AI models. These include Muse Image, a powerful image generation model, and Muse Glimmer, a 30-billion-parameter open-source agentic LLM. This flurry of activity creates a clear context for new foundational model releases.

Crucially, Meta explicitly announced a video model was 'forthcoming' alongside the Muse Image release. The sudden appearance of Polaris and Vega on the Artificial Analysis leaderboards, without official attribution, aligns suspiciously with this previously stated intention. This timing suggests Meta may be quietly rolling out its promised generative video capabilities.

This move reinforces Meta's aggressive strategy in the generative media space. The company consistently builds and releases both open and closed-source foundational models, positioning itself to compete directly with other major players in advanced AI. The unconfirmed emergence of these new video models would be a logical next step in their announced roadmap.

Polaris: A Master of Consistency and Physics

Polaris quickly distinguishes itself with remarkable multi-shot consistency. One striking output features a pilot writing numbers on his hand; these digits remain stable and legible across multiple camera angle changes, a significant feat for generative video models. This capability suggests a robust internal representation of objects and their persistent state.

Further, Polaris demonstrates an impressive understanding of physics and complex interactions. A clip depicting a large banner unfurling vividly illustrates this: the wind catches the fabric, transforming it into a sail that realistically billows and nearly lifts a worker off their feet. This dynamic simulation marks a clear advancement over earlier models, which often struggled with such nuanced environmental responses.

Despite these strengths, Polaris exhibits minor imperfections, such as rendering clean text on clothing patches, which often appears garbled. However, its prompt adherence is exceptional, consistently generating specific details like a dog's breath fogging nostrils. The model even produces recognizable intellectual property, including an endearing 'Batbaby' character. For more on Meta’s broader generative AI initiatives, including earlier work on image and video models, see Introducing Muse Image and Muse Video - Meta AI.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Vega's Multilingual Magic and What's Next

Vega, the second emerging model, showcases distinct capabilities that set it apart. It demonstrates remarkable multilingual lip-syncing, a critical component for realistic character animation. In one striking example, an Italian speaker in a hall of mirrors sees all reflections precisely synchronize with their spoken words, underscoring Vega's advanced spatio-temporal and auditory coherence. This nuanced performance hints at a sophisticated understanding of human speech and its visual manifestation.

While Vega proves notably adept at rendering on-screen text and signage, maintaining impressive font consistency and spelling across frames, it is not without its imperfections. The model can introduce unprompted artifacts, such as "pop-up text" appearing with a "magic paintbrush" effect. This phenomenon suggests the model sometimes hallucinates visual elements, generating content not explicitly specified in the input prompt.

The clear divergence in strengths and weaknesses between Polaris's physics-based consistency and Vega's linguistic and textual finesse strongly indicates their origin from separate model families. This points to a multi-pronged development approach — suggesting a major lab, likely Meta's Superintelligence Labs, is pursuing distinct architectural paths for AI video models. This parallel innovation strategy significantly intensifies pressure on all competitors, forcing them to accelerate their own research and expand their capabilities to keep pace.

Frequently Asked Questions

What are Polaris and Vega?

Polaris and Vega are code names for two new, unannounced text-to-video AI models that appeared on the Artificial Analysis leaderboards, showcasing advanced capabilities in consistency, physics, and multilingual lip-syncing.

Is Meta confirmed to be behind Polaris or Vega?

There is no official confirmation, but strong circumstantial evidence, including recent activity from Meta's Superintelligence Labs and a previous promise of a video model, points to their involvement.

How do Polaris and Vega compare to other AI video models?

Initial outputs show they are highly competitive. Polaris excels at multi-shot consistency and physics simulation, while Vega demonstrates impressive text handling and multilingual lip-syncing, positioning them as strong contenders in the field.

What are the current limitations of these models?

The public demos are limited to 10 seconds at 1080p. Polaris shows minor issues with text generation, while Vega can produce unwanted pop-up text artifacts despite its overall strength in text rendering.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only