Skip to content
ai agents

Alibaba's AI Agent Codes for 10 Days

A new open-weight AI just completed a 10-day coding marathon, building a complex app from nothing. This isn't just another chatbot; it's a direct challenge to the closed-source giants and a glimpse into a future of autonomous developers.

Sol Aguirre
Alibaba's AI Agent Codes for 10 Days

An AI Just Coded for 10 Days Straight

Alibaba just pushed the boundaries of autonomous AI agents. Their new flagship, Qwen 3.8 Max, a 2.4-trillion-parameter open-weight model, recently demonstrated an unprecedented feat: it coded for over 10 days straight. Handed an empty folder, the model autonomously built "O my CLI," a fully functional code agent now live on GitHub, showcasing its capacity for sustained, independent development.

This groundbreaking demo highlights Qwen 3.8 Max’s core strength: agentic coding. Built to stay on task for days, this capability allows the model to handle long-horizon projects requiring complex planning, iterative execution, and continuous self-correction without constant human intervention. It excels in navigating multi-day programming challenges, leveraging its 1 million token context window to maintain project coherence.

Further enhancing its autonomy, Qwen 3.8 Max boasts native multimodal capabilities. Its integrated vision system acts as a continuous feedback loop, enabling the model to "see" its progress, analyze outputs, and refine its plan dynamically. This visual insight allows for sophisticated self-correction and adaptation, crucial for navigating the iterative nature of complex software development, from initial concept to deployment.

Inside the 2.4 Trillion Parameter Beast

Alibaba’s Qwen 3.8 Max is a monumental achievement in model architecture: a 2.4 trillion parameter sparse Mixture-of-Experts (MoE) system. This design cleverly activates only about 95 billion parameters per token during inference, striking a critical balance between raw computational power and operational efficiency. The model also boasts a 1 million token context window and multimodal capabilities, processing images, video, and text for comprehensive understanding.

This release represents a profound strategic shift for Alibaba, marking the first time they have released the weights for their most powerful, "Max" tier model. This aggressive 'open-weight' move signals a significant challenge to the proprietary AI giants, fundamentally reshaping the competitive landscape for cutting-edge agentic systems. It suggests a future where top-tier AI capabilities become more broadly available.

Despite its colossal scale, Alibaba has ensured broad accessibility. Developers can access Qwen 3.8 Max via API, with competitive pricing at $2 per million input tokens and $6 per million output tokens. Furthermore, for local experimentation and development, Alibaba released a smaller, more manageable 27 billion parameter variant, designed to run efficiently on personal machines. This dual approach democratizes access to advanced AI agentic capabilities.

The Benchmark Gauntlet vs. GPT-5 & Claude

Qwen 3.8 Max doesn't just code autonomously; it benchmarks competitively against leading closed-source models. For agentic tasks, it shines, scoring 86.1 on OSWorld-Verified, surpassing Claude Fable 5 (approx. 85.0) and GPT-5.6 Sol (83.2). It also leads PaperBench for research analysis at 93.0, ahead of GPT-5.6 Sol (90.5) and Claude Fable 5 (88.8).

On Terminal-Bench 2.1, Qwen 3.8 Max scores 86.6, outperforming both Claude Opus 4.8 and Claude Fable 5 (both at 84.6), though it trails GPT-5.6 Sol (88.8). However, for complex engineering on SWE-bench Pro, it scores 67.7, trailing Claude Fable 5 (80.0) and Claude Opus 4.8 (69.2).

This 2.4 trillion parameter system presents a powerful, balanced profile, directly challenging the long-held dominance of closed-source models from Anthropic and OpenAI. It particularly excels in agentic use cases and research reproduction, positioning itself as a robust all-rounder.

While the 10-day demo building 'O my CLI' proved groundbreaking, it operated within a controlled environment, a crucial nuance for real-world deployment. Consider its API pricing: $2 per million input tokens and $6 per million output tokens. This open-weight model can run for days, offering a compelling, cost-aware alternative in the evolving AI landscape.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Is This the End of Closed-Source Dominance?

Alibaba's release of Qwen 3.8 Max fundamentally shifts the global AI power dynamic. A top-tier 2.4 trillion parameter MoE model emerging from China, now with open weights, directly challenges the established closed-source dominance of Western tech giants. This intensifies the worldwide race to build and distribute frontier AI, creating a more diversified competitive landscape.

Crucially, "open-weight" does not equate to fully "open-source." Alibaba provides access to the model's parameters, enabling broad experimentation and development. However, potential future licensing restrictions could still govern commercial deployment, defining the true boundaries of its accessibility for enterprise applications.

For the wider industry, this democratizes access to highly capable agentic AI. Developers gain a powerful new tool, designed for multi-day autonomous coding, which could dramatically accelerate innovation cycles. Yet, the widespread adoption of such potent models from diverse geopolitical origins also raises new questions about commercial usage, data governance, and ethical deployment standards globally. The frontier of AI development now feels decidedly more open, and more complex.

Frequently Asked Questions

What is Qwen 3.8 Max?

Qwen 3.8 Max is a 2.4 trillion parameter, open-weight AI model from Alibaba, specifically designed for long-duration, autonomous tasks like agentic coding.

Can I run Qwen 3.8 Max on my laptop?

No, the 2.4T parameter model requires significant computational resources. However, Alibaba has also released a 27 billion parameter variant, Qwen 3.8-27B, designed to run on personal machines.

How does Qwen 3.8 Max compare to models like GPT or Claude?

It's highly competitive. Qwen 3.8 Max outperforms models like Claude Fable 5 and GPT-5.6 Sol on agentic computer use benchmarks (OSWorld-Verified) but trails on some complex software engineering tests (SWE-bench Pro).

What did Qwen 3.8 Max build in its 10-day demo?

In its most impressive demonstration, the model was given an empty folder and, over 10 days, it autonomously built 'O my CLI,' a complete, working code agent now available on GitHub.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only