Skip to content
research

The 975B-Param AI That Hears Raw Audio

Ex-OpenAI CTO Mira Murati just released Inkling, a 975B-parameter open-weight model. Its unique ability to process raw audio and its focus on hyper-specific fine-tuning could change how enterprises build AI.

Aki Tanaka
The 975B-Param AI That Hears Raw Audio

OpenAI's Ex-CTO Drops a 975B-Param Behemoth

Former OpenAI CTO Mira Murati has launched Thinking Machines, a $12 billion startup poised to challenge the landscape of closed AI development. Her venture champions powerful open-weight models, aiming to democratize advanced AI capabilities and provide enterprises with unparalleled control and customization over their intelligent systems.

Thinking Machines' inaugural offering, Inkling, is a formidable presence with 975 billion parameters. This impressive scale is managed through a sophisticated Mixture-of-Experts (MoE) architecture, where only a lean 41 billion parameters are actively engaged for any given token, balancing power with computational efficiency. Released under a permissive Apache 2.0 license, Inkling's full weights are openly accessible on Hugging Face.

Crucially, Thinking Machines positions Inkling not as a general-purpose AI competitor, but as a strategic foundation model. While it may not outperform every broad-spectrum benchmark, its design excels as a starting point for specialized applications. Inkling's true potential is unlocked through custom fine-tuning, exclusively facilitated via the Tinker platform, empowering developers to craft highly tailored solutions for specific, high-volume tasks.

This AI Hears Sound, It Doesn't Read Transcripts

Inkling diverges sharply from conventional multimodal models by bypassing pre-processing for non-textual inputs. Rather than transcribing audio into text or describing images with a separate vision encoder, it ingests raw audio as spectrograms and images as 40x40 pixel patches directly. These are then processed alongside text tokens in a unified representational space.

This novel architecture theoretically minimizes the data loss and compression inherent in traditional methods, where information is inevitably lost during the conversion of rich sensory data into a textual intermediary. It allows Inkling to perceive the nuances of sound, such as speaking style, not just its semantic content.

Thinking Machines is transparent about Inkling's general performance. While it outperforms GLM 5.2 in instruction-following (79.8 to 73.3), it significantly underperforms on Terminal-Bench agentic tasks (63.8 to 82.7). This indicates proficiency in executing direct commands but weaker autonomous problem-solving.

Such benchmark results reinforce Inkling's design as a fine-tuning specialist, not a generalist. Its strength lies in providing a robust, raw-data-aware foundation for highly specific, high-volume tasks, particularly those involving nuanced audio or visual inputs, rather than broad, unsupervised reasoning.

Why Fine-Tuning Is Inkling's Real Superpower

Inkling’s true power emerges through fine-tuning, not raw general performance. This process teaches models how to respond—shaping tone, format, and adherence to specific output structures, rather than imparting new factual knowledge.

Consider Harvey AI, which fine-tuned a smaller model to extract legal citations. This specialized model saw its F1 score improve from 0.56 to 0.68. It also matched or outperformed GPT-4.0 in 93% of cases, all while running faster and more cost-effectively.

Thinking Machines' Tinker platform streamlines this customization, enabling rapid, task-specific adaptation. A compelling demo showcased Inkling fine-tuning itself to exhibit 'epsilonphobia'—never using the letter 'e'—by prompting the model to automatically correct its training data and execute the entire fine-tuning process. The resulting responses, despite the linguistic constraint, still made remarkable sense.

Underpinning this efficiency is Low-Rank Adaptation (LoRA). LoRA freezes the original 975 billion parameters and injects lightweight, smaller matrices to capture new, specific data. This approach drastically reduces the cost and speed of customization compared to retraining an entire model, making tailored AI accessible for niche applications.

This targeted adaptation means enterprises can craft highly specialized agents for high-volume, narrow tasks. For those interested in exploring Inkling's open-weight architecture and further details, its resources are available on thinkingmachines/Inkling - Hugging Face.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

When Should You Actually Use Inkling?

Inkling isn't designed as a generalist workhorse; its creators at Thinking Machines openly acknowledge its performance on broad benchmarks like Terminal-Bench trails frontier models. For wide-ranging, unspecialized tasks, established closed-source AIs remain the go-to. However, for a narrow, high-volume job supported by robust labeled data, Inkling emerges as a formidable contender, purpose-built for deep fine-tuning.

On the Tinker platform, Inkling carves out two unique niches. It is the sole model capable of processing raw audio directly as spectrograms, bypassing the data loss inherent in text transcription. This direct approach allows for unprecedented analysis of sound's intrinsic properties. Additionally, users can finely tune its thinking effort with a granular 0-1 slider, offering control far beyond the typical low, medium, or high settings found in other models.

Choosing Inkling hinges on your application's specificity. It excels in specialized audio tasks such as medical dictation, adeptly learning complex drug names, or performing sophisticated speech style classification to discern how something was said. When deep customization, direct audio analysis, and fine-tuning a model for a precise, repeatable function are primary goals, Inkling offers an unparalleled open-weight solution.

Frequently Asked Questions

What is the Inkling AI model?

Inkling is a 975-billion-parameter, open-weight Mixture-of-Experts (MoE) model from Thinking Machines, a startup founded by ex-OpenAI CTO Mira Murati. It's designed primarily for fine-tuning on specific tasks.

How does Inkling process audio and images differently?

Unlike models that first convert audio/images to text descriptions, Inkling processes raw data directly. It ingests audio as spectrograms and images as pixel patches, mixing them with text tokens to reduce data loss and improve fidelity.

Is Inkling better than GPT-4?

Inkling is not designed to be a stronger general model than GPT-4. Its strength lies in being a highly effective base model for fine-tuning on narrow, specific tasks, where a customized version can outperform general models in both accuracy and efficiency.

What is the Tinker platform?

Tinker is Thinking Machines' platform designed for customizing and fine-tuning AI models like Inkling. It simplifies the process of adapting a model to a highly specific task, using techniques like LoRA for cost-effective training.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only