Skip to content
ai tools

Your Entire AI Stack in One Install

Developers waste countless hours configuring separate runtimes for local speech, text, and embedding models. A new SDK from a surprising player in the crypto space aims to make that entire complex setup obsolete with a single command.

Theo Brandt
Your Entire AI Stack in One Install

From Runtime Hell to One Command

Runtime hell is a real bottleneck for local AI development. Juggling separate, complex runtimes—Ollama for LLMs, Whisper for speech-to-text, plus dedicated services for embeddings and text-to-speech—creates unnecessary configuration overhead. This fragmented approach wastes developer cycles and introduces fragility into any local AI system.

QVAC by Tether finally unifies this mess. A single npm install @qvac/sdk provisions a complete, multi-modal AI system out of the box, literally. It’s the ultimate local AI SDK, offering a full suite of capabilities from LLMs and RAG to image and video generation, transcription, and text-to-speech within one platform, ready to take on any task.

Cole Medin's video demo showcases the immediate impact. He effortlessly builds a voice-driven RAG pipeline using just a few lines of code. This powerful example leverages:

  • Whisper for speech-to-text
  • Gemma for embeddings
  • Qwen3 as the LLM
  • Supertonic for text-to-speech

The advantage is clear: import these popular models directly from the SDK. No individual runtime setup or configuration is needed for each component. They simply run, entirely locally, without rate limits, cloud dependencies, or the usual integration headaches.

More Than a Wrapper: QVAC's Core Engine

QVAC isn't just a basic wrapper; it's a complete local AI system in a single install. This SDK ships with over a dozen core capabilities, giving you everything from foundational models to advanced workflows. Expect robust LLMs, complete with on-device LoRA fine-tuning, for personalized model adaptations.

It handles Retrieval-Augmented Generation (RAG) natively, alongside comprehensive image and video generation. Transcription and text-to-speech are built-in, complemented by OCR, text embeddings, and multimodal LLM inference. This powerful suite eliminates the runtime hell of juggling Whisper, Ollama, and separate Stable Diffusion installs.

Under the hood, QVAC prioritizes raw performance, wiring together highly optimized native C++ engines. Expect its custom fork of llama.cpp for blazing-fast LLM inference, whisper.cpp for efficient transcription, and direct integration with Stable Diffusion for image/video generation. This pushes local hardware to its limits, avoiding abstraction layer bottlenecks.

Crucially, this isn't platform-locked. QVAC offers true cross-platform compatibility. The same simple API works seamlessly across Node.js, mobile applications built with Expo (for both iOS and Android), and all major desktop operating systems including Linux, macOS, and Windows. One SDK, one API, zero compromises.

Declaring Independence from the Cloud

Local-first execution with QVAC ensures total data privacy. Your sensitive information, proprietary code, or personal data never leaves your machine, eliminating third-party data exposure risks inherent in cloud-based solutions. This also guarantees full offline functionality, making your AI applications resilient to internet disruptions and immune from cloud provider outages or unannounced service deprecations.

Ditch the unpredictable costs and arbitrary restrictions of cloud APIs. QVAC operates entirely on your hardware, eliminating per-request fees, bandwidth charges, and monthly subscriptions. It's free, open-source under the Apache 2.0 license, ensuring you own your models and completely avoid vendor lock-in. You face no rate limits, ever; run your models as hard as your local hardware allows.

This independence aligns directly with Tether's vision for decentralized AI. QVAC transforms AI models from rented utilities into portable capital assets you fully own. This empowers developers, securing workflows and shifting computational power back to the edge. For more on this paradigm shift, explore QVAC - Decentralized, Local AI in a Single API.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Building the Peer-to-Peer AI Network

Beyond single-machine local, QVAC pushes into distributed AI. Its peer-to-peer (P2P) communication layer enables decentralized model sharing and delegated inference. Imagine offloading a heavy VisionPsy-Nano VLM task to a peer with beefier hardware, or seamlessly sharing a custom LoRA fine-tune across devices without a central server. This is workflow optimization at its core.

Ecosystem integration is non-negotiable. QVAC provides a built-in OpenAI-compatible server, allowing existing tools to plug-and-play without modification. While currently a JavaScript powerhouse via npm install, the roadmap includes a crucial Python SDK, broadening its appeal to the wider ML community and complex data pipelines.

QVAC also relentlessly tracks the bleeding edge. Recent updates include local support for NVIDIA's GR00T robotics model, unlocking on-device intelligence for advanced robotics applications. The high-performance VisionPsy-Nano VLM further demonstrates a commitment to multimodal capabilities, delivering cutting-edge vision processing directly on your hardware. This isn't just an SDK; it's a platform for the next generation of edge AI.

Frequently Asked Questions

What is QVAC by Tether?

QVAC is a free, open-source Software Development Kit (SDK) by Tether that allows developers to run a wide range of AI models (LLMs, speech, image generation) locally on any device with a single installation.

How is QVAC different from tools like Ollama?

While Ollama primarily focuses on running large language models, QVAC is a comprehensive system that integrates runtimes for LLMs, speech-to-text (like Whisper), text-to-speech, embeddings, and more into a single, cross-platform API.

Is QVAC free to use?

Yes, QVAC is completely free and open-source, licensed under the Apache License 2.0. There are no API fees, subscriptions, or usage costs since all models run on your own hardware.

What programming languages does QVAC support?

Currently, QVAC primarily supports JavaScript and TypeScript via its NPM package. A Python client is planned for future development, along with SDKs for other languages.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only