Skip to content
AI Tool

Hy4 preview Review

Hy4 preview is a 770B-parameter AI model with a 1M-token context window, designed for software engineering, office work, game development, and scientific research.

shipped Aug 29, 2026writingpaid
Domain rating43Monthly visits1/mo
writingresearch
Hy4 preview — product screenshot

Why it matters

1Features a 770 billion total parameter Mixture-of-Experts (MoE) architecture with 49 billion active parameters.
2Offers an extended context window exceeding 1 million tokens for complex tasks.
3Achieved a 31.8% increase in end-to-end throughput through a recursive self-improvement loop.
4Pricing includes $0.834 per million input tokens and $2.501 per million output tokens.

About Hy4 preview

Business Model
Usage-Based (Pay Per Use)
Usage Pricing
$0.042 (cached input) / $0.834 (input) / $2.501 (output) per 1M tokens
Headquarters
Shenzhen, China

Pricing Plans

Hy4 preview API Pricing
$0.042 per 1M tokens (cached input) / per-request
  • $0.042 for cached input
  • $0.834 for input
  • $2.501 for output

Cost Examples

  • 1M cached input tokens: $0.042
  • 1M input tokens: $0.834
  • 1M output tokens: $2.501
GitHubOpen Source

Specs

API Available

Yes, public API

overview

What is Hy4 preview?

Hy4 preview is a large language model (LLM) tool developed by Tencent that enables developers, engineers, office workers, and researchers to perform complex productivity tasks. It is a 770B-parameter Mixture-of-Experts (MoE) model with a 1M-token context window, released on August 28, 2026. This model is designed to enhance capabilities in areas such as software engineering, scientific research, financial analysis, and game development. The architecture includes 49 billion active parameters per token, contributing to its efficiency in processing extensive data and complex workflows. Tencent has also released a lower-precision Hy4 preview-FP8 variant for deployments requiring a smaller footprint.

features

Key Features of Hy4 preview

Hy4 preview incorporates several technical features designed to support advanced AI applications across various professional domains. Its architecture and training methodologies contribute to its performance in handling complex, long-horizon tasks.

  • 770B-parameter Mixture-of-Experts (MoE) architecture with 49 billion active parameters per token.
  • 1M-token context window, enabling processing of extensive documents and codebases.
  • Recursive self-improvement loop, where the model optimized its own training methods, leading to a 31.8% increase in end-to-end throughput.
  • High-quality training data co-created with Tencent's internal domain experts in software engineering, gaming, finance, and security.
  • Enhanced capabilities for understanding, planning, debugging, and validation in software engineering tasks.
  • Improved handling of financial analysis workflows, cross-document synthesis, and data analysis.
  • Stronger capabilities in understanding, reasoning, and solving complex problems in scientific research, including AI development and molecular dynamics.
  • Ability to generate playable prototypes and iterate on game projects through natural-language interactions.

use cases

Who Should Use Hy4 preview?

Hy4 preview is specifically engineered for professionals requiring advanced AI assistance in complex, data-intensive, and reasoning-heavy tasks. Its design caters to specific industry needs, offering specialized support.

  • Software Engineers and Developers: For enhanced understanding, planning, debugging, and validation of long-horizon development tasks, including front-end development.
  • Scientific Researchers: For complex problem-solving, reasoning, and understanding across fields like AI research, molecular dynamics, condensed-matter physics, and pure mathematics.
  • Financial Analysts and Office Workers: For improved handling of financial analysis workflows, cross-document synthesis, data analysis, and the creation of documents, spreadsheets, and presentations.
  • Game Developers: For generating playable prototypes from natural-language requests and iterating on complex game projects through multi-turn interactions.

how to use

How to Use Hy4 preview

Hy4 preview is accessible via API through Tencent Cloud TokenHub and OpenRouter, allowing integration into existing workflows. A limited free evaluation period is available on Tencent's WorkBuddy and CodeBuddy platforms.

  • 1Access the Hy4 preview API through Tencent Cloud TokenHub or OpenRouter.
  • 2Integrate the API into software engineering, research, or office applications.
  • 3Utilize the 1M-token context window for processing large datasets or extensive codebases.
  • 4Leverage the model for tasks such as code generation, document summarization, or scientific problem-solving.
  • 5Monitor input and output token usage for cost management based on the usage-based pricing model.
  • 6For initial evaluation, use the free access period on Tencent's WorkBuddy and CodeBuddy platforms.

pricing

Hy4 preview Pricing & Plans

Hy4 preview operates on a usage-based pricing model, with distinct costs for input, output, and cached tokens. This structure allows users to pay only for the resources consumed during API interactions.

  • Input Tokens: $0.834 per million tokens.
  • Output Tokens: $2.501 per million tokens.
  • Cache Read Tokens: $0.042 per million tokens.
  • Free Evaluation: Available for a limited two-week period on Tencent's WorkBuddy and CodeBuddy platforms from its launch date.

Pros

  • +Features a substantial 1M-token context window, suitable for extensive documents and codebases.
  • +Utilizes a 770B-parameter Mixture-of-Experts (MoE) architecture for efficient processing.
  • +Demonstrated a 31.8% increase in end-to-end throughput due to a recursive self-improvement loop.
  • +Offers competitive pricing, with output tokens 82% cheaper than Kimi K3 and 36% cheaper than GLM-5.3.
  • +Strong performance in productivity-focused engineering tasks, as indicated by internal evaluations (2.99/4.00 score).

Cons

  • Known limitations include taking longer than necessary on complex questions, potentially increasing latency.
  • Exhibits a tendency to over-verify its own answers, which can also contribute to increased latency.
  • Reportedly lags behind Zhipu's GLM-5.3 in public benchmarks for code and cybersecurity tests (e.g., DeepSWE, CyberGym).
  • As a recent release (August 28, 2026), extensive third-party technical reviews are still emerging.

Similar Tools

Hy4 preview vs Competitors

Hy4 preview competes in the open-source and hosted LLM market, particularly against models from other major Chinese developers and global AI providers. Its competitive edge often lies in its specific technical specifications and pricing strategy.

1
Llama 3.1

An open-weight model family from Meta, offering strong performance across various tasks, with Llama 3.1 specifically expanding the context window to 128K tokens and a 405B parameter variant.

While Llama 3.1's 128K token context window is smaller than Hy4 preview's 1M, it offers the flexibility and cost-effectiveness of an open-source model that can be run locally or on various platforms, making it a powerful free alternative.

2

Offers a range of high-performance models, including specialized ones like Codestral for software engineering, with competitive pricing and a focus on efficiency and reasoning.

Mistral AI provides a strong, cost-effective API alternative with models specifically tuned for coding and complex tasks, though its largest context windows might not consistently reach Hy4 preview's 1M tokens across all models.

3

Known for strong reasoning, safety, and a large context window, with several models offering a 1M token context window at standard pricing.

Claude directly competes with Hy4 preview on context window size (1M tokens) and advanced reasoning capabilities, offering a robust alternative for complex tasks, though its per-token pricing for top-tier models can be higher.

4

Specializes in search-augmented generation, providing real-time web retrieval and citations, making it highly effective for research and information-intensive tasks.

Perplexity AI excels at grounding responses with real-time web data, a feature not explicitly highlighted for Hy4 preview, making it superior for tasks requiring up-to-date information, but its context window might be smaller than Hy4 preview's 1M tokens for some models.

5

Enables users to download and run open-source large language models (like Llama, Mistral, DeepSeek, Qwen) locally on their own machines, providing complete control and privacy.

Ollama offers a truly free and private alternative by allowing local execution of various open-source models, but it requires the user to manage hardware and model deployment, and the performance will be limited by local machine specifications, unlike a hosted API like Hy4 preview.

More on Stork

Related AI Tools

Other tools in this category, matched by shared tags