Hy3
Tencent's Open-Source Hybrid Reasoning Model — Apache 2.0
Benchmarks & Evaluation Results
vs. Models 2-5x Larger
Matches or exceeds on complex reasoning, instruction-following, in-context learning
Code Generation
Significant improvement over Hy2 across programming benchmarks
Agent Capabilities
Enhanced tool-use and multi-step task execution vs Hy2
Hybrid Reasoning
Dynamically switches between fast intuitive and slow deliberate reasoning
Full Analysis of Hy3
Hy3 represents Tencent's commitment to open-source AI development. After falling behind competitors in the 2024-2025 AI race, Tencent made the strategic decision to rebuild its AI infrastructure from scratch in January 2026. The Hy3 generation is the result: a completely new architecture that replaces the previous Hy2 foundation. The preview launched in April demonstrated the potential, and the July 6 official release delivers on that promise with improved stability, cost efficiency, and perform
ance. The hybrid fast-and-slow-thinking architecture is Hy3's most innovative feature. For simple queries — translation, summarization, basic Q&A — the model uses a fast inference path that delivers responses with minimal latency. For complex tasks — multi-step reasoning, code generation, mathematical problem-solving — it engages a slower, more deliberate reasoning process that produces higher-quality outputs. This dynamic adaptation means Hy3 uses compute efficiently, only spending extr
a resources when the task demands it. The Apache 2.0 license makes it commercially friendly for businesses wanting to build on open-source AI without legal complications. Tencent has integrated Hy3 across WeChat for conversational AI, Tencent Cloud for enterprise customers, and Tencent Studio for developer tools.
Strengths & Considerations
Strengths
- + Apache 2.0 license — most permissive for commercial use with no restrictions
- + Hybrid reasoning architecture adapts compute to task complexity
- + Matches models 2-5x larger on key benchmarks
- + Massive deployment footprint via WeChat (1.3B+ users) and Tencent Cloud
- + Available on multiple global inference platforms (OpenRouter, etc.)
- + Comprehensive evaluation and product feedback incorporated from preview
- + Open source with full weights available for download
Considerations
- − 256K context is smaller than 1M-token competitors like K3 and Fable 5
- − Primarily optimized for Chinese-language and Tencent ecosystem use cases
- − Smaller global community compared to Llama, Qwen, or DeepSeek
- − Limited independent benchmark data outside Tencent's evaluations
- − Western developer tooling and integration still developing
Best Use Cases for Hy3
Where this model excels and the types of workloads it is best suited for.
Use Case 1
Enterprise AI deployments needing commercial-friendly Apache 2.0 license
Use Case 2
WeChat ecosystem integration and Chinese-language applications
Use Case 3
Cost-sensitive production workloads at 21B active parameters
Use Case 4
Applications benefiting from hybrid fast/slow reasoning adaptation
Use Case 5
Self-hosted deployments for data-sensitive enterprise workloads
Use Case 6
Agent pipelines integrated with Tencent Cloud services
How to Use Hy3
Quick Start Guide
Download Hy3 from Hugging Face (huggingface.co/Tencent) or ModelScope. For hosted inference, access via OpenRouter at openrouter.ai — search for 'Tencent/Hy3'. For Tencent Cloud integration, visit cloud.tencent.com. The model runs efficiently at 21B active parameters, requiring approximately 40-60GB VRAM for self-hosting in FP16. For WeChat integration, contact Tencent's enterprise team.
Other Models from July 2026
Compare Hy3 with other recent releases.