Moonshot AI Financial Momentum
Moonshot AI reached a $35 billion valuation following a $3.5 billion funding round on July 29, 2026. This investment tripled the company’s value in six months. China’s National Artificial Intelligence Industry Investment Fund anchored the round. The National Artificial Intelligence Industry Investment Fund, which established itself in January 2025 with $8.8 billion in capital, is backed by the government’s semiconductor investment vehicle and negotiates to lead DeepSeek’s first-ever external funding round. Moonshot also secured $2 billion in May 2026 at an $18 billion valuation from Meituan. Alibaba and Tencent participated in earlier rounds. Moonshot closed a $700 million round in February 2026 at a $10 billion valuation with Alibaba as a co-investor. The company also considers a Hong Kong IPO for late 2026 and discusses new funds at a $50 billion pre-money valuation. The company’s rapid growth follows the release of Kimi K3. This 2.8 trillion parameter model drove a sixfold surge in daily sales. The Bureau of Industry and Security currently investigates the company over allegations regarding NVIDIA GB300 chips and the distillation of Anthropic’s Fable model. The National Development and Reform Commission instructed Moonshot, ByteDance, and StepFun to reject US-origin capital without government approval in April 2026. This policy follows Meta’s deal with Manus AI and aligns with US bans on backing Chinese AI firms.
Kimi K3 Architecture and Benchmarks
Kimi K3 is a 2.8 trillion parameter model. It activates 16 of 896 experts per token. This design keeps inference costs low because only 50 billion parameters participate in each forward pass. The model supports a 1 million token context window. Kimi K1.5 used reinforcement learning to improve reasoning in math and coding. K3 also performs well on AIME 2025, MATH-500, and MMLU benchmarks. I find the coding performance impressive. K3 scored 1,679 Elo on the Frontend Code Arena and beat Claude Fable 5. It also leads in Program Bench and SWE Marathon. K3 achieves 93.5% on GPQA Diamond. The architecture includes Kimi Delta Attention and Attention Residuals. These components improve how information flows across long sequences. K3 provides a 2.5x improvement in scaling efficiency compared to Kimi K2. K3 uses MXFP4 weights and MXFP8 activations for hardware compatibility.
| Metric | Kimi K3 Specification |
|---|---|
| Total Parameters | 2.8 Trillion |
| Active Parameters | 50 Billion |
| Context Window | 1 Million Tokens |
| Input Price (Cache-hit) | $0.30 per 1M tokens |
| Input Price (Cache-miss) | $3.00 per 1M tokens |
| Output Price | $15.00 per 1M tokens |
If you manage API budgets, you know that prefix caching changes the math. K3 shows a cache hit rate above 90% for coding workloads. This makes the $0.30 input price a reality for many users. K3 successfully built a Triton-style GPU compiler and designed a chip in 48 hours using open-source EDA tools. It also produced 3,000 lines of Python to reproduce astrophysics research in two hours. K3 provides integration with VS Code, Cursor, and Zed via the Kimi Code CLI.
Comparing K3 to Proprietary Models
Enterprise buyers choose between Moonshot’s K3 and models like Claude Fable 5 or GPT-5.6 Sol. Claude Fable 5 wins on user experience and ecosystem maturity. K3’s hallucination rate climbed to 51 percent. I would skip K3 for standard SaaS tasks where you need high reliability or a polished interface. K3 excels in long-horizon coding and research. The model handles large codebases because the 1 million token window removes the need for RAG in many cases. K3 remains behind Claude Fable 5 on the GDPval-AA agentic evaluation. The model shows sensitivity to thinking history. Quality becomes unstable if an agent harness does not pass all thinking content back. K3 also exhibits excess proactivity. This leads to unexpected decisions in ambiguous scenarios. K3 is an open-weight model. You can fine-tune it on proprietary data for specific domains. K3 is a direct competitor to DeepSeek-V3 and Meta’s Llama 4. For those requiring deep file analysis, Claude Fable 5 and GPT-5.6 still lead. Does the high cost of specialized GPU infrastructure for such a massive model eventually negate the savings from its efficient MoE architecture?




