Was this newsletter forwarded to you? Sign up to get it in your inbox.
1. The AI Race Before Kimi K3
Prior to the introduction of Kimi K3, the landscape of frontier artificial intelligence was defined by a stark bifurcation between closed-source proprietary APIs and open-weight architectures. Closed-source models such as OpenAI’s GPT-5.5 (and preview variants like GPT-5.6 Sol), Anthropic’s Claude Fable 5 and Mythos 5, and Google’s Gemini 1.5/2.0 Pro maintained a consistent performance lead in multi-file software engineering, mathematical formalization, and complex multi-step reasoning.
While open-weight models like Meta’s Llama 3.3 405B, Alibaba’s Qwen 2.5 series, and DeepSeek’s V3/R1 architectures drastically lowered inference costs and enabled self-hosted enterprise deployments, a gap persisted in long-horizon task execution. Open-weight models frequently struggled with multi-turn tool loops, maintaining long-context coherence beyond 128,000 tokens, and surviving high-step software engineering benchmarks without state degradation.
2. What is Kimi K3?
Kimi K3 is Moonshot AI’s flagship third-generation foundation model. Built on a sparse Mixture-of-Experts transformer framework, K3 integrates 2.8 trillion total parameters with an active parameter footprint of 380 billion per token pass.
| Specification | Details |
|---|---|
| Core Parameter Count | 2.8 Trillion Total / 380 Billion Active |
| Architecture | Sparse MoE + Delta Attention + Latent Routing |
| Context Window Length | 1,048,576 Tokens (1M Native) |
| Expert Topology | 512 Total Experts / 128 Active per Token |
| Training Corpus | 18.2 Trillion Multi-Modal Tokens |
3. Company Behind Kimi: Moonshot AI
Founded in March 2023 in Beijing by Yang Zhilin and a team of AI researchers from Tsinghua University, Google Brain, and Meta AI, Moonshot AI established itself as a pioneer in long-context language modeling. Yang Zhilin’s prior research contributions—specifically Transformer-XL and XLNet—focused heavily on scaling sequence lengths and capturing long-range dependencies.
4. Architectural Breakdown
The architecture of Kimi K3 resolves two fundamental bottlenecks: computational throughput during inference and memory bandwidth saturation during 1M long-context processing.
5. Core Research Innovations & Kimi Delta Attention
Kimi Delta Attention compresses historical key-value states into dynamic state-space deltas, cutting KV-cache memory usage by 68% at 1 million tokens while preserving single-needle retrieval accuracy at 99.7%.
7. Benchmark Evaluation & Internal Work Benches
In complex enterprise knowledge tasks—such as financial modeling, presentation slide generation, and research synthesis—Kimi K3 achieves top scores on internal and public benchmark suites.

Figure 1: Internal Knowledge Work Bench results comparing Kimi K3 against GPT 5.5 and Claude Opus 4.8 across Online Exp Bench, DECK-Bench, and Finance-Bench.
8. Software Engineering & Long-Horizon Coding
Kimi K3 scores 42.0% on SWE Marathon, 88.3% on Terminal Bench 2.1, and 77.8% on Program Bench, demonstrating long-horizon execution endurance across multi-file repositories.
9. General & Visual Agent Benchmarks
Across agent benchmarks evaluating multi-step tool invocation, browser control, spreadsheet parsing, and visual diagram inspection, K3 demonstrates frontier-class performance.

Figure 2: General Agents and Visual Agents benchmark scores (GDPval-AA v2 Elo, JobBench, AA-Briefcase Elo, SpreadsheetBench 2, Automation Bench, BrowseComp, CharXiv, Zerobench).
10. Roofline & Hardware Efficiency
The chart below details MiniTriton CUDA-core roofline performance on NVIDIA hardware, showing achieved compute performance (GFLOP/s) vs. arithmetic intensity (FLOP/byte).

Figure 3: MiniTriton CUDA-core roofline analysis on NVIDIA L20 (sm_89), fp32.
11. The Open-Weight Strategy
By releasing K3's weights publicly, Moonshot AI enables enterprise data sovereignty, allowing self-hosted deployments without API data privacy concerns.
Expert Analysis
Kimi K3 shifts the balance of open AI architecture. By combining fine-grained Top-128 MoE routing with Delta Attention, Moonshot AI proves that open-weight foundation models can match closed proprietary APIs on complex agentic execution.
Key Takeaways
- ► 2.8T Parameter open-weight MoE with 380B active parameters per forward pass.
- ► Native 1-Million token context window powered by Kimi Delta Attention.
- ► Leading benchmark performance across General Agents, Visual Agents, and Internal Knowledge Benches.
Share this research article
Frequently Asked Questions (FAQ)
What is Kimi K3?
Kimi K3 is a 2.8-trillion parameter open-weight Mixture-of-Experts (MoE) foundation model created by Moonshot AI, designed for advanced coding, autonomous agent execution, and 1M context processing.
Is Kimi K3 open-source or open-weight?
Kimi K3 is an open-weight model, meaning its neural network weights and architecture specifications are publicly accessible for self-hosting and research.
How many active parameters does Kimi K3 use per token?
Out of 2.8 trillion total parameters across 512 experts, K3 routes tokens to 128 active experts per pass, utilizing approximately 380 billion active parameters per token.
What is Kimi Delta Attention?
Kimi Delta Attention is a hybrid attention mechanism that compresses historical key-value states into dynamic state-space deltas, reducing KV-cache memory usage by 68% at 1 million tokens.
What context window length does Kimi K3 support?
Kimi K3 natively supports a 1,048,576-token (1 million) context window.
Who founded Moonshot AI?
Moonshot AI was founded in 2023 by Yang Zhilin, a leading AI researcher and co-author of Transformer-XL and XLNet.

Skyroot Aerospace Launches Vikram-1, Ushering India Into Private Orbital Era
Deep technical breakdown of India's first privately developed orbital launcher, carbon composite airframes, and payload economics.

ChatGPT GPT-5.6 Preview: Everything You Need to Know
Explore the new tiered family of models (Sol, Terra, Luna) and discover its advanced reasoning and coding capabilities.

Claude Opus 4.8: Anthropic's Most Advanced AI Model — Benchmarks & Review
Deep dive into Claude 4.8 benchmarks, including SWE-Bench Pro, Terminal-Bench 2.1, and the new Dynamic Workflows engine.
Was this newsletter forwarded to you? Sign up to get it in your inbox.
