Changes confirmed high confidence

Moonshot AI Launches Kimi K3, a 2.8-Trillion-Parameter 'Open Frontier' Flagship — Weights Promised, Not Yet Downloadable

The strongest open-weight-class model yet shipped July 16 as a hosted service, with full weights committed for July 27 and no license published — buyers should treat 'open' as a dated promise, not an artifact.

Moonshot AI launched Kimi K3 on July 16, 2026, a 2.8-trillion-total-parameter Stable LatentMoE model with 1M-token context, native image and video input, and always-on reasoning, per launch coverage quoting Moonshot's announcement and technical blog.

Context

K3 is the closest an open-weight-class model has come to the frontier since DeepSeek R1, and the first open-announced model in the 3-trillion-parameter class. It landed in the middle of July's release wave, one day after Thinking Machines' Inkling and five days before Google's Gemini 3.6 Flash.

What changed

  • Architecture: Stable LatentMoE with 16 of 896 experts active per token; Kimi Delta Attention with Attention Residuals; 1M-token context.
  • Day-one availability: Kimi app and kimi.com (free tier), the `kimi-k3` OpenAI-compatible API, Kimi Work desktop, Kimi Code CLI, and OpenRouter.
  • Pricing: $3 per million input tokens on cache miss, $0.30 on cache hit, $15 output.
  • Variants: K3 Max and K3 Swarm Max.
  • Weights: committed to public release by July 27, 2026 — but not downloadable at launch, and no license text, model card repo or checksums had been published as of July 22. K3 is API-only today.

Why it matters

K3 repriced the frontier-adjacent tier and demonstrated that a Chinese lab can ship near-frontier capability with open distribution intent. Moonshot's own technical blog concedes overall quality "still trails the strongest proprietary Claude Fable 5 and GPT-5.6 Sol models" while beating Opus 4.8 and GPT-5.5, per BreachRoad's enterprise analysis. Independent testing by Artificial Analysis placed K3 fourth of 189 models (score 57) and first on the Frontend Code Arena at 1,679 Elo — ahead of Fable 5 (1,631) and GPT-5.6 Sol (1,618) — with Vals.ai measuring 93.4% on SWE-bench Verified, per Vectrel's roundup.

Details

The launch also showcased agentic engineering feats — including a chip-design task completed in 48 hours at over 8,700 tokens/second throughput — as vendor demonstrations. Analyst Nathan Lambert called K3 "the closest open models have been to the frontier since DeepSeek R1." The gap between announcement and artifact matters: as aireiter's analysis notes, Moonshot's Hugging Face org still tops out at Kimi K2.7 Code.

Limitations and caveats

Do not cite a license. Press reports listing "MIT" or "Modified MIT" are inference from the K2-family precedent; no license text exists as of July 22. A Modified MIT license is widely expected based on K2/K2.7 precedent — whose attribution clause triggers at 100M monthly active users or $20M monthly revenue — but is unconfirmed. Launch-table benchmarks are vendor-reported; independent figures above come from third-party evaluators. Self-hosting will not relieve hosted demand soon: K3 requires 64+-accelerator supernodes.

Sources

*Update note: This post was last reviewed on 2026-07-22. Next checkpoint: July 27 weights release — verify the repo on Moonshot's Hugging Face org, the license file text, and any MAU/revenue attribution clause.*

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Changes coverage