Moonshot AI · Kimi
Kimi K3
Kimi’s hosted multimodal flagship for long-horizon coding, knowledge work and tool-using tasks, with a one-million-token context window.
A serious candidate, not a new default. Test long-horizon quality, total reasoning cost, history handling and whether its proactiveness crosses the user’s boundary.
Use it for
- Large document and knowledge-work packs
- Long-running coding and terminal work
- Tasks combining visual evidence with code or tools
Do not use it as a quick-answer default, rely on its web search in production or give it broad authority when an unexpected decision would matter.
Effort without the jargon
The task is simple enough that a long reasoning pass would mostly buy tokens.
The work needs real analysis but does not justify the deepest and most expensive pass.
The task is long-horizon or difficult enough to justify the deepest reasoning pass and its token cost.
Advanced recordCost, limits, privacy and operations
K3 always thinks and defaults to its deepest reasoning setting, so the effort level is the cost control; its million-token capacity is not a reason to send everything, because reasoning and output can dominate cost.
- API cost
- Officially disclosedKimi uses flat pay-as-you-go token pricing with separate cache-hit, cache-miss and output rates and no context-length tiering; check the live table for current amounts.
- Context window
- Officially disclosed1,000,000 tokens
- Maximum output
- Officially disclosed1,048,576 completion tokens (131,072 default)
- Knowledge cutoff
- Not disclosedNot disclosed in the cited record
- Recorded inputs
- Text · Images · Video
- Recorded output
- Text · Structured JSON
Data handling is a product decision
Not disclosedThe cited model and pricing records do not state a portable retention or training rule. Confirm Kimi API and product terms before sending sensitive material.
Operational constraints
- Thinking is always on and cannot be disabled; reasoning effort selects low, high or max, and defaults to max.
- Sampling values are fixed, and complete assistant messages must be preserved across tool turns.
- Public image URLs are unsupported; use base64 or an ms:// file ID.
- Kimi says its web-search tool is being updated and is not recommended for production in the near term.
Unknown means the cited official record does not disclose a safe value. It is not an estimate. Prices are provider list prices in USD where stated and can change before this record's review date.
Keep the evidence separate
Moonshot calls K3 its most capable model, with 2.8 trillion parameters, native understanding of images and video, and a one-million-token context. Its quickstart documents reasoning effort at low, high or max. Hosted access is live and requires a minimum top-up to unlock, while full weights are promised by 27 July 2026.
His first hands-on run confirmed working vision and valid SVG output, but the reasoning pass burned 13,241 tokens to produce 3,417 tokens of answer, making one simple test cost 25 cents. Writing on 2026-07-16 he found only one reasoning effort available, max; Kimi's quickstart now documents low, high and max, so the cost trap he measured is now partly avoidable, and his token-volume warning is the part that still holds.
Read the independent source ↗No Human Bit result is claimed; use this as evaluation guidance only.
Not scheduled
No Human Bit result is claimed for this model yet.