Skip to main content
The Human Bit

Moonshot AI · Kimi

Kimi K3

Kimi’s hosted multimodal flagship for long-horizon coding, knowledge work and tool-using tasks, with a one-million-token context window.

The Human Bit position

A serious candidate, not a new default. Test long-horizon quality, total reasoning cost, history handling and whether its proactiveness crosses the user’s boundary.

Use it for

  • Large document and knowledge-work packs
  • Long-running coding and terminal work
  • Tasks combining visual evidence with code or tools
Leave it alone when

Do not use it as a quick-answer default, rely on its web search in production or give it broad authority when an unexpected decision would matter.

Effort without the jargon

Low

The task is simple enough that a long reasoning pass would mostly buy tokens.

High

The work needs real analysis but does not justify the deepest and most expensive pass.

Max

The task is long-horizon or difficult enough to justify the deepest reasoning pass and its token cost.

Advanced recordCost, limits, privacy and operations

K3 always thinks and defaults to its deepest reasoning setting, so the effort level is the cost control; its million-token capacity is not a reason to send everything, because reasoning and output can dominate cost.

API cost
Officially disclosedKimi uses flat pay-as-you-go token pricing with separate cache-hit, cache-miss and output rates and no context-length tiering; check the live table for current amounts.
Context window
Officially disclosed1,000,000 tokens
Maximum output
Officially disclosed1,048,576 completion tokens (131,072 default)
Knowledge cutoff
Not disclosedNot disclosed in the cited record
Recorded inputs
Text · Images · Video
Recorded output
Text · Structured JSON

Data handling is a product decision

Not disclosedThe cited model and pricing records do not state a portable retention or training rule. Confirm Kimi API and product terms before sending sensitive material.

Operational constraints

  • Thinking is always on and cannot be disabled; reasoning effort selects low, high or max, and defaults to max.
  • Sampling values are fixed, and complete assistant messages must be preserved across tool turns.
  • Public image URLs are unsupported; use base64 or an ms:// file ID.
  • Kimi says its web-search tool is being updated and is not recommended for production in the near term.

Unknown means the cited official record does not disclose a safe value. It is not an estimate. Prices are provider list prices in USD where stated and can change before this record's review date.

Keep the evidence separate

Canonical provider claim

Moonshot calls K3 its most capable model, with 2.8 trillion parameters, native understanding of images and video, and a one-million-token context. Its quickstart documents reasoning effort at low, high or max. Hosted access is live and requires a minimum top-up to unlock, while full weights are promised by 27 July 2026.

Independent observation · Simon Willison

His first hands-on run confirmed working vision and valid SVG output, but the reasoning pass burned 13,241 tokens to produce 3,417 tokens of answer, making one simple test cost 25 cents. Writing on 2026-07-16 he found only one reasoning effort available, max; Kimi's quickstart now documents low, high and max, so the cost trap he measured is now partly avoidable, and his token-volume warning is the part that still holds.

Read the independent source ↗
Readiness · informational

No Human Bit result is claimed; use this as evaluation guidance only.

Human Bit test

Not scheduled

No Human Bit result is claimed for this model yet.

Sources