Skip to main content
The Human Bit

The Lab · 13 current model entries

Know the model. Choose the job.

Names change faster than people can learn them. The Lab keeps official positions, independent observations and Human Bit tests separate, then lets you compare only the fields relevant to your job.

Coverage rule

Major current text and reasoning models people can choose or deploy.

We include a model when it is currently accessible, materially different from another tier and relevant to a real user decision. The catalogue order is not a rank.

OpenAI · Anthropic · Google · xAI · Moonshot AI

Filter the bench

Compare recorded fit, not a universal score.

Start with the work surface or provider you can actually use. Select up to three entries to compare their stated fit, limits, evidence coverage and review dates side by side.

More filters
13 of 13 entries · catalogue order, not rank
OpenAI · GPT-5Available

GPT-5.5 Instant

The everyday ChatGPT default: fast enough for routine questions, drafting, rewriting and low-consequence transformations.

Useful clue
Quick explanations and first drafts
Access
ChatGPT
Official record onlyinformational · gpt-55-instant@v1
Reviewed 16 July 2026Review by 30 July 2026
OpenAI · GPT-5.6Rolling out

GPT-5.6 Sol

OpenAI’s main reasoning model for complex knowledge work, research, coding and tasks that need more deliberate analysis.

Useful clue
Conflicting evidence and multi-step analysis
Access
ChatGPT · Work · Codex · API
Official + independent noteinformational · gpt-56-sol@v1
Reviewed 16 July 2026Review by 30 July 2026
OpenAI · GPT-5.6Rolling out

GPT-5.6 Sol Pro

The highest-capability GPT-5.6 option for difficult and longer-running work where a standard reasoning pass is not enough.

Useful clue
Long-running professional workflows
Access
ChatGPT Pro · API
Official + independent noteinformational · gpt-56-sol-pro@v1
Reviewed 16 July 2026Review by 30 July 2026
Anthropic · Claude SonnetAvailable

Claude Sonnet 5

Anthropic’s current Sonnet balance for everyday professional work, coding and agentic tasks, with adaptive thinking by default.

Useful clue
Professional writing and analysis
Access
Claude · Claude API · Cloud platforms
Official + independent noteinformational · claude-sonnet-5@v1
Reviewed 25 July 2026Review by 5 Sept 2026
Anthropic · Claude OpusAvailable

Claude Opus 5

Anthropic’s current model for complex agentic coding and enterprise work, and the model it points Opus 4.8 users at. Thinking is on by default, so the effort setting, not a thinking switch, is how you control depth and cost.

Useful clue
Complex agentic coding and long-horizon tasks
Access
Claude · Claude API · Cloud platforms
Official record onlyinformational · claude-opus-5@v1
Reviewed 25 July 2026Review by 5 Sept 2026
Anthropic · Claude OpusAvailable

Claude Opus 4.8

A high-capability Claude model for complex agentic coding and enterprise work, now superseded by Claude Opus 5 but still available on every platform that carried it.

Useful clue
Existing integrations not yet migrated to Opus 5
Access
Claude · Claude API · Cloud platforms
Official record onlyinformational · claude-opus-48@v1
Reviewed 25 July 2026Review by 22 Aug 2026
Anthropic · Claude FableAvailable

Claude Fable 5

Anthropic’s most capable widely released model for demanding reasoning and long-horizon agentic work, with stricter safety classifiers than Mythos.

Useful clue
Very long and difficult agentic projects
Access
Claude · Claude API · Cloud platforms
Official + independent noteinformational · claude-fable-5@v1
Reviewed 25 July 2026Review by 5 Sept 2026
Google · GeminiAvailable

Gemini 3.5 Flash

Google’s fast current model for multimodal, coding and agentic work, deployed broadly across consumer, developer and enterprise products.

Useful clue
Fast multimodal document work
Access
Gemini app · AI Mode in Google Search · Google AI Studio · Gemini API · Gemini Enterprise
Official + independent noteinformational · gemini-35-flash@v1
Reviewed 25 July 2026Review by 22 Aug 2026
Google · GeminiPreview

Gemini 3.1 Pro Preview

A preview model for precise multi-step reasoning, software engineering and tool use across text, images, video, audio and PDFs.

Useful clue
Complex multimodal analysis
Access
Gemini API · Google AI Studio · Vertex AI
Official record onlywatch · gemini-31-pro-preview@v1
Reviewed 25 July 2026Review by 8 Aug 2026
Google · GeminiAvailable

Gemini 3.1 Flash-Lite

Google’s lower-latency, lower-cost model for high-volume extraction and lightweight multimodal jobs.

Useful clue
Classification and simple extraction
Access
Gemini API · Google AI Studio · Vertex AI
Official record onlyinformational · gemini-31-flash-lite@v1
Reviewed 25 July 2026Review by 22 Aug 2026
xAI · GrokAvailable

Grok 4.5

xAI’s current flagship for coding, agentic tasks and knowledge work, with configurable reasoning and an emphasis on speed.

Useful clue
Engineering and coding evaluation
Access
Grok Build · Cursor · xAI API
Official record onlyinformational · grok-45@v1
Reviewed 16 July 2026Review by 30 July 2026
Moonshot AI · KimiAvailable

Kimi K3

Kimi’s hosted multimodal flagship for long-horizon coding, knowledge work and tool-using tasks, with a one-million-token context window.

Useful clue
Large document and knowledge-work packs
Access
Kimi · Kimi Work · Kimi Code · Kimi API
Official + independent noteinformational · kimi-k3@v1
Reviewed 25 July 2026Review by 29 July 2026
Moonshot AI · Kimi CodeAvailable

Kimi K2.7 Code

Kimi’s dedicated coding model for long-context, tool-using software work, with a separate high-speed endpoint for the same model.

Useful clue
Long-context repository work
Access
Kimi API · OpenAI-compatible clients
Official record onlyinformational · kimi-k27-code@v1
Reviewed 25 July 2026Review by 5 Sept 2026

The model list can grow. The standard cannot loosen.

Each entry needs an official record, a practical decision, an evidence state and a review owner. Missing independent evidence remains visible rather than being filled with a guess.

Read the evidence standard →