Kimi K3 evaluation

Kimi K3 Benchmark Rankings

Read Kimi K3 benchmark rankings without overfitting to one leaderboard. Compare coding, agentic, reasoning, vision, and long-context tasks.

Kimi K3 benchmark dashboard with coding and agentic categories
Comparisonpage shape
Kimi K3topic
6further reading links
2026-07-29updated

Kimi K3 benchmark snapshot

BenchmarkKimi K3 scoreHow to read it
GPQA Diamond93.5Reasoning and knowledge screening; close to the highest listed model in the official table
DeepSWE67.5Coding benchmark; competitive, but not the top score in the listed set
Terminal-Bench 2.188.3Terminal-heavy coding; near the top of the official comparison
FrontierSWE81.2Long-horizon software work; ahead of several listed frontier baselines
SWE-Marathon42.0Extended software tasks; strongest listed score in that official row
BrowseComp91.2Agentic browsing and long-context work; strongest listed score in that row
MCPMark-Verified94.5Tool and MCP-style agent tasks; strongest listed score in that row
Artificial Analysis Intelligence Index57External comparison point; DeepSeek V3.2 Reasoning is listed at 32 on the same page

How to choose

Kimi K3 ranks as a frontier-class open-weight model in Moonshot's official benchmark table. It is especially strong on long-horizon coding and agentic work: 88.3 on Terminal-Bench 2.1, 81.2 on FrontierSWE, 42.0 on SWE-Marathon, 91.2 on BrowseComp, and 94.5 on MCPMark-Verified.

The ranking is not a single universal number. Kimi K3 is near the top on several coding and agentic tasks, but different harnesses produce different winners, so use the scores below as a shortlist and then run your own repeatable task.

AI handoff prompt and permissions

Copy this prompt into Kimi Code, K3Nova, or another AI agent. Keep approval manual for file writes, shell commands, account actions, and secrets.

Copyable AI prompt

Use this to give the agent the task and safety boundary in one message.

You are my AI agent for this task: verify Kimi K3 benchmark claims from official and external sources, then turn them into a short model-evaluation memo.
Start by restating the goal and the permissions you need.
Use official docs or the files I provide before making claims.
Give me a direct answer first, then a short table or checklist.
If commands are needed, show exact copyable commands without a shell prompt.
Ask before writing files, running shell commands, deleting or moving data, logging in, spending money, changing account settings, or handling API keys.
Stop and ask me when a step requires secrets, payment, account access, destructive cleanup, or a permission broader than the task.

Recommended AI permissions

PermissionGive AIWhy
Public researchAllow official docs, cited public pages, and web search results.Needed to refresh facts without touching private data.
Local filesDo not allow local file access unless you name a folder for inspection.Most research pages do not need your machine.
Write or shellAsk before Write, Edit, Bash, or any command execution.Keeps the task reviewable.
Sensitive actionsDeny login, payment, API keys, account changes, and private screenshots.Those actions need a human decision.

Test sequence

Start with the official score table

Use the Moonshot Kimi K3 README for the current score snapshot, model scale, benchmark categories, and footnotes about harnesses.

Check one outside comparison

Artificial Analysis lists Kimi K3 at 57 on its Intelligence Index versus 32 for DeepSeek V3.2 Reasoning, with a much larger context window.

Run your own fixture

Save a coding task, research task, and long-document task. Keep the prompt, expected result, model setting, and observed failure mode.

Compare by job type

A model that wins a coding benchmark may not be the best fit for a document-heavy operating workflow, and the reverse can also be true.

Decision criteria

Coding ranking signal

Kimi K3 scores 88.3 on Terminal-Bench 2.1 and 81.2 on FrontierSWE in the official table, which makes it worth testing for repository and terminal-heavy work.

Agentic ranking signal

Kimi K3 scores 91.2 on BrowseComp, 94.5 on MCPMark-Verified, and 30.8 on AutomationBench, so agent workflows deserve a separate evaluation path.

Context and cost caveat

External comparisons can rank Kimi K3 strongly on intelligence while still showing different cost, speed, and context tradeoffs. Do not choose from scores alone.

FAQ

Are Kimi K3 benchmark rankings enough to choose a model?

No. Use rankings for screening, then test your own repeatable workflow.

Which Kimi K3 ranking matters for coding?

Look for benchmarks that include long-horizon coding, terminal use, repository navigation, and bug-fix quality.

Should I compare Kimi K3 with DeepSeek only by scores?

No. Compare scores, context, cost, provider path, tool behavior, and the actual work you plan to run.

Further reading