Gemma 4 Guides

Kimi K3 Benchmarks: How It Actually Stacks Up

9 min read
kimi k3benchmarksllm evaluationcoding llmmoonshot ai
Kimi K3 Benchmarks: How It Actually Stacks Up

Kimi K3 Benchmarks: How It Actually Stacks Up

Within hours of Kimi K3's July 16, 2026 release, it had jumped from #18 to #1 on LMArena's Frontend Code Arena, overtaking Claude Fable 5. It also posted a higher score than Claude Opus 4.8 on the Artificial Analysis Intelligence Index. Those two facts alone explain most of the launch-week noise — but a single leaderboard position is a poor substitute for actually understanding where a model is strong, where it isn't, and how much of that gap is real versus benchmark-harness noise.

This is a breakdown of Kimi K3's benchmark performance by category, sourced from Moonshot's own materials and Artificial Analysis's independent evaluation, with the methodology caveats that matter before you make a decision based on any of it.

Kimi K3 benchmark dashboard illustration showing intelligence index bars, a coding leaderboard, and a speed/latency comparison chart

Quick answer

  • Artificial Analysis Intelligence Index v4.1: Kimi K3 scores 57, ahead of Claude Opus 4.8 (56), behind Claude Fable 5 (60).
  • Frontend coding (LMArena Frontend Code Arena): Kimi K3 is #1, first in 6 of 7 frontend domains, having overtaken Claude Fable 5.
  • Agentic reasoning (GDPval-AA v2): Kimi K3 scores ~1,687, ahead of Claude Opus 4.8 (1,600) but behind Claude Fable 5 Max (~1,815) and GPT-5.6 Sol Max (~1,748).
  • Head-to-head vs. Claude Fable 5: across the benchmarks both vendors report, Fable 5 wins roughly 8 of 14, K3 wins roughly 6 — including long-horizon agentic coding, BrowseComp, and Terminal-Bench 2.1.
  • Head-to-head vs. Claude Opus 4.8: Kimi K3 leads on 7 of 10 compared benchmarks (including BrowseComp, FrontierSWE, GDPval-AA, and Toolathlon); Opus 4.8 still leads on GPQA Diamond, Humanity's Last Exam, and OfficeQA Pro.
  • Speed: K3 generates roughly 62 tokens/sec with a ~2-second time-to-first-token, notably faster than Opus 4.8's max-effort reasoning mode.
  • Caveat: these numbers are days old, mostly vendor- or single-aggregator-reported, and run under specific evaluation settings (reasoning effort set to "max," temperature 1.0, top-p 1.0). Treat them as a strong early signal, not a settled verdict.

The headline number: Artificial Analysis Intelligence Index

Artificial Analysis combines nine evaluations — including GDPval-AA v2, Terminal-Bench 2.1, SciCode, Humanity's Last Exam, GPQA Diamond, and long-context reasoning tests — into a single Intelligence Index score.

Model AA Intelligence Index v4.1
Claude Fable 5 (Adaptive, Max effort) 60
Kimi K3 (max reasoning effort) 57
Claude Opus 4.8 (Adaptive, Max effort) 56

The takeaway that actually matters: K3 is not the smartest model available, but it's within a few points of the frontier while costing a fraction of what Fable 5 or Opus 4.8 charge per task. For most production use cases, "close to frontier, much cheaper" beats "slightly ahead of frontier, expensive" — but not always, which is why the category breakdown below matters more than the single composite number.

Agentic and long-horizon benchmarks

This is where K3 is genuinely strong, and where Moonshot clearly optimized.

GDPval-AA v2

A private agentic benchmark measuring how well a model performs realistic, economically valuable knowledge work.

Model GDPval-AA v2 (Elo)
Claude Fable 5 Max ~1,815
GPT-5.6 Sol Max ~1,748
Kimi K3 ~1,687
Claude Opus 4.8 1,600

K3 places third overall — behind the two most expensive closed-source frontier models, but clearly ahead of Opus 4.8.

AA-Briefcase (agentic knowledge work)

Model AA-Briefcase (Elo)
Claude Fable 5 Max ~1,587
Kimi K3 ~1,527
GPT-5.6 Sol Max ~1,495

Here K3 climbs to second place, ahead of GPT-5.6 Sol Max.

BrowseComp (agentic web research)

K3 leads Claude Fable 5 on BrowseComp — one of the clearest wins in Moonshot's favor, and a meaningful one if your workload involves autonomous research or multi-step web navigation.

Coding benchmarks

Benchmark Kimi K3
Terminal-Bench 2.1 88.3
FrontierSWE 81.2
Program Bench 77.8
DeepSWE 67.5
SWE Marathon 42.0

A note on SWE Marathon: a low absolute score (42.0) doesn't mean the model is weak — it's one of the hardest long-horizon coding benchmarks available, where every frontier model scores low in absolute terms. What matters is the comparison: K3 leads Claude Fable 5 on SWE Marathon by roughly 7 points, which is a meaningful margin at this difficulty level.

LMArena Frontend Code Arena

This is the benchmark that generated the most launch-day attention, and unlike the vendor-reported numbers above, it's a crowd-voted, independently-run leaderboard. Kimi K3 went from #18 to #1 within hours of release, scoring 1,679 and placing first in 6 of 7 frontend domains, overtaking Claude Fable 5.

Because this is community-voted rather than a fixed test set, it's harder to game with training-set overlap — which makes it one of the more credible signals in this entire launch, not just the loudest one.

Head-to-head: Kimi K3 vs. Claude Opus 4.8

Across the benchmarks both are commonly evaluated on:

Kimi K3 leads on: BrowseComp, CharXiv-R, DeepSearchQA, FrontierSWE, GDPval-AA, MCP Atlas, Toolathlon (7 benchmarks)

Claude Opus 4.8 leads on: GPQA Diamond, Humanity's Last Exam, OfficeQA Pro (3 benchmarks)

The pattern is consistent: K3 wins on agentic, tool-using, and coding-adjacent tasks; Opus 4.8 wins on pure academic reasoning and knowledge-heavy tasks. If your product looks more like an agent than a quiz-taker, that pattern favors K3.

Head-to-head: Kimi K3 vs. Claude Fable 5

Fable 5 is the newer, stronger Claude model, and it shows: across 14 shared benchmarks, Fable 5 wins about 8, including a 5.4-point win on FrontierSWE and a sweep of both visual-reasoning suites. Kimi K3 wins about 6, including SWE Marathon (by ~7 points), BrowseComp, and Terminal-Bench 2.1.

The full comparison, including pricing and a workload-by-workload recommendation, is in Kimi K3 vs Claude: Fable 5 and Opus 4.8 compared.

Speed and latency

Benchmarks measure correctness; for many real applications, latency matters just as much.

Metric Kimi K3 Claude Opus 4.8 (Max effort)
Output speed ~62 tokens/sec ~56 tokens/sec
Time to first token ~1.99s ~34.49s

The time-to-first-token gap is the more consequential number for user-facing products — Opus 4.8's max-reasoning-effort mode spends a long time "thinking" before the first token appears, while K3's always-on thinking mode is tuned to start streaming much sooner. If you're building anything with a visible loading state, this difference is felt by users, not just measured in a spreadsheet.

Methodology caveats — read this before you decide anything

A few things worth knowing before treating any of the above as gospel:

  • Reasoning effort matters. Kimi's reported results use reasoning effort set to "max," temperature 1.0, and top-p 1.0. A different configuration — including whatever your production deployment defaults to — can move scores meaningfully.
  • Harness choice matters. Agentic benchmarks are run through a harness (Kimi Code, Claude Code, Codex, etc.), and the harness itself — how it handles retries, context management, and tool orchestration — can shift scores by several points independent of the underlying model.
  • These are days-old numbers. K3 launched July 16, 2026. Some figures circulating (including a few above) come from a single aggregator or vendor claim rather than multiple independent replications. Expect some numbers to be revised as more of the community runs its own evaluations, and especially once the public weights ship on July 27, 2026 and independent groups can test the actual model rather than the hosted API.
  • Composite indices hide category performance. A single Intelligence Index number is convenient but flattens real differences — a model that's mediocre at everything and a model that's brilliant at coding but weak at trivia can land on the same composite score. Always check the category breakdown for your actual use case.

What this means by workload

Workload How K3 benchmarks Verdict
Frontend / UI code generation #1 on LMArena Frontend Code Arena Strong pick
Long-horizon autonomous coding agents Leads on SWE Marathon, Terminal-Bench 2.1 Strong pick
Agentic web research Leads Fable 5 on BrowseComp Strong pick
Frontier software engineering (hardest tier) Trails Fable 5 on FrontierSWE by 5.4 pts Evaluate Fable 5 too
Academic reasoning / science QA Trails Opus 4.8 on GPQA, HLE Evaluate Opus 4.8 / Fable 5
Cost-sensitive agent workloads at scale Near-frontier score at a fraction of the price Strong pick

Bottom line

Kimi K3's benchmark story is genuinely good, not just loud: a #1 community-voted coding leaderboard result, a composite score ahead of Claude Opus 4.8, and clear wins on the agentic and tool-use categories that matter most for anyone building autonomous coding or research agents. It is not the smartest model on the market — Claude Fable 5 still leads on frontier engineering and academic reasoning — but it's close enough, at a low enough price, that "test K3 first" is a reasonable default for most agentic and coding workloads as of this launch window.

For the pricing side of that equation, see Kimi K3 pricing. For the full model-vs-model breakdown, see Kimi K3 vs Claude.

FAQ

Is Kimi K3 better than Claude Opus 4.8? On most shared benchmarks, yes — K3 leads on 7 of 10 compared evaluations, including BrowseComp, FrontierSWE, and GDPval-AA, and scores higher on the Artificial Analysis Intelligence Index (57 vs. 56). Opus 4.8 still leads on GPQA Diamond, Humanity's Last Exam, and OfficeQA Pro.

Is Kimi K3 better than Claude Fable 5? Not overall — Fable 5 wins roughly 8 of 14 shared benchmarks and scores higher on the Intelligence Index (60 vs. 57). K3 wins on long-horizon agentic coding, BrowseComp, and Terminal-Bench 2.1, and costs substantially less per task.

What is Kimi K3's score on the Artificial Analysis Intelligence Index? 57, as of the index's v4.1 methodology, placing it between Claude Opus 4.8 (56) and Claude Fable 5 (60).

Did Kimi K3 really hit #1 on a coding leaderboard? Yes — LMArena's Frontend Code Arena, a crowd-voted leaderboard, ranked K3 #1 within hours of its release, up from #18, overtaking Claude Fable 5 and placing first in 6 of 7 frontend domains.

Are Kimi K3's benchmark numbers reliable? They're a reasonable early signal but not a final verdict. Most figures are vendor-reported or come from a single independent aggregator (Artificial Analysis), evaluated under specific settings (max reasoning effort). Expect refinement as more groups test the model independently, particularly after the public weights ship.

Related guides

Related guides

Continue through the Gemma 4 cluster with the next guide that matches your current decision.

Still deciding what to read next?

Go back to the guide hub to browse model comparisons, setup walkthroughs, and hardware planning pages.