Site navigation

Model intelligenceLive snapshot

Verdicts by the work, not a leaderboard.

HiveBase routes each work class to a preferred model — and for coding, to a preferred harness × model × role. This page publishes that opinion, with the evidence we are allowed to show.

Tracked sources are named and linked. Artificial Analysis is internal-only. Sample sizes and confidence intervals are shown when we have them. There is no “#1.”

LAST REFRESHED
Sep 14, 2026, 06:30 AM UTC

Live snapshot from the public Model Intel publication (slug current), regenerated at most hourly. The ledger API returns these same numbers.

Classes
11
Auto-reverts
12
Feed
live
WORK CLASSES

Eleven jobs. Eleven verdicts.

Jump a class, or read them in order. Coding rows always name the harness. n is the first-party sample; intervals are shown when present.

code.agentic

Agentic coding

Multi-file repository work: implement, recover, verify. Ranked as (harness, model, role).

n unpublished
Unpublished default · harness × model
No published default

No published HiveBase default for this class.

Currently serving · reverted
kimi-k2.7-code

Applicable: Multi-file repository work: implement, recover, verify. Ranked as (harness, model, role). · Harness × model × role (planner, executor, reviewer)

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • The same model in a different harness is a different row.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

Attributed sources for this class. No qualified recommendation published.

code.frontend

Front-end / UI

Generated websites, components, and full-stack UI. Taste plus deployable quality.

n unpublished
Unpublished default · harness × model
No published default

No published HiveBase default for this class.

Applicable: Generated websites, components, and full-stack UI. Taste plus deployable quality. · Harness × model × role (executor)

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • The same model in a different harness is a different row.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

Attributed sources for this class. No qualified recommendation published.

  • LMArena WebDev / Code Arena

    See LMArena WebDev / Code Arena (numbers not republished).

    Source
code.review

Code review

Finding real defects without comment fatigue. Cross-family review is required.

n unpublished
Unpublished default · harness × model
No published default

No published HiveBase default for this class.

Applicable: Finding real defects without comment fatigue. Cross-family review is required. · Harness × model × role (reviewer)

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • The same model in a different harness is a different row.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

No attributed public sources in this snapshot.

No attributed evidence published for this class yet.

reasoning.deep

Deep reasoning

Strategic, second-order, long-horizon analysis.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Currently serving · reverted
gpt-5.6-sol

Applicable: Strategic, second-order, long-horizon analysis.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

Attributed sources for this class. No qualified recommendation published.

agentic.tools

Agentic tool use

Multi-step tool calling with schema fidelity and recovery.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Applicable: Multi-step tool calling with schema fidelity and recovery.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

Attributed sources for this class. No qualified recommendation published.

writing.narrative

Long-form writing

User-visible prose, brand voice, and durable narratives.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Applicable: User-visible prose, brand voice, and durable narratives.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

Attributed sources for this class. No qualified recommendation published.

  • LMArena Text

    See LMArena Text (numbers not republished).

    Source
extraction.structured

Structured extraction

Schema-valid JSON / Output.object without silent field loss.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Currently serving · reverted
gpt-oss-120b

Applicable: Schema-valid JSON / Output.object without silent field loss.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

No attributed public sources in this snapshot.

No attributed evidence published for this class yet.

longctx

Long context

Documents, transcripts, and large codebases. Effective window, not marketed size.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Currently serving · reverted
qwen3-6-plus

Applicable: Documents, transcripts, and large codebases. Effective window, not marketed size.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

No attributed public sources in this snapshot.

No attributed evidence published for this class yet.

vision.doc

Document / vision

PDF, OCR, diagram, and image understanding.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Currently serving · reverted
gemini-3.5-flash

Applicable: PDF, OCR, diagram, and image understanding.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

Attributed sources for this class. No qualified recommendation published.

  • LMArena Vision

    See LMArena Vision (numbers not republished).

    Source
speed.classify

Classification / gates

Binary and cheap structured gates. Recoverable if wrong.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Applicable: Binary and cheap structured gates. Recoverable if wrong.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

No attributed public sources in this snapshot.

No attributed evidence published for this class yet.

intel.x

X / community intel

Native X search and community recency. Discovery, not deciding.

n unpublished
Unpublished default · model
No published default

No published HiveBase default for this class.

Applicable: Native X search and community recency. Discovery, not deciding.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.
Research basis · Sep 14, 2026, 06:30 AM UTC

No attributed public sources in this snapshot.

No attributed evidence published for this class yet.

METHODOLOGY

How a verdict is allowed to exist.

Work classes are the spine. External benchmarks, community consensus, and HiveBase first-party outcomes all hang off the same eleven jobs so they stay commensurable. Fusion never promotes a default from community recency alone — that channel is capped and decays.

Coding is harness-keyed. The ranked unit is (harness, model, role). An executor on Claude Code is not comparable to the same weights inside Cursor or Codex. Review must be cross-family.

Outcomes stay separate. Research-based defaults can exist before HiveBase outcome data accumulates. Dogfood statistics appear in their own block with population, dates, denominator, and uncertainty. n < 20 is below the reporting threshold, not a causal override. Durable selection guidance lives in /docs/guides/model-intelligence. The same snapshot is at GET /api/models/ledger.

Artificial Analysis is not rehosted. AA indexes inform internal routing. This page will only say they are tracked and point at artificialanalysis.ai. Attributed sources (Terminal-Bench, SWE-Rebench, Design Arena, LMArena, BFCL, Epoch) appear as name + outbound link.

Saturated suites (SWE-bench Verified, HumanEval, retired open LLM leaderboards) are weight-zero. A canary that fails the quality floor auto-reverts; those events are public when the live feed includes them.

CHANGELOG

What moved — including auto-reverts.

Promotions that did not hold stay on the record. Auto-reverts are not hidden behind a successful later default.

  1. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: claude-opus-4-8-1m → claude-opus-4-8-1m

  2. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  3. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: kimi-k2.6 → kimi-k2.6

  4. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  5. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.1-pro → gemini-3.1-pro

  6. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  7. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on extraction.structured: gpt-oss-120b → gpt-oss-120b

  8. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: qwen3-6-plus → qwen3-6-plus

  9. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  10. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  11. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on reasoning.deep: gpt-5.6-sol → gpt-5.6-sol

  12. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: claude-opus-4-8-1m → claude-opus-4-8-1m

  13. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  14. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.6 → kimi-k2.6

  15. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  16. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.1-pro → gemini-3.1-pro

  17. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  18. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on extraction.structured: gpt-oss-120b → gpt-oss-120b

  19. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: qwen3-6-plus → qwen3-6-plus

  20. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  21. Sep 7, 2026, 08:00 AM UTCauto revert

    Rollback on longctx: qwen3-6-plus → qwen3-6-plus

  22. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  23. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on reasoning.deep: gpt-5.6-sol → gpt-5.6-sol

  24. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: claude-opus-4-8-1m → claude-opus-4-8-1m

  25. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  26. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.6 → kimi-k2.6

  27. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  28. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.1-pro → gemini-3.1-pro

  29. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  30. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on extraction.structured: gpt-oss-120b → gpt-oss-120b

  31. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: qwen3-6-plus → qwen3-6-plus

  32. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  33. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on longctx: qwen3-6-plus → qwen3-6-plus

  34. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  35. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on reasoning.deep: gpt-5.6-sol → gpt-5.6-sol

SOURCE RECEIPT

Last refresh by source.

Ingest health for this snapshot. A skipped or empty refresh is shown, not implied successful. Artificial Analysis appears here as a tracked source, never as a copied score.

Last refresh time and status for each Model Intel source
SourceFinishedStatus
fusionSep 14, 2026, 06:00 AM UTCok
research_pollSep 14, 2026, 06:00 AM UTCok
overlaySep 14, 2026, 06:00 AM UTCok
research_pollSep 14, 2026, 05:30 AM UTCok
overlaySep 14, 2026, 05:00 AM UTCok
first_partySep 14, 2026, 05:00 AM UTCok
swe_rebenchSep 14, 2026, 05:00 AM UTCskipped
hf_hubSep 14, 2026, 05:00 AM UTCskipped
design_arenaSep 14, 2026, 05:00 AM UTCskipped
wulong_arenaSep 14, 2026, 05:00 AM UTCok
bfclSep 14, 2026, 05:00 AM UTCok
epoch_eciSep 14, 2026, 05:00 AM UTCok
terminal_benchSep 14, 2026, 05:00 AM UTCok
openrouterSep 14, 2026, 05:00 AM UTCok
openrouter_dataSep 14, 2026, 05:01 AM UTCok
lmarena_textSep 14, 2026, 05:00 AM UTCok
artificial_analysisTracked internally — artificialanalysis.aiSep 14, 2026, 05:00 AM UTCok
research_pollSep 14, 2026, 05:00 AM UTCok
research_pollSep 14, 2026, 04:30 AM UTCok
overlaySep 14, 2026, 04:00 AM UTCok
research_pollSep 14, 2026, 04:00 AM UTCok
research_pollSep 14, 2026, 03:30 AM UTCok
overlaySep 14, 2026, 03:00 AM UTCok
research_pollSep 14, 2026, 03:00 AM UTCok
research_pollSep 14, 2026, 02:30 AM UTCok
research_pollSep 14, 2026, 02:00 AM UTCok
overlaySep 14, 2026, 02:00 AM UTCok
research_pollSep 14, 2026, 02:00 AM UTCok
overlaySep 14, 2026, 02:00 AM UTCok
research_pollSep 14, 2026, 01:00 AM UTCok
overlaySep 14, 2026, 01:00 AM UTCok
research_pollSep 14, 2026, 12:30 AM UTCok
overlaySep 14, 2026, 12:00 AM UTCok
research_pollSep 14, 2026, 12:00 AM UTCok
research_pollSep 13, 2026, 11:30 PM UTCok
research_pollSep 13, 2026, 11:00 PM UTCok
overlaySep 13, 2026, 11:00 PM UTCok
research_pollSep 13, 2026, 10:30 PM UTCok
research_pollSep 13, 2026, 10:00 PM UTCok
overlaySep 13, 2026, 10:00 PM UTCok
research_pollSep 13, 2026, 09:30 PM UTCok
research_pollSep 13, 2026, 09:00 PM UTCok
overlaySep 13, 2026, 09:00 PM UTCok
research_pollSep 13, 2026, 08:30 PM UTCok
research_pollSep 13, 2026, 08:00 PM UTCok
overlaySep 13, 2026, 08:00 PM UTCok
research_pollSep 13, 2026, 07:30 PM UTCok
research_pollSep 13, 2026, 07:00 PM UTCok
overlaySep 13, 2026, 07:00 PM UTCok
research_pollSep 13, 2026, 06:30 PM UTCok
overlaySep 13, 2026, 06:00 PM UTCok
research_pollSep 13, 2026, 06:00 PM UTCok
research_pollSep 13, 2026, 05:30 PM UTCok
research_pollSep 13, 2026, 05:00 PM UTCok
overlaySep 13, 2026, 05:00 PM UTCok
research_pollSep 13, 2026, 04:30 PM UTCok
research_pollSep 13, 2026, 04:00 PM UTCok
overlaySep 13, 2026, 04:00 PM UTCok
research_pollSep 13, 2026, 03:30 PM UTCok
overlaySep 13, 2026, 03:00 PM UTCok
research_pollSep 13, 2026, 03:00 PM UTCok
research_pollSep 13, 2026, 02:30 PM UTCok
research_pollSep 13, 2026, 02:00 PM UTCok
overlaySep 13, 2026, 02:00 PM UTCok
research_pollSep 13, 2026, 01:30 PM UTCok
overlaySep 13, 2026, 01:00 PM UTCok
research_pollSep 13, 2026, 01:00 PM UTCok
research_pollSep 13, 2026, 12:30 PM UTCok
overlaySep 13, 2026, 12:00 PM UTCok
research_pollSep 13, 2026, 11:30 AM UTCok
overlaySep 13, 2026, 11:00 AM UTCok
research_pollSep 13, 2026, 11:00 AM UTCerror
research_pollSep 13, 2026, 10:30 AM UTCok
overlaySep 13, 2026, 10:00 AM UTCok
research_pollSep 13, 2026, 10:00 AM UTCok
research_pollSep 13, 2026, 09:30 AM UTCok
research_pollSep 13, 2026, 09:00 AM UTCok
overlaySep 13, 2026, 09:00 AM UTCok
research_pollSep 13, 2026, 08:30 AM UTCok
research_pollSep 13, 2026, 08:00 AM UTCok
overlaySep 13, 2026, 08:00 AM UTCok
research_pollSep 13, 2026, 07:30 AM UTCok
research_pollSep 13, 2026, 07:00 AM UTCok
overlaySep 13, 2026, 07:00 AM UTCok
research_pollSep 13, 2026, 06:30 AM UTCok
publicationSep 13, 2026, 06:30 AM UTCok
research_pollSep 13, 2026, 06:00 AM UTCok
overlaySep 13, 2026, 06:00 AM UTCok
fusionSep 13, 2026, 06:00 AM UTCok
research_pollSep 13, 2026, 05:30 AM UTCok
hf_hubSep 13, 2026, 05:00 AM UTCskipped
wulong_arenaSep 13, 2026, 05:00 AM UTCok
first_partySep 13, 2026, 05:00 AM UTCok
design_arenaSep 13, 2026, 05:00 AM UTCskipped
swe_rebenchSep 13, 2026, 05:00 AM UTCskipped
epoch_eciSep 13, 2026, 05:00 AM UTCok
terminal_benchSep 13, 2026, 05:00 AM UTCok
bfclSep 13, 2026, 05:00 AM UTCok
openrouter_dataSep 13, 2026, 05:01 AM UTCok
openrouterSep 13, 2026, 05:00 AM UTCok
lmarena_textSep 13, 2026, 05:00 AM UTCok
artificial_analysisTracked internally — artificialanalysis.aiSep 13, 2026, 05:01 AM UTCok
overlaySep 13, 2026, 05:00 AM UTCok
research_pollSep 13, 2026, 05:00 AM UTCok
research_pollSep 13, 2026, 04:30 AM UTCok
research_pollSep 13, 2026, 04:00 AM UTCok
overlaySep 13, 2026, 04:00 AM UTCok
research_pollSep 13, 2026, 03:30 AM UTCok
research_pollSep 13, 2026, 03:00 AM UTCok
overlaySep 13, 2026, 03:00 AM UTCok
research_pollSep 13, 2026, 02:30 AM UTCok
research_pollSep 13, 2026, 02:00 AM UTCok
overlaySep 13, 2026, 02:00 AM UTCok
research_pollSep 13, 2026, 01:30 AM UTCok
overlaySep 13, 2026, 01:00 AM UTCok
research_pollSep 13, 2026, 01:00 AM UTCok
research_pollSep 13, 2026, 12:30 AM UTCok
overlaySep 13, 2026, 12:00 AM UTCok
research_pollSep 13, 2026, 12:00 AM UTCok
research_pollSep 12, 2026, 11:30 PM UTCok
research_pollSep 12, 2026, 11:00 PM UTCok
overlaySep 12, 2026, 11:00 PM UTCok
research_pollSep 12, 2026, 10:30 PM UTCok
overlaySep 12, 2026, 10:00 PM UTCok
research_pollSep 12, 2026, 10:00 PM UTCok
research_pollSep 12, 2026, 09:30 PM UTCok
overlaySep 12, 2026, 09:00 PM UTCok
research_pollSep 12, 2026, 09:00 PM UTCok
research_pollSep 12, 2026, 08:30 PM UTCok
research_pollSep 12, 2026, 08:00 PM UTCok
overlaySep 12, 2026, 08:00 PM UTCok
research_pollSep 12, 2026, 07:30 PM UTCok
research_pollSep 12, 2026, 07:00 PM UTCok
overlaySep 12, 2026, 07:00 PM UTCok
research_pollSep 12, 2026, 06:30 PM UTCok
overlaySep 12, 2026, 06:00 PM UTCok
research_pollSep 12, 2026, 06:00 PM UTCok
research_pollSep 12, 2026, 05:30 PM UTCok
overlaySep 12, 2026, 05:00 PM UTCok
research_pollSep 12, 2026, 05:00 PM UTCok
research_pollSep 12, 2026, 04:30 PM UTCok
overlaySep 12, 2026, 04:00 PM UTCok
research_pollSep 12, 2026, 04:00 PM UTCok
overlaySep 12, 2026, 03:00 PM UTCok
research_pollSep 12, 2026, 03:00 PM UTCok
research_pollSep 12, 2026, 02:30 PM UTCok
research_pollSep 12, 2026, 02:00 PM UTCok
overlaySep 12, 2026, 02:00 PM UTCok
research_pollSep 12, 2026, 01:30 PM UTCok
overlaySep 12, 2026, 01:00 PM UTCok
research_pollSep 12, 2026, 01:00 PM UTCok
research_pollSep 12, 2026, 12:30 PM UTCok
overlaySep 12, 2026, 12:00 PM UTCok
research_pollSep 12, 2026, 12:00 PM UTCok
research_pollSep 12, 2026, 11:30 AM UTCok
research_pollSep 12, 2026, 11:00 AM UTCok
overlaySep 12, 2026, 11:00 AM UTCok
research_pollSep 12, 2026, 10:30 AM UTCok
research_pollSep 12, 2026, 10:00 AM UTCok
overlaySep 12, 2026, 10:00 AM UTCok
research_pollSep 12, 2026, 09:30 AM UTCok
overlaySep 12, 2026, 09:00 AM UTCok
research_pollSep 12, 2026, 09:00 AM UTCok
research_pollSep 12, 2026, 08:30 AM UTCok
overlaySep 12, 2026, 08:00 AM UTCok
research_pollSep 12, 2026, 08:00 AM UTCok
research_pollSep 12, 2026, 07:30 AM UTCok
overlaySep 12, 2026, 07:00 AM UTCok
research_pollSep 12, 2026, 07:00 AM UTCok
research_pollSep 12, 2026, 06:30 AM UTCok
publicationSep 12, 2026, 06:30 AM UTCok
overlaySep 12, 2026, 06:00 AM UTCok
research_pollSep 12, 2026, 06:00 AM UTCok
fusionSep 12, 2026, 06:00 AM UTCok
research_pollSep 12, 2026, 05:30 AM UTCok
overlaySep 12, 2026, 05:00 AM UTCok
swe_rebenchSep 12, 2026, 05:00 AM UTCskipped
first_partySep 12, 2026, 05:00 AM UTCok
wulong_arenaSep 12, 2026, 05:00 AM UTCok
epoch_eciSep 12, 2026, 05:00 AM UTCok
design_arenaSep 12, 2026, 05:00 AM UTCskipped
hf_hubSep 12, 2026, 05:00 AM UTCskipped
openrouter_dataSep 12, 2026, 05:01 AM UTCok
bfclSep 12, 2026, 05:00 AM UTCok
terminal_benchSep 12, 2026, 05:00 AM UTCok
openrouterSep 12, 2026, 05:00 AM UTCok
lmarena_textSep 12, 2026, 05:00 AM UTCok
artificial_analysisTracked internally — artificialanalysis.aiSep 12, 2026, 05:00 AM UTCok
research_pollSep 12, 2026, 05:00 AM UTCok
research_pollSep 12, 2026, 04:30 AM UTCok
overlaySep 12, 2026, 04:00 AM UTCok
research_pollSep 12, 2026, 04:00 AM UTCok
research_pollSep 12, 2026, 03:30 AM UTCok
research_pollSep 12, 2026, 03:00 AM UTCok
overlaySep 12, 2026, 03:00 AM UTCok
research_pollSep 12, 2026, 02:30 AM UTCok
research_pollSep 12, 2026, 02:00 AM UTCok
overlaySep 12, 2026, 02:00 AM UTCok
research_pollSep 12, 2026, 01:30 AM UTCok
research_pollSep 12, 2026, 01:00 AM UTCok

Frequently asked questions

Why not one ranking?

A model that is right for cheap classification is the wrong unit for multi-file coding. HiveBase publishes a verdict per work class. For code.agentic, code.frontend, and code.review the unit is (harness, model, role) — the same model on Claude Code versus Cursor is not the same system.

Why are Artificial Analysis scores missing?

Artificial Analysis is tracked internally for routing. Their terms do not allow us to rehost numeric scores. When AA is a source, the page says so and links to artificialanalysis.ai.

What does an auto-revert mean?

If a canary drops below HiveBase's quality floor, routing returns to the previous default. Public auto-reverts are listed in the changelog so a promotion that did not hold is visible, not buried.

Is this live production data?

When the public snapshot is reachable, the receipt reads live. If environment, table, or query fail, this page renders a labeled preview snapshot so the marketing site never depends on production being up.

Are first-party numbers a ranking?

No. HiveBase-owned dogfood outcomes appear separately with population, dates, denominator, and uncertainty. They are observational. Research-based defaults remain useful before that evidence accumulates. Existing eval consent does not authorize publishing customer traces.

Routing you can inspect.

HiveBase uses these work-class verdicts when it dispatches coding missions and the rest of the company. The same boundaries, receipts, and undo as everywhere else.