Live · Measured · Sourced

SkyMind Benchmarks

Three kinds of numbers, clearly separated: what we measure on our own live API, what our routing-pool backends publish, and where the paid frontier sits for honest context. No vibes, sources on every row.

SkyMind 2.0 — measured on the live API

CentroSky Internal Eval — 36 original items (reasoning · math · code), exact-match graded, run by tools/benchmark.php against production. Tracks our quality over time; not comparable to MMLU/GPQA.
Loading measured results…

How SkyMind 2.0 compares — GPQA Diamond

GPQA Diamond: 198 graduate-level science questions, the industry's standard hard-reasoning benchmark. One chart, three kinds of bars, each labeled: measured = we ran the public benchmark on our own live API. inherited = published score of the strongest backend in our routing pool. published = that vendor's own reported result.
Loading comparison…

The full picture — capability across benchmarks

One benchmark hides the point. SkyMind 2.0 routes per request, so on each axis its ceiling is the strongest backend for that axis — and different backends win different axes. Each SkyMind bar below is labeled with the pool model it inherits from. inherited until we publish a measured run for that axis.
Loading capability envelope…
Each axis: SkyMind 2.0's bar is the best published score among its routing-pool backends on that benchmark (backend named on the bar), compared against the frontier and well-known models on the same benchmark. Inherited from vendor-published numbers — the pool table below lists every backend and source. Where a measured SkyMind run exists for an axis, it replaces the inherited bar.

Live network — last 7 days

Real production traffic. The resolved-backend distribution is our routing audit, published — see exactly which models actually served requests.
Loading live metrics…

What SkyMind 2.0 routes to — published scores

SkyMind 2.0 is a health-aware router: its capability ceiling is the strongest live backend in its pool. These are the vendors' published results on their own benchmark runs, attributed per row — they are inherited context, not CentroSky measurements.
BackendMMLU-ProGPQA DiamondAIME25LiveCodeBenchSource

The paid frontier — for context

Published scores of the current frontier leaders. We show these because honesty ranks higher than ranking.
ModelHeadline scores (published)Source
Where we honestly stand: SkyMind 2.0 routes across frontier-class open models — published GPQA-class scores in the ~79–87 range — below the paid frontier's ~92–95, and at $0 in provider fees. What no one above publishes: the exact backend behind every response ("resolved" field), a public routing audit, and live measured latency. Transparency is our benchmark.

Methodology

Internal Eval — 36 original items written in-house (12 reasoning, 12 math, 12 code), answerable with a single letter/number/word and graded by exact match at temperature 0. Every item logs the resolved backend, so each run doubles as a routing audit. Run weekly via cron; history retained one year.

GPQA Diamond (measured) — the 198-question public benchmark run by tools/benchmark.php --gpqa against the live production API: exact-match graded, temperature 0, deterministic answer-order shuffle so runs are comparable. This is the one number on this page directly comparable to other models' published GPQA results. Until the first run publishes, the comparison chart shows an inherited bar instead — the published score of the strongest routing-pool backend, labeled as such.

Live metrics — aggregated from the last 7 days of production runs: request count, average latency, success rate, resolved-backend distribution, provider health.

Published scores — taken from vendor technical reports and independent leaderboards, with the source named on every row; rows marked are approximate and pending re-verification. We never present published backend scores as SkyMind's own results.