Live · Measured · Sourced
SkyMind Benchmarks
Three kinds of numbers, clearly separated: what we measure on our own live API, what our routing-pool backends publish, and where the paid frontier sits for honest context. No vibes, sources on every row.
SkyMind 2.0 — measured on the live API
CentroSky Internal Eval — 36 original items (reasoning · math · code), exact-match graded, run by tools/benchmark.php against production. Tracks our quality over time; not comparable to MMLU/GPQA.
Loading measured results…
How SkyMind 2.0 compares — GPQA Diamond
GPQA Diamond: 198 graduate-level science questions, the industry's standard hard-reasoning benchmark. One chart, three kinds of bars, each labeled: measured = we ran the public benchmark on our own live API. inherited = published score of the strongest backend in our routing pool. published = that vendor's own reported result.
The full picture — capability across benchmarks
One benchmark hides the point. SkyMind 2.0 routes per request, so on each axis its ceiling is the strongest backend for that axis — and different backends win different axes. Each SkyMind bar below is labeled with the pool model it inherits from. inherited until we publish a measured run for that axis.
Loading capability envelope…
Each axis: SkyMind 2.0's bar is the best published score among its routing-pool backends on that benchmark (backend named on the bar), compared against the frontier and well-known models on the same benchmark. Inherited from vendor-published numbers — the pool table below lists every backend and source. Where a measured SkyMind run exists for an axis, it replaces the inherited bar.
Live network — last 7 days
Real production traffic. The resolved-backend distribution is our routing audit, published — see exactly which models actually served requests.
What SkyMind 2.0 routes to — published scores
SkyMind 2.0 is a health-aware router: its capability ceiling is the strongest live backend in its pool. These are the vendors' published results on their own benchmark runs, attributed per row — they are inherited context, not CentroSky measurements.
| Backend | MMLU-Pro | GPQA Diamond | AIME25 | LiveCodeBench | Source |
The paid frontier — for context
Published scores of the current frontier leaders. We show these because honesty ranks higher than ranking.
| Model | Headline scores (published) | Source |
Where we honestly stand: SkyMind 2.0 routes across frontier-class open models — published GPQA-class scores in the ~79–87 range — below the paid frontier's ~92–95, and at $0 in provider fees. What no one above publishes: the exact backend behind every response ("resolved" field), a public routing audit, and live measured latency. Transparency is our benchmark.
Methodology
Internal Eval — 36 original items written in-house (12 reasoning, 12 math, 12 code), answerable with a single letter/number/word and graded by exact match at temperature 0. Every item logs the resolved backend, so each run doubles as a routing audit. Run weekly via cron; history retained one year.
GPQA Diamond (measured) — the 198-question public benchmark run by tools/benchmark.php --gpqa against the live production API: exact-match graded, temperature 0, deterministic answer-order shuffle so runs are comparable. This is the one number on this page directly comparable to other models' published GPQA results. Until the first run publishes, the comparison chart shows an inherited bar instead — the published score of the strongest routing-pool backend, labeled as such.
Live metrics — aggregated from the last 7 days of production runs: request count, average latency, success rate, resolved-backend distribution, provider health.
Published scores — taken from vendor technical reports and independent leaderboards, with the source named on every row; rows marked ≈ are approximate and pending re-verification. We never present published backend scores as SkyMind's own results.