The current board — every model run on the same 32 tasks, graded the same way, and costed on what it actually spent to get work accepted. The 6 models at the top are doing work we cannot tell apart, at a 73× spread in price.
Price is what you pay for tokens. Cost is what you pay for mistakes.
The benchmark measures what an attempt costs and how often it is accepted. Everything else that decides your bill is yours: volume, review time, what a wrong answer costs when it reaches a customer. Supply those and the ranking can change — which is the point.
Retries assume a failed task fails again, because that is what we measured — asking a second time usually returns the same wrong answer at twice the price. Why that is, and when it is not true →
Success rates and cost per attempt are measured on the benchmark tasks shown below. Volume, review time, error cost, detection rate and retry behaviour are assumptions you supplied. Totals combine both and are estimates, not quotes.
Everything below is how that recommendation was reached: what each model cost, how often a grader accepted its output, and where the numbers are too thin to act on. The Octane score explains the ranking; it is not the product.
These cards reflect performance on the specific benchmark tasks listed under “What we test” — they are starting points for a model evaluation, not blanket recommendations for an entire industry. The badge is the observed success rate on those tasks, with the run counts beneath it; not a claim about production reliability. Each card leads with the highest-scoring model above the 80% threshold, then names a lower-cost option that also cleared it. Where no model reached the threshold, the card says so instead of recommending one.
A pricing page sells you a rate per million tokens. That rate is one of three terms, and it is the only one published. A model can advertise the lowest rate, write two or three times as many tokens to answer the same question, fail a share of the time — and finish as the most expensive option here.
Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs
Each bar adds one thing the price page does not tell you. The faintest is the advertised rate on this exact workload, priced as if the model were as concise as the leanest one measured. Then the tokens it really wrote. Then the attempts you paid for and could not use. Then your actual work mix, since the index weights every vertical equally instead of letting a model's weak vertical be diluted by its strong one. Only the faintest bar is visible when you pick a model.
The leading model in each panel sets 100; every other grade is its share of that leader's successful work per dollar. Log axis — a real panel spans orders of magnitude. Whiskers are 95% clustered bootstrap intervals, resampled by vertical, then task, then repetition. Raw baseline-relative scores are in the CSV and JSON exports.
gpt-oss 120B ranks first overall because its lead in Customer Service outweighs DeepSeek V3.2's lead in Banking & Finance, Software Engineering under the benchmark's equal-weight methodology — each vertical counts the same regardless of how many tasks or attempts it holds. Weight the verticals to your own workload in the calculator above and the ranking can change.
Every panel shares one axis, so they are directly comparable. This is the cut that matters — cost-effectiveness is a property of a model on a kind of work, not of a model alone.
A hollow, struck-through dot marks a model that recorded below 80% observed success on that vertical. It is excluded from the recommendation, not judged unusable — cost per successful task already prices retries, but it assumes a failure is free to detect and that retrying eventually works. Neither held in these measurements: most model-task pairs returned byte-identical text on every repetition, so a failed attempt repeated rather than varied.
Observed success against cost per successful task. Bottom-right is the good corner: high observed success at low cost. Position carries the reading, not colour — every point is directly labelled.
Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs
Total billed cost includes input tokens, output tokens, cached tokens where applicable, reasoning tokens where separately billed, failed attempts, and retries where benchmarked. Work you paid for and could not use stays in the numerator.
Full methodology, grader definitions, confidence rules and limitations →
Every task behind these numbers, with how it is graded and what it weighs. Full prompts are deliberately unpublished — printed prompts end up in training data and the benchmark decays — but per-task results are in the CSV and JSON exports, so any figure on this page traces to graded attempts.
| Task | Graded by | Weight | Version |
|---|---|---|---|
| Accrue interest under a stated day-count convention | exact answer | 2 | v1 |
| Convert a nominal rate to an effective annual rate | structured output, field-checked | 1.5 | v1 |
| Refuse to settle an FX conversion at the wrong date's rate | exact answer | 2 | v1 |
| Assign a KYC risk tier | exact answer | 1 | v1 |
| Allocate a partial payment down a contractual waterfall | structured output, field-checked | 1.5 | v1 |
| Extract structured fields from a bank statement line | structured output, field-checked | 1 | v1 |
| Task | Graded by | Weight | Version |
|---|---|---|---|
| Apply a return policy where the exception overrides the deadline | exact answer | 2 | v1 |
| Apply a multi-condition escalation rule | exact answer | 1.5 | v1 |
| Compute a prorated refund on cancellation | structured output, field-checked | 1.5 | v1 |
| Decide refund eligibility against a policy | structured output, field-checked | 1 | v1 |
| Classify a support ticket into a routing intent | exact answer | 1 | v1 |
| Task | Graded by | Weight | Version |
|---|---|---|---|
| Resolve conflicting governing-law clauses via a precedence rule | structured output, field-checked | 1.5 | v1 |
| Extract governing law and venue from a contract clause | structured output, field-checked | 1 | v1 |
| Apply a liability cap with a carve-out and an absolute exclusion | exact answer | 2 | v1 |
| Compute a notice expiry with exclusion and weekend roll-forward | structured output, field-checked | 1.5 | v1 |
| Earliest termination date under an initial-term bar | exact answer | 2 | v1 |
| Task | Graded by | Weight | Version |
|---|---|---|---|
| Apply a drug interaction only when its stated condition holds | exact answer | 2 | v1 |
| Identify the ICD-10 code for a described condition | exact answer | 1 | v2 |
| Convert a weight-based infusion order into a pump rate | exact answer | 2 | v1 |
| Extract a drug interaction into structured form | structured output, field-checked | 1 | v1 |
| Apply a maximum-dose ceiling before converting to volume | exact answer | 2 | v1 |
| Refuse to fill a weight-based order when the weight is missing | exact answer | 2 | v1 |
| Convert a weight-based dose into a dispensing volume | structured output, field-checked | 1.5 | v1 |
| Interpret a lab value against a sex-specific reference range | exact answer | 1.5 | v1 |
| Adjust a dose for renal function using Cockcroft-Gault | exact answer | 2 | v1 |
| Task | Graded by | Weight | Version |
|---|---|---|---|
| Add months to a date, clamping to the end of the month | code executed against tests | 2 | v1 |
| Implement order-preserving deduplication | code executed against tests | 1.5 | v1 |
| Predict the output of an order-preserving deduplication | exact answer | 1 | v1 |
| Merge overlapping intervals, including touching endpoints | code executed against tests | 1.5 | v1 |
| Count divisible pairs at a scale that punishes brute force | code executed against tests | 2 | v1 |
| Parse durations, honouring a buried error-handling contract | code executed against tests | 1.5 | v1 |
| Subtract one set of half-open intervals from another | code executed against tests | 2 | v1 |
This public basket is 23 short, single-shot tasks. Your work is not these tasks. The same machinery runs against representative samples of your own workload — extraction from your documents, your support tickets, your code — and returns model comparison, total-cost analysis under your review and error costs, observed success by task type, routing recommendations, and the caveats that apply.
Useful for model selection, vendor evaluation, cost optimisation, routing design and procurement support. Get in touch to benchmark your workload →
Task-level rows are omitted here for length. The CSV and JSON exports carry them in full.
| Scope | Vertical | Model | Observed success | Cost per successful task | Cost 95% CI | Octane | Octane 95% CI | Total billed cost |
|---|---|---|---|---|---|---|---|---|
| overall | — | gpt-oss 120B | 91/96 · 95% | $0.0000677 | $0.0000299–$0.00013 | 100 | — | $0.00540 |
| overall | — | DeepSeek V3.2 | 81/96 · 84% | $0.0000716 | — | 94.6 | 29.5–178 | $0.00460 |
| overall | — | GPT-5.6 Luna | 95/96 · 99% | $0.0000775 | — | 87.4 | 54.5–147 | $0.00718 |
| overall | — | Gemini 2.5 Flash | 87/96 · 91% | $0.00029 | — | 23.4 | 14.0–56.0 | $0.0246 |
| overall | — | Llama 4 Maverick | 50/96 · 52% | $0.00041 | — | 16.6 | 4.27–41.6 | $0.0137 |
| overall | — | DeepSeek V4 Flash | 84/96 · 88% | $0.00043 | — | 15.6 | 4.91–108 | $0.0277 |
| overall | — | GPT-5.4 mini | 80/96 · 83% | $0.00043 | — | 15.6 | 7.07–34.9 | $0.0249 |
| overall | — | GPT-5.6 Luna Pro | 96/96 · 100% | $0.00047 | — | 14.5 | 9.71–24.0 | $0.0439 |
| overall | — | GPT-5.6 Terra | 93/96 · 97% | $0.00068 | — | 9.89 | 5.78–17.6 | $0.0618 |
| overall | — | Claude Haiku 4.5 | 75/96 · 78% | $0.00081 | — | 8.32 | 2.60–16.6 | $0.0466 |
| overall | — | GLM-4.7 Flash | 77/96 · 80% | $0.00145 | — | 4.66 | 1.63–15.9 | $0.0884 |
| overall | — | Claude Sonnet 4.6 | 92/96 · 96% | $0.00190 | — | 3.56 | 1.97–5.74 | $0.1609 |
| overall | — | GLM-5.2 | 93/96 · 97% | $0.00233 | — | 2.90 | 1.51–5.70 | $0.2078 |
| overall | — | Qwen3.5 Flash | 53/96 · 55% | $0.00291 | — | 2.33 | 1.30–20.1 | $0.0514 |
| overall | — | Claude Opus 4.6 | 92/96 · 96% | $0.00335 | — | 2.02 | 1.11–3.44 | $0.3005 |
| overall | — | GPT-5.6 Sol | 93/96 · 97% | $0.00342 | — | 1.98 | 1.11–3.82 | $0.3080 |
| overall | — | Kimi K2.5 | 90/96 · 94% | $0.00438 | — | 1.55 | 1.02–2.72 | $0.3362 |
| overall | — | Kimi K3 | 95/96 · 99% | $0.00565 | — | 1.20 | 0.78–1.88 | $0.5145 |
| overall | — | Qwen3.8 Max | 91/96 · 95% | $0.00899 | — | 0.75 | 0.26–2.90 | $0.6465 |
| overall | — | GLM-5 | 80/96 · 83% | $0.0105 | — | 0.64 | 0.20–2.39 | $0.5287 |
| vertical | banking | Claude Haiku 4.5 | 18/18 · 100% | $0.00020 | — | 16.5 | 8.94–29.1 | $0.00364 |
| vertical | banking | Claude Opus 4.6 | 15/18 · 83% | $0.00289 | — | 1.12 | 0.46–1.83 | $0.0496 |
| vertical | banking | Claude Sonnet 4.6 | 18/18 · 100% | $0.00086 | — | 3.75 | 2.94–4.97 | $0.0158 |
| vertical | banking | DeepSeek V3.2 | 18/18 · 100% | $0.0000322 | — | 100 | 55.6–190 | $0.00058 |
| vertical | banking | DeepSeek V4 Flash | 18/18 · 100% | $0.0000358 | — | 90.1 | 52.6–142 | $0.00068 |
| vertical | banking | Gemini 2.5 Flash | 18/18 · 100% | $0.0000827 | — | 39.0 | 18.9–77.6 | $0.00160 |
| vertical | banking | GLM-4.7 Flash | 17/18 · 94% | $0.00051 | — | 6.28 | 1.67–17.0 | $0.0111 |
| vertical | banking | GLM-5 | 18/18 · 100% | $0.00205 | — | 1.57 | 0.93–2.64 | $0.0373 |
| vertical | banking | GLM-5.2 | 18/18 · 100% | $0.00093 | — | 3.46 | 1.99–5.19 | $0.0182 |
| vertical | banking | GPT-5.4 mini | 18/18 · 100% | $0.00013 | — | 24.6 | 13.7–43.3 | $0.00244 |
| vertical | banking | GPT-5.6 Luna | 18/18 · 100% | $0.0000352 | — | 91.5 | 58.0–137 | $0.00065 |
| vertical | banking | GPT-5.6 Luna Pro | 18/18 · 100% | $0.00030 | — | 10.8 | 6.99–16.7 | $0.00539 |
| vertical | banking | GPT-5.6 Sol | 18/18 · 100% | $0.00090 | — | 3.58 | 1.95–6.49 | $0.0167 |
| vertical | banking | GPT-5.6 Terra | 18/18 · 100% | $0.00026 | — | 12.6 | 9.22–17.2 | $0.00454 |
| vertical | banking | gpt-oss 120B | 18/18 · 100% | $0.0000421 | $0.0000262–$0.0000689 | 76.6 | — | $0.00075 |
| vertical | banking | Kimi K2.5 | 18/18 · 100% | $0.00147 | — | 2.20 | 1.11–4.00 | $0.0277 |
| vertical | banking | Kimi K3 | 18/18 · 100% | $0.00364 | — | 0.89 | 0.60–1.27 | $0.0684 |
| vertical | banking | Llama 4 Maverick | 9/18 · 50% | $0.00044 | — | 7.34 | 1.58–22.8 | $0.00273 |
| vertical | banking | Qwen3.5 Flash | 11/18 · 61% | $0.00048 | — | 6.76 | 2.09–12.8 | $0.00586 |
| vertical | banking | Qwen3.8 Max | 18/18 · 100% | $0.00151 | — | 2.14 | 1.70–2.73 | $0.0266 |
| vertical | coding | Claude Haiku 4.5 | 21/21 · 100% | $0.00121 | — | 8.08 | 5.39–13.7 | $0.0233 |
| vertical | coding | Claude Opus 4.6 | 21/21 · 100% | $0.00588 | — | 1.66 | 1.14–2.80 | $0.1123 |
| vertical | coding | Claude Sonnet 4.6 | 21/21 · 100% | $0.00339 | — | 2.88 | 1.99–5.55 | $0.0647 |
| vertical | coding | DeepSeek V3.2 | 21/21 · 100% | $0.0000978 | — | 100 | 67.8–194 | $0.00189 |
| vertical | coding | DeepSeek V4 Flash | 15/21 · 71% | $0.00178 | — | 5.48 | 1.81–25.0 | $0.0231 |
| vertical | coding | Gemini 2.5 Flash | 21/21 · 100% | $0.00090 | — | 10.9 | 6.80–18.8 | $0.0173 |
| vertical | coding | GLM-4.7 Flash | 14/21 · 67% | $0.00458 | — | 2.13 | 0.67–7.90 | $0.0575 |
| vertical | coding | GLM-5 | 11/21 · 52% | $0.0408 | — | 0.24 | 0.08–0.64 | $0.3577 |
| vertical | coding | GLM-5.2 | 21/21 · 100% | $0.00746 | — | 1.31 | 0.73–3.45 | $0.1374 |
| vertical | coding | GPT-5.4 mini | 21/21 · 100% | $0.00071 | — | 13.8 | 9.18–23.3 | $0.0137 |
| vertical | coding | GPT-5.6 Luna | 21/21 · 100% | $0.00021 | — | 47.0 | 24.0–98.9 | $0.00413 |
| vertical | coding | GPT-5.6 Luna Pro | 21/21 · 100% | $0.00106 | — | 9.26 | 5.39–18.3 | $0.0208 |
| vertical | coding | GPT-5.6 Sol | 21/21 · 100% | $0.0102 | — | 0.96 | 0.53–2.74 | $0.1991 |
| vertical | coding | GPT-5.6 Terra | 21/21 · 100% | $0.00197 | — | 4.97 | 2.83–12.6 | $0.0383 |
| vertical | coding | gpt-oss 120B | 19/21 · 90% | $0.00013 | $0.0000673–$0.00023 | 72.6 | — | $0.00225 |
| vertical | coding | Kimi K2.5 | 21/21 · 100% | $0.0102 | — | 0.96 | 0.52–2.20 | $0.1844 |
| vertical | coding | Kimi K3 | 21/21 · 100% | $0.0121 | — | 0.81 | 0.36–1.89 | $0.2484 |
| vertical | coding | Llama 4 Maverick | 16/21 · 76% | $0.00016 | — | 62.1 | 29.1–115 | $0.00236 |
| vertical | coding | Qwen3.5 Flash | 3/21 · 14% | $0.0116 | — | 0.84 | 0.48–4.42 | $0.0202 |
| vertical | coding | Qwen3.8 Max | 18/21 · 86% | $0.0291 | — | 0.34 | 0.10–2.22 | $0.4374 |
| vertical | customer_service | Claude Haiku 4.5 | 6/15 · 40% | $0.00147 | — | 2.05 | 0.65–6.90 | $0.00604 |
| vertical | customer_service | Claude Opus 4.6 | 15/15 · 100% | $0.00290 | — | 1.03 | 0.71–2.60 | $0.0412 |
| vertical | customer_service | Claude Sonnet 4.6 | 15/15 · 100% | $0.00154 | — | 1.95 | 1.39–4.13 | $0.0219 |
| vertical | customer_service | DeepSeek V3.2 | 15/15 · 100% | $0.0000372 | — | 80.6 | 61.5–122 | $0.00054 |
| vertical | customer_service | DeepSeek V4 Flash | 15/15 · 100% | $0.0000383 | — | 78.2 | 68.4–92.4 | $0.00054 |
| vertical | customer_service | Gemini 2.5 Flash | 15/15 · 100% | $0.0000604 | — | 49.7 | 36.8–59.0 | $0.00090 |
| vertical | customer_service | GLM-4.7 Flash | 15/15 · 100% | $0.00020 | — | 14.8 | 12.9–17.3 | $0.00294 |
| vertical | customer_service | GLM-5 | 15/15 · 100% | $0.00179 | — | 1.68 | 1.24–2.08 | $0.0259 |
| vertical | customer_service | GLM-5.2 | 15/15 · 100% | $0.00072 | — | 4.19 | 2.99–5.41 | $0.0111 |
| vertical | customer_service | GPT-5.4 mini | 13/15 · 87% | $0.00016 | — | 18.3 | 13.1–21.4 | $0.00206 |
| vertical | customer_service | GPT-5.6 Luna | 15/15 · 100% | $0.000032 | — | 93.8 | 66.8–143 | $0.00047 |
| vertical | customer_service | GPT-5.6 Luna Pro | 15/15 · 100% | $0.00025 | — | 12.0 | 9.04–16.9 | $0.00369 |
| vertical | customer_service | GPT-5.6 Sol | 15/15 · 100% | $0.00139 | — | 2.16 | 1.50–3.13 | $0.0201 |
| vertical | customer_service | GPT-5.6 Terra | 15/15 · 100% | $0.00030 | — | 9.94 | 6.71–15.7 | $0.00435 |
| vertical | customer_service | gpt-oss 120B | 15/15 · 100% | $0.00003 | $0.0000222–$0.0000387 | 100 | — | $0.00043 |
| vertical | customer_service | Kimi K2.5 | 15/15 · 100% | $0.00169 | — | 1.77 | 1.19–2.76 | $0.0243 |
| vertical | customer_service | Kimi K3 | 15/15 · 100% | $0.00231 | — | 1.30 | 1.07–1.61 | $0.0334 |
| vertical | customer_service | Llama 4 Maverick | 9/15 · 60% | $0.00041 | — | 7.23 | 1.57–18.6 | $0.00283 |
| vertical | customer_service | Qwen3.5 Flash | 9/15 · 60% | $0.00048 | — | 6.24 | 1.46–17.0 | $0.00442 |
| vertical | customer_service | Qwen3.8 Max | 15/15 · 100% | $0.00128 | — | 2.34 | 1.83–2.95 | $0.0183 |
| vertical | legal | Claude Haiku 4.5 | 6/15 · 40% | $0.00077 | — | 9.84 | 3.68–15.3 | $0.00448 |
| vertical | legal | Claude Opus 4.6 | 14/15 · 93% | $0.00318 | — | 2.39 | 1.32–5.20 | $0.0480 |
| vertical | legal | Claude Sonnet 4.6 | 11/15 · 73% | $0.00257 | — | 2.96 | 0.99–6.59 | $0.0299 |
| vertical | legal | DeepSeek V3.2 | 5/15 · 33% | $0.00015 | — | 50.6 | 9.29–78.8 | $0.00075 |
| vertical | legal | DeepSeek V4 Flash | 9/15 · 60% | $0.00028 | — | 26.8 | 9.48–76.1 | $0.00269 |
| vertical | legal | Gemini 2.5 Flash | 6/15 · 40% | $0.00029 | — | 26.6 | 10.1–42.9 | $0.00169 |
| vertical | legal | GLM-4.7 Flash | 6/15 · 40% | $0.00172 | — | 4.42 | 0.80–16.6 | $0.0107 |
| vertical | legal | GLM-5 | 10/15 · 67% | $0.00612 | — | 1.24 | 0.56–2.52 | $0.0634 |
| vertical | legal | GLM-5.2 | 12/15 · 80% | $0.00187 | — | 4.07 | 2.75–5.20 | $0.0223 |
| vertical | legal | GPT-5.4 mini | 4/15 · 27% | $0.00102 | — | 7.49 | 2.64–15.2 | $0.00322 |
| vertical | legal | GPT-5.6 Luna | 14/15 · 93% | $0.000076 | — | 100 | 69.2–155 | $0.00101 |
| vertical | legal | GPT-5.6 Luna Pro | 15/15 · 100% | $0.00044 | — | 17.4 | 10.5–28.6 | $0.00629 |
| vertical | legal | GPT-5.6 Sol | 12/15 · 80% | $0.00326 | — | 2.33 | 1.77–2.96 | $0.0385 |
| vertical | legal | GPT-5.6 Terra | 12/15 · 80% | $0.00060 | — | 12.7 | 8.21–16.8 | $0.00704 |
| vertical | legal | gpt-oss 120B | 12/15 · 80% | $0.00010 | $0.0000486–$0.00019 | 74.2 | — | $0.00121 |
| vertical | legal | Kimi K2.5 | 10/15 · 67% | $0.00681 | — | 1.12 | 0.70–2.99 | $0.0603 |
| vertical | legal | Kimi K3 | 14/15 · 93% | $0.00825 | — | 0.92 | 0.56–2.02 | $0.1109 |
| vertical | legal | Llama 4 Maverick | 4/15 · 27% | $0.00063 | — | 12.1 | 3.03–23.7 | $0.00237 |
| vertical | legal | Qwen3.5 Flash | 9/15 · 60% | $0.00163 | — | 4.67 | 2.37–8.17 | $0.0143 |
| vertical | legal | Qwen3.8 Max | 13/15 · 87% | $0.0119 | — | 0.64 | 0.30–3.03 | $0.1341 |
| vertical | medical | Claude Haiku 4.5 | 24/27 · 89% | $0.00042 | — | 6.44 | 3.68–13.6 | $0.00919 |
| vertical | medical | Claude Opus 4.6 | 27/27 · 100% | $0.00191 | — | 1.42 | 0.82–3.17 | $0.0493 |
| vertical | medical | Claude Sonnet 4.6 | 27/27 · 100% | $0.00115 | — | 2.37 | 1.70–4.21 | $0.0286 |
| vertical | medical | DeepSeek V3.2 | 22/27 · 81% | $0.0000404 | — | 67.3 | 46.0–89.3 | $0.00084 |
| vertical | medical | DeepSeek V4 Flash | 27/27 · 100% | $0.0000272 | — | 100 | 77.9–125 | $0.00072 |
| vertical | medical | Gemini 2.5 Flash | 27/27 · 100% | $0.00012 | — | 22.4 | 11.3–59.7 | $0.00309 |
| vertical | medical | GLM-4.7 Flash | 25/27 · 93% | $0.00025 | — | 10.9 | 8.35–12.6 | $0.00625 |
| vertical | medical | GLM-5 | 26/27 · 96% | $0.00178 | — | 1.53 | 1.26–1.88 | $0.0445 |
| vertical | medical | GLM-5.2 | 27/27 · 100% | $0.00068 | — | 3.99 | 3.05–5.20 | $0.0188 |
| vertical | medical | GPT-5.4 mini | 24/27 · 89% | $0.00015 | — | 18.2 | 10.9–25.5 | $0.00349 |
| vertical | medical | GPT-5.6 Luna | 27/27 · 100% | $0.0000358 | — | 76.1 | 63.3–90.1 | $0.00092 |
| vertical | medical | GPT-5.6 Luna Pro | 27/27 · 100% | $0.00030 | — | 9.18 | 7.23–11.2 | $0.00772 |
| vertical | medical | GPT-5.6 Sol | 27/27 · 100% | $0.00131 | — | 2.07 | 1.77–2.49 | $0.0336 |
| vertical | medical | GPT-5.6 Terra | 27/27 · 100% | $0.00030 | — | 9.09 | 7.43–11.2 | $0.00764 |
| vertical | medical | gpt-oss 120B | 27/27 · 100% | $0.0000293 | $0.0000209–$0.0000394 | 92.8 | — | $0.00076 |
| vertical | medical | Kimi K2.5 | 26/27 · 96% | $0.00170 | — | 1.60 | 1.05–2.73 | $0.0395 |
| vertical | medical | Kimi K3 | 27/27 · 100% | $0.00195 | — | 1.39 | 1.09–1.76 | $0.0534 |
| vertical | medical | Llama 4 Maverick | 12/27 · 44% | $0.00040 | — | 6.86 | 1.58–25.7 | $0.00345 |
| vertical | medical | Qwen3.5 Flash | 21/27 · 78% | $0.00031 | — | 8.64 | 5.40–13.1 | $0.00667 |
| vertical | medical | Qwen3.8 Max | 27/27 · 100% | $0.00115 | — | 2.37 | 1.98–2.75 | $0.0301 |