Find the right model for the evidence, the workload, and the way you deploy.
Updated October 10, 2026 · Saved measurements · No pooled ranking across workloads
All saved models remain available. Earlier versions and clearly dominated API results may be hidden in the recommended view. Your selection overrides that default.
0measurements shown
No matching measurements. Change your filters or select models to compare.
EXPLORE THE TRADE-OFF
What matters most?
Each point is a saved measurement. Missing costs are omitted, never plotted as zero. Hover or focus a point for details.
SHARED EVIDENCE
One request. More questions.
1 · 10 · 25 · 50 · 100
Separate scaling workload: timings and quality here are not the headline accuracy benchmark. Native request points are connected; versions also use distinct line patterns. Split workflows use separate diamonds.
Measured 100-question workflows
Clef's native curve stops at 50. Its measured 100-question workflow uses two sequential calls. Only selected models with saved scaling measurements appear. Historical Ling runs are not pooled here. Costs show reported charges, excluding uncertain failed-call charges. Microsoft had three successfully retried failures; its run reserved $0.01813 for uncertain charges.
Cost is USD. Local operating cost is unmeasured. Open weights and a local runtime are separate properties; capabilities reflect saved research, not a fresh catalog audit.
LATEST ADDITIONS
METHODOLOGY & HISTORY
The comparison has boundaries.
Compare the same questions
Image Inputs shows measured image workloads. General Text includes choices, ordinal scores and booleans. Yes/No compares identical boolean questions from general and binary models. Workloads keep separate denominators.
Evidence comes before action
Models surface evidence. Procedural underwriting rules determine the resulting action. A lack of supporting evidence does not establish an exposure's absence.
Keep failures visible
Missing outputs, abstentions and errors have different meanings. Confidence is not automatically calibrated. Reported charges may omit unknown failed-call charges, which stay reserved in original runs.
Preserve every version
Earlier versions remain in the model picker. Perplexity v1 and v1.1 returned identical observed answers and distributions in saved paired tests; the serving cause is unresolved. Distinct weights are not independently confirmed.
Sources, hardware, and preview limitations
This comparison uses saved aggregate summaries with SHA-256 source hashes in the download. A full raw-ledger audit remains pending. Local smoke results use AMD RX 6800 XT; precision appears in model labels where recorded. Memory is peak PyTorch allocation. Decider's eager reference runtime is not optimized serving. Image labels and development labels have the qualifications shown for each workload.
October 9: combined workload views, model selection, availability filters and version history. Includes the October 9 three-model accuracy evaluation and Microsoft scaling run, plus October 10 paid Mercury accuracy and scaling. Insurance evidence benchmarks.