InsureBenchInsurance evidence benchmarksMethodology ↗
INSURANCE EVIDENCE BENCHMARKS

Decision Models.

Find the right model for the evidence, the workload, and the way you deploy.

Updated October 10, 2026 · Saved measurements · No pooled ranking across workloads

0measurements shown

EXPLORE THE TRADE-OFF

What matters most?

Each point is a saved measurement. Missing costs are omitted, never plotted as zero. Hover or focus a point for details.

SHARED EVIDENCE

One request. More questions.

1 · 10 · 25 · 50 · 100

Separate scaling workload: timings and quality here are not the headline accuracy benchmark. Native request points are connected; versions also use distinct line patterns. Split workflows use separate diamonds.

Measured 100-question workflows

Clef's native curve stops at 50. Its measured 100-question workflow uses two sequential calls. Only selected models with saved scaling measurements appear. Historical Ling runs are not pooled here. Costs show reported charges, excluding uncertain failed-call charges. Microsoft had three successfully retried failures; its run reserved $0.01813 for uncertain charges.

THE NUMBERS

Results you can inspect

Download aggregates ↓
Request / deploymentNotes

Cost is USD. Local operating cost is unmeasured. Open weights and a local runtime are separate properties; capabilities reflect saved research, not a fresh catalog audit.

LATEST ADDITIONS
METHODOLOGY & HISTORY

The comparison has boundaries.

Compare the same questions

Image Inputs shows measured image workloads. General Text includes choices, ordinal scores and booleans. Yes/No compares identical boolean questions from general and binary models. Workloads keep separate denominators.

Evidence comes before action

Models surface evidence. Procedural underwriting rules determine the resulting action. A lack of supporting evidence does not establish an exposure's absence.

Keep failures visible

Missing outputs, abstentions and errors have different meanings. Confidence is not automatically calibrated. Reported charges may omit unknown failed-call charges, which stay reserved in original runs.

Preserve every version

Earlier versions remain in the model picker. Perplexity v1 and v1.1 returned identical observed answers and distributions in saved paired tests; the serving cause is unresolved. Distinct weights are not independently confirmed.

Sources, hardware, and preview limitations

This comparison uses saved aggregate summaries with SHA-256 source hashes in the download. A full raw-ledger audit remains pending. Local smoke results use AMD RX 6800 XT; precision appears in model labels where recorded. Memory is peak PyTorch allocation. Decider's eager reference runtime is not optimized serving. Image labels and development labels have the qualifications shown for each workload.

October 9: combined workload views, model selection, availability filters and version history. Includes the October 9 three-model accuracy evaluation and Microsoft scaling run, plus October 10 paid Mercury accuracy and scaling. Insurance evidence benchmarks.