LLM ontology matching - benchmark results (human-mouse)

Generated 2026-07-13 · focus subset human-mouse · metrics & error rates recomputed from raw predictions

86
models
15
families
3
datasets / 9 subsets
3792
experiment runs
0.774
best Youden's index (human-mouse)

What this is

We benchmark LLMs as binary judges for ontology entity matching. Each model is run over candidate pairs under several completion configurations parsed from the experiment name (reasoning effort, thinking budget, few-shot count, prompt). The leaderboard is regenerated directly from the results on disk, so every table and figure below stays in sync with the data.

Leaderboard - ranked by Youden's index (best config per model)

Core metric: Youden's index (sensitivity + specificity − 1). Reliable = fail_rate ≤ 5% and ≥ 30 predictions. fail_rate = fraction of calls that errored; mean_num_retries = average retries per call (shown where the provider logged it). Drill-down in the explorer at the bottom.

model family config prompt Youden's index F1 Score fail_rate mean_num_retries cost_total_$
google|gemini-2.5-pro google thinking_budget=1024 direct_entity 0.774 0.924 0.000 0.000 1.602
qwen|qwen3.5-35b-a3b qwen base direct_entity 0.753 0.923 0.000
deepseek|deepseek-v3.2 deepseek thinking_budget=0 direct_entity 0.752 0.919 0.000 0.000 0.064
google|gemini-3.1-flash-lite-preview google reasoning_effort=medium direct_entity 0.750 0.927 0.000 2.811 0.076
qwen|qwen3.5-122b-a10b qwen base direct_entity 0.742 0.920 0.000
google|gemini-2.5-flash google thinking_budget=1024 direct_entity 0.741 0.916 0.000 0.000 0.223
mistralai|mistral-medium-3-5 mistralai reasoning_effort=high direct_entity 0.737 0.920 0.000 4.000 0.700
deepseek|deepseek-v4-flash deepseek thinking_budget=1024 direct_entity 0.737 0.920 0.000 0.075 0.007
z-ai|glm-5 z-ai base direct_entity 0.735 0.913 0.000
anthropic|claude-sonnet-4.6 anthropic base direct_entity 0.734 0.924 0.000
openai|gpt-5-nano openai reasoning_effort=medium direct_entity 0.730 0.897 0.000
openai|gpt-5.5 openai reasoning_effort=medium direct_entity_with_synonyms 0.728 0.921 0.000 0.004 1.197
google|gemini-2.5-flash-lite google thinking_budget=512 direct_entity 0.726 0.918 0.000 0.012 0.032
bytedance-seed|seed-2.0-mini bytedance-seed base direct_entity 0.726 0.885 0.000 0.000 0.059
x-ai|grok-4.3 x-ai base direct_entity 0.724 0.910 0.000 0.000 0.446

Cost is per full run of the subset, from real token usage × live OpenRouter pricing.

Recommended model selection

Ranked by Youden's index (balanced sensitivity + specificity). Ten model-config pairs balancing quality, value (quality per $), reasoning need and provider diversity - not just the top scores. gpt-5.5 excluded as too expensive for the quality gained.

Cost vs performance for the proposed models (gold, with their Pareto front) against all other models (grey, with a secondary dashed front). Use the input to change the fail-rate reliability threshold and watch both frontiers shift.

Max fail rate (reliability threshold):
model family config prompt Youden's index F1 Score fail_rate mean_num_retries cost_total_$ rationale
google|gemini-2.5-pro google thinking_budget=1024 direct_entity 0.774 0.924 0.000 0.000 1.602 Top Youden's index
qwen|qwen3.5-35b-a3b qwen base direct_entity 0.753 0.923 0.000 2nd best Youden's index
deepseek|deepseek-v4-flash deepseek thinking_budget=1024 direct_entity 0.737 0.920 0.000 0.075 0.007 Cheapest among above-median quality (best value)
mistralai|mistral-medium-3-5 mistralai reasoning_effort=high direct_entity 0.737 0.920 0.000 4.000 0.700 Family diversity (mistralai)
z-ai|glm-5 z-ai base direct_entity 0.735 0.913 0.000 Family diversity (z-ai)
anthropic|claude-sonnet-4.6 anthropic base direct_entity 0.734 0.924 0.000 Family diversity (anthropic)
openai|gpt-5-nano openai reasoning_effort=medium direct_entity 0.730 0.897 0.000 Family diversity (openai)
bytedance-seed|seed-2.0-mini bytedance-seed base direct_entity 0.726 0.885 0.000 0.000 0.059 Family diversity (bytedance-seed)
x-ai|grok-4.3 x-ai base direct_entity 0.724 0.910 0.000 0.000 0.446 Family diversity (x-ai)
minimax|minimax-m2.5 minimax base direct_entity 0.691 0.902 0.000 0.008 0.085 Family diversity (minimax)

Honorable mentions (also usable / backups)

Strong reliable models just outside the selection - handy as cheaper alternatives or drop-in backups.

model family config prompt Youden's index F1 Score fail_rate mean_num_retries cost_total_$ note
deepseek|deepseek-v3.2 deepseek thinking_budget=0 direct_entity 0.752 0.919 0.000 0.000 0.064 strong backup
google|gemini-3.1-flash-lite-preview google reasoning_effort=medium direct_entity 0.750 0.927 0.000 2.811 0.076 strong backup
qwen|qwen3.5-122b-a10b qwen base direct_entity 0.742 0.920 0.000 strong backup
google|gemini-2.5-flash google thinking_budget=1024 direct_entity 0.741 0.916 0.000 0.000 0.223 strong backup
google|gemini-2.5-flash-lite google thinking_budget=512 direct_entity 0.726 0.918 0.000 0.012 0.032 strong backup
x-ai|grok-4.1-fast x-ai base direct_entity 0.721 0.918 0.000 strong backup

Best model per provider family

Cost, speed & reliability trade-offs

Failure rate per model (red = unreliable, > 5%). These are the models that need re-running.

Average retries per call (logged for OpenRouter runs) - high retries flag flaky providers.

Quality (Youden's index) vs cost and vs speed for reliable models - the dotted Pareto front marks the best achievable trade-offs (labelled); everything below-left is dominated.

Pareto-optimal models (best quality for their cost):

model family config Youden's index F1 Score fail_rate mean_num_retries mean_proc_time cost_total_$
ibm-granite|granite-4.1-8b ibm-granite base 0.551 0.846 0.000 0.004 5.188 0.002
google|gemma-3-27b-it google base 0.555 0.821 0.000 1.618 4.298 0.003
qwen|qwen3-14b qwen base 0.590 0.857 0.000 3.390 4.470 0.003
qwen|qwen3-235b-a22b-2507 qwen base 0.620 0.870 0.000 2.441 10.608 0.003
deepseek|deepseek-v4-flash deepseek thinking_budget=1024 0.737 0.920 0.000 0.075 7.107 0.007
deepseek|deepseek-v3.2 deepseek thinking_budget=0 0.752 0.919 0.000 0.000 25.829 0.064
google|gemini-2.5-pro google thinking_budget=1024 0.774 0.924 0.000 0.000 15.717 1.602

Effect of completion configuration

How each model responds to its configuration knobs (cell = best Youden's index across prompts; empty = not run). The failure-rate grid shows which configs are unreliable.

Prompt & few-shot effects

Coverage & what to re-run

11 model(s) produced only failures on human-mouse (grouped by root cause):

New candidate models added to the registry but not yet run: claude-fable-5, gemini-3.5-flash, qwen3.7-plus, xiaomi/mimo-v2.5, gpt-5.5-pro.

Interactive explorer

Pick family → model → subset → metric to see every configuration actually run for a model (recommended config starred, unreliable runs flagged) - all in this page.

★ = recommended (best reliable config)  |  red border / ⚠ = unreliable run  |  only configurations actually run are shown.
Cost vs performance across models for the current family / subset (best reliable config per model; dotted Pareto front labelled - top-left is best value):
Per-model configurations - pick a model above to see its runs:

Everything on this page regenerates from src.report.build_report.