Generated 2026-07-13 · focus subset human-mouse · metrics & error rates recomputed from raw predictions
We benchmark LLMs as binary judges for ontology entity matching. Each model is run over candidate pairs under several completion configurations parsed from the experiment name (reasoning effort, thinking budget, few-shot count, prompt). The leaderboard is regenerated directly from the results on disk, so every table and figure below stays in sync with the data.
Core metric: Youden's index (sensitivity + specificity − 1). Reliable = fail_rate ≤ 5% and ≥ 30 predictions. fail_rate = fraction of calls that errored; mean_num_retries = average retries per call (shown where the provider logged it). Drill-down in the explorer at the bottom.
| model | family | config | prompt | Youden's index | F1 Score | fail_rate | mean_num_retries | cost_total_$ |
|---|---|---|---|---|---|---|---|---|
| google|gemini-2.5-pro | thinking_budget=1024 | direct_entity | 0.774 | 0.924 | 0.000 | 0.000 | 1.602 | |
| qwen|qwen3.5-35b-a3b | qwen | base | direct_entity | 0.753 | 0.923 | 0.000 | — | — |
| deepseek|deepseek-v3.2 | deepseek | thinking_budget=0 | direct_entity | 0.752 | 0.919 | 0.000 | 0.000 | 0.064 |
| google|gemini-3.1-flash-lite-preview | reasoning_effort=medium | direct_entity | 0.750 | 0.927 | 0.000 | 2.811 | 0.076 | |
| qwen|qwen3.5-122b-a10b | qwen | base | direct_entity | 0.742 | 0.920 | 0.000 | — | — |
| google|gemini-2.5-flash | thinking_budget=1024 | direct_entity | 0.741 | 0.916 | 0.000 | 0.000 | 0.223 | |
| mistralai|mistral-medium-3-5 | mistralai | reasoning_effort=high | direct_entity | 0.737 | 0.920 | 0.000 | 4.000 | 0.700 |
| deepseek|deepseek-v4-flash | deepseek | thinking_budget=1024 | direct_entity | 0.737 | 0.920 | 0.000 | 0.075 | 0.007 |
| z-ai|glm-5 | z-ai | base | direct_entity | 0.735 | 0.913 | 0.000 | — | — |
| anthropic|claude-sonnet-4.6 | anthropic | base | direct_entity | 0.734 | 0.924 | 0.000 | — | — |
| openai|gpt-5-nano | openai | reasoning_effort=medium | direct_entity | 0.730 | 0.897 | 0.000 | — | — |
| openai|gpt-5.5 | openai | reasoning_effort=medium | direct_entity_with_synonyms | 0.728 | 0.921 | 0.000 | 0.004 | 1.197 |
| google|gemini-2.5-flash-lite | thinking_budget=512 | direct_entity | 0.726 | 0.918 | 0.000 | 0.012 | 0.032 | |
| bytedance-seed|seed-2.0-mini | bytedance-seed | base | direct_entity | 0.726 | 0.885 | 0.000 | 0.000 | 0.059 |
| x-ai|grok-4.3 | x-ai | base | direct_entity | 0.724 | 0.910 | 0.000 | 0.000 | 0.446 |
Cost is per full run of the subset, from real token usage × live OpenRouter pricing.
Ranked by Youden's index (balanced sensitivity + specificity). Ten model-config pairs balancing quality, value (quality per $), reasoning need and provider diversity - not just the top scores. gpt-5.5 excluded as too expensive for the quality gained.
Cost vs performance for the proposed models (gold, with their Pareto front) against all other models (grey, with a secondary dashed front). Use the input to change the fail-rate reliability threshold and watch both frontiers shift.
| model | family | config | prompt | Youden's index | F1 Score | fail_rate | mean_num_retries | cost_total_$ | rationale |
|---|---|---|---|---|---|---|---|---|---|
| google|gemini-2.5-pro | thinking_budget=1024 | direct_entity | 0.774 | 0.924 | 0.000 | 0.000 | 1.602 | Top Youden's index | |
| qwen|qwen3.5-35b-a3b | qwen | base | direct_entity | 0.753 | 0.923 | 0.000 | — | — | 2nd best Youden's index |
| deepseek|deepseek-v4-flash | deepseek | thinking_budget=1024 | direct_entity | 0.737 | 0.920 | 0.000 | 0.075 | 0.007 | Cheapest among above-median quality (best value) |
| mistralai|mistral-medium-3-5 | mistralai | reasoning_effort=high | direct_entity | 0.737 | 0.920 | 0.000 | 4.000 | 0.700 | Family diversity (mistralai) |
| z-ai|glm-5 | z-ai | base | direct_entity | 0.735 | 0.913 | 0.000 | — | — | Family diversity (z-ai) |
| anthropic|claude-sonnet-4.6 | anthropic | base | direct_entity | 0.734 | 0.924 | 0.000 | — | — | Family diversity (anthropic) |
| openai|gpt-5-nano | openai | reasoning_effort=medium | direct_entity | 0.730 | 0.897 | 0.000 | — | — | Family diversity (openai) |
| bytedance-seed|seed-2.0-mini | bytedance-seed | base | direct_entity | 0.726 | 0.885 | 0.000 | 0.000 | 0.059 | Family diversity (bytedance-seed) |
| x-ai|grok-4.3 | x-ai | base | direct_entity | 0.724 | 0.910 | 0.000 | 0.000 | 0.446 | Family diversity (x-ai) |
| minimax|minimax-m2.5 | minimax | base | direct_entity | 0.691 | 0.902 | 0.000 | 0.008 | 0.085 | Family diversity (minimax) |
Strong reliable models just outside the selection - handy as cheaper alternatives or drop-in backups.
| model | family | config | prompt | Youden's index | F1 Score | fail_rate | mean_num_retries | cost_total_$ | note |
|---|---|---|---|---|---|---|---|---|---|
| deepseek|deepseek-v3.2 | deepseek | thinking_budget=0 | direct_entity | 0.752 | 0.919 | 0.000 | 0.000 | 0.064 | strong backup |
| google|gemini-3.1-flash-lite-preview | reasoning_effort=medium | direct_entity | 0.750 | 0.927 | 0.000 | 2.811 | 0.076 | strong backup | |
| qwen|qwen3.5-122b-a10b | qwen | base | direct_entity | 0.742 | 0.920 | 0.000 | — | — | strong backup |
| google|gemini-2.5-flash | thinking_budget=1024 | direct_entity | 0.741 | 0.916 | 0.000 | 0.000 | 0.223 | strong backup | |
| google|gemini-2.5-flash-lite | thinking_budget=512 | direct_entity | 0.726 | 0.918 | 0.000 | 0.012 | 0.032 | strong backup | |
| x-ai|grok-4.1-fast | x-ai | base | direct_entity | 0.721 | 0.918 | 0.000 | — | — | strong backup |
Failure rate per model (red = unreliable, > 5%). These are the models that need re-running.
Average retries per call (logged for OpenRouter runs) - high retries flag flaky providers.
Quality (Youden's index) vs cost and vs speed for reliable models - the dotted Pareto front marks the best achievable trade-offs (labelled); everything below-left is dominated.
Pareto-optimal models (best quality for their cost):
| model | family | config | Youden's index | F1 Score | fail_rate | mean_num_retries | mean_proc_time | cost_total_$ |
|---|---|---|---|---|---|---|---|---|
| ibm-granite|granite-4.1-8b | ibm-granite | base | 0.551 | 0.846 | 0.000 | 0.004 | 5.188 | 0.002 |
| google|gemma-3-27b-it | base | 0.555 | 0.821 | 0.000 | 1.618 | 4.298 | 0.003 | |
| qwen|qwen3-14b | qwen | base | 0.590 | 0.857 | 0.000 | 3.390 | 4.470 | 0.003 |
| qwen|qwen3-235b-a22b-2507 | qwen | base | 0.620 | 0.870 | 0.000 | 2.441 | 10.608 | 0.003 |
| deepseek|deepseek-v4-flash | deepseek | thinking_budget=1024 | 0.737 | 0.920 | 0.000 | 0.075 | 7.107 | 0.007 |
| deepseek|deepseek-v3.2 | deepseek | thinking_budget=0 | 0.752 | 0.919 | 0.000 | 0.000 | 25.829 | 0.064 |
| google|gemini-2.5-pro | thinking_budget=1024 | 0.774 | 0.924 | 0.000 | 0.000 | 15.717 | 1.602 |
How each model responds to its configuration knobs (cell = best Youden's index across prompts; empty = not run). The failure-rate grid shows which configs are unreliable.
11 model(s) produced only failures on human-mouse (grouped by root cause):
routes_registry: bytedance-seed|seed-2.0-lite, moonshotai|kimi-k2.6, nvidia|nemotron-3-nano-omni-30b-a3b-reasoning, nvidia|nemotron-3-super-120b-a12b, qwen|qwen3.6-35b-a3b, qwen|qwen3.6-flash, qwen|qwen3.6-max-previewcustom_parsing=True (no structured output): minimax|minimax-m2.7, z-ai|glm-5.1New candidate models added to the registry but not yet run: claude-fable-5, gemini-3.5-flash, qwen3.7-plus, xiaomi/mimo-v2.5, gpt-5.5-pro.
Pick family → model → subset → metric to see every configuration actually run for a model (recommended config starred, unreliable runs flagged) - all in this page.
Everything on this page regenerates from src.report.build_report.