★ = recommended (best reliable config) |
red border / ⚠ = unreliable run | only configurations actually run are shown.
Cost vs performance across models for the current family / subset (best reliable config
per model; dotted Pareto front labelled - top-left is best value):
Per-model configurations - pick a model above to see its runs: