Candidate Models Fan-out runs all selected models in parallel
Evaluate candidate models in parallel: rubric-based scoring, independent evals, full trace logging.
Prompt
Sent to candidate models
Evaluation
Used by the judge to score outputs
Hello! Select a model in the panel and start typing.
Trace Explorer

All runs append full outputs, token usage, served_by backends, and independent judge verdicts to JSONL.

Timestamp Query Goal Candidates Achieved Goal Judge Model Status
Loading trace logs…