Commands
run
Execute the full benchmark pipeline.
See Supported Models for all available judge and answering models.
compare
Run benchmark across multiple providers in parallel.test
Evaluate a single question for debugging.status
Check progress of a run.show-failures
Debug failed questions with full context.list-questions
Browse benchmark questions.Random Sampling
Sample N questions per category with optional randomization.serve
Start the web UI.help
Get help on providers, models, or benchmarks.Checkpointing
Runs are saved todata/runs/{runId}/ and automatically resume from the last successful phase. Use --force to restart.