Skip to main content
Bench connects your prompts, code and business context to repeatable evaluations. Start with a repository or uploaded prompts. Inspect discovered AI systems, run a bench, and review the failed checks and proposed improvements.
These guides include the system-first staging preview. SDK monitoring, system-wide Free allowances and the new context tools require the matching preview deployment. They are not a production release announcement. Use the setup instructions and endpoint shown inside your Bench environment.

Run your first bench

Connect a source and evaluate a prompt.

Connect your coding agent

Use Claude Code, Cursor or Codex.

Define success

Add business rules and examples when useful.

Build a test library

Keep trusted cases and customize criteria.

What a bench does

Bench creates criteria and scenarios, measures the current prompt, then tests candidate prompt and model changes on that suite. You can inspect the cases, scores, explanations and comparison before accepting a recommendation.
A code scan is not a runtime evaluation. Current prompt-level evaluations do not execute an entire application, agent graph or production tool side effect. The result page distinguishes a suggested improvement from an applied change.