1
Connect your repository
Connect GitHub and select a repository. Bench scans its code for AI systems,
prompts and tools. You can also begin with SDK events without GitHub.
2
Check your system’s business outcome
Open an AI system’s Understanding page. Add or correct what success means:
for example, “refund an eligible order exactly once” and the relevant refund
policy. Bench uses this context to make evaluation criteria relevant.
3
Confirm your SDK connection
Choose an environment, install the SDK and run a synthetic request. The step
completes after Bench receives an event. Creating a key alone is not evidence
that instrumentation works. Follow the SDK quickstart.
4
Run and inspect a Bench
Select Start benching on a system to compare prompts and models. Inspect
failed cases and the evidence behind recommendations. Use Application tests
to inspect tests of application behavior and tool effects. See
application testing.
5
Connect your coding agent, optionally
Connect MCP with sign-in, then ask your agent to list your Bench
systems. The checklist confirms a successful authenticated MCP request. An API
key created for an agent does not by itself confirm MCP connection.

