Add a case
Choose Add case, enter the input and expected result, then save. If the system has several prompts, choose the target. Cases are prompt-specific so an expectation for one agent does not become a policy for every agent.Edit a case
Choose Edit to open the case beside the table. Update Input and Expected result. Open More details for notes and any saved conversation, variables or context. Structured values use individual fields rather than a read-only JSON dump. Save for future benches creates a new context revision. Editing a Bench-retained case replaces its future-use copy with your correction. Bench does not run both versions, change previous scores, or allow a later automatic import to overwrite your correction. Use Pause to leave a case in the library without including it in future benches. Use Remove for a custom or corrected case you no longer want stored. Removing evidence can also erase the saved details of evaluations that used it; their scores remain unchanged.Import a dataset
1
Upload
Choose Upload cases. CSV, XLSX, JSON and JSONL are supported, up to 4 MB and 5,000 rows. Select a worksheet for a multi-sheet workbook.
2
Map and preview
Map the input and expected result columns. You can also map history, actual outputs, variables, resolved prompts and notes. Extra columns and repeated inputs are preserved after redaction.
3
Save
Save the mapped dataset. Comparing recorded answers is an exact, case-sensitive comparison, not a semantic LLM judgment.
4
Include in a bench
Choose Use a dataset, select the saved dataset and target prompt, then Add golden cases. This copies that dataset version into the library. Uploading alone does not start an evaluation.