Assist
An AI agent built into AI Evals that reads your prompts, datasets, outputs, and scores, then proposes edits you review and apply in place.
Assist
Assist is an AI agent built into AI Evals that reads the context you're working in — your prompts, datasets, model outputs, and evaluation scores — and reasons about what's working and what isn't. Work with it conversationally in a docked chat panel: type an instruction in plain language, and it proposes concrete edits without leaving the page, whether that's refining a prompt or building out your golden dataset from a task description or existing examples. Every proposal — a prompt edit, generated rows, or edits to existing items — is staged for review; nothing is written until you accept it.
What Assist needs
- A model. Assist runs on a model you choose from the LLM connections configured for your project. If none is set up, add one in Settings → LLM Connections.
- Something to work with. To improve a prompt, Assist is most useful once you have outputs and evaluation scores for it to reason over. To generate dataset items, it needs at least a task description — or existing traces, documents, or dataset items to ground on.
Where you can use Assist
Playground
Analyze prompts, outputs, and scores, then review and apply proposed prompt edits as a diff — across one or many windows.
Datasets
Build and maintain datasets conversationally — generate items from a description, traces, or documents, and edit existing items in bulk.
Scenarios
Draft persona-driven conversation scenarios grounded on your agent, then review and commit them.