Autonomous synthetic user testing

Agentic Evals DASHboard

A synthetic user plans its own scenarios and drives the real authenticated Snow Dev widget on its own, with no scripted replies. An independent panel of three judges then grades the conversation it produced.

Testing CaseRun byValidationModelStatusProgressEvaluationCostStarted (US East)