Replay one eval and grade it against the recorded answer
Runs the real generation path, then grades the reply against preferred_output on substance rather than wording…
POST
Runs the real generation path, then grades the reply against preferred_output on substance rather than wording.
passed is null — never false — when grading could not run, so an unavailable judge is not mistaken for a failing agent.
Path parameters
string (uuid)
required
The eval id.
Response
Returnsobject.