Run an agent's eval suite (regression check)
Each eval costs a real generation plus a judge call, so runs are sequential and capped per call; has_more reports when the suite was longer than the cap.
Each eval costs a real generation plus a judge call, so runs are sequential and capped per call; has_more reports when the suite was longer than the cap.