Evaluating LLM Agents in CI
Thursday, May 13, 2027 · 9:30 AM – 10:00 AM · Main Stage
Agent evaluations are flaky, expensive and easy to fool. This talk covers the harness we run on every pull request: deterministic seeds, judge calibration, cost budgets, and the tr…
- IOIris Oyelaran — Staff Machine Learning Engineer, Corvid AI