← Back to sessions
Evaluating LLM Agents in CI
Thursday, May 13, 2027 · 9:30 AM – 10:00 AMMain Stage
AI EngineeringTalk (30 min)Advanced
Agent evaluations are flaky, expensive and easy to fool. This talk covers the harness we run on every pull request: deterministic seeds, judge calibration, cost budgets, and the tripwires that catch a regression before it reaches a customer.