Eval: From Computer Systems to Agents
As agents take on complex tasks, evaluation provides evidence that they meet specified requirements under deployment conditions. With growing adoption of open-weight models, robust Eval design becomes increasingly important: it guides improvements to narrow the gap with frontier models on a company’s target workloads. Assessing these improvements requires distinguishing the contributions of model capability, harness design, and execution budget, then verifying that the benefits persist in production.
Interestingly, there are parallels with Computer Systems Eval, where profiling helps investigate performance limits across interacting components, and the measurements themselves require validation. Agents add a further complication: the grader can be consistently wrong, and the criteria can reward behavior that misses the intended goal. Drawing on my experience across the two domains, I penned down these parallels, where they hold, and where Agent Eval introduces new challenges. Feedback welcome!
Thanks Bikash for the feedback on the early versions of the essay.