# How to Build Custom Evals and Regression Tests for AI Agents

DevFeed: [How to Build Custom Evals and Regression Tests for AI Agents](<https://devfeed.tech/articles/i-built-a-benchmark-for-my-agent-the-smaller-model-won-57823.md>)

Original publisher: [Read original article](<https://www.decodingai.com/p/evaluate-ai-agents-benchmarks-regression-tests>)

Author: Paul Iusztin

Published: 2026-09-22T05:01:45Z

Content type: tutorial

Language: en

Sources: [Decoding ML](<https://devfeed.tech/sources/decoding-ml.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Agent Harness](<https://devfeed.tech/topics/agent-harness.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Code](<https://devfeed.tech/topics/code.md>), [gpt-oss](<https://devfeed.tech/topics/gpt-oss.md>), [Python](<https://devfeed.tech/topics/python.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [LangChain](<https://devfeed.tech/topics/langchain.md>)

Tags: [agent-evals](<https://devfeed.tech/tags/agent-evals.md>), [agent-harness](<https://devfeed.tech/tags/agent-harness.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [context-engineering](<https://devfeed.tech/tags/context-engineering.md>), [cost](<https://devfeed.tech/tags/cost.md>), [latency](<https://devfeed.tech/tags/latency.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [tests](<https://devfeed.tech/tags/tests.md>)

## AI overview

A practical guide to building custom evaluation harnesses and regression tests for AI agents. It argues that application-specific evals can reveal meaningful differences between models, including cases where a smaller model outperforms a larger one, while measuring performance, cost, and latency.

## Source excerpt

A 35B beat a 120B by 42 points: the benchmark and regression tests behind it