# AI evals

Published articles for AI evals.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Learn Claude Code, evals, AI systems, and more: ByteByteGo Live is here

DevFeed: [Learn Claude Code, evals, AI systems, and more: ByteByteGo Live is here](<https://devfeed.tech/articles/learn-claude-code-evals-ai-systems-and-more-bytebytego-live-is-here-17996.md>)

Original publisher: [Read original article](<https://blog.bytebytego.com/p/learn-claude-code-evals-ai-systems>)

Author: ByteByteGo

Published: 2026-09-11T15:32:16Z

Content type: release

Language: en

Sources: [ByteByteGo](<https://devfeed.tech/sources/bytebytego.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-development](<https://devfeed.tech/tags/ai-development.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [development](<https://devfeed.tech/tags/development.md>)

### AI overview

ByteByteGo announces ByteByteGo Live, a membership offering live courses on Claude Code, production AI systems, AI engineering, AI evaluations, cost optimization, and related topics. The announcement cites higher completion rates for live cohorts and says the membership covers courses offered over the next 12 months.

### Source excerpt

Most online courses never get finished (~4% completion). Live cohorts get ~40%, roughly 10x higher. Live courses are the only courses people actually finish. So we're launching ByteByteGo Live.

## Catch AI Regressions Before They Ship with AI Evals in CI/CD

DevFeed: [Catch AI Regressions Before They Ship with AI Evals in CI/CD](<https://devfeed.tech/articles/catch-ai-regressions-before-they-ship-with-ai-evals-in-ci-cd-13376.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/catch-ai-regressions-before-they-ship-with-ai-evals-in-ci-cd>)

Author: Shibam Dhar

Published: 2026-09-02T00:00:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [production](<https://devfeed.tech/tags/production.md>), [quality](<https://devfeed.tech/tags/quality.md>), [tests](<https://devfeed.tech/tags/tests.md>)

### AI overview

Harness AI Evals uses golden datasets, response-quality metrics, and blocking quality gates in CI/CD to catch AI agent regressions before production. In an e-commerce support-agent test, an early run passed about 65% of cases, below the 70% deployment threshold, revealing incorrect, missing, or incomplete answers.

### Source excerpt

Harness AI Evals tests AI agent quality in CI/CD, using golden datasets and quality gates to catch behavioral regressions before production. | Blog

## Building Trust in AI DevOps: Validating the Harness Knowledge Graph

DevFeed: [Building Trust in AI DevOps: Validating the Harness Knowledge Graph](<https://devfeed.tech/articles/building-trust-in-ai-devops-validating-the-harness-knowledge-graph-13374.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/building-trust-in-our-knowledge-graph>)

Author: Vikram Sahu

Published: 2026-08-31T18:37:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [DevOps](<https://devfeed.tech/topics/devops.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Software](<https://devfeed.tech/topics/software.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [api](<https://devfeed.tech/tags/api.md>), [automated](<https://devfeed.tech/tags/automated.md>), [data](<https://devfeed.tech/tags/data.md>), [devops](<https://devfeed.tech/tags/devops.md>), [evals](<https://devfeed.tech/tags/evals.md>), [graph](<https://devfeed.tech/tags/graph.md>), [knowledge-graph](<https://devfeed.tech/tags/knowledge-graph.md>), [lifecycle](<https://devfeed.tech/tags/lifecycle.md>), [operational](<https://devfeed.tech/tags/operational.md>), [other](<https://devfeed.tech/tags/other.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [schema](<https://devfeed.tech/tags/schema.md>), [sdlc](<https://devfeed.tech/tags/sdlc.md>), [security](<https://devfeed.tech/tags/security.md>), [services](<https://devfeed.tech/tags/services.md>), [software](<https://devfeed.tech/tags/software.md>), [software-delivery](<https://devfeed.tech/tags/software-delivery.md>), [validation](<https://devfeed.tech/tags/validation.md>)

### AI overview

This article explains how Harness validates answers from its SDLC Knowledge Graph. Its multi-layered approach combines AI evaluations, schema traversal, API checks, direct product verification, production data, and shift-left testing to improve reliability.

### Source excerpt

Discover our multi-layered validation approach combining AI evals to ensure reliable AI-powered software delivery insights. | Blog

## Artificial Intelligence: Glossary

DevFeed: [Artificial Intelligence: Glossary](<https://devfeed.tech/articles/artificial-intelligence-glossary-9033.md>)

Original publisher: [Read original article](<https://www.nngroup.com/articles/artificial-intelligence-glossary/>)

Author: Caleb Sponheim

Published: 2026-08-21T17:00:00Z

Content type: article

Language: en

Sources: [NN/g latest articles and announcements](<https://devfeed.tech/sources/nn-g-latest-articles-and-announcements.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Algorithms](<https://devfeed.tech/topics/algorithms.md>), [Prompt Engineering](<https://devfeed.tech/topics/prompt-engineering.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-definitions](<https://devfeed.tech/tags/ai-definitions.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [ai-glossary](<https://devfeed.tech/tags/ai-glossary.md>), [ai-glossary-for-ux](<https://devfeed.tech/tags/ai-glossary-for-ux.md>), [ai-hallucination](<https://devfeed.tech/tags/ai-hallucination.md>), [ai-terminology](<https://devfeed.tech/tags/ai-terminology.md>), [ai-terminology-for-product-teams](<https://devfeed.tech/tags/ai-terminology-for-product-teams.md>), [ai-terms](<https://devfeed.tech/tags/ai-terms.md>), [ai-terms-for-designers](<https://devfeed.tech/tags/ai-terms-for-designers.md>), [ai-vocabulary](<https://devfeed.tech/tags/ai-vocabulary.md>), [article](<https://devfeed.tech/tags/article.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [artificial-intelligence-glossary](<https://devfeed.tech/tags/artificial-intelligence-glossary.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [genai-glossary](<https://devfeed.tech/tags/genai-glossary.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [generative-ui](<https://devfeed.tech/tags/generative-ui.md>), [glossary](<https://devfeed.tech/tags/glossary.md>), [knowledge-cutoff](<https://devfeed.tech/tags/knowledge-cutoff.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [model-context-protocol](<https://devfeed.tech/tags/model-context-protocol.md>), [prompt-engineering](<https://devfeed.tech/tags/prompt-engineering.md>), [prompt-injection](<https://devfeed.tech/tags/prompt-injection.md>), [rag](<https://devfeed.tech/tags/rag.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [ux](<https://devfeed.tech/tags/ux.md>), [vibe-coding](<https://devfeed.tech/tags/vibe-coding.md>)

### AI overview

A plain-language glossary of artificial-intelligence terminology used in products and UX work. It explains concepts including agents, agentic systems, AI development, algorithms, AI-generated content, and AI-related claims, while noting that terminology can vary among vendors and researchers.

### Source excerpt

Plain-language definitions of the AI terms that come up in product and design work, from tokens and context windows to agents, evals, and prompt injection.

## From a Raw Shell to a Sandboxed Coding Agent

DevFeed: [From a Raw Shell to a Sandboxed Coding Agent](<https://devfeed.tech/articles/from-a-raw-shell-to-a-sandboxed-coding-agent-18302.md>)

Original publisher: [Read original article](<https://www.decodingai.com/p/run-coding-agents-safely>)

Author: Paul Iusztin

Published: 2026-08-18T11:02:59Z

Content type: tutorial

Language: en

Sources: [Decoding ML](<https://devfeed.tech/sources/decoding-ml.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [LangChain](<https://devfeed.tech/topics/langchain.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [Docker](<https://devfeed.tech/topics/docker.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>), [codex](<https://devfeed.tech/topics/codex.md>), [Python](<https://devfeed.tech/topics/python.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [claude](<https://devfeed.tech/tags/claude.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [computer-use](<https://devfeed.tech/tags/computer-use.md>), [docker](<https://devfeed.tech/tags/docker.md>), [guide](<https://devfeed.tech/tags/guide.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [python](<https://devfeed.tech/tags/python.md>), [terminal](<https://devfeed.tech/tags/terminal.md>)

### AI overview

A tutorial on isolating coding-agent tools inside local Docker or remote Modal sandboxes. It explains how to build a Python harness that safely executes commands and supports remote, parallel agent workflows.

### Source excerpt

The guide to isolating your harness and safely executing its commands, locally or remotely.

## Building a Coding Agent From Scratch

DevFeed: [Building a Coding Agent From Scratch](<https://devfeed.tech/articles/building-a-coding-agent-from-scratch-18293.md>)

Original publisher: [Read original article](<https://www.decodingai.com/p/building-a-coding-agent-from-scratch-system-design>)

Author: Paul Iusztin

Published: 2026-07-22T11:04:24Z

Content type: tutorial

Language: en

Sources: [Decoding ML](<https://devfeed.tech/sources/decoding-ml.md>)

Topics: [Agent Harness](<https://devfeed.tech/topics/agent-harness.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [coding](<https://devfeed.tech/topics/coding.md>), [LangChain](<https://devfeed.tech/topics/langchain.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [codex](<https://devfeed.tech/topics/codex.md>)

Tags: [agent-harness](<https://devfeed.tech/tags/agent-harness.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [langchain](<https://devfeed.tech/tags/langchain.md>)

### AI overview

This tutorial explains how to build a coding-agent harness from scratch in Python. It covers the agent loop, shell execution, context engineering, subagents, remote parallel agents, and evaluation workflows, using the project Decode as the practical example.

### Source excerpt

Designing the harness around the model, from the agent loop to a remote swarm.

## Harness AI Evals adds CI/CD quality gates for testing and monitoring AI agents

DevFeed: [Harness AI Evals adds CI/CD quality gates for testing and monitoring AI agents](<https://devfeed.tech/articles/ship-ai-agents-you-can-trust-introducing-ai-evals-13441.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/introducing-ai-evals>)

Author: Shibam Dhar Uri Scheiner

Published: 2026-07-21T00:00:00Z

Content type: release

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [evals](<https://devfeed.tech/tags/evals.md>), [monitor](<https://devfeed.tech/tags/monitor.md>), [production](<https://devfeed.tech/tags/production.md>), [quality](<https://devfeed.tech/tags/quality.md>), [releases](<https://devfeed.tech/tags/releases.md>), [test](<https://devfeed.tech/tags/test.md>)

### AI overview

Harness introduces AI Evals, a tool for testing, scoring, and monitoring AI agents before and after deployment. It provides a native CI/CD pipeline step, evaluation metrics, and blocking or advisory pass strategies intended to prevent poor releases from reaching production.

### Source excerpt

Harness AI Evals helps you test, score, and monitor AI agents with native CI/CD quality gates, blocking poor releases before production. | Blog

## Design AI Products for Verification Before Building Evals

DevFeed: [Design AI Products for Verification Before Building Evals](<https://devfeed.tech/articles/it-s-hard-to-eval-is-a-product-smell-18787.md>)

Original publisher: [Read original article](<https://hamel.dev/blog/posts/eval-smell/>)

Author: Hamel Husain

Published: 2026-06-29T07:00:00Z

Content type: opinion

Language: en

Sources: [Hamel Husain](<https://devfeed.tech/sources/hamel-husain.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [data](<https://devfeed.tech/topics/data.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [SQL](<https://devfeed.tech/topics/sql.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [data-agents](<https://devfeed.tech/tags/data-agents.md>), [evals](<https://devfeed.tech/tags/evals.md>), [interface](<https://devfeed.tech/tags/interface.md>), [llms](<https://devfeed.tech/tags/llms.md>), [techniques](<https://devfeed.tech/tags/techniques.md>), [verification](<https://devfeed.tech/tags/verification.md>)

### AI overview

The article argues that products described as difficult to evaluate often make their outputs difficult for users to verify. Using AI data agents as an example, it recommends providing checkable artifacts--such as source comparisons, precise metric definitions, breakdowns, SQL, and uncertainty notes--before focusing on eval design.

### Source excerpt

For the past 3 years, AI evals have been my professional focus.1 The most common objection I hear to evals is "our product is hard to eval". This objection is a product smell. Artifacts that are hard for you to verify are often hard for users too. In the worst case, users have to redo the work from scratch to verify the output. More importantly, designing your product for ease of verification should come before building evals. In this post, I'll walk through three products I advised on that faced this issue. I'll also show before and after sketches to demonstrate design principles. After these examples, I'll discuss how to apply this general pattern to your product. Example 1: the AI data agent Almost every company I've worked with builds an internal AI data agent. You ask it a business question, like what was net revenue for Product A last quarter, and it finds relevant data sources, runs the queries, and provides an answer. The goal of this agent is to reduce dependency on data analysts. A common mistake when building AI data agents is to make the answer the only output, as illustrated below. Data Agent What was net revenue for Product A last quarter? Net revenue for Product A last quarter was $4.21M. Ask anything about your business...➤ Since the only output is the answer, there is nothing here to check. In the sketch above, the user has no way to verify the answer beyond redoing work.2 A better design is to provide the user with checkable artifacts, informed by how a domain expert might validate the output. Here are techniques I use to validate metrics as a data scientist: Compare the quantity and any intermediate calculations against a trusted source, like a vetted dashboard or report, or a similar analysis a colleague has already vetted.3 Confirm the metric definition precisely. A number like net revenue can include or exclude things like returns and discounts. Sanity-check a related quantity. If I can't verify the number directly, I pull a related number that s

## How Evaluation-Driven Development (EDD) Works

DevFeed: [How Evaluation-Driven Development (EDD) Works](<https://devfeed.tech/articles/how-evaluation-driven-development-edd-works-18296.md>)

Original publisher: [Read original article](<https://www.decodingai.com/p/how-evaluation-driven-development-works>)

Author: Paul Iusztin

Published: 2026-06-23T08:57:02Z

Content type: tutorial

Language: en

Sources: [Decoding ML](<https://devfeed.tech/sources/decoding-ml.md>)

Topics: [Development](<https://devfeed.tech/topics/development.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [dataset](<https://devfeed.tech/topics/dataset.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [case-study](<https://devfeed.tech/tags/case-study.md>), [development](<https://devfeed.tech/tags/development.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [saas](<https://devfeed.tech/tags/saas.md>), [test](<https://devfeed.tech/tags/test.md>), [tests](<https://devfeed.tech/tags/tests.md>)

### AI overview

This case study explains Evaluation-Driven Development (EDD) for AI agents: measure a new feature, compare results before and after changes, and detect regressions before merging. It also discusses generating realistic test data when historical datasets, traces, or ground truth are unavailable.

### Source excerpt

Turn every AI agent change into a measured experiment you compare before and after to detect regressions and measure performance.

## Key takeaways from the PyAI conference on AI evaluation, software design, and open-source maintenance

DevFeed: [Key takeaways from the PyAI conference on AI evaluation, software design, and open-source maintenance](<https://devfeed.tech/articles/learnings-from-the-pyai-conference-21748.md>)

Original publisher: [Read original article](<http://blog.pamelafox.org/2026/03/learnings-from-pyai-conference.html>)

Author: Pamela Fox (noreply@blogger.com)

Published: 2026-03-12T06:40:00Z

Content type: article

Language: en

Sources: [Pamela Fox](<https://devfeed.tech/sources/pamela-fox.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Python](<https://devfeed.tech/topics/python.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [MCP](<https://devfeed.tech/topics/mcp.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [Pydantic](<https://devfeed.tech/topics/pydantic.md>), [FastAPI](<https://devfeed.tech/topics/fastapi.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [code](<https://devfeed.tech/tags/code.md>), [coding](<https://devfeed.tech/tags/coding.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [conference](<https://devfeed.tech/tags/conference.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [fastapi](<https://devfeed.tech/tags/fastapi.md>), [github](<https://devfeed.tech/tags/github.md>), [maintainers](<https://devfeed.tech/tags/maintainers.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [python](<https://devfeed.tech/tags/python.md>), [sdk](<https://devfeed.tech/tags/sdk.md>)

### AI overview

The article summarizes lessons from PyAI conference sessions on evaluating AI systems, designing Python software for maintainability by coding agents, and handling AI-generated pull requests in open-source projects. It recommends validating LLM judges with labeled data and conventional evaluation metrics, using clearer software abstractions, and developing systems to triage low-quality contributions.

### Source excerpt

I recently spoke at the PyAI conference, put on by the good folks at Prefect and Pydantic, and I learnt so much from the talks I attended. Here are my top takeaways from the sessions that I watched: AI Evals Pitfalls Hamel Husain 📺 Watch the video recording | 📊 View slides Hamel cautioned against blindly using automated evaluation frameworks and built-in evaluators (like helpfulness and coherence). Instead, we should adopt a data science approach to evaluation: explore the data, discover what's actually breaking, identify the most important metric, and iterate as new data comes in. We shouldn't just trust an LLM-as-a-judge to be given accurate scores. Instead, we should validate it like we would validate a ML classifier- with labeled data, train/dev/test splits, and precision/recall metrics. LLM-judges should always give pass/fail results, instead of 1-5 scores, so that there's no ambiguity in their judgment. When generating synthetic data, first come up with dimensions (such as persona), generate combinations based off dimensions, and convert those into realistic queries. Hamel created evals-skills, a collection of skills for coding agents that can be run against evaluation pipelines to find issues like poorly designed LLM-judges. Build Reasonable Software Jeremiah Lowin (FastMCP/Prefect) 📺 Watch the video recording Write your Python programs in a way that coding agents can reason about them, so that they can more easily maintain and build them. For example, FastMCP v2 SDK was not well designed (bad abstractions) so a new CodeMod feature required 4,000 lines of code. In the new FastMCP v3 SDK (same functional API, different abstractions backing it), the same feature only required 500 lines of code. To make Python FastMCP servers more Pythonic, Jeremiah is developing a new package for MCP apps which includes the most common UIs (forms/tables/charts), called PreFab: https://github.com/PrefectHQ/prefab Panel: Open Source in the Age of AI Guido van Rossum (CPython), Sa

## Evals Skills for Coding Agents

DevFeed: [Evals Skills for Coding Agents](<https://devfeed.tech/articles/evals-skills-for-coding-agents-18790.md>)

Original publisher: [Read original article](<https://hamel.dev/blog/posts/evals-skills/>)

Author: Shreya Shankar

Published: 2026-03-02T08:00:00Z

Content type: release

Language: en

Sources: [Hamel Husain](<https://devfeed.tech/sources/hamel-husain.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [coding](<https://devfeed.tech/topics/coding.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [audit](<https://devfeed.tech/tags/audit.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [evals](<https://devfeed.tech/tags/evals.md>), [review](<https://devfeed.tech/tags/review.md>), [skills](<https://devfeed.tech/tags/skills.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

Shreya Shankar and the author publish evals skills for AI product evaluations. The collection routes users to skills for auditing eval pipelines, discovering errors from traces, generating synthetic test inputs, designing and validating LLM judges, and evaluating RAG quality.

### Source excerpt

Today, Shreya Shankar and I are publishing evals skills, a set of skills for AI product evals1. Eval tools often get in the way. They nudge you toward generic off-the-shelf metrics and fully automated evals before you've looked at your data. These skills help you avoid common mistakes we've seen helping 50+ companies and teaching students in our AI Evals course. Why skills for evals There are many easily avoidable footguns in evals. These skills help you avoid them. evals-start is the entry point. It looks at your situation and routes you to the right skill. Most of the time it will send you to one of these two: eval-audit, if you already have an eval pipeline. It inspects your setup and recommends next steps. 2 error-discovery, if you have traces but haven't analyzed them yet. It builds a customized annotation interface and helps you sample traces intelligently. Shreya does a live walkthrough of using this skill here. The skills Install the skills: npx skills add https://github.com/ai-evals-course/evals-skills Then give your agent this prompt: Run the evals-start skill from the evals plugin and follow the skill it picks. If it picks eval-audit, investigate each diagnostic area using a separate subagent in parallel, then synthesize the findings into a single report. If you're experienced with evals, skip the router and pick the skill you need: Skill What it does evals-start Entry point. Routes to the skill that matches your situation eval-audit Audit an eval pipeline and surface problems with prioritized severity error-discovery Build a review app, select diverse samples, and organize your notes into failure modes generate-synthetic-data Create diverse synthetic test inputs using dimension-based tuple generation write-judge-prompt Design LLM-as-Judge evaluators for subjective quality criteria validate-evaluator Calibrate LLM judges against human labels using data splits, TPR/TNR, and bias correction evaluate-rag Evaluate retrieval and generation quality in RAG pipel

## Selecting The Right AI Evals Tool

DevFeed: [Selecting The Right AI Evals Tool](<https://devfeed.tech/articles/selecting-the-right-ai-evals-tool-18788.md>)

Original publisher: [Read original article](<https://hamel.dev/blog/posts/eval-tools/>)

Author: Hamel Husain

Published: 2025-10-01T07:00:00Z

Content type: article

Language: en

Sources: [Hamel Husain](<https://devfeed.tech/sources/hamel-husain.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>), [Developer experience](<https://devfeed.tech/topics/developer-experience.md>), [LangChain](<https://devfeed.tech/topics/langchain.md>), [phoenix](<https://devfeed.tech/topics/phoenix.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [developer-experience](<https://devfeed.tech/tags/developer-experience.md>), [evals](<https://devfeed.tech/tags/evals.md>), [langchain](<https://devfeed.tech/tags/langchain.md>), [phoenix](<https://devfeed.tech/tags/phoenix.md>), [quality](<https://devfeed.tech/tags/quality.md>)

### AI overview

The article examines how to select AI evaluation tools, arguing that no single tool is best for every team. It compares approaches from LangSmith, Braintrust, and Arize Phoenix through a shared assignment and highlights workflow, developer experience, SDK ergonomics, documentation, integrations, and human-in-the-loop support as selection criteria.

### Source excerpt

Over the past year, I've focused heavily on AI Evals, both in my consulting work and teaching. A question I get constantly is, "What's the best tool for evals?". I've always resisted answering directly for two reasons. First, people focus too much on tools instead of the process, thinking the tool will be an off-the-shelf solution when it rarely is. Second, the tools change so quickly that comparisons become outdated immediately. Having used many of the popular eval tools, I can genuinely say that no single one is superior in every dimension. The "best" tool depends on your team's skillset, technical stack, and maturity. Instead of a feature-by-feature comparison, I think it's more valuable to show you how a panel of data scientists skilled in evals assesses these tools. As part of my AI Evals course, we had three of the most dominant vendors--Langsmith, Braintrust, and Arize Phoenix complete the same homework assignment. This gave us a unique opportunity to see how they tackle the exact same challenge. We recorded the entire process and live commentary, which is available below. We think this might be helpful in learning about the kinds of things you should consider when selecting a tool for your team. Thanks to Shreya Shankar and Bryan Bischof for serving as the panelists (alongside me). Langsmith With Harrison Chase, CEO of LangChain. Braintrust With Wayde Gilliam, former developer relations at Braintrust. Arize Phoenix With SallyAnn DeLucia, Technical AI Product Leader at Arize. Criteria for Assessing AI Evals Tools Here are themes that consistently surfaced during our review. 1. Workflow and Developer Experience Reducing friction is more important than any single feature. Concretely, you should be mindful of the time it takes to go from observing a failure to iterating on a solution. For example, we appreciated the ability to go from viewing a single trace to experimenting with that same trace in a playground. For some teams with data-science backgrounds, a note

## AI Evals: Common Questions About Model and Product Evaluation

DevFeed: [AI Evals: Common Questions About Model and Product Evaluation](<https://devfeed.tech/articles/ai-evals-everything-you-need-to-know-18789.md>)

Original publisher: [Read original article](<https://hamel.dev/blog/posts/evals-faq/>)

Author: Shreya Shankar

Published: 2025-05-28T07:00:00Z

Content type: tutorial

Language: en

Sources: [Hamel Husain](<https://devfeed.tech/sources/hamel-husain.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [evals](<https://devfeed.tech/tags/evals.md>), [llms](<https://devfeed.tech/tags/llms.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

This FAQ explains AI evaluations as tests for determining whether an AI system meets user and business goals. It distinguishes model benchmarks from product evaluations, which measure a specific product across its model, prompts, retrieval, tools, and application code.

### Source excerpt

This document curates the most common questions Shreya and I received while teaching 700+ engineers & PMs AI Evals. Warning: These are sharp opinions about what works in most cases. They are not universal truths. Use your judgment. For a guided path through the rest of our evals work, use the AI evals topic hub. 👉 Want to learn more about AI Evals? Check out our AI Evals course. It's a live cohort with hands on exercises and office hours. Here is a 25% discount code for readers. 👈 Getting Started & Fundamentals Q: What are AI Evals? AI evals are tests that tell you whether an AI system is doing what you want. They give your team feedback when the product drifts from user needs or business goals. The failures they catch also become data you can use to improve the system. More formally, evaluation is the systematic measurement of quality. Each eval checks one behavior on relevant examples and returns a score or structured review. Most AI products need several evals because they can fail in different ways. When you hear the word "evals," it usually refers to one of two things: model benchmarks or product evals. Model benchmarks Model benchmarks compare general-purpose models on shared tasks. Model providers publish these benchmark results when they release new models. Common examples include GPQA Diamond for graduate-level science reasoning, Terminal-Bench for agents doing complex work in command-line environments, and MMLU for knowledge and reasoning across a wide range of subjects. These scores can help you choose a promising model as a starting point. To assess quality on your own tasks you need product evals, which we discuss next. Product evals Product evals measure whether your specific AI product does what you want it to do. They turn your judgment about what a good product experience looks like into metrics you can track. Product evals encompass all components of your product, including the model, prompts, retrieval, tools, and application code. This flavor of