# Selecting The Right AI Evals Tool

DevFeed: [Selecting The Right AI Evals Tool](<https://devfeed.tech/articles/selecting-the-right-ai-evals-tool-18788.md>)

Original publisher: [Read original article](<https://hamel.dev/blog/posts/eval-tools/>)

Author: Hamel Husain

Published: 2025-10-01T07:00:00Z

Content type: article

Language: en

Sources: [Hamel Husain](<https://devfeed.tech/sources/hamel-husain.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>), [Developer experience](<https://devfeed.tech/topics/developer-experience.md>), [LangChain](<https://devfeed.tech/topics/langchain.md>), [phoenix](<https://devfeed.tech/topics/phoenix.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-evals](<https://devfeed.tech/tags/ai-evals.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [developer-experience](<https://devfeed.tech/tags/developer-experience.md>), [evals](<https://devfeed.tech/tags/evals.md>), [langchain](<https://devfeed.tech/tags/langchain.md>), [phoenix](<https://devfeed.tech/tags/phoenix.md>), [quality](<https://devfeed.tech/tags/quality.md>)

## AI overview

The article examines how to select AI evaluation tools, arguing that no single tool is best for every team. It compares approaches from LangSmith, Braintrust, and Arize Phoenix through a shared assignment and highlights workflow, developer experience, SDK ergonomics, documentation, integrations, and human-in-the-loop support as selection criteria.

## Source excerpt

Over the past year, I've focused heavily on AI Evals, both in my consulting work and teaching. A question I get constantly is, "What's the best tool for evals?". I've always resisted answering directly for two reasons. First, people focus too much on tools instead of the process, thinking the tool will be an off-the-shelf solution when it rarely is. Second, the tools change so quickly that comparisons become outdated immediately. Having used many of the popular eval tools, I can genuinely say that no single one is superior in every dimension. The "best" tool depends on your team's skillset, technical stack, and maturity. Instead of a feature-by-feature comparison, I think it's more valuable to show you how a panel of data scientists skilled in evals assesses these tools. As part of my AI Evals course, we had three of the most dominant vendors--Langsmith, Braintrust, and Arize Phoenix complete the same homework assignment. This gave us a unique opportunity to see how they tackle the exact same challenge. We recorded the entire process and live commentary, which is available below. We think this might be helpful in learning about the kinds of things you should consider when selecting a tool for your team. Thanks to Shreya Shankar and Bryan Bischof for serving as the panelists (alongside me). Langsmith With Harrison Chase, CEO of LangChain. Braintrust With Wayde Gilliam, former developer relations at Braintrust. Arize Phoenix With SallyAnn DeLucia, Technical AI Product Leader at Arize. Criteria for Assessing AI Evals Tools Here are themes that consistently surfaced during our review. 1. Workflow and Developer Experience Reducing friction is more important than any single feature. Concretely, you should be mindful of the time it takes to go from observing a failure to iterating on a solution. For example, we appreciated the ability to go from viewing a single trace to experimenting with that same trace in a playground. For some teams with data-science backgrounds, a note