# Evaluating AI at Scale: How Thumbtack Approaches Reliability, Safety, and Quality in GenAI

DevFeed: [Evaluating AI at Scale: How Thumbtack Approaches Reliability, Safety, and Quality in GenAI](<https://devfeed.tech/articles/evaluating-ai-at-scale-how-thumbtack-approaches-reliability-safety-and-quality-in-genai-24724.md>)

Original publisher: [Read original article](<https://medium.com/thumbtack-engineering/evaluating-ai-at-scale-how-thumbtack-approaches-reliability-safety-and-quality-in-genai-f75d0211ac54?source=rss----1199c607a13f---4>)

Author: Thumbtack Engineering

Published: 2026-04-29T00:16:16Z

Content type: article

Language: en

Sources: [Thumbtack Engineering - Medium](<https://devfeed.tech/sources/thumbtack-engineering-medium.md>)

Topics: [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [genai](<https://devfeed.tech/topics/genai.md>), [trust & safety](<https://devfeed.tech/topics/trust-safety.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-evaluation](<https://devfeed.tech/tags/ai-evaluation.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [genai](<https://devfeed.tech/tags/genai.md>), [research](<https://devfeed.tech/tags/research.md>), [safety](<https://devfeed.tech/tags/safety.md>), [technical](<https://devfeed.tech/tags/technical.md>)

## AI overview

Thumbtack describes a learning-driven, exploratory approach to evaluating generative AI experiences. Its strategy combines cross-functional insights with a parallel-path MVP evaluation system to address probabilistic outputs, unsupported claims, harmful content, changing model behavior, and trust-related risks.

## Source excerpt

A practical look at how Thumbtack navigates evaluation for emerging AI experiences and what we've learned along the way. By: Shishir Dash, Director of Applied Science & Teja Venkat Kolli, Senior Applied Scientist Evaluating AI at ScaleIntroduction AI is reshaping how people interact with products, and Thumbtack is no exception. We're introducing AI into more aspects of our customer and local service professional (pro) experiences -- from helping customers articulate what they need, to generating helpful summaries, to offering clearer explanations of how pros may fit those needs. But evaluating generative AI is uniquely challenging. Unlike traditional software, its outputs are probabilistic, wide-ranging, and capable of subtle errors: mistakes in tone, inaccuracies, unsupported claims, or harmful assumptions. Rather than attempt to formalize a single rigid evaluation framework, we've taken a learning-driven, exploratory approach, pairing cross-functional insights with a parallel-path MVP evaluation system. This balanced strategy allows us to move quickly while staying grounded in safety, responsibility, and quality. Why AI Evaluation Matters Evaluation is essential because generative AI can produce unsupported or overly strong claims. Sometimes it can misinterpret user intent or vary in style or tone from one version to the next. It can sometimes generate harmful, biased, or inappropriate content. It can also drift over time due to model updates or prompt changes. For a marketplace built on trust, these challenges matter. Customers need accurate guidance; pros need fair, clear representation. Evaluation helps ensure every AI interaction strengthens and not undermines that trust. Our Approach: Exploration, Learning, and MVP Paths The landscape of AI evaluation is still evolving. New research, tooling, and patterns emerge every month. Rather than over-commit to a single approach, we've adopted a mixed strategy rooted in: Exploration and fast learning across multiple pro