# How we build and evaluate our MCP server for SRE agents

DevFeed: [How we build and evaluate our MCP server for SRE agents](<https://devfeed.tech/articles/how-we-build-and-evaluate-our-mcp-server-for-sre-agents-4984.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/benchmarking-the-clickstack-mcp-server-with-hdx-evals>)

Author: Brandon Pereira

Published: 2026-07-22T00:00:00Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [incident](<https://devfeed.tech/topics/incident.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [claude](<https://devfeed.tech/tags/claude.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [framework](<https://devfeed.tech/tags/framework.md>), [incident](<https://devfeed.tech/tags/incident.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [model-context-protocol](<https://devfeed.tech/tags/model-context-protocol.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [sql](<https://devfeed.tech/tags/sql.md>), [sre](<https://devfeed.tech/tags/sre.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

## AI overview

ClickHouse explains its open-source hdx-evals framework for benchmarking the ClickStack MCP server against direct ClickHouse SQL access for AI agents investigating production incidents.

## Source excerpt

A behind-the-scenes look at hdx-evals, the open-source framework we built to benchmark the ClickStack MCP server against a raw SQL baseline using deterministic synthetic incidents, sandboxed Claude agents, and blind LLM grading -- and why the MCP scored hi