# DABStep: Data Agent Benchmark for Multi-step Reasoning

DevFeed: [DABStep: Data Agent Benchmark for Multi-step Reasoning](<https://devfeed.tech/articles/dabstep-data-agent-benchmark-for-multi-step-reasoning-7155.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/dabstep>)

Author: Alex Egg; Martin Iglesias Goyanes; Friso Kingma; Andreu Mora; Leandro von Werra; Thomas Wolf; Aymeric Roucher

Published: 2025-02-04T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Data analysis](<https://devfeed.tech/topics/data-analysis.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [data](<https://devfeed.tech/tags/data.md>), [data-analysis](<https://devfeed.tech/tags/data-analysis.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [llms](<https://devfeed.tech/tags/llms.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>)

## AI overview

Adyen and Hugging Face introduce DABstep, a benchmark of more than 450 data analysis tasks for evaluating state-of-the-art LLMs and AI agents. The article reports that the strongest reasoning-based agents reached only 16% accuracy, revealing a substantial gap in current systems' ability to handle rigorous, context-rich, real-world data analysis.

## Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.