# Data Engineering Is the Cache Invalidation Problem ft. Josh Wills

DevFeed: [Data Engineering Is the Cache Invalidation Problem ft. Josh Wills](<https://devfeed.tech/articles/data-engineering-is-the-cache-invalidation-problem-ft-josh-wills-84268.md>)

Original publisher: [Read original article](<https://motherduck.com/blog/cache-invalidation-problem-data-engineering-josh-wills>)

Author: Simon Späti

Published: 2026-10-09T00:00:00Z

Content type: article

Language: en

Sources: [MotherDuck](<https://devfeed.tech/sources/motherduck.md>)

Topics: [data observability](<https://devfeed.tech/topics/data-observability.md>), [cache-invalidation](<https://devfeed.tech/topics/cache-invalidation.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [engineering-leadership](<https://devfeed.tech/topics/engineering-leadership.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [dbt](<https://devfeed.tech/tags/dbt.md>), [duckdb](<https://devfeed.tech/tags/duckdb.md>)

## AI overview

An interview with data engineer Josh Wills explores how AI agents affect large-scale data engineering. He says agents can handle routine tasks such as flaky tests and database migrations, but long-running Spark and Ray pipelines lack the fast feedback loops that help agents produce reliable, efficient work. Wills argues that architectural decisions, communicating trade-offs, and reviewing agent output remain central to engineering, and describes data engineering as the cache invalidation problem in a batch-processing context.

## Source excerpt

Josh Wills, creator of dbt-duckdb and former head of data engineering at Slack, on why the model barely matters, why agents still write ugly and expensive data pipelines, and why data engineering is still the cache invalidation problem.