# DataExpert.io Newsletter

A newsletter dedicated to talking about data engineering, AI, and data science trends

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## The Complete Databricks Learning Roadmap for 2026

DevFeed: [The Complete Databricks Learning Roadmap for 2026](<https://devfeed.tech/articles/the-complete-databricks-learning-roadmap-for-2026-27257.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/the-2026-mastering-databricks-roadmap>)

Author: Zach Wilson

Published: 2026-08-14T20:11:45Z

Content type: tutorial

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [databricks](<https://devfeed.tech/topics/databricks.md>), [Learning](<https://devfeed.tech/topics/learning.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [learn](<https://devfeed.tech/tags/learn.md>), [learning](<https://devfeed.tech/tags/learning.md>), [master](<https://devfeed.tech/tags/master.md>), [skip](<https://devfeed.tech/tags/skip.md>)

### AI overview

A 2026 learning roadmap for Databricks covering what to learn, what to skip, and the recommended order for mastering the platform.

### Source excerpt

What to learn, what to skip, and the right order to master Databricks.

## Junior data engineers build pipelines. Seniors build trust

DevFeed: [Junior data engineers build pipelines. Seniors build trust](<https://devfeed.tech/articles/junior-data-engineers-build-pipelines-seniors-build-trust-27241.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/data-engineering-was-never-about>)

Author: Sahar Massachi

Published: 2026-06-30T13:02:18Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [junior](<https://devfeed.tech/tags/junior.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [product-owner](<https://devfeed.tech/tags/product-owner.md>), [series](<https://devfeed.tech/tags/series.md>), [trust](<https://devfeed.tech/tags/trust.md>)

### AI overview

The article contrasts the work associated with junior and senior data engineers, framing pipeline building as junior work and trust building as senior work. It is identified as Part 3 of a data engineering essentials series.

### Source excerpt

You are the product owner for data. (Part 3 of our DE essentials series)

## RAG and note-taking tools are not sufficient to save a business

DevFeed: [RAG and note-taking tools are not sufficient to save a business](<https://devfeed.tech/articles/a-well-architected-secretary-is-76-agents-in-a-trenchcoat-27240.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/ai-architecture-is-misunderstood>)

Author: Sahar Massachi

Published: 2026-05-11T16:02:23Z

Content type: opinion

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>)

Tags: [rag](<https://devfeed.tech/tags/rag.md>)

### AI overview

The article argues that RAG and note-taking tools alone will not save a business.

### Source excerpt

RAG and note takers won't save your business

## How do One Big Table and AI fit together

DevFeed: [How do One Big Table and AI fit together](<https://devfeed.tech/articles/how-do-one-big-table-and-ai-fit-together-27248.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/how-to-data-model-for-your-ai-context>)

Author: Zach Wilson

Published: 2026-04-06T20:54:07Z

Content type: tutorial

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>)

### AI overview

The article discusses how One Big Table relates to AI and states that dimensional modeling is dead.

### Source excerpt

Dimensional Modeling is dead

## Understanding Parquet Format for beginners

DevFeed: [Understanding Parquet Format for beginners](<https://devfeed.tech/articles/understanding-parquet-format-for-beginners-27252.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/parquet-can-shrink-your-data-100x>)

Author: Zach Wilson

Published: 2026-03-03T21:52:56Z

Content type: tutorial

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [parquet](<https://devfeed.tech/topics/parquet.md>)

Tags: [beginners](<https://devfeed.tech/tags/beginners.md>), [format](<https://devfeed.tech/tags/format.md>), [parquet](<https://devfeed.tech/tags/parquet.md>)

### AI overview

A beginner-oriented walkthrough of the Parquet file format.

### Source excerpt

A walk through of the most important file format to ever exist

## Databricks is abstracting away physical data engineering controls

DevFeed: [Databricks is abstracting away physical data engineering controls](<https://devfeed.tech/articles/databricks-is-no-longer-about-tuning-knobs-27242.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/databricks-is-for-data-analysts-not>)

Author: Zach Wilson

Published: 2026-02-24T01:03:11Z

Content type: opinion

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [databricks](<https://devfeed.tech/topics/databricks.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partition](<https://devfeed.tech/tags/partition.md>), [partitioning](<https://devfeed.tech/tags/partitioning.md>), [sorting](<https://devfeed.tech/tags/sorting.md>), [spark](<https://devfeed.tech/tags/spark.md>)

### AI overview

This opinion article argues that Databricks is shifting away from hands-on data engineering by abstracting physical data modeling through features such as liquid clustering and predictive optimization. It also criticizes Databricks' support for managed Apache Iceberg tables after acquiring Tabular.

### Source excerpt

Databricks abstracts away almost all of the data engineering skills. Liquid clustering is the first place where things will get messy!

## The 2026 AI Data Engineer Roadmap

DevFeed: [The 2026 AI Data Engineer Roadmap](<https://devfeed.tech/articles/the-2026-ai-data-engineer-roadmap-27256.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/the-2026-ai-data-engineer-roadmap>)

Author: Zach Wilson

Published: 2026-02-05T20:26:37Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [flink](<https://devfeed.tech/topics/flink.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-coding-agents](<https://devfeed.tech/tags/ai-coding-agents.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>)

### AI overview

This article presents a 2026 roadmap for data engineering as AI automates more pipeline, SQL, Spark, testing, migration, and operational work. It compares responsibilities that are becoming more automated with areas such as system design, deep technical debt, performance tradeoffs, and organizational constraints where human expertise remains important.

### Source excerpt

And how to avoid getting replaced

## Processing 1 TB with DuckDB in less than 30 seconds

DevFeed: [Processing 1 TB with DuckDB in less than 30 seconds](<https://devfeed.tech/articles/processing-1-tb-with-duckdb-in-less-than-30-seconds-27250.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/i-processed-1-tb-with-duckdb-in-30>)

Author: Matt Martin

Published: 2025-12-23T19:58:37Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [DuckDB](<https://devfeed.tech/topics/duckdb.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [data](<https://devfeed.tech/topics/data.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Python](<https://devfeed.tech/topics/python.md>), [Processes](<https://devfeed.tech/topics/processes.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [data](<https://devfeed.tech/tags/data.md>), [duckdb](<https://devfeed.tech/tags/duckdb.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [process](<https://devfeed.tech/tags/process.md>), [python](<https://devfeed.tech/tags/python.md>)

### AI overview

This article benchmarks DuckDB on progressively larger datasets, reporting that it read approximately 200 GB in under 10 seconds and 500 GB in about 40 seconds. It also describes generating a 1 TB dataset in roughly 70 minutes on an M2 Pro Mac with 16 GB of RAM using 10 parallel workers.

### Source excerpt

And so can you

## Data security shouldn't be an afterthought

DevFeed: [Data security shouldn't be an afterthought](<https://devfeed.tech/articles/data-security-shouldn-t-be-an-afterthought-27249.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/how-to-secure-your-data-a-practical>)

Author: Zach Wilson

Published: 2025-11-26T16:36:35Z

Content type: tutorial

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Security](<https://devfeed.tech/topics/security.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [guide](<https://devfeed.tech/tags/guide.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

A practical guide for data engineers about treating data security as a core consideration rather than an afterthought.

### Source excerpt

A practical guide for Data Engineers

## A Datestamp-Based Alternative to Slowly Changing Dimensions

DevFeed: [A Datestamp-Based Alternative to Slowly Changing Dimensions](<https://devfeed.tech/articles/scd-2-considered-harmful-part-2-27254.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/stop-using-slowly-changing-dimensions>)

Author: Sahar Massachi

Published: 2025-11-04T20:44:28Z

Content type: tutorial

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [Structured-data](<https://devfeed.tech/topics/structured-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [data](<https://devfeed.tech/tags/data.md>), [data-recovery](<https://devfeed.tech/tags/data-recovery.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [sql](<https://devfeed.tech/tags/sql.md>), [time](<https://devfeed.tech/tags/time.md>)

### AI overview

This second installment in a data warehousing series explains the difficulty of querying historical state across multiple tables with slowly changing dimensions. It proposes datestamps and idempotent pipelines as a simpler approach for historical queries, backfills, data recovery, and operational alerts.

### Source excerpt

Date stamp your data!

## A Date-Stamped Data Warehouse Setup

DevFeed: [A Date-Stamped Data Warehouse Setup](<https://devfeed.tech/articles/the-data-warehouse-setup-no-one-taught-you-27258.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/the-data-warehouse-setup-no-one-taught>)

Author: Sahar Massachi

Published: 2025-10-24T21:03:30Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [data](<https://devfeed.tech/topics/data.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>)

Tags: [ab-testing](<https://devfeed.tech/tags/ab-testing.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [storage](<https://devfeed.tech/tags/storage.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article presents date-stamping data as a simple organizing principle for building resilient data warehouses and pipelines. It explains how this approach can manage changes over time, support experimentation and metrics, and work with systems including Hive metastore, Iceberg, Delta, and Hudi.

### Source excerpt

Storage is cheap, your time is not!

## The 2025 AI + Data Engineering Roadmap

DevFeed: [The 2025 AI + Data Engineering Roadmap](<https://devfeed.tech/articles/the-2025-ai-data-engineering-roadmap-27255.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/the-2025-breaking-into-data-engineering-roadmap>)

Author: Zach Wilson

Published: 2025-10-17T22:35:45Z

Content type: tutorial

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Python](<https://devfeed.tech/topics/python.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [data-modeling](<https://devfeed.tech/topics/data-modeling.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [airflow](<https://devfeed.tech/tags/airflow.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [count](<https://devfeed.tech/tags/count.md>), [course](<https://devfeed.tech/tags/course.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-modeling](<https://devfeed.tech/tags/data-modeling.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [framer](<https://devfeed.tech/tags/framer.md>), [newsletter](<https://devfeed.tech/tags/newsletter.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [python](<https://devfeed.tech/tags/python.md>), [rag](<https://devfeed.tech/tags/rag.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [right-join](<https://devfeed.tech/tags/right-join.md>), [spark](<https://devfeed.tech/tags/spark.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

A 2025 roadmap for entering data engineering, covering foundational SQL and Python skills, distributed computing, orchestration, data modeling, data quality, AI and data integrations, portfolio projects, and personal branding.

### Source excerpt

Getting a data engineering job is complicated.

## DuckDB benchmarked against Spark

DevFeed: [DuckDB benchmarked against Spark](<https://devfeed.tech/articles/duckdb-benchmarked-against-spark-27243.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/duckdb-can-be-100x-faster-than-spark>)

Author: Matt Martin

Published: 2025-09-22T20:13:34Z

Content type: comparison

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [DuckDB](<https://devfeed.tech/topics/duckdb.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>)

Tags: [duckdb](<https://devfeed.tech/tags/duckdb.md>), [spark](<https://devfeed.tech/tags/spark.md>)

### AI overview

A comparison article about benchmarking DuckDB against Apache Spark.

### Source excerpt

You Don't Always Need A Sledgehammer

## Migrating 13,000 Iceberg Tables in 4 hours to Glue Catalog

DevFeed: [Migrating 13,000 Iceberg Tables in 4 hours to Glue Catalog](<https://devfeed.tech/articles/migrating-13-000-iceberg-tables-in-4-hours-to-glue-catalog-27247.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/how-i-migrated-13000-iceberg-tables>)

Author: Zach Wilson

Published: 2025-09-17T22:12:17Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>)

### AI overview

The article describes migrating 13,000 Apache Iceberg tables to the Glue Catalog in four hours. The supplied text also reports a message saying that Tabular would be sunsetted within 24 hours.

### Source excerpt

At midnight on September 16th, Jason Reid (data engineering advocate at Databricks) messages me on LinkedIn saying, "Tabular will be sunsetted in 24 hours.

## Three Free Tech Bootcamps Covering Solutions Architecture, Cybersecurity, and Data Engineering

DevFeed: [Three Free Tech Bootcamps Covering Solutions Architecture, Cybersecurity, and Data Engineering](<https://devfeed.tech/articles/three-free-tech-bootcamps-that-could-change-your-career-27259.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/three-free-tech-bootcamps-that-could>)

Author: Zach Wilson

Published: 2025-08-27T15:02:55Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Bootcamp](<https://devfeed.tech/topics/bootcamp.md>), [Cybersecurity](<https://devfeed.tech/topics/cybersecurity.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>)

Tags: [career](<https://devfeed.tech/tags/career.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [free](<https://devfeed.tech/tags/free.md>), [tech](<https://devfeed.tech/tags/tech.md>)

### AI overview

The article presents three free technology bootcamps covering solutions architecture, cybersecurity foundations, and data engineering basics.

### Source excerpt

How to become a Solutions Architect, Mastering the Foundations of Cybersecurity and the Absolute Basics in Data Engineering

## Building, Scaling and Thinking in the Age of AI with Chip Huyen

DevFeed: [Building, Scaling and Thinking in the Age of AI with Chip Huyen](<https://devfeed.tech/articles/navigating-ai-s-new-frontier-with-chip-huyen-27251.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/navigating-ais-new-frontier-with>)

Author: Zach Wilson

Published: 2025-08-25T13:03:24Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [scaling](<https://devfeed.tech/topics/scaling.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [building](<https://devfeed.tech/tags/building.md>), [scaling](<https://devfeed.tech/tags/scaling.md>)

### AI overview

An article about building, scaling, and thinking in the age of AI, featuring Chip Huyen.

### Source excerpt

Building, Scaling and Thinking in the Age of AI

## Stopping Silent Failures for Meta's Fake Accounts Pipeline

DevFeed: [Stopping Silent Failures for Meta's Fake Accounts Pipeline](<https://devfeed.tech/articles/stopping-silent-failures-for-meta-s-fake-accounts-pipeline-27253.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/saving-metas-fake-accounts-pipeline>)

Author: Zach Wilson

Published: 2025-08-12T19:38:14Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Meta](<https://devfeed.tech/topics/meta.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>)

Tags: [fake-accounts](<https://devfeed.tech/tags/fake-accounts.md>), [meta](<https://devfeed.tech/tags/meta.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>)

### AI overview

An article about stopping silent failures in Meta's fake accounts pipeline.

### Source excerpt

Data Orchestration Challenges I Faced at Airbnb, Netflix & Facebook - Part IV

## How I got a 12x speed up in a 50 TB pipeline at Meta

DevFeed: [How I got a 12x speed up in a 50 TB pipeline at Meta](<https://devfeed.tech/articles/how-i-got-a-12x-speed-up-in-a-50-tb-pipeline-at-meta-27246.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/how-i-got-a-12x-speed-up-in-a-50>)

Author: Zach Wilson

Published: 2025-08-04T19:39:26Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Meta](<https://devfeed.tech/topics/meta.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [data](<https://devfeed.tech/topics/data.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [meta](<https://devfeed.tech/tags/meta.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [speed](<https://devfeed.tech/tags/speed.md>)

### AI overview

The article describes a 12x speed improvement in a 50 TB data pipeline at Meta and places the topic within data orchestration challenges discussed across Airbnb, Netflix, and Facebook.

### Source excerpt

Data Orchestration Challenges I Faced at Airbnb, Netflix & Facebook - Part III