# The structured data lake: How schema evolution enables the next generation of data platforms

DevFeed: [The structured data lake: How schema evolution enables the next generation of data platforms](<https://devfeed.tech/articles/the-structured-data-lake-how-schema-evolution-enables-the-next-generation-of-data-platforms-80357.md>)

Original publisher: [Read original article](<https://dlthub.com/blog/next-generation-data-platform>)

Author: Adrian Brudaru

Published: 2023-05-26T00:00:00Z

Content type: article

Language: en

Sources: [DLT Hub](<https://devfeed.tech/sources/dlt-hub.md>)

Topics: [schema-evolution](<https://devfeed.tech/topics/schema-evolution.md>), [data-platforms](<https://devfeed.tech/topics/data-platforms.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [Structured-data](<https://devfeed.tech/topics/structured-data.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [database-optimization](<https://devfeed.tech/topics/database-optimization.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-management](<https://devfeed.tech/tags/data-management.md>), [data-mesh](<https://devfeed.tech/tags/data-mesh.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [product](<https://devfeed.tech/tags/product.md>), [schema-evolution](<https://devfeed.tech/tags/schema-evolution.md>), [structured-data](<https://devfeed.tech/tags/structured-data.md>), [tutorials](<https://devfeed.tech/tags/tutorials.md>)

## AI overview

The article compares data warehouses, data lakes, and lakehouses, then argues for structuring data during ingestion through schema evolution. It describes how the open-source Python library dlt applies schemas as data is loaded to improve consistency, governance, maintenance, and query performance.

## Source excerpt

## What is schema evolution? In the fast-paced world of data, the only constant is change, and it usually comes unannounced. ### **Schema on read** Schema on read means your data does not have a schema, but your consumer expects one. So when they read, they define the schema, and if the unstructured data does not have the same schema, issues happen. ### **Schema on write** So, to avoid things breaking on running, you would want to define a schema upfront - hence you would structure the data.