# data lake

A data lake is a storage repository that stores large volumes of raw data in its original formats and supports data management, analytics, and processing at scale.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## How Partition Access Visualizations Reduced our Data Lake S3 Cost by 33%

DevFeed: [How Partition Access Visualizations Reduced our Data Lake S3 Cost by 33%](<https://devfeed.tech/articles/how-partition-access-visualizations-reduced-our-data-lake-s3-cost-by-33-27427.md>)

Original publisher: [Read original article](<https://engineeringblog.yelp.com/2026/05/partition-access-visualizations.html>)

Author: Nick Del Nano, Data Streaming

Published: 2026-05-21T00:00:00Z

Content type: article

Language: en

Sources: [Yelp](<https://devfeed.tech/sources/yelp.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [AWS IAM](<https://devfeed.tech/topics/aws-iam.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [aws](<https://devfeed.tech/tags/aws.md>), [data](<https://devfeed.tech/tags/data.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [iam](<https://devfeed.tech/tags/iam.md>), [partition](<https://devfeed.tech/tags/partition.md>), [s3](<https://devfeed.tech/tags/s3.md>)

### AI overview

Yelp describes visualizations that map partition keys against access-event timestamps to reveal daily batch jobs, backfills, and ad hoc queries. The resulting usage attribution supported Apache Iceberg migration and storage-efficiency work that reduced the cost of its petabyte-scale data lake by 33%.

### Source excerpt

Introduction In large analytics environments, data teams often struggle to answer deceptively simple questions, like who their stakeholders are and how their data is being used. At Yelp, we address this by visualizing access patterns, plotting time-based partition key values against access event timestamps. These visualizations reveal distinct usage signatures - ad hoc queries, daily batch jobs, and periodic backfills - allowing data owners to understand their stakeholders and use cases. This deeper insight into data usage has enabled high-impact platform initiatives including migrating thousands of tables to Apache Iceberg format and identifying storage efficiencies which reduced the cost of...

## ClickHouse Announces Integration with Microsoft OneLake for Federated Analytics

DevFeed: [ClickHouse Announces Integration with Microsoft OneLake for Federated Analytics](<https://devfeed.tech/articles/clickhouse-strengthens-collaboration-with-microsoft-through-microsoft-onelake-integration-for-seamless-data-interoperability-5418.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/microsoft-collaboration-onelake>)

Author: Melvyn Peignon

Published: 2025-11-18T00:00:00Z

Content type: release

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [interoperability](<https://devfeed.tech/topics/interoperability.md>), [Microsoft](<https://devfeed.tech/topics/microsoft.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>)

Tags: [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [databases](<https://devfeed.tech/tags/databases.md>), [fabric](<https://devfeed.tech/tags/fabric.md>), [integration](<https://devfeed.tech/tags/integration.md>), [interoperability](<https://devfeed.tech/tags/interoperability.md>), [latency](<https://devfeed.tech/tags/latency.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

ClickHouse announced an integration with Microsoft OneLake, the unified data lake service within Microsoft Fabric. The integration exposes OneLake Iceberg tables to ClickHouse so users can query and analyze data across ClickHouse and Fabric, supporting real-time analytical workloads.

### Source excerpt

ClickHouse today announced the availability of a powerful new integration with Microsoft OneLake, the unified data lake service within Microsoft Fabric.

## Project Teleport: Cost-Effective and Scalable Kafka Data Processing at Block

DevFeed: [Project Teleport: Cost-Effective and Scalable Kafka Data Processing at Block](<https://devfeed.tech/articles/project-teleport-cost-effective-and-scalable-kafka-data-processing-at-block-29015.md>)

Original publisher: [Read original article](<https://code.cash.app/project-teleport>)

Author: Unni Krishnan

Published: 2025-03-20T00:00:00Z

Content type: article

Language: en

Sources: [Cash App Code Blog](<https://devfeed.tech/sources/cash-app-code-blog.md>)

Topics: [data-processing](<https://devfeed.tech/topics/data-processing.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [migration](<https://devfeed.tech/topics/migration.md>), [databricks](<https://devfeed.tech/topics/databricks.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>)

Tags: [acquisition](<https://devfeed.tech/tags/acquisition.md>), [amazon](<https://devfeed.tech/tags/amazon.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [delta-lake](<https://devfeed.tech/tags/delta-lake.md>), [emr](<https://devfeed.tech/tags/emr.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [redshift](<https://devfeed.tech/tags/redshift.md>), [s3](<https://devfeed.tech/tags/s3.md>), [scale](<https://devfeed.tech/tags/scale.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

Project Teleport is Block's cross-region Kafka data-processing system for integrating Afterpay's Sydney-hosted data lake into Block's US-based ecosystem. Built with Delta Lake, Spark on Databricks, and object storage, it supports migration of legacy pipelines and reduced cloud egress costs by USD 540,000 per year.

### Source excerpt

Teleport achieves efficient and reliable cross-region Kafka data processing at scale. Using this approach, Afterpay data team reduced cloud egress costs by USD 540,000 per year.

## Data Quality at Udemy -- Part 1

DevFeed: [Data Quality at Udemy -- Part 1](<https://devfeed.tech/articles/data-quality-at-udemy-part-1-26351.md>)

Original publisher: [Read original article](<https://medium.com/udemy-engineering/data-quality-at-udemy-part-1-63e3b099ff81?source=rss----19c6d3367ed4---4>)

Author: Murat Migdisoglu

Published: 2023-09-06T22:01:16Z

Content type: article

Language: en

Sources: [Udemy Engineering](<https://devfeed.tech/sources/udemy-engineering.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data-architecture](<https://devfeed.tech/topics/data-architecture.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [data-architecture](<https://devfeed.tech/tags/data-architecture.md>), [data-catalog](<https://devfeed.tech/tags/data-catalog.md>), [data-governance](<https://devfeed.tech/tags/data-governance.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [data-lineage](<https://devfeed.tech/tags/data-lineage.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [data-quality-management](<https://devfeed.tech/tags/data-quality-management.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [principal-engineer](<https://devfeed.tech/tags/principal-engineer.md>), [quality](<https://devfeed.tech/tags/quality.md>), [spark](<https://devfeed.tech/tags/spark.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

This article describes Udemy's efforts to improve data quality by establishing an end-to-end data lineage solution. It explains how distributed data ownership and self-service analytics make lineage important for impact analysis, change management, and identifying unused columns or orphan tables.

### Source excerpt

Data Quality at Udemy -- Part 1Data Lineage Demystified- Why it Matters and How to Leverage its Magic for Informed Business Success! In late 2020, upon joining Udemy as a principal engineer for the data platform team, my focus shifted toward enhancing data quality within the organization. My journey began with conducting a comprehensive poll across the data organization, aimed at identifying the key pain points of data users. The results of the poll were eye-opening, revealing that 78% of users considered the absence of data provenance/lineage as a data quality issue. Furthermore, it was obvious that for a vast majority of the users, the inability to track data lineage was an important problem in impact analysis and detecting unused columns or orphan tables in the system. Inspired by these insights, I took the initiative to propose and launch two transformative projects. The first one, which is the subject of this article, is an ambitious initiative to establish a comprehensive end-to-end data lineage solution that will revolutionize our data ecosystem. The second project centers around data monitoring, which will be explored in another post. Udemy's sophisticated data architecture revolves around a data lake fed by diverse pipelines: system logs, streaming data from services, CDC listeners for replicated service databases, and more. The backbone of data transformations lies in Hive and Spark, while Airflow takes charge of orchestrating thousands of these pipelines. Unraveling Data Flow Complexity: Data Lineage in Growing Data Driven Organizations In the early stages of an organization's data-driven journey, data lineage may not be deemed crucial. With just a few pipelines managed by a centralized team, the dependency tree of the workflow orchestration typically suffices to comprehend the relationships between data entities. However, as the business scales up, relying on a single centralized team for all data flows becomes a bottleneck. Consequently, data organizatio

## The unnecessary hype strategy behind Microsoft Fabric

DevFeed: [The unnecessary hype strategy behind Microsoft Fabric](<https://devfeed.tech/articles/the-unnecessary-hype-strategy-behind-microsoft-fabric-40838.md>)

Original publisher: [Read original article](<https://mutto.fyi/posts/2023/05/the-unnecessary-hype-fabric/>)

Published: 2023-05-29T00:00:00Z

Content type: opinion

Language: en

Sources: [Mutt0-ds Notes](<https://devfeed.tech/sources/mutt0-ds-notes.md>)

Topics: [Microsoft](<https://devfeed.tech/topics/microsoft.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [etl](<https://devfeed.tech/topics/etl.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [azure](<https://devfeed.tech/tags/azure.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [copilot](<https://devfeed.tech/tags/copilot.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [etl](<https://devfeed.tech/tags/etl.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [microsoft-azure](<https://devfeed.tech/tags/microsoft-azure.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

An opinion article examines Microsoft Fabric, a unified data platform announced at Microsoft Build. It describes Fabric's integration of data storage, processing, ETL, analytics, and business intelligence tools, while criticizing Microsoft's secrecy and hype-oriented launch strategy and noting that the platform was still in beta.

### Source excerpt

If you are into Data Engineering in Microsoft Azure Cloud Environment, you probaly heard about Microsoft Fabric being announced last week...

## Trino Summit 2022: Sessions, speakers, and event details

DevFeed: [Trino Summit 2022: Sessions, speakers, and event details](<https://devfeed.tech/articles/trino-summit-2022-will-be-legendary-8688.md>)

Original publisher: [Read original article](<https://trino.io/blog/2022/09/22/trino-summit-2022-teaser.html>)

Author: Brian Olsen, Dain Sundstrom

Published: 2022-09-22T00:00:00Z

Content type: news

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [data lake](<https://devfeed.tech/topics/data-lake.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>)

Tags: [architectures](<https://devfeed.tech/tags/architectures.md>), [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [conference](<https://devfeed.tech/tags/conference.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [event](<https://devfeed.tech/tags/event.md>), [lyft](<https://devfeed.tech/tags/lyft.md>), [mpp](<https://devfeed.tech/tags/mpp.md>), [routing](<https://devfeed.tech/tags/routing.md>), [summit](<https://devfeed.tech/tags/summit.md>)

### AI overview

This article announces Trino Summit 2022, a free hybrid conference taking place on November 10th, and previews selected sessions and speakers. Topics include Trino's open source project, query federation, large-scale ETL at Lyft, autoscaling, and fault-tolerant execution.

### Source excerpt

Commander Bun Bun is back and this year we have an exciting lineup of speakers. Topics range from architectures like data mesh and data lakehouse, to running Trino at scale with fault-tolerant execution, and query federation. This conference is free and takes place on November 10th. The summit is a hybrid event for in-person and virtual attendance. Find out more details below!

## Integrating Confluent Schema Registry with Apache Spark applications

DevFeed: [Integrating Confluent Schema Registry with Apache Spark applications](<https://devfeed.tech/articles/integrating-confluent-schema-registry-with-apache-spark-applications-24745.md>)

Original publisher: [Read original article](<https://medium.com/yazio-engineering/integrating-confluent-schema-registry-with-apache-spark-applications-d3426e33bc51?source=rss----65bd178b00af---4>)

Author: Dominik Liebler

Published: 2022-01-24T08:04:19Z

Content type: tutorial

Language: en

Sources: [YAZIO Engineering - Medium](<https://devfeed.tech/sources/yazio-engineering-medium.md>)

Topics: [Kafka](<https://devfeed.tech/topics/kafka.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Kotlin](<https://devfeed.tech/topics/kotlin.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [ceph](<https://devfeed.tech/topics/ceph.md>), [JSON Schema](<https://devfeed.tech/topics/json-schema.md>)

Tags: [apache-spark](<https://devfeed.tech/tags/apache-spark.md>), [backpressure](<https://devfeed.tech/tags/backpressure.md>), [ceph](<https://devfeed.tech/tags/ceph.md>), [confluent](<https://devfeed.tech/tags/confluent.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [data-pipeline](<https://devfeed.tech/tags/data-pipeline.md>), [json](<https://devfeed.tech/tags/json.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [kotlin](<https://devfeed.tech/tags/kotlin.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [payload](<https://devfeed.tech/tags/payload.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [schema](<https://devfeed.tech/tags/schema.md>), [schemaregistry](<https://devfeed.tech/tags/schemaregistry.md>), [serialization](<https://devfeed.tech/tags/serialization.md>), [spark](<https://devfeed.tech/tags/spark.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

This engineering article explains YAZIO's data pipeline from mobile and web applications through Kafka and Spark Structured Streaming into a Ceph-based data lake. It discusses why schemas matter and describes replacing JSON with Apache Avro and Confluent Schema Registry to reduce message size while keeping schema information externally stored and cached.

### Source excerpt

At YAZIO, we believe in making decisions backed by data to help people live healthier lives through better nutrition. For each new and existing feature we want to evaluate how well it performs and how our users interact with it. In order to do so, we need a lot of data and we need to handle backpressure in our systems. To cope with that we use a Kafka cluster managed by Strimzi operators running in Kubernetes. The data itself is being ingested from our mobile and web apps via HTTP or TCP endpoints serialized into JSON and stored in Kafka by a small application written in Kotlin/JVM. Overview of our data pipeline architecture At the other end of the pipeline, different Spark Structured Streaming applications (also written in Kotlin) dump this information into our data lake residing in a Ceph bucket. They read data from Kafka, deserialize it, transform some of the fields and write Parquet files into the data lake using a new schema. Why schemas? Schemas play an important role in data pipelines because they give meaning and context to data. In a world without schemas we would still do random interpretations about the context and meaning of data every now and then when using it. As you might have guessed already this would lead to a lot of bugs and misunderstandings. Photo by EJ Strat https://unsplash.com/photos/VjWi56AWQ9k Similar to a legal contract that binds you to certain limits, a schema binds the data to certain limits and meaning which narrow down the need of interpretation. Choice of serialization formats At the time of writing, Confluent Schema Registry supports these three serialization formats: Apache Avro Protocol Buffers (protobuf) JSON Schema From those choices, only two really provide more than just validation of the data that is ingested and transmitted through our data pipelines. Avro and Protobuf also allow us to shrink the sizes of our topics because only the payload is contained in a message, while the repeating schema will not be stored. In the cas

## Building a Data Lake with Spark and Iceberg at Home to over-complicate shopping for a House

DevFeed: [Building a Data Lake with Spark and Iceberg at Home to over-complicate shopping for a House](<https://devfeed.tech/articles/building-a-data-lake-with-spark-and-iceberg-at-home-to-over-complicate-shopping-for-a-house-41493.md>)

Original publisher: [Read original article](<https://chollinger.com/blog/2021/12/building-a-data-lake-with-spark-and-iceberg-at-home-to-over-complicate-shopping-for-a-house/>)

Author: Christian Hollinger

Published: 2021-12-03T00:00:00Z

Content type: tutorial

Language: en

Sources: [Christian Hollinger](<https://devfeed.tech/sources/christian-hollinger.md>)

Topics: [data lake](<https://devfeed.tech/topics/data-lake.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [Python](<https://devfeed.tech/topics/python.md>), [geospatial](<https://devfeed.tech/topics/geospatial.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [data](<https://devfeed.tech/tags/data.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [geopandas](<https://devfeed.tech/tags/geopandas.md>), [geospatial](<https://devfeed.tech/tags/geospatial.md>), [geospatial-data](<https://devfeed.tech/tags/geospatial-data.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [hive](<https://devfeed.tech/tags/hive.md>), [iceberg](<https://devfeed.tech/tags/iceberg.md>), [presto](<https://devfeed.tech/tags/presto.md>), [python](<https://devfeed.tech/tags/python.md>), [scala](<https://devfeed.tech/tags/scala.md>), [spark](<https://devfeed.tech/tags/spark.md>), [sql](<https://devfeed.tech/tags/sql.md>), [trino](<https://devfeed.tech/tags/trino.md>)

### AI overview

A tutorial about building a local data lake with Spark, Iceberg, and Python to query demographic and geographic data while searching for a new neighborhood. It discusses data lakes, geospatial inconsistencies, and local systems design.

### Source excerpt

How I build what is essentially a self-service Data Lake at home to narrow down the search area for a new house, instead of using Zillow like a normal person, using Spark, Iceberg, and Python.

## Pinion -- The Load Framework Part-2

DevFeed: [Pinion -- The Load Framework Part-2](<https://devfeed.tech/articles/pinion-the-load-framework-part-2-26224.md>)

Original publisher: [Read original article](<https://medium.com/groupon-eng/pinion-the-load-framework-part-2-e6a47586e7be?source=rss----5c13a88f9872---4>)

Author: Saurabh Jain

Published: 2021-10-29T16:50:24Z

Content type: article

Language: en

Sources: [Groupon Engineering -- Medium](<https://devfeed.tech/sources/groupon-engineering-medium.md>)

Topics: [data lake](<https://devfeed.tech/topics/data-lake.md>), [Framework](<https://devfeed.tech/topics/framework.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [Transactions](<https://devfeed.tech/topics/transactions.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [acid](<https://devfeed.tech/tags/acid.md>), [audit](<https://devfeed.tech/tags/audit.md>), [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data-validation](<https://devfeed.tech/tags/data-validation.md>), [delta-lake](<https://devfeed.tech/tags/delta-lake.md>), [deltalake](<https://devfeed.tech/tags/deltalake.md>), [hdfs](<https://devfeed.tech/tags/hdfs.md>), [logging](<https://devfeed.tech/tags/logging.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [reporting](<https://devfeed.tech/tags/reporting.md>), [s3](<https://devfeed.tech/tags/s3.md>), [schema](<https://devfeed.tech/tags/schema.md>), [science](<https://devfeed.tech/tags/science.md>), [spark](<https://devfeed.tech/tags/spark.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transactions](<https://devfeed.tech/tags/transactions.md>)

### AI overview

This second post in the Pinion -- The Load Framework series explains how Pinion extends Apache Delta Lake APIs for slowly changing dimension operations. It covers data validation, compaction, auditing, streamlined logging, and chained APIs, and introduces Delta Lake capabilities such as ACID transactions, schema enforcement, batch and streaming interfaces, and time travel.

### Source excerpt

Pinion -- The Load Framework Part-2 This post is the 2nd part of the "Pinion -- The Load Framework" series. In case you have not read the 1st post, you can read it here. In this post, we are going to cover the following topics. How does Pinion use Delta Lake for SCD operations? Small file problem with Delta Lake and its resolution. Before we dive into the topics of this post, let's look at the definition of DeltaLake to set the context right. Apache Delta Lake - Apache Delta Lake is an open-source framework that enables the addition of ACID transactions support to a new data lake or an existing data lake created on top of S3, GCS, and HDFS. In addition to this, it provides other features such as scalable metadata handling, unified interface for both batch and streaming application, schema enforcement, time travel, and a rich interface of APIs to enable complex use cases like change-data-capture (CDC) and slowly-changing-dimension (SCD) operations. To keep the post concise and to the point, I won't go into much detail here about Delta Lake, since there is already great documentation available about it, that you can read it here. How does Pinion use Delta Lake for SCD operations? - Apache Delta Lake provides a rich set of APIs to handle slowly-changing dimensions, however, those APIs were not enough alone to build the features that we want to have in The Pinion Framework. So, we decided to enrich the APIs provided by Delta Lake by adding the following features to it: Data Validation Compaction Audit Streamlined logging infrastructure to make the data engineer's life easier during debugging of a failed job Chained APIs Let's dive a little further into the features that we had listed above. Data Validation -- By default, schema enforcement is enabled in Pinion for all the APIs where we have a need of inserting rows from source data(LRFs) into the target table. In case of a schema mismatch, Pinion raises an error and stops processing of further stages. It ensures the data a

## How we build the Image Gallery on trivago

DevFeed: [How we build the Image Gallery on trivago](<https://devfeed.tech/articles/how-we-build-the-image-gallery-on-trivago-28012.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2021-07-07-image-gallery-pipeline/>)

Author: Praneeth Peiris I want

Published: 2021-07-07T00:00:00Z

Content type: article

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>), [etl](<https://devfeed.tech/topics/etl.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Front end](<https://devfeed.tech/topics/frontend.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [amazon-web-services-aws](<https://devfeed.tech/tags/amazon-web-services-aws.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [backend](<https://devfeed.tech/tags/backend.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [etl](<https://devfeed.tech/tags/etl.md>), [frontend](<https://devfeed.tech/tags/frontend.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [platforms](<https://devfeed.tech/tags/platforms.md>)

### AI overview

trivago describes migrating its hotel image-gallery ETL pipeline from Amazon Web Services to Google Cloud Platform. The redesigned architecture uses Dataflow jobs to snapshot images and tags, fetch changed accommodations, build and validate sorted galleries, and stream approved changes to frontend teams through Kafka.

### Source excerpt

When was the last time you booked accommodation without checking its photos? Most probably never! Because having imagery information makes our decision-making process much easier and faster. How...

## Prophecy: Teamwork's Data Lake

DevFeed: [Prophecy: Teamwork's Data Lake](<https://devfeed.tech/articles/prophecy-teamwork-s-data-lake-35101.md>)

Original publisher: [Read original article](<https://engineroom.teamwork.com/prophecy-teamworks-data-lake-1a8ebb6dd3ae?source=rss----cea4eecd5960---4>)

Author: Joe Minichino

Published: 2020-07-17T11:41:10Z

Content type: opinion

Language: en

Sources: [Teamwork](<https://devfeed.tech/sources/teamwork.md>)

Topics: [data lake](<https://devfeed.tech/topics/data-lake.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [data](<https://devfeed.tech/topics/data.md>), [decision-making](<https://devfeed.tech/topics/decision-making.md>), [big-data](<https://devfeed.tech/topics/big-data.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Google Analytics](<https://devfeed.tech/topics/google-analytics.md>), [stripe](<https://devfeed.tech/topics/stripe.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [aws](<https://devfeed.tech/tags/aws.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [databases](<https://devfeed.tech/tags/databases.md>), [decision-making](<https://devfeed.tech/tags/decision-making.md>), [google-analytics](<https://devfeed.tech/tags/google-analytics.md>), [stripe](<https://devfeed.tech/tags/stripe.md>)

### AI overview

Teamwork describes reevaluating its analytics approach because data was distributed across product databases and third-party services, making it difficult to connect leads, product usage, and revenue. The article argues that a data lake can make analytics and data-informed decision-making more practical, while clarifying that data lakes and big data are not limited to enterprise-scale environments.

### Source excerpt

You need a Data Lake. The Context Teamwork has been around for more than 10 years. Starting out as a project management and work collaboration platform and later expanding into other areas, such as help-desk, chat, document management and CRM software. As the company has grown and evolved, data has grown, changed, expanded, diversified, fragmented, then changed again. Analytics in this landscape are not for the faint of heart. The Problem The issue at the base of the decision to start re-evaluating our analytics approach, at Teamwork, was because we have way too many data sources to make sense of all the data we collect. Much of the data ends up being "dark", underutilized, unleveraged. We have multiple databases shards, for each product, along with platform databases containing global customer information; and then we have 3rd parties such as Stripe, Marketo, Chart Mogul and Google Analytics to name the most important ones. We realized that making a connection between sales leads in Marketo, their usage of our products (GA and our DBs) and their revenue (our DBs, Stripe, Chart Mogul) was impossible. Well, it was possible, but incredibly painful and poorly automated, making it slow and error prone. And by addressing this problem we are getting a large number of benefits back, practically for free. You need Analytics Let's cut to the chase: you need analytics. No matter how big or small your business or enterprise is, you need analytics to take more informed business decisions. I do not doubt there are individuals with great gut-driven decision-making skills, but by and large, it's better if you look at data to make decisions. That's why you need a Data Lake. It could be a really small lake, it could be a pond or a puddle. But you need it for better decision making. Data Lakes, Big Data and other buzzwords Let's clear the air on a couple of misconceptions. The terms "Big Data" and "Data Lake" absolutely scream of corporate, of enterprise, of governance, compliance, r

## Data Lakes: Some thoughts on Hadoop, Hive, HBase, and Spark

DevFeed: [Data Lakes: Some thoughts on Hadoop, Hive, HBase, and Spark](<https://devfeed.tech/articles/data-lakes-some-thoughts-on-hadoop-hive-hbase-and-spark-41475.md>)

Original publisher: [Read original article](<https://chollinger.com/blog/2017/11/data-lakes-some-thoughts-on-hadoop-hive-hbase-and-spark/>)

Author: Christian Hollinger

Published: 2017-11-04T00:00:00Z

Content type: opinion

Language: en

Sources: [Christian Hollinger](<https://devfeed.tech/sources/christian-hollinger.md>)

Topics: [data lake](<https://devfeed.tech/topics/data-lake.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [data](<https://devfeed.tech/topics/data.md>), [etl](<https://devfeed.tech/topics/etl.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>)

Tags: [big-data](<https://devfeed.tech/tags/big-data.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [etl](<https://devfeed.tech/tags/etl.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [hbase](<https://devfeed.tech/tags/hbase.md>), [hive](<https://devfeed.tech/tags/hive.md>), [phoenix](<https://devfeed.tech/tags/phoenix.md>), [scala](<https://devfeed.tech/tags/scala.md>), [spark](<https://devfeed.tech/tags/spark.md>)

### AI overview

An overview of data lakes and the organizational and technical questions involved in building and using them. It discusses data storage, metadata, connected business systems, data models, ETL requirements, regulations, and tools including Hadoop, BigQuery, Hive, HBase, and Spark.

### Source excerpt

This article will talk about how organizations can make use of the wonderful thing that is commonly referred to as "Data Lake" - what constitutes a Data Lake, how probably should (and shouldn't) use it to gather insights and why evaluating technologies is just as important as understanding your data...

## A Journey Towards a Custom Data Warehouse Solution Part 2: We Need Storage

DevFeed: [A Journey Towards a Custom Data Warehouse Solution Part 2: We Need Storage](<https://devfeed.tech/articles/a-journey-towards-a-custom-data-warehouse-solution-part-2-we-need-storage-35108.md>)

Original publisher: [Read original article](<https://upday.github.io/blog/dwh-part2-we-need-storage/>)

Author: Robert Bordo (robert@upday.com)

Published: 2017-08-22T04:39:55Z

Content type: article

Language: en

Sources: [Upday](<https://devfeed.tech/sources/upday.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [business-intelligence](<https://devfeed.tech/tags/business-intelligence.md>), [data](<https://devfeed.tech/tags/data.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [s3](<https://devfeed.tech/tags/s3.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [storage](<https://devfeed.tech/tags/storage.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

This article examines storage choices for a custom data warehouse. It describes application log data, the limitations of time-series databases for additional master and historical data, business intelligence access needs, and the use of Amazon S3 as a data lake while considering other storage options including AWS Redshift.

### Source excerpt

In the beginning we created a cluster. And the cluster was without form, and void; and nulls were upon the face of the storage. As we learned in part 1 of our series, a data warehouse consists of several components. The key component is the storage. All the others group around it. But how can one draw a decision on which storage solution to adopt? What is out there anyway? Preface In a perfect world there would be only one kind of storage that fits all the needs of current DWH development and analysis. But since we are not living in that kind of place, we have several options. And the number of options increase the deeper one dives into the topic. There seem to be solutions for every use case you can think of. That might be a good starting point. What is our most common use case? What are we going to store? And how would we like to access our data in the end? Our major source is a massive amount of log data coming from our app. Everything the user does (e.g swiping through articles, selecting categories, leaving the app) is tracked, enriched with metadata (e.g. the user's location, app version, article identifier) and stored by a third-party service in big, semi-structured log files. Having only this source, a time series database like Graphite or InfluxDB could do the job. But also having slow changing master data, like user profiles, article metadata and maybe even to keep a history of data, this solution would not satisfy our current and future needs. Another thing that comes to my mind is how the data will be accessed by our final consumer (namely: Business Intelligence). Usually they use tools like Jasper Reports or Tableau for generating reports. For analyses we have to pre-aggregate the data to make queries more performant and translate raw information into a digestible format. What else is on the market? Storage good at bad at Example S3/Flat Files scalability, easy to use, data lake querying S3 Time Series DB handling time series data non time series data G