# Cloud-native data ingestion architecture using AWS, Databricks, and open-source tools

DevFeed: [Cloud-native data ingestion architecture using AWS, Databricks, and open-source tools](<https://devfeed.tech/articles/let-s-save-tons-of-money-with-cloud-native-data-ingestion-22560.md>)

Original publisher: [Read original article](<https://tech.scribd.com/blog/2025/cloud-native-data-ingestion.html>)

Author: R Tyler Croy

Published: 2025-08-01T00:00:00Z

Content type: tutorial

Language: en

Sources: [Scribd Tech](<https://devfeed.tech/sources/scribd-tech.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [databricks](<https://devfeed.tech/topics/databricks.md>), [event driven](<https://devfeed.tech/topics/event-driven.md>), [Amazon Simple Queue Service (SQS)](<https://devfeed.tech/topics/amazon-simple-queue-service-sqs.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [catalog](<https://devfeed.tech/tags/catalog.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [data](<https://devfeed.tech/tags/data.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [deltalake](<https://devfeed.tech/tags/deltalake.md>), [event-driven](<https://devfeed.tech/tags/event-driven.md>), [featured](<https://devfeed.tech/tags/featured.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [rust](<https://devfeed.tech/tags/rust.md>), [sqs](<https://devfeed.tech/tags/sqs.md>)

## AI overview

This article presents Scribd's cloud-native data-ingestion architecture for building large datasets for Delta Lake. It describes using AWS services and open-source tools such as kafka-delta-ingest, oxbow, and Airbyte in a more event-driven and reliable platform, with Databricks and Unity Catalog. The approach can also be adapted to Azure, Google Cloud Platform, or on-premises environments.

## Source excerpt

Delta Lake is a fantastic technology for quickly querying massive data sets, but first you need those massive data sets! In this talk from Data and AI Summit 2025 I dive into the cloud-native architecture Scribd has adopted to ingest data from AWS Aurora, SQS, Kinesis Data Firehose and more!