# AWS Glue

AWS Glue is a serverless data integration service for discovering, preparing, combining, and processing data, including ETL pipelines built with Apache Spark, Python, and Scala.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Tableflow: Turn Kafka Topics into Iceberg Tables

DevFeed: [Tableflow: Turn Kafka Topics into Iceberg Tables](<https://devfeed.tech/articles/tableflow-turn-kafka-topics-into-iceberg-tables-11555.md>)

Original publisher: [Read original article](<https://www.confluent.io/blog/tableflow-kafka-iceberg/>)

Author: Mohtasham Sayeed Mohiuddin

Published: 2026-07-10T15:36:14Z

Content type: tutorial

Language: en

Sources: [Confluent: Data in motion](<https://devfeed.tech/sources/confluent-data-in-motion.md>)

Topics: [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [schema-evolution](<https://devfeed.tech/topics/schema-evolution.md>), [Confluent Cloud](<https://devfeed.tech/topics/confluent-cloud.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>)

Tags: [amazon-redshift](<https://devfeed.tech/tags/amazon-redshift.md>), [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [confluent-cloud](<https://devfeed.tech/tags/confluent-cloud.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [schema-evolution](<https://devfeed.tech/tags/schema-evolution.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [technologies](<https://devfeed.tech/tags/technologies.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>)

### AI overview

This tutorial explains how Confluent Cloud Tableflow continuously materializes Apache Kafka topics as Apache Iceberg or Delta Lake tables. It covers automatic schema handling, type conversion, schema evolution, Parquet conversion, catalog publishing, and table maintenance for querying streaming data with analytics engines and warehouses.

### Source excerpt

Learn how Confluent Tableflow turns Kafka topics into Iceberg tables for zero-ETL analytics with automatic schema evolution and open catalog access.

## If it's in your catalog, you can query it: The DataLakeCatalog engine in ClickHouse Cloud

DevFeed: [If it's in your catalog, you can query it: The DataLakeCatalog engine in ClickHouse Cloud](<https://devfeed.tech/articles/if-it-s-in-your-catalog-you-can-query-it-the-datalakecatalog-engine-in-clickhouse-cloud-5538.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/query-your-catalog-clickhouse-cloud>)

Author: Tom Schreiber

Published: 2025-10-15T00:00:00Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [databricks](<https://devfeed.tech/topics/databricks.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [caching](<https://devfeed.tech/tags/caching.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [files](<https://devfeed.tech/tags/files.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [performance](<https://devfeed.tech/tags/performance.md>)

### AI overview

ClickHouse Cloud can query Iceberg and Delta Lake tables directly through the DataLakeCatalog engine. It integrates with catalogs such as AWS Glue Catalog and Databricks Unity Catalog, automatically discovers table formats, and supports federated queries across catalogs.

### Source excerpt

ClickHouse Cloud can now query Iceberg and Delta Lake tables directly through the DataLakeCatalog engine. Connect to Glue or Unity catalogs, discover tables automatically, and query your Lakehouse data instantly, all at ClickHouse speed.

## Announcing the ClickHouse Connector for AWS Glue

DevFeed: [Announcing the ClickHouse Connector for AWS Glue](<https://devfeed.tech/articles/announcing-the-clickhouse-connector-for-aws-glue-5087.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/clickhouse-connector-aws-glue>)

Author: Luke Gannon

Published: 2025-08-21T00:00:00Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [aws-marketplace](<https://devfeed.tech/topics/aws-marketplace.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>)

Tags: [apache](<https://devfeed.tech/tags/apache.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [aws-marketplace](<https://devfeed.tech/tags/aws-marketplace.md>), [blog](<https://devfeed.tech/tags/blog.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [launch](<https://devfeed.tech/tags/launch.md>), [python](<https://devfeed.tech/tags/python.md>), [scala](<https://devfeed.tech/tags/scala.md>), [spark](<https://devfeed.tech/tags/spark.md>)

### AI overview

This article announces the official ClickHouse Connector for AWS Glue, a serverless Apache Spark-based ETL integration available through AWS Marketplace. It explains how the connector supports PySpark and Scala, simplifies setup, and enables production-ready Spark jobs that connect AWS Glue with ClickHouse.

### Source excerpt

Today, we're announcing the launch of the official ClickHouse Connector for AWS Glue, which utilizes their Apache Spark-based serverless ETL engine.

## Redpanda 25.2: Advancing Iceberg integrations for the real-time lakehouse

DevFeed: [Redpanda 25.2: Advancing Iceberg integrations for the real-time lakehouse](<https://devfeed.tech/articles/redpanda-25-2-advancing-iceberg-integrations-for-the-real-time-lakehouse-12663.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/25-2-advancing-iceberg-integration>)

Author: Matt Schumpert

Published: 2025-08-05T00:00:00Z

Content type: article

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [databricks](<https://devfeed.tech/topics/databricks.md>), [Event-Streaming](<https://devfeed.tech/topics/event-streaming.md>), [interoperability](<https://devfeed.tech/topics/interoperability.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [migration](<https://devfeed.tech/topics/migration.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [aws-glue-catalog-integration](<https://devfeed.tech/tags/aws-glue-catalog-integration.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [databricks-unity-catalog-integration](<https://devfeed.tech/tags/databricks-unity-catalog-integration.md>), [event-streaming](<https://devfeed.tech/tags/event-streaming.md>), [iceberg-capabilities](<https://devfeed.tech/tags/iceberg-capabilities.md>), [integration](<https://devfeed.tech/tags/integration.md>), [interoperability](<https://devfeed.tech/tags/interoperability.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [kafka-json-schema-mapping](<https://devfeed.tech/tags/kafka-json-schema-mapping.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [migration](<https://devfeed.tech/tags/migration.md>), [open-source-ai-connectors](<https://devfeed.tech/tags/open-source-ai-connectors.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [real-time-data-lakes](<https://devfeed.tech/tags/real-time-data-lakes.md>), [real-time-lakehouse-architecture](<https://devfeed.tech/tags/real-time-lakehouse-architecture.md>), [redpanda-25-2-release](<https://devfeed.tech/tags/redpanda-25-2-release.md>), [redpanda-iceberg-integration](<https://devfeed.tech/tags/redpanda-iceberg-integration.md>), [schema-registry-authorization](<https://devfeed.tech/tags/schema-registry-authorization.md>), [streaming-data-analytics](<https://devfeed.tech/tags/streaming-data-analytics.md>), [streaming-platform-updates](<https://devfeed.tech/tags/streaming-platform-updates.md>), [vendor-lock-in](<https://devfeed.tech/tags/vendor-lock-in.md>)

### AI overview

Redpanda 25.2 expands Iceberg integration for a real-time open lakehouse architecture. It adds integrations with Databricks Unity Catalog and AWS Glue Catalog, supports AWS Glue as an Iceberg REST Catalog for Redpanda Iceberg Topics, and expands schema support. The release also includes migration improvements from Confluent and Schema Registry Authorization updates, while emphasizing Kafka client compatibility, low-latency ingestion, open interoperability, and reduced vendor lock-in.

### Source excerpt

Check what's new in Redpanda 25.2 as we expand our Iceberg capabilities, delivering on the promise of a real-time open lakehouse architecture.

## ClickHouse Release 25.3

DevFeed: [ClickHouse Release 25.3](<https://devfeed.tech/articles/clickhouse-release-25-3-5120.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/clickhouse-release-25-03>)

Author: ClickHouse

Published: 2025-03-27T00:00:00Z

Content type: release

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [JSON](<https://devfeed.tech/topics/json.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [data](<https://devfeed.tech/topics/data.md>), [elasticsearch](<https://devfeed.tech/topics/elasticsearch.md>), [MongoDB](<https://devfeed.tech/topics/mongodb.md>)

Tags: [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [apache-iceberg-tables](<https://devfeed.tech/tags/apache-iceberg-tables.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [cache](<https://devfeed.tech/tags/cache.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [compression](<https://devfeed.tech/tags/compression.md>), [database](<https://devfeed.tech/tags/database.md>), [elasticsearch](<https://devfeed.tech/tags/elasticsearch.md>), [json](<https://devfeed.tech/tags/json.md>), [mongodb](<https://devfeed.tech/tags/mongodb.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

ClickHouse 25.3 introduces 18 features, 13 performance optimizations, and 48 bug fixes. The release adds AWS Glue and Unity catalog support, a query condition cache, automatic S3 query parallelization, new array functions, and a production-ready JSON data type.

### Source excerpt

ClickHouse 25.3 is out! In this post, we highlight expanded Lakehouse catalog support with AWS Glue and Unity, the new query condition cache, automatic parallelization for external data sources, two new handy functions--and the GA of our new JSON type.

## Terraform module to manage Oxbow Lambda and its components

DevFeed: [Terraform module to manage Oxbow Lambda and its components](<https://devfeed.tech/articles/terraform-module-to-manage-oxbow-lambda-and-its-components-22561.md>)

Original publisher: [Read original article](<https://tech.scribd.com/blog/2025/terraform-oxbow-module.html>)

Author: Oleh Motrunych

Published: 2025-03-14T00:00:00Z

Content type: release

Language: en

Sources: [Scribd Tech](<https://devfeed.tech/sources/scribd-tech.md>)

Topics: [Terraform](<https://devfeed.tech/topics/terraform.md>), [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [infrastructure as code (IAC)](<https://devfeed.tech/topics/infrastructure-as-code-iac.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [Amazon Simple Queue Service (SQS)](<https://devfeed.tech/topics/amazon-simple-queue-service-sqs.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [event driven](<https://devfeed.tech/topics/event-driven.md>), [DynamoDB](<https://devfeed.tech/topics/dynamodb.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [aws-s3](<https://devfeed.tech/tags/aws-s3.md>), [deltalake](<https://devfeed.tech/tags/deltalake.md>), [dynamodb](<https://devfeed.tech/tags/dynamodb.md>), [event-driven](<https://devfeed.tech/tags/event-driven.md>), [iac](<https://devfeed.tech/tags/iac.md>), [oxbow](<https://devfeed.tech/tags/oxbow.md>), [rust](<https://devfeed.tech/tags/rust.md>), [security](<https://devfeed.tech/tags/security.md>), [sqs](<https://devfeed.tech/tags/sqs.md>), [terraform](<https://devfeed.tech/tags/terraform.md>)

### AI overview

The article introduces terraform-oxbow, an open-source Terraform module for deploying and managing an Oxbow AWS Lambda workflow and its supporting components. It describes configurable integrations including AWS Glue, Kinesis Data Firehose, SQS, DynamoDB, IAM policies, and S3 notifications, while noting AWS notification limits and least-privilege considerations.

### Source excerpt

Oxbow is a project to take an existing storage location which contains Apache Parquet files into a Delta Lake table. It is intended to run both as an AWS Lambda or as a command line application. We are excited to introduce terraform-oxbow, an open-source Terraform module that simplifies the deployment and management of AWS Lambda and its supporting components. Whether you're working with AWS Glue, Kinesis Data Firehose, SQS, or DynamoDB, this module provides a streamlined approach to infrastructure as code (IaC) in AWS.

## Trino on ice IV: Deep dive into Iceberg internals

DevFeed: [Trino on ice IV: Deep dive into Iceberg internals](<https://devfeed.tech/articles/trino-on-ice-iv-deep-dive-into-iceberg-internals-8667.md>)

Original publisher: [Read original article](<https://trino.io/blog/2021/08/12/deep-dive-into-iceberg-internals.html>)

Author: Brian Olsen

Published: 2021-08-12T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [Instrumentation](<https://devfeed.tech/topics/instrumentation.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [blog](<https://devfeed.tech/tags/blog.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [data](<https://devfeed.tech/tags/data.md>), [deep-dive](<https://devfeed.tech/tags/deep-dive.md>), [instrumentation](<https://devfeed.tech/tags/instrumentation.md>), [minio](<https://devfeed.tech/tags/minio.md>)

### AI overview

This deep-dive article explains Iceberg internals in the context of the Trino query engine. It examines metadata and files produced by Trino operations, using tools such as Avro tools, the MinIO client, and Iceberg's core library, with emphasis on understanding, troubleshooting, and querying Iceberg tables.

### Source excerpt

Welcome to the Trino on ice series, covering the details around how the Iceberg table format works with the Trino query engine. The examples build on each previous post, so it's recommended to read the posts sequentially and reference them as needed later. Here are links to the posts in this series: Trino on ice I: A gentle introduction to Iceberg Trino on ice II: In-place table evolution and cloud compatibility with Iceberg Trino on ice III: Iceberg concurrency model, snapshots, and the Iceberg spec Trino on ice IV: Deep dive into Iceberg internals So far, this series has covered some very interesting user level concepts of the Iceberg model, and how you can take advantage of them using the Trino query engine. This blog post dives into some implementation details of Iceberg by dissecting some files that result from various operations carried out using Trino. To dissect you must use some surgical instrumentation, namely Trino, Avro tools, the MinIO client tool and Iceberg's core library. It's useful to dissect how these files work, not only to help understand how Iceberg works, but also to aid in troubleshooting issues, should you have any issues during ingestion or querying of your Iceberg table. I like to think of this type of debugging much like a fun game of operation, and you're looking to see what causes the red errors to fly by on your screen.

## A Report about Presto Conference Tokyo 2020 Online

DevFeed: [A Report about Presto Conference Tokyo 2020 Online](<https://devfeed.tech/articles/a-report-about-presto-conference-tokyo-2020-online-8656.md>)

Original publisher: [Read original article](<https://trino.io/blog/2020/11/21/a-report-about-presto-conference-tokyo-2020.html>)

Author: Toru Takahashi, Treasure Data

Published: 2020-11-21T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>)

Tags: [article](<https://devfeed.tech/tags/article.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [conference](<https://devfeed.tech/tags/conference.md>), [github](<https://devfeed.tech/tags/github.md>), [japan](<https://devfeed.tech/tags/japan.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [performance](<https://devfeed.tech/tags/performance.md>), [report](<https://devfeed.tech/tags/report.md>), [saas](<https://devfeed.tech/tags/saas.md>), [spark](<https://devfeed.tech/tags/spark.md>), [technical](<https://devfeed.tech/tags/technical.md>), [youtube](<https://devfeed.tech/tags/youtube.md>)

### AI overview

This English report summarizes the 2nd Presto Conference held by the Japan Presto Community in 2020. It covers recent Presto community updates, contribution guidance, open-source principles, Treasure Data's SaaS support for Presto, and ways to use and compare Presto on AWS services including EC2, EMR, Athena, and AWS Glue. The supplied text ends mid-sentence.

### Source excerpt

On Nov 11th, 2020, Japan Presto Community held the 2nd Presto Conference welcoming Martin Traverso and Brian Olsen. The conference was hosted at Youtube Live. This article is the summary of the conference aiming to share their great talks.

## Continuous Deployment for AWS Glue

DevFeed: [Continuous Deployment for AWS Glue](<https://devfeed.tech/articles/continuous-deployment-for-aws-glue-22993.md>)

Original publisher: [Read original article](<https://bravenewgeek.com/continuous-deployment-for-aws-glue/>)

Author: Mohammed

Published: 2020-10-15T15:51:25Z

Content type: tutorial

Language: en

Sources: [Brave New Geek](<https://devfeed.tech/sources/brave-new-geek.md>)

Topics: [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [Continuous Deployment (CD)](<https://devfeed.tech/topics/continuous-deployment.md>), [GitHub Actions](<https://devfeed.tech/topics/github-actions.md>), [Jupyter Notebook](<https://devfeed.tech/topics/jupyter-notebook.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [analytics-pipeline](<https://devfeed.tech/tags/analytics-pipeline.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [continuous-delivery](<https://devfeed.tech/tags/continuous-delivery.md>), [continuous-deployment](<https://devfeed.tech/tags/continuous-deployment.md>), [etl](<https://devfeed.tech/tags/etl.md>), [github](<https://devfeed.tech/tags/github.md>), [github-actions](<https://devfeed.tech/tags/github-actions.md>), [jupyter](<https://devfeed.tech/tags/jupyter.md>), [jupyter-notebook](<https://devfeed.tech/tags/jupyter-notebook.md>), [s3](<https://devfeed.tech/tags/s3.md>), [serverless](<https://devfeed.tech/tags/serverless.md>)

### AI overview

A tutorial for automating continuous deployment of AWS Glue ETL jobs. It uses GitHub Actions to generate a Python script from a Jupyter notebook, copy it to Amazon S3, and update the Glue job to use the new script.

### Source excerpt

AWS Glue is a managed service for building ETL (Extract-Transform-Load) jobs. It's a useful tool for implementing analytics pipelines in AWS without having to manage server infrastructure. Jobs are implemented using Apache Spark and, with the help of Development Endpoints, can be built using Jupyter notebooks. This makes it reasonably easy to write ETL processes in an interactive, iterative fashion. Once finished, the Jupyter notebook is converted into a Python script, uploaded to S3, and then run as a Glue job.

## Accommodation Consolidation: How we created an ETL pipeline on cloud

DevFeed: [Accommodation Consolidation: How we created an ETL pipeline on cloud](<https://devfeed.tech/articles/accommodation-consolidation-how-we-created-an-etl-pipeline-on-cloud-27990.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2020-03-26-accommodationconsolidationhowwecreatedan/>)

Author: Praneeth Peiris I want

Published: 2020-03-26T00:00:00Z

Content type: article

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>), [AWS Step Functions](<https://devfeed.tech/topics/aws-step-functions.md>), [etl](<https://devfeed.tech/topics/etl.md>), [data](<https://devfeed.tech/topics/data.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [aws-step-functions](<https://devfeed.tech/tags/aws-step-functions.md>), [backend](<https://devfeed.tech/tags/backend.md>), [batch](<https://devfeed.tech/tags/batch.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [engineering-culture](<https://devfeed.tech/tags/engineering-culture.md>), [etl](<https://devfeed.tech/tags/etl.md>), [overhead](<https://devfeed.tech/tags/overhead.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>)

### AI overview

trivago describes a hybrid AWS architecture using AWS Glue and AWS Step Functions to build ETL pipelines for consolidating frequently changing hotel information from hundreds of partners. The approach batches updates to reduce computational overhead and supports separately tested consolidation models and sandbox environments.

### Source excerpt

Imagine you go to your hotel for check-in and they say that your dog is not allowed even though the website clearly states that it is!trivago gets information about millions of accommodat...