# Data pipelines

Published articles for Data pipelines.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## The official ClickHouse provider for Apache Airflow is now available

DevFeed: [The official ClickHouse provider for Apache Airflow is now available](<https://devfeed.tech/articles/the-official-clickhouse-provider-for-apache-airflow-is-now-available-42157.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/clickhouse-airflow-provider>)

Author: Aditya Chidurala; Bentsi Leviav; Alex Francoeur

Published: 2026-09-17T18:06:01Z

Content type: release

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [airflow](<https://devfeed.tech/topics/airflow.md>), [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Python](<https://devfeed.tech/topics/python.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [integration](<https://devfeed.tech/tags/integration.md>), [python](<https://devfeed.tech/tags/python.md>), [release](<https://devfeed.tech/tags/release.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

ClickHouse has released an officially maintained Apache Airflow provider for orchestrating ClickHouse data workflows. The provider uses ClickHouse Connect over HTTP(S), supports Airflow's common SQL operators, and includes a hook for bulk and client-specific operations.

### Source excerpt

The official ClickHouse provider for Apache Airflow simplifies data workflows with standard SQL operators, bulk inserts, and shared setup across self-managed Airflow and Astronomer.

## AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity

DevFeed: [AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity](<https://devfeed.tech/articles/ai-changed-how-spotify-builds-what-we-learned-and-fixed-about-quality-at-higher-velocity-41282.md>)

Original publisher: [Read original article](<https://engineering.atspotify.com/2026/9/ai-changed-how-spotify-builds-what-we-learned-and-fixed-about-quality-at-higher-velocity/>)

Author: Spotify Engineering

Published: 2026-09-16T19:13:53Z

Content type: article

Language: en

Sources: [Spotify Engineering](<https://devfeed.tech/sources/spotify-engineering.md>), [Spotify Engineering Blog](<https://devfeed.tech/sources/spotify-engineering-blog.md>)

Topics: [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [reliability](<https://devfeed.tech/topics/reliability.md>), [Microservices](<https://devfeed.tech/topics/microservices.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Job](<https://devfeed.tech/topics/job.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [bug](<https://devfeed.tech/tags/bug.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [reliability](<https://devfeed.tech/tags/reliability.md>)

### AI overview

Spotify describes how rapid change, content-processing weaknesses, capacity limits, and a scheduling bug contributed to delays in publishing episodes. It reports adding end-to-end monitoring, fixing the scheduler, lowering batch-job priority, and increasing capacity.

### Source excerpt

Quality and reliability have always been a point of pride for Spotify. We run an extraordinarily complex... The post AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity appeared first on Spotify Engineering.

## Dropbox Evolves Riviera Content Processing Platform to Support AI Workloads

DevFeed: [Dropbox Evolves Riviera Content Processing Platform to Support AI Workloads](<https://devfeed.tech/articles/dropbox-evolves-riviera-content-processing-platform-to-support-ai-workloads-31517.md>)

Original publisher: [Read original article](<https://www.infoq.com/news/2026/09/dropbox-riviera-ai-platform/>)

Author: Leela Kumili

Published: 2026-09-16T14:42:00Z

Content type: news

Language: en

Sources: [InfoQ](<https://devfeed.tech/sources/infoq.md>)

Topics: [dropbox](<https://devfeed.tech/topics/dropbox.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [API](<https://devfeed.tech/topics/api.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-ml-data-engineering](<https://devfeed.tech/tags/ai-ml-data-engineering.md>), [apache](<https://devfeed.tech/tags/apache.md>), [apis](<https://devfeed.tech/tags/apis.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [architecture-design](<https://devfeed.tech/tags/architecture-design.md>), [asynchronous-architecture](<https://devfeed.tech/tags/asynchronous-architecture.md>), [backend](<https://devfeed.tech/tags/backend.md>), [caching](<https://devfeed.tech/tags/caching.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [development](<https://devfeed.tech/tags/development.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [dropbox](<https://devfeed.tech/tags/dropbox.md>), [dropbox-riviera-ai-platform](<https://devfeed.tech/tags/dropbox-riviera-ai-platform.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [enterprise-content-management](<https://devfeed.tech/tags/enterprise-content-management.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [ml-data-engineering](<https://devfeed.tech/tags/ml-data-engineering.md>), [model-context-protocol-mcp](<https://devfeed.tech/tags/model-context-protocol-mcp.md>), [news](<https://devfeed.tech/tags/news.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [platform-engineering](<https://devfeed.tech/tags/platform-engineering.md>), [plugins](<https://devfeed.tech/tags/plugins.md>), [rag](<https://devfeed.tech/tags/rag.md>), [tika](<https://devfeed.tech/tags/tika.md>)

### AI overview

Dropbox has expanded Riviera from an internal file-preview service into a content-processing platform supporting more than 300 file formats and over 100 transformation capabilities. The platform supports Dropbox products including Search, Replay, Sign, and Dash, and provides APIs for asynchronous document conversion, media transcription, and structured metadata extraction for AI and RAG workflows.

### Source excerpt

Dropbox has evolved Riviera from a file preview service into a universal content processing platform supporting more than 300 file formats and over 100 transformation capabilities. Processing hundreds of thousands of transformations per second, Riviera now supports Search, Replay, Sign, and Dash, while its APIs enable asynchronous content extraction for AI and RAG workflows. By Leela Kumili

## New data pipeline management platform at Khan Academy

DevFeed: [New data pipeline management platform at Khan Academy](<https://devfeed.tech/articles/new-data-pipeline-management-platform-at-khan-academy-27388.md>)

Original publisher: [Read original article](<http://engineering.khanacademy.org/posts/khanalytics.htm>)

Author: Khan Academy

Published: 2018-04-30T22:00:00Z

Content type: article

Language: en

Sources: [Khan Academy](<https://devfeed.tech/sources/khan-academy.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [batch](<https://devfeed.tech/tags/batch.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [cloud-dataflow](<https://devfeed.tech/tags/cloud-dataflow.md>), [data](<https://devfeed.tech/tags/data.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [news](<https://devfeed.tech/tags/news.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [scheduling](<https://devfeed.tech/tags/scheduling.md>)

### AI overview

Khan Academy developed Khanalytics to manage its growing collection of data pipelines. The platform provides a sandboxed environment for batch jobs, a web interface, automatic parallelization, centralized logs, and pipeline scheduling with dependencies.

### Source excerpt

By Ragini Gupta Data is very crucial to Khan Academy and is itself an internal product for the ... Read more

## How dbt works, and why orchestrators shouldn't split it into tasks

DevFeed: [How dbt works, and why orchestrators shouldn't split it into tasks](<https://devfeed.tech/articles/how-dbt-works-and-why-orchestrators-shouldn-t-split-it-into-tasks-30714.md>)

Original publisher: [Read original article](<https://www.windmill.dev/blog/how-dbt-works-and-its-orchestrators>)

Author: Ruben Fiszel

Published: 2026-08-20T00:00:00Z

Content type: article

Language: en

Sources: [Windmill Blog](<https://devfeed.tech/sources/windmill-blog.md>)

Topics: [Compiler](<https://devfeed.tech/topics/compiler.md>), [Job](<https://devfeed.tech/topics/job.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [data](<https://devfeed.tech/topics/data.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [cli](<https://devfeed.tech/tags/cli.md>), [dagster](<https://devfeed.tech/tags/dagster.md>), [data](<https://devfeed.tech/tags/data.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [dbt](<https://devfeed.tech/tags/dbt.md>), [dbt-data-pipelines-airflow-dagster-orchestration](<https://devfeed.tech/tags/dbt-data-pipelines-airflow-dagster-orchestration.md>), [job](<https://devfeed.tech/tags/job.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [scheduler](<https://devfeed.tech/tags/scheduler.md>), [series](<https://devfeed.tech/tags/series.md>)

### AI overview

This primer explains dbt's architecture and examines five ways teams orchestrate it in production. It argues that running a project as one dbt command, while reading dbt's per-model execution state, is generally more efficient and less fragile than creating one orchestrator task per model. It discusses Dagster and astronomer-cosmos's convergence on this approach, including a reported cost comparison.

### Source excerpt

How is dbt actually orchestrated, and why does running it as one job beat one task per model? A primer on dbt as a compiler with a scheduler attached, the five ways teams wrap it, and why both Dagster and astronomer-cosmos converged on a single dbt invocation with per-model state projected out of it.

## The 5 Silent Failures in Data Pipelines

DevFeed: [The 5 Silent Failures in Data Pipelines](<https://devfeed.tech/articles/the-5-silent-failures-in-data-pipelines-37149.md>)

Original publisher: [Read original article](<https://seattledataguy.substack.com/p/the-5-silent-failures-in-data-pipelines>)

Author: SeattleDataGuy

Published: 2026-04-24T19:06:02Z

Content type: article

Language: en

Sources: [SeattleDataGuy's Newsletter](<https://devfeed.tech/sources/seattledataguy-s-newsletter.md>)

Topics: [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [data](<https://devfeed.tech/topics/data.md>), [CSV](<https://devfeed.tech/topics/csv.md>)

Tags: [csv](<https://devfeed.tech/tags/csv.md>), [dashboard](<https://devfeed.tech/tags/dashboard.md>), [data-pipeline](<https://devfeed.tech/tags/data-pipeline.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [reports](<https://devfeed.tech/tags/reports.md>), [schema](<https://devfeed.tech/tags/schema.md>)

### AI overview

The article explains how data pipelines can fail silently without triggering errors or obvious warnings, causing stale or incorrect data to reach dashboards and reports. It introduces schema drift as one failure mode, including unexpected changes to CSV or XML files loaded from SFTP.

### Source excerpt

How Your Pipelines Lie to You Without Throwing a Single Error

## Data Pipeline Foundations - Everything You Need To Know About Data Pipelines

DevFeed: [Data Pipeline Foundations - Everything You Need To Know About Data Pipelines](<https://devfeed.tech/articles/data-pipeline-foundations-everything-you-need-to-know-about-data-pipelines-37141.md>)

Original publisher: [Read original article](<https://seattledataguy.substack.com/p/data-pipeline-foundations-everything>)

Author: SeattleDataGuy

Published: 2026-04-18T15:50:56Z

Content type: article

Language: en

Sources: [SeattleDataGuy's Newsletter](<https://devfeed.tech/sources/seattledataguy-s-newsletter.md>)

Topics: [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [data](<https://devfeed.tech/topics/data.md>), [API](<https://devfeed.tech/topics/api.md>), [Database](<https://devfeed.tech/topics/database.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [data](<https://devfeed.tech/tags/data.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [database](<https://devfeed.tech/tags/database.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [s3](<https://devfeed.tech/tags/s3.md>), [sftp](<https://devfeed.tech/tags/sftp.md>)

### AI overview

This article serves as a central collection of introductory resources about data pipelines. It explains that pipelines move data from sources to destinations and highlights sources such as S3 buckets, SFTP files, APIs, and databases.

### Source excerpt

Hi, fellow future and current Data Leaders; Ben here 👋

## Daily Tasks With Data Pipelines - Data Quality Checks And The Problem With Noisy Checks

DevFeed: [Daily Tasks With Data Pipelines - Data Quality Checks And The Problem With Noisy Checks](<https://devfeed.tech/articles/daily-tasks-with-data-pipelines-data-quality-checks-and-the-problem-with-noisy-checks-37140.md>)

Original publisher: [Read original article](<https://seattledataguy.substack.com/p/daily-tasks-with-data-pipelines-data>)

Author: SeattleDataGuy

Published: 2026-04-07T22:24:38Z

Content type: article

Language: en

Sources: [SeattleDataGuy's Newsletter](<https://devfeed.tech/sources/seattledataguy-s-newsletter.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [quality](<https://devfeed.tech/tags/quality.md>)

### AI overview

The article discusses data quality checks in data pipelines and notes that teams may receive 137 data quality alerts every morning.

### Source excerpt

Every morning, your team wakes up to 137 data quality alerts.

## Introducing the Apache Airflow Registry

DevFeed: [Introducing the Apache Airflow Registry](<https://devfeed.tech/articles/introducing-the-apache-airflow-registry-32543.md>)

Original publisher: [Read original article](<https://airflow.apache.org/blog/airflow-registry/>)

Author: Apache Airflow

Published: 2026-03-19T00:00:00Z

Content type: release

Language: en

Sources: [Apache Airflow Blog](<https://devfeed.tech/sources/apache-airflow-blog.md>)

Topics: [airflow](<https://devfeed.tech/topics/airflow.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [API](<https://devfeed.tech/topics/api.md>), [JSON](<https://devfeed.tech/topics/json.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [apache](<https://devfeed.tech/tags/apache.md>), [apache-airflow](<https://devfeed.tech/tags/apache-airflow.md>), [community](<https://devfeed.tech/tags/community.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [messaging](<https://devfeed.tech/tags/messaging.md>), [notifications](<https://devfeed.tech/tags/notifications.md>), [openai](<https://devfeed.tech/tags/openai.md>), [registry](<https://devfeed.tech/tags/registry.md>)

### AI overview

Apache Airflow launches the Airflow Registry, a searchable catalog of official providers and modules. It includes provider and module search, installation and compatibility details, connection generation in URI, JSON, and environment-variable formats, ecosystem statistics, and a structured JSON API.

### Source excerpt

Today we're launching the Apache Airflow Registry -- a searchable catalog of every official Airflow provider and its modules, live at airflow.apache.org/registry/. Need an S3 operator? A Snowflake hook? An OpenAI sensor? The Registry helps you find, compare, and configure the right components for your data pipelines -- without digging through docs or PyPI pages. By the Numbers 98 Official providers 1,602 Modules (operators, hooks, sensors, triggers, transfers, and more) 329M+ Monthly PyPI downloads across all providers 125+ Integrations with cloud platforms, databases, ML tools, and messaging services Search Everything Hit Cmd+K from any page and start typing. Results show up instantly, grouped by Providers and Modules, with type badges so you can tell a hook from an operator at a glance. Provider Pages Each provider gets a dedicated page with everything in one place: install command with copy-to-clipboard, version selector, extras dropdown, compatibility info, connection types, and the full module listing organized by type. The Amazon provider, for example, has 372 modules across operators, hooks, sensors, triggers, transfers, and more. Module type tabs let you filter to exactly what you're looking for, and a category sidebar groups modules by AWS service (S3, Lambda, Glue, Step Functions, etc.). Connection Builder Click any connection type badge on a provider page, fill in the fields, and the builder generates the connection in three formats -- URI, JSON, and Env Var -- ready to copy into your configuration. No more guessing URI encoding or JSON structure. Explore by Category Not sure which provider you need? The Explore page organizes providers into categories: Cloud Platforms, Databases, Data Warehouses, Messaging & Notifications, AI & Machine Learning, Data Processing, and more. Statistics The Stats page breaks down the ecosystem: 848 operators, 298 hooks, 164 triggers, 157 sensors, 83 transfers, and more -- plus top providers by downloads and module count. JSON API

## System Migration: Minimize Downtime, Maximize Efficiency

DevFeed: [System Migration: Minimize Downtime, Maximize Efficiency](<https://devfeed.tech/articles/system-migration-minimize-downtime-maximize-efficiency-39555.md>)

Original publisher: [Read original article](<https://ankit-rana.com/logs/03-system-migration/>)

Author: hello@ankit-rana.com

Published: 2026-03-16T00:00:00Z

Content type: tutorial

Language: en

Sources: [Ankit Rana | Mechanical Sympathy](<https://devfeed.tech/sources/ankit-rana-mechanical-sympathy.md>)

Topics: [migration](<https://devfeed.tech/topics/migration.md>), [systems](<https://devfeed.tech/topics/systems.md>), [async](<https://devfeed.tech/topics/async.md>), [client](<https://devfeed.tech/topics/client.md>), [event driven](<https://devfeed.tech/topics/event-driven.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [apache-kafka](<https://devfeed.tech/tags/apache-kafka.md>), [api](<https://devfeed.tech/tags/api.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [async](<https://devfeed.tech/tags/async.md>), [bridge-layer](<https://devfeed.tech/tags/bridge-layer.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [downtime](<https://devfeed.tech/tags/downtime.md>), [event-driven-architecture](<https://devfeed.tech/tags/event-driven-architecture.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [migration](<https://devfeed.tech/tags/migration.md>), [observability](<https://devfeed.tech/tags/observability.md>), [production](<https://devfeed.tech/tags/production.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [rollback](<https://devfeed.tech/tags/rollback.md>), [system-migration](<https://devfeed.tech/tags/system-migration.md>), [traffic-leakage](<https://devfeed.tech/tags/traffic-leakage.md>)

### AI overview

A practical guide to migrating an existing system with minimal disruption. It recommends isolated-environment testing, load testing, adapters for incompatible contracts, synchronized asynchronous pipelines, a Kafka-based shared stream, a bridge layer, staged traffic switching, monitoring, and rollback preparation.

### Source excerpt

Migrate behind a bridge layer that routes all client traffic and supports three modes: old-only, dual, and new-only. Run dual mode to compare responses without user impact, keep a back-sync pipeline so the old system stays current for rollback, and shift traffic in stages while watching metrics at each step.

## Backfills - The Necessary Evil of Data Engineering

DevFeed: [Backfills - The Necessary Evil of Data Engineering](<https://devfeed.tech/articles/backfills-the-necessary-evil-of-data-engineering-37138.md>)

Original publisher: [Read original article](<https://seattledataguy.substack.com/p/backfills-the-necessary-evil-of-data>)

Author: SeattleDataGuy

Published: 2026-02-23T23:37:52Z

Content type: tutorial

Language: en

Sources: [SeattleDataGuy's Newsletter](<https://devfeed.tech/sources/seattledataguy-s-newsletter.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [bug](<https://devfeed.tech/tags/bug.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-type](<https://devfeed.tech/tags/data-type.md>), [databases](<https://devfeed.tech/tags/databases.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [schema](<https://devfeed.tech/tags/schema.md>)

### AI overview

This article explains why data teams perform backfills and why data engineers often dislike them. It describes backfills as rerunning or rebuilding tables and pipelines to account for corrected source data, pipeline bugs, schema changes, logic changes, or required data-type conversions.

### Source excerpt

Why backfills happen, why we hate them, and how to handle them without breaking trust

## How we Use Dagster Automations in our Data Pipeline

DevFeed: [How we Use Dagster Automations in our Data Pipeline](<https://devfeed.tech/articles/how-we-use-dagster-automations-in-our-data-pipeline-29999.md>)

Original publisher: [Read original article](<https://engineering.freeagent.com/2025/12/10/how-we-use-dagster-automations-in-our-data-pipeline/>)

Author: Delphine Rabiller

Published: 2025-12-10T11:32:42Z

Content type: article

Language: en

Sources: [FreeAgent](<https://devfeed.tech/sources/freeagent.md>)

Topics: [Automation](<https://devfeed.tech/topics/automation.md>), [data](<https://devfeed.tech/topics/data.md>), [migration](<https://devfeed.tech/topics/migration.md>)

Tags: [analytics-engineering](<https://devfeed.tech/tags/analytics-engineering.md>), [automation](<https://devfeed.tech/tags/automation.md>), [dagster](<https://devfeed.tech/tags/dagster.md>), [data-ml](<https://devfeed.tech/tags/data-ml.md>), [data-pipeline](<https://devfeed.tech/tags/data-pipeline.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [declarative](<https://devfeed.tech/tags/declarative.md>), [migration](<https://devfeed.tech/tags/migration.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [scheduled](<https://devfeed.tech/tags/scheduled.md>), [sensors](<https://devfeed.tech/tags/sensors.md>), [tooling](<https://devfeed.tech/tags/tooling.md>), [triggers](<https://devfeed.tech/tags/triggers.md>)

### AI overview

This engineering post explains how FreeAgent is migrating data pipelines to Dagster and re-architecting its automation logic. It describes three approaches to automating asset materialization: schedules, declarative automation, and asset sensors, including when schedules are appropriate and how declarative automation uses asset dependencies and materialization status.

### Source excerpt

Introduction The heart of a reliable data platform are robust and automated data pipelines. As our team migrates our data pipelines to Dagster, re-architecting our automation logic is a crucial task. Dagster offers condition-based approaches to creating or updating a data asset (table or file), moving us toward a modern, asset-centric view of data. This [...]

## How Can I Be An AI Engineer?

DevFeed: [How Can I Be An AI Engineer?](<https://devfeed.tech/articles/how-can-i-be-an-ai-engineer-33446.md>)

Original publisher: [Read original article](<https://timkellogg.me/blog/2024/12/09/ai-engineer>)

Published: 2024-12-09T00:00:00Z

Content type: tutorial

Language: en

Sources: [Tim Kellogg](<https://devfeed.tech/sources/tim-kellogg.md>)

Topics: [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Retrieval-Augmented Generation](<https://devfeed.tech/topics/retrieval-augmented-generation.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [Front end](<https://devfeed.tech/topics/frontend.md>), [React](<https://devfeed.tech/topics/react.md>)

Tags: [ai-engineer](<https://devfeed.tech/tags/ai-engineer.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [frontend](<https://devfeed.tech/tags/frontend.md>), [genai](<https://devfeed.tech/tags/genai.md>), [llms](<https://devfeed.tech/tags/llms.md>)

### AI overview

The article explains what AI engineers do, emphasizing that the role involves integrating generative AI models into applications and may include building user interfaces, APIs, and data pipelines. It describes several archetypes, including data-pipeline and UX-focused AI engineers, and discusses skills and experience that may be useful.

### Source excerpt

You want to be an AI Engineer? Do you even have the right skills? What do they do? All great questions. I've had this same conversation several times, so I figured it would be best to write it down. Here I answer all those, and break down the job into archetypes that should help you understand how you'll contribute.

## Office Hours with Engineering Managing Director Mae Santos

DevFeed: [Office Hours with Engineering Managing Director Mae Santos](<https://devfeed.tech/articles/office-hours-with-engineering-managing-director-mae-santos-39480.md>)

Original publisher: [Read original article](<https://www.twosigma.com/articles/office-hours-with-engineering-managing-director-mae-santos/>)

Author: Emily Majewski

Published: 2024-06-26T17:34:49Z

Content type: article

Language: en

Sources: [Two Sigma Engineering](<https://devfeed.tech/sources/two-sigma-engineering.md>)

Topics: [reliability](<https://devfeed.tech/topics/reliability.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Automation](<https://devfeed.tech/topics/automation.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [integrity](<https://devfeed.tech/tags/integrity.md>), [reliability-engineering](<https://devfeed.tech/tags/reliability-engineering.md>)

### AI overview

In this Office Hours interview, Two Sigma engineering leader Mae Santos discusses her career, leadership values, and responsibilities overseeing critical applications and data reliability engineering. She emphasizes integrity, collaboration, resilience, and designing reliability into applications and data systems from the beginning.

### Source excerpt

The post Office Hours with Engineering Managing Director Mae Santos appeared first on Two Sigma.

## Image Migration to Google Cloud Platform

DevFeed: [Image Migration to Google Cloud Platform](<https://devfeed.tech/articles/image-migration-to-google-cloud-platform-28045.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2024-05-14-image-migration-to-gcp/>)

Author: Praneeth Peiris I want

Published: 2024-05-14T00:00:00Z

Content type: article

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>), [migration](<https://devfeed.tech/topics/migration.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [data](<https://devfeed.tech/topics/data.md>), [Image processing](<https://devfeed.tech/topics/image-processing.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [amazon-web-services](<https://devfeed.tech/tags/amazon-web-services.md>), [backend](<https://devfeed.tech/tags/backend.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [google-cloud-platform](<https://devfeed.tech/tags/google-cloud-platform.md>), [image-processing](<https://devfeed.tech/tags/image-processing.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [migration](<https://devfeed.tech/tags/migration.md>)

### AI overview

This article explains how trivago migrated its image infrastructure from AWS and an on-premises data centre to Google Cloud. It describes the image-processing and machine-learning pipelines involved, the limitations of a lift-and-shift migration, and the need to redesign the pipelines for a single cloud environment.

### Source excerpt

Migration projects can be hard, especially when we were not around when the original projects were built. We migrated our images infrastructure to Google Cloud which was spread across multiple environments and here is how we did that.

## Insights from Workflow History Export on Temporal Cloud

DevFeed: [Insights from Workflow History Export on Temporal Cloud](<https://devfeed.tech/articles/insights-from-workflow-history-export-on-temporal-cloud-35839.md>)

Original publisher: [Read original article](<https://temporal.io/blog/get-insights-from-workflow-histories-export-on-temporal-cloud>)

Author: Alice Yin

Published: 2024-04-24T06:00:00Z

Content type: tutorial

Language: en

Sources: [Temporal Blog](<https://devfeed.tech/sources/temporal-blog.md>)

Topics: [export](<https://devfeed.tech/topics/export.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [Data analysis](<https://devfeed.tech/topics/data-analysis.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [AWS Glue](<https://devfeed.tech/topics/aws-glue.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [api](<https://devfeed.tech/tags/api.md>), [aws-glue](<https://devfeed.tech/tags/aws-glue.md>), [case-study](<https://devfeed.tech/tags/case-study.md>), [data](<https://devfeed.tech/tags/data.md>), [data-analysis](<https://devfeed.tech/tags/data-analysis.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [export](<https://devfeed.tech/tags/export.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [product-news](<https://devfeed.tech/tags/product-news.md>)

### AI overview

This tutorial explains how to export closed workflow histories from Temporal Cloud to Amazon S3, convert Protocol Buffer files to Parquet, and build data pipelines for analyzing execution metadata, operational efficiency, and bottlenecks.

### Source excerpt

Explore how exporting workflow histories in Temporal Cloud gives you deep insights into execution metadata, events, and performance over time.

## Leveraging Spark 3 and NVIDIA's GPUs to Reduce Cloud Cost by up to 70% for Big Data Pipelines

DevFeed: [Leveraging Spark 3 and NVIDIA's GPUs to Reduce Cloud Cost by up to 70% for Big Data Pipelines](<https://devfeed.tech/articles/leveraging-spark-3-and-nvidia-s-gpus-to-reduce-cloud-cost-by-up-to-70-for-big-data-pipelines-31935.md>)

Original publisher: [Read original article](<https://medium.com/paypal-tech/leveraging-spark-3-and-nvidias-gpus-to-reduce-cloud-cost-by-up-to-70-for-big-data-pipelines-e0bc02ec4f88?source=rss----6423323524ba---4>)

Author: Ilay Chen

Published: 2024-02-21T16:42:14Z

Content type: tutorial

Language: en

Sources: [PayPal Technology](<https://devfeed.tech/sources/paypal-technology.md>)

Topics: [Apache Spark](<https://devfeed.tech/topics/spark.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [RAPIDS](<https://devfeed.tech/topics/rapids.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [migration](<https://devfeed.tech/topics/migration.md>), [upgrade](<https://devfeed.tech/topics/upgrade.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [apache-spark](<https://devfeed.tech/tags/apache-spark.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-computing](<https://devfeed.tech/tags/cloud-computing.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [migration](<https://devfeed.tech/tags/migration.md>), [rapids](<https://devfeed.tech/tags/rapids.md>), [upgrade](<https://devfeed.tech/tags/upgrade.md>)

### AI overview

A PayPal engineering blog explains how upgrading from Apache Spark 2 to Spark 3 and migrating workloads to GPU clusters with NVIDIA Spark RAPIDS can accelerate selected big-data processing tasks and potentially reduce cloud costs by up to 70%. It covers the migration, parameter tuning, challenges, and reported benefits.

### Source excerpt

By Ilay Chen and Tomer Akirav At PayPal, hundreds of thousands of Apache Spark jobs run on an hourly basis, processing petabytes of data and requiring a high volume of resources. To handle the growth of machine learning solutions, PayPal requires scalable environments, cost awareness and constant innovation. This blog explains how Apache Spark 3 and GPUs can help enterprises potentially reduce Apache Spark's jobs cloud costs by up to 70% for big data processing and AI applications. Our journey will begin with a brief introduction of Spark RAPIDS -- Apache Spark's accelerator that leverages GPUs to accelerate processing via the RAPIDS libraries. We will then review PayPal's CPU-based Spark 2 application, our upgrade to Spark 3 and its new capabilities, explore the migration of our Apache Spark application to a GPU cluster, and how we tuned Spark RAPIDS parameters. We will then discuss some challenges we encountered and the benefits of the updates. Libra scales in the cloud, generated by AIBackground GPUs are everywhere, and their parallelism characteristics are perfect for processing AI and graphics applications, among other things. For those unfamiliar: what makes GPUs different from CPUs, computation-wise, is that CPUs have a limited amount of very strong cores, whereas GPUs have thousands, or even tens of thousands or more, relatively weak cores that work together very well. PayPal has been leveraging GPUs to train models for some time now, and so we decided to evaluate if the parallelism of the GPU can be helpful with processing big data applications based on Apache Spark. In our research, we encountered NVIDIA's Spark RAPIDS open-source project. It has many purposes, however we focused on Spark RAPIDS's cost reduction potential, because enterprises like PayPal spend lots of money on running Spark jobs in the cloud. Using Spark with GPUs isn't common in the industry yet, but according to our findings as described in this blog, the potential benefits could be enorm

## Introducing Setup and Teardown tasks

DevFeed: [Introducing Setup and Teardown tasks](<https://devfeed.tech/articles/introducing-setup-and-teardown-tasks-32564.md>)

Original publisher: [Read original article](<https://airflow.apache.org/blog/introducing_setup_teardown/>)

Author: Apache Airflow

Published: 2023-08-18T00:00:00Z

Content type: article

Language: en

Sources: [Apache Airflow Blog](<https://devfeed.tech/sources/apache-airflow-blog.md>)

Topics: [airflow](<https://devfeed.tech/topics/airflow.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [cleanup](<https://devfeed.tech/tags/cleanup.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [dependencies](<https://devfeed.tech/tags/dependencies.md>), [setup](<https://devfeed.tech/tags/setup.md>), [tasks](<https://devfeed.tech/tags/tasks.md>)

### AI overview

This article introduces setup and teardown tasks in Airflow 2.7 for managing infrastructure around work in data pipelines. It explains their dependency semantics, cleanup behavior, DAG run state handling, and behavior within task groups.

### Source excerpt

In data pipelines, commonly we need to create infrastructure resources, like a cluster or GPU nodes in an existing cluster, before doing the actual "work" and delete them after the work is done. Airflow 2.7 adds "setup" and "teardown" tasks to better support this type of pipeline. This blog post aims to highlight the key features so you know what's possible. For full documentation on how to use setup and teardown tasks, see the setup and teardown docs. Why setup and teardown? Before we dig into examples, let me state at high level what setup and teardown bring to the table. More expressive dependencies Before setup and teardown, upstream and downstream relationships could only mean one thing: "this comes before that". With setup and teardown, in effect we can say "this requires that". And what it means in practice is, if you clear your task, and it requires a setup, that setup will be cleared too. And if that setup has a teardown, that will run again as well. Separating the work from the infra Sometimes the part of the dag you care about is not, say, the cleanup task. For example, suppose you have a dag that loads some data and then deletes temp files. As long as the data loads, you want your dag to be marked successful. By default, this is how teardown tasks work; that is, they are ignored when determining dag run state. Simple case A simple example is one setup / teardown pair, and one normal or "work" task. Setups and teardowns are indicated by the up and down arrows, respectively. From that we can see that .create_cluster is a setup task and delete_cluster is a teardown. The link between a setup and a teardown is always dotted to highlight the special relationship. Some things to observe: If create_cluster fails, neither run_query nor delete_cluster will run. If create_cluster succeeds and run_query fails, then delete_cluster will still run. If create_cluster is skipped, run_query and delete_cluster will be skipped By default, if run_query succeeds, and delete_c

## Algo Hour - Large Scale Data & ML Monitoring with whylogs | Alessya Visnjic

DevFeed: [Algo Hour - Large Scale Data & ML Monitoring with whylogs | Alessya Visnjic](<https://devfeed.tech/articles/algo-hour-large-scale-data-ml-monitoring-with-whylogs-alessya-visnjic-29338.md>)

Original publisher: [Read original article](<https://multithreaded.stitchfix.com/blog/2022/09/29/alessya-algo-hour-announcement/>)

Published: 2022-09-29T09:00:00Z

Content type: article

Language: en

Sources: [Stitch Fix](<https://devfeed.tech/sources/stitch-fix.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-ml](<https://devfeed.tech/tags/data-ml.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

### AI overview

This talk explains how the open-source whylogs library supports end-to-end data quality and monitoring across machine learning pipelines. It covers whylogs' lightweight statistical data collection, language- and platform-agnostic approach, architecture, and application to existing data and ML pipelines.

### Source excerpt

Title: Large Scale Data & ML Monitoring with whylogs Talk Abstract: In the era of microservices, decentralized ML architectures and complex data pipelines, data quality has become a bigger challenge than ever. When data is involved in complex business processes and decisions, bad data can, and will, affect the bottom line. As a result, ensuring data quality across the entire ML pipeline is both costly, and cumbersome while data monitoring is often fragmented and performed ad hoc. An open source library called whylogs is built to address these challenges. It is a lightweight data profiling library that enables end-to-end data monitoring across the entire software stack. The library implements a language and platform agnostic approach to data quality and data monitoring. It's been deployed at massive-scale data environments, on structured and unstructured data modalities, and across a range of points in the ML lifecycle. In this talk, we will provide an overview of the whylogs architecture, including its lightweight statistical data collection approach and we will show how users can apply this library to existing data and ML pipelines. Date and Time: The talk will be held on Tuesday, October 11th at 1:00PM PDT. Recording Info: This talk was recorded live and is viewable below: Speaker Info: Alessya Visnjic is the CEO of WhyLabs, the AI Observability company building tools that power robust and responsible AI deployment. Prior to WhyLabs, Alessya was a CTO-in-residence at the Allen Institute for AI, where she evaluated commercial potential for the latest AI research. Earlier, Alessya spent 9 years at Amazon leading ML initiatives, including forecasting and data science platforms. Alessya is also the founder of Rsqrd AI, a global community of 1,000+ AI practitioners who are committed to making enterprise AI technology responsible.

## Airflow Summit 2022

DevFeed: [Airflow Summit 2022](<https://devfeed.tech/articles/airflow-summit-2022-32553.md>)

Original publisher: [Read original article](<https://airflow.apache.org/blog/airflow_summit_2022/>)

Author: Apache Airflow

Published: 2022-05-16T00:00:00Z

Content type: news

Language: en

Sources: [Apache Airflow Blog](<https://devfeed.tech/sources/apache-airflow-blog.md>)

Topics: [airflow](<https://devfeed.tech/topics/airflow.md>), [Data pipelines](<https://devfeed.tech/topics/data-pipelines.md>), [data-governance](<https://devfeed.tech/topics/data-governance.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [airflow-summit](<https://devfeed.tech/tags/airflow-summit.md>), [apache-airflow](<https://devfeed.tech/tags/apache-airflow.md>), [community](<https://devfeed.tech/tags/community.md>), [data-governance](<https://devfeed.tech/tags/data-governance.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [events](<https://devfeed.tech/tags/events.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [network](<https://devfeed.tech/tags/network.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [summit](<https://devfeed.tech/tags/summit.md>)

### AI overview

Airflow Summit 2022 was scheduled for May 23-27 as a free conference for Apache Airflow practitioners and data leaders. The program covered Airflow practices, data pipelines, data governance, machine learning, the project's future, and non-code open-source contributions.

### Source excerpt

The biggest Airflow Event of the Year returns May 23-27! Airflow Summit 2022 will bring together the global community of Apache Airflow practitioners and data leaders. What's on the Agenda During the free conference, you will hear about Apache Airflow best practices, trends in building data pipelines, data governance, Airflow and machine learning, and the future of Airflow. There will also be a series of presentations on non-code contributions driving the open-source project. How to Attend This year's edition will include a variety of online sessions across different time zones. Additionally, you can take part in local in-person events organized worldwide for data communities to watch the event and network. Interested? 🪶 Register for Airflow Summit 2022 today 🤝 Check out the in-person events planned for Airflow Summit 2022.

## Why Data Pipelines Should Avoid Monolithic Queue Processors

DevFeed: [Why Data Pipelines Should Avoid Monolithic Queue Processors](<https://devfeed.tech/articles/building-data-pipelines-learning-number-01-die-monolith-die-28152.md>)

Original publisher: [Read original article](<http://fuzzyblog.io/blog/data_pipeline/2020/07/20/building-data-pipelines-die-monolith-die.html>)

Author: Fuzzygroup

Published: 2020-07-20T00:00:00Z

Content type: article

Language: en

Sources: [Scott Johnson](<https://devfeed.tech/sources/scott-johnson.md>)

Topics: [data observability](<https://devfeed.tech/topics/data-observability.md>), [Amazon Simple Queue Service (SQS)](<https://devfeed.tech/topics/amazon-simple-queue-service-sqs.md>), [debugging](<https://devfeed.tech/topics/debugging.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [Tensorflow](<https://devfeed.tech/topics/tensorflow.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [data-pipeline](<https://devfeed.tech/tags/data-pipeline.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [process](<https://devfeed.tech/tags/process.md>), [sqs](<https://devfeed.tech/tags/sqs.md>), [tensorflow](<https://devfeed.tech/tags/tensorflow.md>)

### AI overview

The article describes lessons from building a high-performance, near-real-time data pipeline on AWS with SQS and machine learning components. It argues that a monolithic queue processor can increase cloud costs when GPU-backed processing is applied to work that does not require a GPU, make debugging harder, and limit independent scaling of different processing routines.

### Source excerpt

I am in the process of wrapping up a year long engagement where I: Built a high performance data pipeline Capable of processing all of Twitter in real time / near real time Applied multiple tools to the data at different stages of the pipeline Applied one or more Machine Learning models at different stages of the pipeline Operated on AWS using SQS as the queueing structure This blog post talks about one of the key things I learned in terms of the data pipeline, specifically: **Do Not Build Data Pipelines Around a Monolithic Queue Processor ** Note: By queue processor I mean the bit of software which pulls the data of the queue, operates on it and then puts it back. This lesson may be obvious to some but we took a meandering approach to this problem where we started with the idea of distributed pipeline components, moved to a monolithic approach and then ended up back at a distributed approach. As with a lot of research endeavors, the obvious conclusion wasn't quite so obvious in the throes of the research. Lesson 01: Monolithic Queue Processing Raises Your Costs At the heart of our queue processing were a number of Machine Learning components (python / tensorflow) that really needed a GPU for efficient data processing. The problem here is that when you have a monolithic queue processor, all your processing happens on a box with the GPU whether or not all that processing needs the GPU. When you are using cloud computing, you pay for the GPU whether not not it is being used for a given operation. And since GPU boxes generally cost at least 5x to 6x more than CPU only boxes, well, our monolithic queue processor proved to be an economic disaster. Lesson 02: Monolithic Queue Processing Is Harder to Debug After realizing Lesson 01, I took our monolithic queue processor apart and broke it down into 8 (ultimately 9) individual queue processors. One thing that I quickly found is that debugging the 8 individual queue processors was dramatically easier than debugging the singl

## Addepar's Migration from Mongo 2.4 to Mongo 3.4

DevFeed: [Addepar's Migration from Mongo 2.4 to Mongo 3.4](<https://devfeed.tech/articles/migrating-mountains-of-mongo-data-30546.md>)

Original publisher: [Read original article](<https://medium.com/build-addepar/migrating-mountains-of-mongo-data-63e530539952?source=rss----596e43e5e150---4>)

Author: Elan Kugelmass

Published: 2017-10-24T13:11:15Z

Content type: article

Language: en

Sources: [Addepar](<https://devfeed.tech/sources/addepar.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [upgrade](<https://devfeed.tech/topics/upgrade.md>), [etl](<https://devfeed.tech/topics/etl.md>), [Dependency management](<https://devfeed.tech/topics/dependency-management.md>), [Replication](<https://devfeed.tech/topics/replication.md>)

Tags: [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [database](<https://devfeed.tech/tags/database.md>), [databases](<https://devfeed.tech/tags/databases.md>), [dependency-management](<https://devfeed.tech/tags/dependency-management.md>), [migration](<https://devfeed.tech/tags/migration.md>), [mongodb](<https://devfeed.tech/tags/mongodb.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [replication](<https://devfeed.tech/tags/replication.md>), [upgrade](<https://devfeed.tech/tags/upgrade.md>)

### AI overview

Addepar describes upgrading its database from Mongo 2.4 (TokuMX 2.0) to Mongo 3.4 as its dataset and stability and performance requirements grew. The article explains Mongo's role in the company's data ingestion and ETL pipelines and why database upgrades are risky because database guarantees and behavior can change between versions.

### Source excerpt

At Addepar, we're building the world's most versatile financial analytics engine. To feed the calculations that give our clients an unprecedented view into their portfolios, we need data -- from as many sources, vendors, and intermediaries as possible. Our market and portfolio data pipelines ingest benchmarks, security terms, accounting, and performance data from hundreds of integration partners. Behind this data pipeline is a database. And like every database, ours requires maintenance and care. Maintaining an obsolete database instance is challenging due to lack of support, inferior performance, and a dwindling developer community. As our dataset grew and we faced increased stability and performance requirements, the engineering group at Addepar decided it was time to upgrade our venerable Mongo 2.4 (TokuMX 2.0) database to the latest and greatest Mongo 3.4. Every organization has at least one database saga, and we're excited to share one of ours. Dependency management extends to databases Upgrading a database is a tricky business. Like other dependencies that support a product, databases have a tendency to fall out of date. Upgrades are deferred until that imagined future where everything is stable, clients have exhausted their feature request lists, and there's not much to do in the office other than play foosball and exchange memes. It's an understandable decision! Databases are complicated, leaky abstractions that inevitably form an implicit extension of our application logic. Their data types, atomicity guarantees, and transactionality semantics define the constraints we place on data and drive how we store application state. And because these guarantees (or lack thereof) tend to change (in ways that are sometimes undocumented!) between database versions, moving to the latest release is a risky proposition. Motivated by our need for an extremely reliable datastore that could handle complex and evolving schemas, we chose Mongo as the sole database for Addepar's