# DataOps

DataOps is a set of principles for developing and delivering analytics, including the orchestration of data, tools, code, and environments.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## New data pipeline management platform at Khan Academy

DevFeed: [New data pipeline management platform at Khan Academy](<https://devfeed.tech/articles/new-data-pipeline-management-platform-at-khan-academy-27388.md>)

Original publisher: [Read original article](<http://engineering.khanacademy.org/posts/khanalytics.htm>)

Author: Khan Academy

Published: 2018-04-30T22:00:00Z

Content type: article

Language: en

Sources: [Khan Academy](<https://devfeed.tech/sources/khan-academy.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [batch](<https://devfeed.tech/tags/batch.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [cloud-dataflow](<https://devfeed.tech/tags/cloud-dataflow.md>), [data](<https://devfeed.tech/tags/data.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [news](<https://devfeed.tech/tags/news.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [scheduling](<https://devfeed.tech/tags/scheduling.md>)

### AI overview

Khan Academy developed Khanalytics to manage its growing collection of data pipelines. The platform provides a sandboxed environment for batch jobs, a web interface, automatic parallelization, centralized logs, and pipeline scheduling with dependencies.

### Source excerpt

By Ragini Gupta Data is very crucial to Khan Academy and is itself an internal product for the ... Read more

## Linux Patched For Silent User-Space Data Loss Bug That's Existed Since 2023

DevFeed: [Linux Patched For Silent User-Space Data Loss Bug That's Existed Since 2023](<https://devfeed.tech/articles/linux-patched-for-silent-user-space-data-loss-bug-that-s-existed-since-2023-17445.md>)

Original publisher: [Read original article](<https://www.phoronix.com/news/Linux-7.3-Fix-Silent-Data-Loss>)

Author: Michael Larabel

Published: 2026-09-14T10:13:45Z

Content type: news

Language: en

Sources: [Phoronix](<https://devfeed.tech/sources/phoronix.md>)

Topics: [Linux](<https://devfeed.tech/topics/linux.md>), [bug](<https://devfeed.tech/topics/bug.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [bug](<https://devfeed.tech/tags/bug.md>), [c](<https://devfeed.tech/tags/c.md>), [code](<https://devfeed.tech/tags/code.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [desktop-linux](<https://devfeed.tech/tags/desktop-linux.md>), [kernel](<https://devfeed.tech/tags/kernel.md>), [linux](<https://devfeed.tech/tags/linux.md>), [linux-benchmarking](<https://devfeed.tech/tags/linux-benchmarking.md>), [linux-hardware-benchmarks](<https://devfeed.tech/tags/linux-hardware-benchmarks.md>), [linux-hardware-reviews](<https://devfeed.tech/tags/linux-hardware-reviews.md>), [linux-how-to](<https://devfeed.tech/tags/linux-how-to.md>), [linux-performance](<https://devfeed.tech/tags/linux-performance.md>), [linux-server-benchmarks](<https://devfeed.tech/tags/linux-server-benchmarks.md>), [open-source-graphics](<https://devfeed.tech/tags/open-source-graphics.md>), [phoronix](<https://devfeed.tech/tags/phoronix.md>), [phoronix-test-suite](<https://devfeed.tech/tags/phoronix-test-suite.md>), [production](<https://devfeed.tech/tags/production.md>), [release](<https://devfeed.tech/tags/release.md>), [ubuntu-benchmarks](<https://devfeed.tech/tags/ubuntu-benchmarks.md>), [ubuntu-hardware](<https://devfeed.tech/tags/ubuntu-hardware.md>)

### AI overview

A Linux kernel bug introduced in July 2023 can silently discard user-space writes when transparent hugepages are enabled, cgroup limits apply, and memory reclaim pressure is high. The issue has caused production data loss, including for users of the Polars data analytics library. A one-line fix was merged into the x86/urgent branch and marked for backporting to supported stable kernels.

### Source excerpt

Being merged after yesterday's Linux 7.3-rc3 release was an important fix for addressing a silent, user-space data loss bug that has existed in the kernel the past three years...

## A practical approach to end-to-end Solvency II reporting in Databricks

DevFeed: [A practical approach to end-to-end Solvency II reporting in Databricks](<https://devfeed.tech/articles/a-practical-approach-to-end-to-end-solvency-ii-reporting-in-databricks-11543.md>)

Original publisher: [Read original article](<https://www.databricks.com/blog/practical-approach-end-end-solvency-ii-reporting-databricks>)

Author: Laurence Ryszka; Jack Yallop

Published: 2026-09-09T16:20:00Z

Content type: article

Language: en

Sources: [Databricks](<https://devfeed.tech/sources/databricks.md>)

Topics: [databricks](<https://devfeed.tech/topics/databricks.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [data](<https://devfeed.tech/topics/data.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [business](<https://devfeed.tech/tags/business.md>), [data](<https://devfeed.tech/tags/data.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [eu](<https://devfeed.tech/tags/eu.md>), [financial](<https://devfeed.tech/tags/financial.md>), [financial-services](<https://devfeed.tech/tags/financial-services.md>), [governance](<https://devfeed.tech/tags/governance.md>), [industries](<https://devfeed.tech/tags/industries.md>), [monitor](<https://devfeed.tech/tags/monitor.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [platform](<https://devfeed.tech/tags/platform.md>), [regulatory](<https://devfeed.tech/tags/regulatory.md>), [solutions](<https://devfeed.tech/tags/solutions.md>), [uk](<https://devfeed.tech/tags/uk.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

The article presents Databricks as a governed data, orchestration, and reporting layer for end-to-end Solvency II reporting. It covers ingestion, data quality, actuarial reserving, capital calculation, QRT production, ORSA drafting, governance, approvals, and disclosure while allowing insurers to retain established actuarial and capital modeling systems.

### Source excerpt

Solvency II reporting is not only a regulatory submission. It is a business process...

## Using Data Contracts to Coordinate Data Evolution at Enterprise Scale

DevFeed: [Using Data Contracts to Coordinate Data Evolution at Enterprise Scale](<https://devfeed.tech/articles/stop-reacting-to-data-problems-here-s-the-architecture-that-prevents-them-22547.md>)

Original publisher: [Read original article](<https://medium.com/walmartglobaltech/stop-reacting-to-data-problems-heres-the-architecture-that-prevents-them-a274d54f624b?source=rss----905ea2b3d4d1---4>)

Author: Keerthipriyan

Published: 2026-08-25T20:22:17Z

Content type: article

Language: en

Sources: [Walmart Global Tech](<https://devfeed.tech/sources/walmart-global-tech.md>)

Topics: [data-platforms](<https://devfeed.tech/topics/data-platforms.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-governance](<https://devfeed.tech/tags/data-governance.md>), [data-platforms](<https://devfeed.tech/tags/data-platforms.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [organizational](<https://devfeed.tech/tags/organizational.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [schema](<https://devfeed.tech/tags/schema.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [teams](<https://devfeed.tech/tags/teams.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

The article explains how data contracts help large enterprises coordinate changes across independently evolving data teams and downstream consumers. It argues that schema validation alone cannot identify ownership, downstream impact, or migration responsibilities, and presents data contracts as machine-enforceable coordination agreements.

### Source excerpt

Coauthored by Satyajeet Coordinating Data Evolution at Enterprise Scale When you operate data platforms on a global enterprise scale, hundreds of engineering teams ship improvements every week, each moving independently to deliver value at the pace of the business demands. This velocity is a competitive advantage. The challenge: How do you enable hundreds of teams to evolve their data products independently while maintaining reliability for thousands of downstream consumers? Traditional coordination methods (messages, wiki updates, shared spreadsheets) work at small scale but break at Walmart scale. A source team ships an enhancement, perfectly valid within their domain, but that change ripples through fifteen downstream pipelines owned by different teams with different release schedules. Without a formal coordination mechanism, you discover the impact after it reaches production. The gap isn't technical debt or fragile systems. It's the absence of machine-enforceable agreements that scale with organizational complexity. Data contracts solve this: enabling teams to move fast independently while maintaining coordinated reliability across organizational boundaries. Here's the architecture we built. Why Schema Validation Alone Isn't Enough When data quality issues surface in production, the first instinct is often added to more schema validation. If a field is missing or has the wrong type, the pipeline catches it. This works for many data quality problems, but not all of them. Consider a scenario where a source team enhances their data model by restructuring field names to support new business capabilities. The schema still validates perfectly: every field exists; every type is correct; the data is well formed. But downstream consumers who depend on the original field names now receive empty results. Schema validation checks whether data has the right shape. It tells you that a field is missing. It does not tell you who owns that field, which downstream teams will bre

## The CI/CD moment for data analytics

DevFeed: [The CI/CD moment for data analytics](<https://devfeed.tech/articles/the-ci-cd-moment-for-data-analytics-12230.md>)

Original publisher: [Read original article](<https://platformengineering.org/blog/the-ci-cd-moment-for-data-analytics>)

Author: Gaurav Nanda

Published: 2026-07-23T05:40:01Z

Content type: article

Language: en

Sources: [Platform Engineering Blog](<https://devfeed.tech/sources/platform-engineering-blog.md>)

Topics: [data analytics](<https://devfeed.tech/topics/data-analytics.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [apache-flink](<https://devfeed.tech/topics/apache-flink.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [apache-flink](<https://devfeed.tech/tags/apache-flink.md>), [architectures](<https://devfeed.tech/tags/architectures.md>), [bridging](<https://devfeed.tech/tags/bridging.md>), [cost](<https://devfeed.tech/tags/cost.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [event](<https://devfeed.tech/tags/event.md>), [fraud](<https://devfeed.tech/tags/fraud.md>), [google](<https://devfeed.tech/tags/google.md>), [insights](<https://devfeed.tech/tags/insights.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [latency](<https://devfeed.tech/tags/latency.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [scale](<https://devfeed.tech/tags/scale.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [systems](<https://devfeed.tech/tags/systems.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

The article argues that data analytics is approaching a CI/CD-like transition toward continuous analytics. It describes how multi-step pipelines introduce latency, operational risk, and maintenance burden, and explains why real-time insight is becoming a baseline platform capability for applications such as personalization, fraud detection, reliability, and AI-driven features.

### Source excerpt

How converging OLTP and OLAP architectures are driving 'Continuous Analytics', the CI/CD moment for data to deliver real-time, unified insights and operational simplicity.

## The continuous validation framework for data pipelines.

DevFeed: [The continuous validation framework for data pipelines.](<https://devfeed.tech/articles/the-continuous-validation-framework-for-data-pipelines-12232.md>)

Original publisher: [Read original article](<https://platformengineering.org/blog/the-continuous-validation-framework-for-data-pipelines>)

Author: Niruta Talwekar

Published: 2026-07-23T05:40:01Z

Content type: article

Language: en

Sources: [Platform Engineering Blog](<https://devfeed.tech/sources/platform-engineering-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [DevOps](<https://devfeed.tech/topics/devops.md>), [Platform Engineering](<https://devfeed.tech/topics/platform-engineering.md>)

Tags: [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [data](<https://devfeed.tech/tags/data.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [devops](<https://devfeed.tech/tags/devops.md>), [platform-engineering](<https://devfeed.tech/tags/platform-engineering.md>), [provenance](<https://devfeed.tech/tags/provenance.md>), [software-testing](<https://devfeed.tech/tags/software-testing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

The article introduces the Continuous Validation Framework (CVF), an end-to-end methodology for validating data pipelines through architectural isolation, configuration-driven data quality management, and continuous automation based on lineage-driven impact analysis. It reports production results including a 50% reduction in incidents and an 80% improvement in detecting data quality issues.

### Source excerpt

A framework for automated, end-to-end data pipeline validation using isolation, declarative quality checks, and lineage-driven impact analysis.

## Building a Centralized Alerting Framework for Data Quality Monitoring and Incident Management

DevFeed: [Building a Centralized Alerting Framework for Data Quality Monitoring and Incident Management](<https://devfeed.tech/articles/building-a-centralized-alerting-framework-for-data-quality-monitoring-and-incident-management-30514.md>)

Original publisher: [Read original article](<https://medium.com/helpshift-engineering/building-a-centralized-alerting-framework-for-data-quality-monitoring-and-incident-management-2f90d93a65b5?source=rss----3229f31ca4f4---4>)

Author: Manav Mehta

Published: 2026-06-18T07:11:45Z

Content type: article

Language: en

Sources: [Helpshift](<https://devfeed.tech/sources/helpshift.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [incident management](<https://devfeed.tech/topics/incident-management.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>)

Tags: [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [email](<https://devfeed.tech/tags/email.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [notifications](<https://devfeed.tech/tags/notifications.md>), [observability](<https://devfeed.tech/tags/observability.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [slack](<https://devfeed.tech/tags/slack.md>), [snowflake](<https://devfeed.tech/tags/snowflake.md>)

### AI overview

This article describes the design of a centralized alerting and incident management framework for data quality and pipeline monitoring. The framework uses Snowflake's native Alerting capabilities with Email, Slack, and Splunk On-Call integrations to detect issues, notify the appropriate engineers, escalate critical incidents, and provide actionable context.

### Source excerpt

Before We Knew Better As our data platform grew, so did the number of pipelines, scheduled tasks, and data quality checks running every day. While Snowflake provided a reliable platform for storing and processing data, operational monitoring was fragmented across multiple systems. Data quality failures were often discovered only after downstream reports showed inconsistencies. Pipeline issues sometimes required engineers to manually inspect logs, query tables, and trace execution paths before identifying the root cause. The challenge wasn't detecting failures -- we already had mechanisms to identify them. The real challenge was ensuring the right people were notified quickly, with enough context to take action. Questions during on-call incidents were often similar: Did the pipeline fail or was data simply delayed? Which validation check triggered the alert? Who should respond to the issue? How can we ensure critical failures don't get missed overnight? As the number of pipelines increased, manually monitoring these failures became increasingly difficult. We needed a centralized alerting framework. What We Actually Needed Our goal wasn't simply to send more notifications. We wanted a system that could: Detect data quality issues automatically Notify engineers through channels they already use Escalate critical incidents to on-call responders Provide actionable context instead of generic failure messages Scale across multiple pipelines and monitoring use cases Most importantly, we wanted to keep the solution as close to the data platform as possible. Since our monitoring logic already lived in Snowflake, it made sense for the alerting framework to live there as well. The Architecture We Chose To address these challenges, we designed a centralized notification and incident management framework using Snowflake's native Alerting capabilities, combined with Email, Slack, and Splunk On-Call integrations. Rather than introducing another monitoring platform, we chose to build

## From SSH to REST: A Security-Driven Modernization of Slack's EMR Data Pipelines

DevFeed: [From SSH to REST: A Security-Driven Modernization of Slack's EMR Data Pipelines](<https://devfeed.tech/articles/from-ssh-to-rest-a-security-driven-modernization-of-slack-s-emr-data-pipelines-146.md>)

Original publisher: [Read original article](<https://slack.engineering/from-ssh-to-rest-a-security-driven-modernization-of-slacks-emr-data-pipelines/>)

Author: Mahendran Vasagam

Published: 2026-05-05T14:00:01Z

Content type: article

Language: en

Sources: [Engineering at Slack](<https://devfeed.tech/sources/engineering-at-slack.md>)

Topics: [Security](<https://devfeed.tech/topics/security.md>), [OpenSSH](<https://devfeed.tech/topics/openssh.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [aws](<https://devfeed.tech/tags/aws.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [modernization](<https://devfeed.tech/tags/modernization.md>), [security](<https://devfeed.tech/tags/security.md>), [ssh](<https://devfeed.tech/tags/ssh.md>), [uncategorized](<https://devfeed.tech/tags/uncategorized.md>)

### AI overview

Slack describes migrating more than 700 SSH-based data pipeline jobs to a REST-based architecture across eight data regions, eliminating SSH access to production AWS EMR clusters without downtime. The article explains the security and operational problems that motivated the modernization, including attack-surface exposure, key-management overhead, resource contention, broken connections, zombie jobs, and unreliable job-status detection.

### Source excerpt

Excerpt By 2024, Slack's data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We're talking daily search indexing that processed terabytes of data, analytics jobs powering business intelligence, the whole shebang. Every single one of these jobs required direct SSH access to production AWS Elastic MapReduce (EMR) clusters. We had a massive security...

## Risk-Based Data Quality Testing for Reliable Finance Pipelines

DevFeed: [Risk-Based Data Quality Testing for Reliable Finance Pipelines](<https://devfeed.tech/articles/test-smarter-not-harder-risk-based-data-quality-without-pipeline-paralysis-20446.md>)

Original publisher: [Read original article](<https://vinted.engineering//2026/03/11/risk-based-testing/>)

Author: Jeremy Chia

Published: 2026-03-11T00:00:00Z

Content type: tutorial

Language: en

Sources: [Vinted](<https://devfeed.tech/sources/vinted.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Finance](<https://devfeed.tech/topics/finance.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [finance](<https://devfeed.tech/tags/finance.md>), [migration](<https://devfeed.tech/tags/migration.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [quality](<https://devfeed.tech/tags/quality.md>), [reporting](<https://devfeed.tech/tags/reporting.md>), [schema](<https://devfeed.tech/tags/schema.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This article explains how Vinted addressed frequent upstream schema and format changes affecting finance reporting pipelines. It describes shifting testing closer to the source, applying materiality-based checks, and balancing data quality with pipeline reliability and timely availability.

### Source excerpt

Upstream schema changes were breaking our finance pipelines daily. With monthly reporting deadlines looming, we needed to balance data quality with pipeline reliability. Here's how we solved it without compromising either.

## Streaming IoT and event data into Snowflake and ClickHouse

DevFeed: [Streaming IoT and event data into Snowflake and ClickHouse](<https://devfeed.tech/articles/streaming-iot-and-event-data-into-snowflake-and-clickhouse-12773.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/stream-iot-snowflake-clickhouse>)

Author: Mdu Sibisi

Published: 2025-12-09T00:00:00Z

Content type: tutorial

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Internet of things](<https://devfeed.tech/topics/iot.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Redpanda-Connect](<https://devfeed.tech/topics/redpanda-connect.md>), [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [data-governance](<https://devfeed.tech/topics/data-governance.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud-storage](<https://devfeed.tech/tags/cloud-storage.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [data](<https://devfeed.tech/tags/data.md>), [data-governance](<https://devfeed.tech/tags/data-governance.md>), [event](<https://devfeed.tech/tags/event.md>), [iot](<https://devfeed.tech/tags/iot.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [redpanda-connect](<https://devfeed.tech/tags/redpanda-connect.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

A guide to streaming IoT and event data through Redpanda and Redpanda Connect into Snowflake and ClickHouse. It compares ClickHouse for real-time analysis with Snowflake for scalable cloud storage, historical reporting, and querying, while discussing governance, compression, data freshness, and robust pipeline design.

### Source excerpt

Learn how to stream IoT and event data into Snowflake and ClickHouse using Redpanda

## Stopping Silent Failures for Meta's Fake Accounts Pipeline

DevFeed: [Stopping Silent Failures for Meta's Fake Accounts Pipeline](<https://devfeed.tech/articles/stopping-silent-failures-for-meta-s-fake-accounts-pipeline-27253.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/saving-metas-fake-accounts-pipeline>)

Author: Zach Wilson

Published: 2025-08-12T19:38:14Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Meta](<https://devfeed.tech/topics/meta.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>)

Tags: [fake-accounts](<https://devfeed.tech/tags/fake-accounts.md>), [meta](<https://devfeed.tech/tags/meta.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>)

### AI overview

An article about stopping silent failures in Meta's fake accounts pipeline.

### Source excerpt

Data Orchestration Challenges I Faced at Airbnb, Netflix & Facebook - Part IV

## How I got a 12x speed up in a 50 TB pipeline at Meta

DevFeed: [How I got a 12x speed up in a 50 TB pipeline at Meta](<https://devfeed.tech/articles/how-i-got-a-12x-speed-up-in-a-50-tb-pipeline-at-meta-27246.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/how-i-got-a-12x-speed-up-in-a-50>)

Author: Zach Wilson

Published: 2025-08-04T19:39:26Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [Meta](<https://devfeed.tech/topics/meta.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [data](<https://devfeed.tech/topics/data.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [meta](<https://devfeed.tech/tags/meta.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [speed](<https://devfeed.tech/tags/speed.md>)

### AI overview

The article describes a 12x speed improvement in a 50 TB data pipeline at Meta and places the topic within data orchestration challenges discussed across Airbnb, Netflix, and Facebook.

### Source excerpt

Data Orchestration Challenges I Faced at Airbnb, Netflix & Facebook - Part III

## From siloed DataOps, MLOps, and LLMOps to a unified data-intelligence platform

DevFeed: [From siloed DataOps, MLOps, and LLMOps to a unified data-intelligence platform](<https://devfeed.tech/articles/from-siloed-dataops-mlops-and-llmops-to-a-unified-data-intelligence-platform-26354.md>)

Original publisher: [Read original article](<https://medium.com/udemy-engineering/from-siloed-dataops-mlops-and-llmops-to-a-unified-data-intelligence-platform-4400be283641?source=rss----19c6d3367ed4---4>)

Author: Rajit Saha

Published: 2025-08-04T18:03:19Z

Content type: opinion

Language: en

Sources: [Udemy Engineering](<https://devfeed.tech/sources/udemy-engineering.md>)

Topics: [DataOps](<https://devfeed.tech/topics/dataops.md>), [MLOps](<https://devfeed.tech/topics/mlops.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [Amazon Bedrock](<https://devfeed.tech/topics/amazon-bedrock.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [Amazon SageMaker](<https://devfeed.tech/topics/amazon-sagemaker.md>), [apache-flink](<https://devfeed.tech/topics/apache-flink.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [aiops](<https://devfeed.tech/tags/aiops.md>), [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [apache-flink](<https://devfeed.tech/tags/apache-flink.md>), [apache-spark](<https://devfeed.tech/tags/apache-spark.md>), [bedrock](<https://devfeed.tech/tags/bedrock.md>), [dataops](<https://devfeed.tech/tags/dataops.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [llmops](<https://devfeed.tech/tags/llmops.md>), [mlops](<https://devfeed.tech/tags/mlops.md>)

### AI overview

The article describes how DataOps, MLOps, and AI/LLM Ops commonly rely on separate systems and teams for data processing, model deployment, inference, evaluation, orchestration, governance, and monitoring. It then introduces Databricks' Data Intelligence Platform as a unified environment intended to bring these domains together.

### Source excerpt

Introduction In modern data-driven businesses, the pace of innovation in analytics and artificial intelligence has outstripped the capacity of many teams. Three distinct disciplines emerged to handle this expansion: Data platform (DataOps) teams built data lakes on cloud storage such as Amazon S3, processed them with Apache Spark and Hive on EMR, ingested streaming data with Spark Structured Streaming or Apache Flink, and loaded tabular copies into MPP warehouses like Redshift for interactive SQL and BI. Cataloguing and governance were offloaded to external tools such as DataHub, and fine-grained access controls required third-party services like Privacera. This architecture worked, but it required separate workflows for batch and streaming, extra systems for lineage and governance, and a mosaic of operational teams. MLOps teams provided an additional layer. Data scientists used notebook environments (for example, Amazon SageMaker) to preprocess data, train, and evaluate models. Deploying models meant writing integration code to move features into a serving layer, to register models in disparate registries and to build custom APIs for inference. Feature stores and model registries were bought from additional vendors. Updates and monitoring were often manual processes. AI/LLM Ops teams are a new addition because generative AI requires specialized components: LLM gateways (e.g., Amazon Bedrock) to proxy access to foundation models; evaluation tooling to compare large language models; orchestration frameworks for agents; vector databases for retrieval augmented generation; and of course another layer of security, access management and cost control. These tools seldom integrate seamlessly with existing data and ML pipelines. This fragmented state makes it difficult to react quickly when product requirements change. Each new capability requires another system, another integration, and another team. Meanwhile, budgets tighten and go-to-market timelines shrink. The questio

## Overclocking dbt: Discord's Custom Solution in Processing Petabytes of Data

DevFeed: [Overclocking dbt: Discord's Custom Solution in Processing Petabytes of Data](<https://devfeed.tech/articles/overclocking-dbt-discord-s-custom-solution-in-processing-petabytes-of-data-273.md>)

Original publisher: [Read original article](<https://discord.com/blog/overclocking-dbt-discords-custom-solution-in-processing-petabytes-of-data>)

Author: Chris Dong

Published: 2025-04-09T00:00:00Z

Content type: article

Language: en

Sources: [Discord Blog](<https://devfeed.tech/sources/discord-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [developer-productivity](<https://devfeed.tech/topics/developer-productivity.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [code](<https://devfeed.tech/tags/code.md>), [data](<https://devfeed.tech/tags/data.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developer-productivity](<https://devfeed.tech/tags/developer-productivity.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [productivity](<https://devfeed.tech/tags/productivity.md>), [quality](<https://devfeed.tech/tags/quality.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [sql](<https://devfeed.tech/tags/sql.md>), [testing](<https://devfeed.tech/tags/testing.md>), [tool](<https://devfeed.tech/tags/tool.md>)

### AI overview

Discord describes scaling dbt to process petabytes of warehouse data while supporting more than 100 developers and over 2,500 models. It explains how custom extensions addressed long recompilation times, inefficient incremental processing, conflicting test tables, breaking changes, and complex calculations, improving performance, developer productivity, and data quality.

### Source excerpt

Explore how Discord supercharged dbt with a tailored solution designed for performance, developer productivity, and data quality.

## Data Quality at Petabyte Scale: Building Trust in the Data Lifecycle

DevFeed: [Data Quality at Petabyte Scale: Building Trust in the Data Lifecycle](<https://devfeed.tech/articles/data-quality-at-petabyte-scale-building-trust-in-the-data-lifecycle-22608.md>)

Original publisher: [Read original article](<https://medium.com/glassdoor-engineering/data-quality-at-petabyte-scale-building-trust-in-the-data-lifecycle-7052361307a4?source=rss----288d984af747---4>)

Author: Zakariah Siyaji

Published: 2025-02-14T15:52:43Z

Content type: article

Language: en

Sources: [Glassdoor Engineering](<https://devfeed.tech/sources/glassdoor-engineering.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Usability](<https://devfeed.tech/topics/usability.md>)

Tags: [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-patterns](<https://devfeed.tech/tags/data-patterns.md>), [data-platform-engineering](<https://devfeed.tech/tags/data-platform-engineering.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [decision-making](<https://devfeed.tech/tags/decision-making.md>), [gable](<https://devfeed.tech/tags/gable.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [shift-left](<https://devfeed.tech/tags/shift-left.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [trust](<https://devfeed.tech/tags/trust.md>), [usability](<https://devfeed.tech/tags/usability.md>)

### AI overview

Glassdoor describes a shift from reactive data engineering to a proactive, trust-centered approach to data quality. The article connects organizational culture with technical checks across the data lifecycle.

### Source excerpt

The data Lifecycle with Data Quality Checks at GlassdoorMotivation Glassdoor has transformed from an employee review site to a community for workplace conversations [1]. As our platform evolves to support content creators, facilitate discussions, and offer rich content, it has become more apparent than ever that adopting a data-driven culture is essential. Businesses rely on accurate, high-quality data to understand their operations and assess strategic outcomes. Flawed or incomplete data results in misguided decisions and undermines trust. Recognizing this risk, we made data quality a foundational principle of our data-driven transformation. Although every company defines data quality differently, there is a universal expectation that data used for decision-making must be trustworthy. Additionally, data quality challenges are not solely technical; a psychological component is closely linked to trust in data. Airbnb recognized this and sought to develop a scoring system that acknowledges the belief that data quality is a multivariate issue, encompassing accuracy, reliability, stewardship, and usability, along with more detailed dimensions within each of these categories [2]. On the other hand, Netflix employs a more technically centered approach to quality: data is initially written to a temporary staging area, audited, and then published to the production location upon passing quality checks [3]. Ultimately, Glassdoor drew inspiration from these lessons and aimed to reinforce trust through a cultural shift and a series of technical solutions. This article demonstrates how a proactive, trust-centered approach that connects data producers and consumers establishes a foundation for more rigorous data quality methods, ultimately bolstering a strong company-wide strategy. Figure 1. Enhancing quality guards at the application code layer.Culture Shift: Reactive to Proactive Historically, Glassdoor's data engineering teams have been reactive, learning about issues only aft

## How to break into data in 2024? With DataCamp's CEO, Jonathan Cornelissen.

DevFeed: [How to break into data in 2024? With DataCamp's CEO, Jonathan Cornelissen.](<https://devfeed.tech/articles/how-to-break-into-data-in-2024-with-datacamp-s-ceo-jonathan-cornelissen-39161.md>)

Original publisher: [Read original article](<https://merinova.substack.com/p/how-to-break-into-data-in-2024-with>)

Author: Meri Nova

Published: 2024-10-09T17:31:09Z

Content type: opinion

Language: en

Sources: [Meri Nova](<https://devfeed.tech/sources/meri-nova.md>)

Topics: [Data Science](<https://devfeed.tech/topics/data-science.md>), [data analytics](<https://devfeed.tech/topics/data-analytics.md>), [genai](<https://devfeed.tech/topics/genai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [beginners](<https://devfeed.tech/tags/beginners.md>), [career-advice](<https://devfeed.tech/tags/career-advice.md>), [data](<https://devfeed.tech/tags/data.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [education](<https://devfeed.tech/tags/education.md>), [future-of-work](<https://devfeed.tech/tags/future-of-work.md>), [genai](<https://devfeed.tech/tags/genai.md>), [podcast](<https://devfeed.tech/tags/podcast.md>), [scaling](<https://devfeed.tech/tags/scaling.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

A Technical Founder podcast episode featuring DataCamp co-founder and CEO Jonathan Cornelissen. The discussion covers building DataCamp, entering data careers, data education, GenAI's impact on edtech, data literacy, and career advice for new graduates.

### Source excerpt

Listen to our first episode of the "Technical Founder" podcast, where we invite AI and Data leaders to learn from their entrepreneurial and technical journey!

## Data quality testing

DevFeed: [Data quality testing](<https://devfeed.tech/articles/data-quality-testing-11744.md>)

Original publisher: [Read original article](<https://incident.io/blog/data-quality-testing>)

Author: Lambert Le Manh

Published: 2024-09-04T16:30:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [ci](<https://devfeed.tech/topics/ci.md>), [Transactions](<https://devfeed.tech/topics/transactions.md>)

Tags: [ci](<https://devfeed.tech/tags/ci.md>), [data-observability](<https://devfeed.tech/tags/data-observability.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [flaky](<https://devfeed.tech/tags/flaky.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [observability](<https://devfeed.tech/tags/observability.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [testing](<https://devfeed.tech/tags/testing.md>), [transactions](<https://devfeed.tech/tags/transactions.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

The article explains how incident.io uses dbt's native testing features for data quality testing and data observability. It describes integrating validation into production and CI workflows, including custom tests and relationship tests, and addresses flaky failures caused by ingestion and transformation pipelines running at different times.

### Source excerpt

Our data observability workflow uses data quality testing to ensure data meets accuracy, consistency, and reliability standards, enabling confident, data-driven decisions. See how we built it, the common challenges we encountered, and the solutions.

## Emerging Kafka Trends in Connectors, Self-Service Environments, and Stream Processing

DevFeed: [Emerging Kafka Trends in Connectors, Self-Service Environments, and Stream Processing](<https://devfeed.tech/articles/three-plus-some-lovely-kafka-trends-18882.md>)

Original publisher: [Read original article](<https://www.morling.dev/blog/three-plus-some-lovely-kafka-trends/>)

Published: 2021-05-28T08:30:00Z

Content type: opinion

Language: en

Sources: [Gunnar Morling](<https://devfeed.tech/sources/gunnar-morling.md>)

Topics: [Kafka](<https://devfeed.tech/topics/kafka.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [kafka](<https://devfeed.tech/tags/kafka.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [review](<https://devfeed.tech/tags/review.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>)

### AI overview

The article identifies recurring themes in Kafka community activity: the expanding connector ecosystem and its operational maturity, self-service Kafka environments, and growing adoption of stream processing. It also mentions related concerns including access control, schema management, data lineage, quality, compliance, privacy, and observability.

### Source excerpt

Table of Contents Cambrian Explosion of Connectors Democratization of Data Pipelines Stream Processing for Everyone Honorable Mentions Over the course of the last few months, I've had the pleasure to serve on the Kafka Summit program committee and review several hundred session abstracts for the three Summits happening this year (Europe, APAC, Americas). That's not only a big honour, but also a unique opportunity to learn what excites people currently in the Kafka eco-system (and yes, it's a fair amount of work, too ;). While voting on the proposals, and also generally aspiring to stay informed of what's going on in the Kafka community at large, I noticed a few repeating themes and topics which I thought would be interesting to share (without touching on any specific talks of course). At first I meant to put this out via a Twitter thread, but then it became a bit too long for that, so I decided to write this quick blog post instead. Here it goes!

## DataOps: How to Develop and Scale Data Intensive Projects

DevFeed: [DataOps: How to Develop and Scale Data Intensive Projects](<https://devfeed.tech/articles/dataops-how-to-develop-and-scale-data-intensive-projects-18467.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/dataops>)

Author: Alberto Romeu

Published: 2021-02-27T00:00:00Z

Content type: tutorial

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [DataOps](<https://devfeed.tech/topics/dataops.md>), [DevOps](<https://devfeed.tech/topics/devops.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [dataops](<https://devfeed.tech/tags/dataops.md>), [develop](<https://devfeed.tech/tags/develop.md>), [devops](<https://devfeed.tech/tags/devops.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [scalable-analytics-architecture](<https://devfeed.tech/tags/scalable-analytics-architecture.md>), [scale](<https://devfeed.tech/tags/scale.md>)

### AI overview

The article presents DataOps practices for developing and scaling data-intensive projects. It describes these workflows as applying DevOps rigor to data pipelines in 2026.

### Source excerpt

DataOps practices separate teams that ship from teams that struggle. These workflows bring DevOps rigor to data pipelines in 2026.

## DataOps: 10 principles to develop data intensive projects

DevFeed: [DataOps: 10 principles to develop data intensive projects](<https://devfeed.tech/articles/dataops-10-principles-to-develop-data-intensive-projects-18468.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/dataops-principles>)

Author: Alberto Romeu

Published: 2021-02-27T00:00:00Z

Content type: article

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [DataOps](<https://devfeed.tech/topics/dataops.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [dataops](<https://devfeed.tech/tags/dataops.md>), [develop](<https://devfeed.tech/tags/develop.md>), [projects](<https://devfeed.tech/tags/projects.md>), [scalable-analytics-architecture](<https://devfeed.tech/tags/scalable-analytics-architecture.md>), [teams](<https://devfeed.tech/tags/teams.md>)

### AI overview

The article presents 10 DataOps principles for developing data-intensive projects and supporting data teams.

### Source excerpt

10 of the principles of DataOps that we make available to data teams.

## Implementing ETL on GCP

DevFeed: [Implementing ETL on GCP](<https://devfeed.tech/articles/implementing-etl-on-gcp-22998.md>)

Original publisher: [Read original article](<https://bravenewgeek.com/implementing-etl-on-gcp/>)

Author: Deepmala

Published: 2020-07-15T20:53:17Z

Content type: tutorial

Language: en

Sources: [Brave New Geek](<https://devfeed.tech/sources/brave-new-geek.md>)

Topics: [DataOps](<https://devfeed.tech/topics/dataops.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [data loss prevention](<https://devfeed.tech/topics/data-loss-prevention.md>), [Low code](<https://devfeed.tech/topics/low-code.md>), [No-code](<https://devfeed.tech/topics/no-code.md>), [olap](<https://devfeed.tech/topics/olap.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [analytics-pipeline](<https://devfeed.tech/tags/analytics-pipeline.md>), [bi-tools](<https://devfeed.tech/tags/bi-tools.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [cdap](<https://devfeed.tech/tags/cdap.md>), [cloud-data-loss-prevention](<https://devfeed.tech/tags/cloud-data-loss-prevention.md>), [cloud-dataflow](<https://devfeed.tech/tags/cloud-dataflow.md>), [cloud-dataprep](<https://devfeed.tech/tags/cloud-dataprep.md>), [cloud-dataproc](<https://devfeed.tech/tags/cloud-dataproc.md>), [cloud-pub-sub](<https://devfeed.tech/tags/cloud-pub-sub.md>), [cloud-storage](<https://devfeed.tech/tags/cloud-storage.md>), [cloud-tasks](<https://devfeed.tech/tags/cloud-tasks.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-fusion](<https://devfeed.tech/tags/data-fusion.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [data-loss-prevention](<https://devfeed.tech/tags/data-loss-prevention.md>), [elt](<https://devfeed.tech/tags/elt.md>), [etl](<https://devfeed.tech/tags/etl.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [no-code](<https://devfeed.tech/tags/no-code.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

A practical guide to building ETL pipelines on Google Cloud Platform using Google-managed services. It explains a two-phase architecture with Cloud Storage as a data lake, Cloud Data Loss Prevention for sensitive-data detection or redaction, and BigQuery as the curated data warehouse, with attention to low-code and no-code approaches.

### Source excerpt

ETL (Extract-Transform-Load) processes are an essential component of any data analytics program. This typically involves loading data from disparate sources, transforming or enriching it, and storing the curated data in a data warehouse for consumption by different users or systems. An example of this would be taking customer data from operational databases, joining it with data from Salesforce and Google Analytics, and writing it to an OLAP database or BI engine.