# data observability

A data-engineering practice for monitoring and improving the health of data in applications and data pipelines.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Data Engineering Weekly #287

DevFeed: [Data Engineering Weekly #287](<https://devfeed.tech/articles/data-engineering-weekly-287-18267.md>)

Original publisher: [Read original article](<https://www.dataengineeringweekly.com/p/data-engineering-weekly-287>)

Author: Ananth Packkildurai

Published: 2026-09-14T02:52:23Z

Content type: article

Language: en

Sources: [Data Engineering Weekly](<https://devfeed.tech/sources/data-engineering-weekly.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [Multi-tenancy](<https://devfeed.tech/topics/multi-tenancy.md>), [Event-Streaming](<https://devfeed.tech/topics/event-streaming.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Library](<https://devfeed.tech/topics/library.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [architectures](<https://devfeed.tech/tags/architectures.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [embedding](<https://devfeed.tech/tags/embedding.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [multi-tenancy](<https://devfeed.tech/tags/multi-tenancy.md>), [object-storage](<https://devfeed.tech/tags/object-storage.md>), [observability](<https://devfeed.tech/tags/observability.md>), [openai](<https://devfeed.tech/tags/openai.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [vector](<https://devfeed.tech/tags/vector.md>)

### AI overview

Data Engineering Weekly #287 covers building data platforms from scratch, including composable architectures, data quality, and observability. It also previews talks on governed machine-executable ontologies for marketing activation and fair, order-preserving Kafka consumption for many tenants. The issue links to OpenAI's storage platform scaling for ChatGPT and Pinterest's embedding retrieval platform.

### Source excerpt

The Weekly Data Engineering Newsletter

## Data pipeline monitoring 101: Tracking health and performance across the data stack

DevFeed: [Data pipeline monitoring 101: Tracking health and performance across the data stack](<https://devfeed.tech/articles/data-pipeline-monitoring-101-tracking-health-and-performance-across-the-data-stack-2253.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/data-pipeline-monitoring/>)

Author: Aaron Kaplan; Ryan Warrier

Published: 2026-08-14T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [observability](<https://devfeed.tech/topics/observability.md>), [data](<https://devfeed.tech/topics/data.md>), [data streams monitoring](<https://devfeed.tech/topics/data-streams-monitoring.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [event driven](<https://devfeed.tech/topics/event-driven.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [data](<https://devfeed.tech/tags/data.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-observability](<https://devfeed.tech/tags/data-observability.md>), [data-streams-monitoring](<https://devfeed.tech/tags/data-streams-monitoring.md>), [event-driven](<https://devfeed.tech/tags/event-driven.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [learn](<https://devfeed.tech/tags/learn.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [observability](<https://devfeed.tech/tags/observability.md>), [performance](<https://devfeed.tech/tags/performance.md>), [spark](<https://devfeed.tech/tags/spark.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

This article introduces end-to-end monitoring for modern data pipelines. It explains how to track pipeline health, performance, data quality, and availability across varied architectures and technology layers, with examples including Kafka, Flink, Apache Spark, data lakes, warehouses, and lakehouses.

### Source excerpt

Learn about monitoring the end-to-end health and performance of modern data pipelines.

## Data Engineering Weekly #282

DevFeed: [Data Engineering Weekly #282](<https://devfeed.tech/articles/data-engineering-weekly-282-18262.md>)

Original publisher: [Read original article](<https://www.dataengineeringweekly.com/p/data-engineering-weekly-282>)

Author: Ananth Packkildurai

Published: 2026-08-10T01:21:26Z

Content type: article

Language: en

Sources: [Data Engineering Weekly](<https://devfeed.tech/sources/data-engineering-weekly.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [data-platforms](<https://devfeed.tech/topics/data-platforms.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>), [semantic-layer](<https://devfeed.tech/topics/semantic-layer.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [gRPC](<https://devfeed.tech/topics/grpc.md>), [Concurrency](<https://devfeed.tech/topics/concurrency.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [caching](<https://devfeed.tech/tags/caching.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [concurrency](<https://devfeed.tech/tags/concurrency.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-platforms](<https://devfeed.tech/tags/data-platforms.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [knowledge-graphs](<https://devfeed.tech/tags/knowledge-graphs.md>), [llms](<https://devfeed.tech/tags/llms.md>), [newsletter](<https://devfeed.tech/tags/newsletter.md>), [observability](<https://devfeed.tech/tags/observability.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [semantic-layer](<https://devfeed.tech/tags/semantic-layer.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Data Engineering Weekly #282 is a newsletter covering data platform fundamentals, semantic layers, ontology-backed knowledge graphs, converged databases, AI modernization, and Netflix's real-time distributed graph query architecture. It highlights composable architectures, data quality and observability, evolving schemas supported by LLM-assisted extraction, Iceberg full-text search, and optimization techniques including concurrency control, streaming filters, and caching.

### Source excerpt

The Weekly Data Engineering Newsletter

## The continuous validation framework for data pipelines.

DevFeed: [The continuous validation framework for data pipelines.](<https://devfeed.tech/articles/the-continuous-validation-framework-for-data-pipelines-12232.md>)

Original publisher: [Read original article](<https://platformengineering.org/blog/the-continuous-validation-framework-for-data-pipelines>)

Author: Niruta Talwekar

Published: 2026-07-23T05:40:01Z

Content type: article

Language: en

Sources: [Platform Engineering Blog](<https://devfeed.tech/sources/platform-engineering-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [DevOps](<https://devfeed.tech/topics/devops.md>), [Platform Engineering](<https://devfeed.tech/topics/platform-engineering.md>)

Tags: [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [data](<https://devfeed.tech/tags/data.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [devops](<https://devfeed.tech/tags/devops.md>), [platform-engineering](<https://devfeed.tech/tags/platform-engineering.md>), [provenance](<https://devfeed.tech/tags/provenance.md>), [software-testing](<https://devfeed.tech/tags/software-testing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

The article introduces the Continuous Validation Framework (CVF), an end-to-end methodology for validating data pipelines through architectural isolation, configuration-driven data quality management, and continuous automation based on lineage-driven impact analysis. It reports production results including a 50% reduction in incidents and an 80% improvement in detecting data quality issues.

### Source excerpt

A framework for automated, end-to-end data pipeline validation using isolation, declarative quality checks, and lineage-driven impact analysis.

## Agentic Data Engineering Is Here -- But Can It Close the Loop?

DevFeed: [Agentic Data Engineering Is Here -- But Can It Close the Loop?](<https://devfeed.tech/articles/agentic-data-engineering-is-here-but-can-it-close-the-loop-38703.md>)

Original publisher: [Read original article](<https://dataengineeringcentral.substack.com/p/agentic-data-engineering-is-here>)

Author: Daniel Beach

Published: 2026-07-22T14:05:27Z

Content type: article

Language: en

Sources: [Data Engineering Central](<https://devfeed.tech/sources/data-engineering-central.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [DuckDB](<https://devfeed.tech/topics/duckdb.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [duckdb](<https://devfeed.tech/tags/duckdb.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>)

### AI overview

A podcast conversation with Hugo Lu about agentic data engineering and the infrastructure needed for data platforms to execute work, observe outcomes, validate changes, and improve pipelines safely. It examines why production data systems remain difficult for AI, including schema changes, realistic testing, business semantics, and secure execution.

### Source excerpt

a conversation with Hugo Lu

## Data Engineering Weekly #276

DevFeed: [Data Engineering Weekly #276](<https://devfeed.tech/articles/data-engineering-weekly-276-18256.md>)

Original publisher: [Read original article](<https://www.dataengineeringweekly.com/p/data-engineering-weekly-276>)

Author: Ananth Packkildurai

Published: 2026-06-29T03:52:17Z

Content type: article

Language: en

Sources: [Data Engineering Weekly](<https://devfeed.tech/sources/data-engineering-weekly.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [schema-evolution](<https://devfeed.tech/topics/schema-evolution.md>), [data-architecture](<https://devfeed.tech/topics/data-architecture.md>), [Amazon OpenSearch Service](<https://devfeed.tech/topics/amazon-opensearch-service.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [compatibility](<https://devfeed.tech/tags/compatibility.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [observability](<https://devfeed.tech/tags/observability.md>), [opensearch](<https://devfeed.tech/tags/opensearch.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [schema-evolution](<https://devfeed.tech/tags/schema-evolution.md>)

### AI overview

Data Engineering Weekly #276 is a newsletter roundup covering data platform fundamentals, storage and workload architecture, schema evolution in Pinterest's ingestion framework, zone-failure-resilient OpenSearch at Uber, AI modernization, and stateful reasoning systems for notebooks.

### Source excerpt

The Weekly Data Engineering Newsletter

## Get reliable answers to business questions with Bits Data Analysis

DevFeed: [Get reliable answers to business questions with Bits Data Analysis](<https://devfeed.tech/articles/get-reliable-answers-to-business-questions-with-bits-data-analysis-2234.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/bits-data-analysis/>)

Author: Jonathan Morin; Jonathan Parisot; Harel Shein

Published: 2026-06-09T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Data analysis](<https://devfeed.tech/topics/data-analysis.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [semantic-layer](<https://devfeed.tech/topics/semantic-layer.md>), [SIEM, Security, Observability](<https://devfeed.tech/topics/siem-security-observability.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [ai-coding](<https://devfeed.tech/topics/ai-coding.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [bits-ai](<https://devfeed.tech/tags/bits-ai.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-observability](<https://devfeed.tech/tags/data-observability.md>), [digital-experience-monitoring](<https://devfeed.tech/tags/digital-experience-monitoring.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [observability](<https://devfeed.tech/tags/observability.md>), [root-cause-analysis](<https://devfeed.tech/tags/root-cause-analysis.md>), [semantic-layer](<https://devfeed.tech/tags/semantic-layer.md>)

### AI overview

Bits Data Analysis, now in Preview, helps teams answer business questions with governed context from their data stack and Datadog telemetry. It uses metric definitions, lineage, freshness, quality signals, application telemetry, and source code to select appropriate data and provide confidence indicators with links to the definitions and tables used.

### Source excerpt

Learn how Bits Data Analysis answers business questions using governed data context from Datadog.

## Malicious Release of elementary-data PyPI Package Steals Cloud Credentials from Data Engineers

DevFeed: [Malicious Release of elementary-data PyPI Package Steals Cloud Credentials from Data Engineers](<https://devfeed.tech/articles/malicious-release-of-elementary-data-pypi-package-steals-cloud-credentials-from-data-engineers-8011.md>)

Original publisher: [Read original article](<https://snyk.io/blog/malicious-release-of-elementary-data-pypi-package-steals-cloud-credentials-from-data-engineers/>)

Author: Liran Tal

Published: 2026-04-27T23:00:00Z

Content type: news

Language: en

Sources: [Blog RSS Feed | Snyk](<https://devfeed.tech/sources/blog-rss-feed-snyk.md>)

Topics: [data observability](<https://devfeed.tech/topics/data-observability.md>), [GitHub Actions](<https://devfeed.tech/topics/github-actions.md>), [Security](<https://devfeed.tech/topics/security.md>), [supply-chain-security](<https://devfeed.tech/topics/supply-chain-security.md>), [vulnerability](<https://devfeed.tech/topics/vulnerability.md>), [backdoor](<https://devfeed.tech/topics/backdoor.md>), [Python](<https://devfeed.tech/topics/python.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [ssh](<https://devfeed.tech/topics/ssh.md>)

Tags: [awareness](<https://devfeed.tech/tags/awareness.md>), [aws](<https://devfeed.tech/tags/aws.md>), [backdoor](<https://devfeed.tech/tags/backdoor.md>), [blog](<https://devfeed.tech/tags/blog.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [developer](<https://devfeed.tech/tags/developer.md>), [devops](<https://devfeed.tech/tags/devops.md>), [devrel](<https://devfeed.tech/tags/devrel.md>), [docker](<https://devfeed.tech/tags/docker.md>), [github-actions](<https://devfeed.tech/tags/github-actions.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [python](<https://devfeed.tech/tags/python.md>), [security](<https://devfeed.tech/tags/security.md>), [snyk-platform](<https://devfeed.tech/tags/snyk-platform.md>), [ssh](<https://devfeed.tech/tags/ssh.md>), [supply-chain](<https://devfeed.tech/tags/supply-chain.md>), [supply-chain-security](<https://devfeed.tech/tags/supply-chain-security.md>), [vulnerability](<https://devfeed.tech/tags/vulnerability.md>)

### AI overview

Attackers compromised the elementary-data Python package's GitHub Actions publication pipeline through script injection and released malicious content that stole cloud credentials and SSH secrets from data engineering environments.

### Source excerpt

Attackers exploited a GitHub Actions script injection vulnerability to publish a malicious version of the elementary-data Python CLI (v0.23.3), embedding a credential-stealing backdoor that targeted dbt profiles, cloud provider keys, and SSH secrets from data engineering environments.

## Chainguard customers safe from elementary-data compromise

DevFeed: [Chainguard customers safe from elementary-data compromise](<https://devfeed.tech/articles/chainguard-customers-safe-from-elementary-data-compromise-12937.md>)

Original publisher: [Read original article](<https://www.chainguard.dev/unchained/chainguard-customers-safe-from-elementary-data-compromise>)

Published: 2026-04-25T00:00:00Z

Content type: article

Language: en

Sources: [Chainguard: Unchained](<https://devfeed.tech/sources/chainguard-unchained.md>)

Topics: [chainguard](<https://devfeed.tech/topics/chainguard.md>), [Malware](<https://devfeed.tech/topics/malware.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [chainguard containers](<https://devfeed.tech/topics/chainguard-containers.md>), [chainguard libraries](<https://devfeed.tech/topics/chainguard-libraries.md>)

Tags: [chainguard](<https://devfeed.tech/tags/chainguard.md>), [chainguard-containers](<https://devfeed.tech/tags/chainguard-containers.md>), [chainguard-customers](<https://devfeed.tech/tags/chainguard-customers.md>), [chainguard-libraries](<https://devfeed.tech/tags/chainguard-libraries.md>), [data-observability](<https://devfeed.tech/tags/data-observability.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [elementary-data-compromise](<https://devfeed.tech/tags/elementary-data-compromise.md>), [malware](<https://devfeed.tech/tags/malware.md>), [pypi](<https://devfeed.tech/tags/pypi.md>), [pypi-malware](<https://devfeed.tech/tags/pypi-malware.md>)

### AI overview

Chainguard reports that customers using its Python Libraries and Container images were unaffected by the compromised elementary-data 0.23.3 package on PyPI. Chainguard detected malicious patterns before building the package, while the compromised release was quarantined and related GitHub and Docker artifacts were removed.

### Source excerpt

Malicious elementary-data version hit PyPI. Chainguard customers stayed protected by detecting malware pre-build and serving only verified safe versions.

## Data-to-Production: Bridging the Gap Between Iceberg and Live Microservices

DevFeed: [Data-to-Production: Bridging the Gap Between Iceberg and Live Microservices](<https://devfeed.tech/articles/data-to-production-bridging-the-gap-between-iceberg-and-live-microservices-22631.md>)

Original publisher: [Read original article](<https://www.wix.engineering/post/data-to-production-bridging-the-gap-between-iceberg-and-live-microservices>)

Author: Wix Engineering

Published: 2026-02-17T11:04:08Z

Content type: article

Language: en

Sources: [Wix Engineering](<https://devfeed.tech/sources/wix-engineering.md>)

Topics: [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Microservice](<https://devfeed.tech/topics/microservice.md>), [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [api](<https://devfeed.tech/tags/api.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [json](<https://devfeed.tech/tags/json.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [spark](<https://devfeed.tech/tags/spark.md>)

### AI overview

Wix describes Data-to-Production, a platform that activates data from Amazon S3 and Apache Iceberg for backend microservices. The system ingests Iceberg data into ClickHouse and serves it through a type-safe JSON API, using metadata governance and an Airflow and Python ingestion engine.

### Source excerpt

At Wix, our Data Warehouse (DWH) is a massive repository of insights. Built on Amazon S3 using Apache Iceberg table formats, and populated by Trino and Spark jobs, it houses petabytes of data--from user segmentation and logs to AI chat analytics. However, storage is only half the battle. The real challenge--and the "holy grail" for many data engineering teams--is Activation : taking that petabyte-scale data and exposing it to backend microservices with millisecond latency, high availability, and...

## Monitoring Temporal Cloud with ClickStack

DevFeed: [Monitoring Temporal Cloud with ClickStack](<https://devfeed.tech/articles/monitoring-temporal-cloud-with-clickstack-5429.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/monitoring-temporal>)

Author: The ClickStack Team

Published: 2026-01-28T14:25:05Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [OpenTelemetry](<https://devfeed.tech/topics/opentelemetry.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>)

Tags: [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [code](<https://devfeed.tech/tags/code.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [opentelemetry](<https://devfeed.tech/tags/opentelemetry.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

ClickStack integrates with Temporal Cloud's OpenMetrics endpoint to provide scalable observability for long-running workflows and high-cardinality telemetry. Built on ClickHouse, it supports logs, metrics, and traces for monitoring, troubleshooting, and performance analysis.

### Source excerpt

Learn how ClickStack integrates with Temporal Cloud to deliver fast, scalable observability for high-cardinality metrics, helping you spot bottlenecks, debug failures, and understand workflow health in minutes.

## Streaming IoT and event data into Snowflake and ClickHouse

DevFeed: [Streaming IoT and event data into Snowflake and ClickHouse](<https://devfeed.tech/articles/streaming-iot-and-event-data-into-snowflake-and-clickhouse-12773.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/stream-iot-snowflake-clickhouse>)

Author: Mdu Sibisi

Published: 2025-12-09T00:00:00Z

Content type: tutorial

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Internet of things](<https://devfeed.tech/topics/iot.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Redpanda-Connect](<https://devfeed.tech/topics/redpanda-connect.md>), [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [data-governance](<https://devfeed.tech/topics/data-governance.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud-storage](<https://devfeed.tech/tags/cloud-storage.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [data](<https://devfeed.tech/tags/data.md>), [data-governance](<https://devfeed.tech/tags/data-governance.md>), [event](<https://devfeed.tech/tags/event.md>), [iot](<https://devfeed.tech/tags/iot.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [redpanda-connect](<https://devfeed.tech/tags/redpanda-connect.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

A guide to streaming IoT and event data through Redpanda and Redpanda Connect into Snowflake and ClickHouse. It compares ClickHouse for real-time analysis with Snowflake for scalable cloud storage, historical reporting, and querying, while discussing governance, compression, data freshness, and robust pipeline design.

### Source excerpt

Learn how to stream IoT and event data into Snowflake and ClickHouse using Redpanda

## No, AI is not replacing Data Engineers.

DevFeed: [No, AI is not replacing Data Engineers.](<https://devfeed.tech/articles/no-ai-is-not-replacing-data-engineers-40853.md>)

Original publisher: [Read original article](<https://mutto.fyi/posts/2025/05/ai-and-data-engineering/>)

Published: 2025-05-29T00:00:00Z

Content type: opinion

Language: en

Sources: [Mutt0-ds Notes](<https://devfeed.tech/sources/mutt0-ds-notes.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [ChatGPT](<https://devfeed.tech/topics/chatgpt.md>), [Python](<https://devfeed.tech/topics/python.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-tools](<https://devfeed.tech/tags/ai-tools.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [claude](<https://devfeed.tech/tags/claude.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [python](<https://devfeed.tech/tags/python.md>)

### AI overview

An opinion article argues that AI is not replacing data engineers, focusing on the continued importance of greenfield prototyping in data engineering. It describes messy, real-world work involving APIs, spreadsheets, data cleaning, and Python scripts, contrasting it with AI demonstrations and synthetic coding benchmarks.

### Source excerpt

Claude 4 landed, and as usual, it kicked off the usual hype cycle. "Best coding model ever" "It writes stories!" "AI is replacing...

## April 2025 Newsletter

DevFeed: [April 2025 Newsletter](<https://devfeed.tech/articles/april-2025-newsletter-4888.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/202504-newsletter>)

Author: Mark Needham

Published: 2025-04-16T08:29:42Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [airflow](<https://devfeed.tech/topics/airflow.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [airflow](<https://devfeed.tech/tags/airflow.md>), [aws](<https://devfeed.tech/tags/aws.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data](<https://devfeed.tech/tags/data.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [data-observability](<https://devfeed.tech/tags/data-observability.md>), [newsletter](<https://devfeed.tech/tags/newsletter.md>), [openai](<https://devfeed.tech/tags/openai.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

The April 2025 ClickHouse newsletter covers CloudQuery's experience with ClickHouse, the new query condition cache in version 25.3, ClickHouse's Rust development, the acquisition of HyperDX, community news, upcoming events, training, and ClickHouse's role in data observability and AI applications.

### Source excerpt

Welcome to the April ClickHouse newsletter, which will round up what's happened in real-time data warehouses over the last month.

## Overclocking dbt: Discord's Custom Solution in Processing Petabytes of Data

DevFeed: [Overclocking dbt: Discord's Custom Solution in Processing Petabytes of Data](<https://devfeed.tech/articles/overclocking-dbt-discord-s-custom-solution-in-processing-petabytes-of-data-273.md>)

Original publisher: [Read original article](<https://discord.com/blog/overclocking-dbt-discords-custom-solution-in-processing-petabytes-of-data>)

Author: Chris Dong

Published: 2025-04-09T00:00:00Z

Content type: article

Language: en

Sources: [Discord Blog](<https://devfeed.tech/sources/discord-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [developer-productivity](<https://devfeed.tech/topics/developer-productivity.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [code](<https://devfeed.tech/tags/code.md>), [data](<https://devfeed.tech/tags/data.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developer-productivity](<https://devfeed.tech/tags/developer-productivity.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [productivity](<https://devfeed.tech/tags/productivity.md>), [quality](<https://devfeed.tech/tags/quality.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [sql](<https://devfeed.tech/tags/sql.md>), [testing](<https://devfeed.tech/tags/testing.md>), [tool](<https://devfeed.tech/tags/tool.md>)

### AI overview

Discord describes scaling dbt to process petabytes of warehouse data while supporting more than 100 developers and over 2,500 models. It explains how custom extensions addressed long recompilation times, inefficient incremental processing, conflicting test tables, breaking changes, and complex calculations, improving performance, developer productivity, and data quality.

### Source excerpt

Explore how Discord supercharged dbt with a tailored solution designed for performance, developer productivity, and data quality.

## Open Preference Dataset for Text-to-Image Generation by the 🤗 Community

DevFeed: [Open Preference Dataset for Text-to-Image Generation by the 🤗 Community](<https://devfeed.tech/articles/open-preference-dataset-for-text-to-image-generation-by-the-community-7274.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/image-preferences>)

Author: David Berenstein; ben burtenshaw; Daniel Vila; Daniel van Strien; Sayak Paul; Ame Vi; Linoy Tsaban

Published: 2024-12-09T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [flux](<https://devfeed.tech/topics/flux.md>), [stable-diffusion](<https://devfeed.tech/topics/stable-diffusion.md>), [argilla](<https://devfeed.tech/topics/argilla.md>), [distilabel](<https://devfeed.tech/topics/distilabel.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>)

Tags: [argilla](<https://devfeed.tech/tags/argilla.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [data](<https://devfeed.tech/tags/data.md>), [data-is-better-together](<https://devfeed.tech/tags/data-is-better-together.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [distilabel](<https://devfeed.tech/tags/distilabel.md>), [flux](<https://devfeed.tech/tags/flux.md>), [generation](<https://devfeed.tech/tags/generation.md>), [github](<https://devfeed.tech/tags/github.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [image](<https://devfeed.tech/tags/image.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [stable-diffusion](<https://devfeed.tech/tags/stable-diffusion.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [text-to-image](<https://devfeed.tech/tags/text-to-image.md>)

### AI overview

The article describes an open community effort to create an image-preference dataset for text-to-image generation. It covers prompt preparation with distilabel, synthetic data generation, image generation with Flux and Stable Diffusion, and filtering with text- and image-based classifiers plus manual review. The resulting dataset and related code are available through the Hugging Face Hub and GitHub.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Data quality testing

DevFeed: [Data quality testing](<https://devfeed.tech/articles/data-quality-testing-11744.md>)

Original publisher: [Read original article](<https://incident.io/blog/data-quality-testing>)

Author: Lambert Le Manh

Published: 2024-09-04T16:30:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [ci](<https://devfeed.tech/topics/ci.md>), [Transactions](<https://devfeed.tech/topics/transactions.md>)

Tags: [ci](<https://devfeed.tech/tags/ci.md>), [data-observability](<https://devfeed.tech/tags/data-observability.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [flaky](<https://devfeed.tech/tags/flaky.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [observability](<https://devfeed.tech/tags/observability.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [testing](<https://devfeed.tech/tags/testing.md>), [transactions](<https://devfeed.tech/tags/transactions.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

The article explains how incident.io uses dbt's native testing features for data quality testing and data observability. It describes integrating validation into production and CI workflows, including custom tests and relationship tests, and addresses flaky failures caused by ingestion and transformation pipelines running at different times.

### Source excerpt

Our data observability workflow uses data quality testing to ensure data meets accuracy, consistency, and reliability standards, enabling confident, data-driven decisions. See how we built it, the common challenges we encountered, and the solutions.

## "The dashboard looks broken!": How should data teams respond to incidents?

DevFeed: ["The dashboard looks broken!": How should data teams respond to incidents?](<https://devfeed.tech/articles/the-dashboard-looks-broken-how-should-data-teams-respond-to-incidents-11829.md>)

Original publisher: [Read original article](<https://incident.io/blog/incident-management-for-data-teams>)

Author: Jack Colsey

Published: 2023-11-21T14:54:00Z

Content type: opinion

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [incident management](<https://devfeed.tech/topics/incident-management.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [observability](<https://devfeed.tech/topics/observability.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Slack](<https://devfeed.tech/topics/slack.md>)

Tags: [dashboards](<https://devfeed.tech/tags/dashboards.md>), [data](<https://devfeed.tech/tags/data.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [observability](<https://devfeed.tech/tags/observability.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack](<https://devfeed.tech/tags/slack.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This article argues that incident management platforms are a good fit for data teams because data pipeline failures are inevitable and require coordinated fixes. It describes how automated testing and observability can identify failing tests, route issues to owners, detect anomalies, and notify tools such as Slack, while incident management supports the subsequent communication and collaboration needed to resolve data incidents.

### Source excerpt

For data teams, the thought of using an incident management tool feels like an odd fit. But if you look deeper, it just makes sense.

## Can Debezium Lose Events?

DevFeed: [Can Debezium Lose Events?](<https://devfeed.tech/articles/can-debezium-lose-events-18805.md>)

Original publisher: [Read original article](<https://www.morling.dev/blog/can-debezium-lose-events/>)

Published: 2023-11-14T14:00:00Z

Content type: article

Language: en

Sources: [Gunnar Morling](<https://devfeed.tech/sources/gunnar-morling.md>)

Topics: [data observability](<https://devfeed.tech/topics/data-observability.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [Database](<https://devfeed.tech/topics/database.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [Amazon RDS](<https://devfeed.tech/topics/amazon-rds.md>)

Tags: [amazon-rds](<https://devfeed.tech/tags/amazon-rds.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [database](<https://devfeed.tech/tags/database.md>), [disk-space](<https://devfeed.tech/tags/disk-space.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [replication](<https://devfeed.tech/tags/replication.md>), [sql-server](<https://devfeed.tech/tags/sql-server.md>)

### AI overview

The article explains that Debezium should generally not miss database change events because its delivery semantics are at-least-once. Event loss can occur when operational issues cause required transaction-log data to be discarded before a connector captures it, such as when retention limits are exceeded during extended downtime. It discusses retention behavior for MySQL, Amazon RDS, SQL Server, and PostgreSQL replication slots.

### Source excerpt

This question came up on the Data Engineering sub-reddit the other day: Can Debezium lose any events? I.e. can there be a situation where a record in a database get inserted, updated, or deleted, but Debezium fails to capture that event from the transaction log and propagate it to downstream consumers?

## Evaluating Equality Predicates with RangeBitmap

DevFeed: [Evaluating Equality Predicates with RangeBitmap](<https://devfeed.tech/articles/evaluating-equality-predicates-25641.md>)

Original publisher: [Read original article](<https://richardstartin.github.io/posts/range-bitmap-equality-queries>)

Author: Richard Startin's Blog

Published: 2022-12-18T00:00:00Z

Content type: article

Language: en

Sources: [Richard Startin's Blog](<https://devfeed.tech/sources/richard-startin-s-blog.md>)

Topics: [Data structures](<https://devfeed.tech/topics/data-structures.md>), [Library](<https://devfeed.tech/topics/library.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>)

Tags: [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [cardinality](<https://devfeed.tech/tags/cardinality.md>), [comparison](<https://devfeed.tech/tags/comparison.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-structure](<https://devfeed.tech/tags/data-structure.md>), [java](<https://devfeed.tech/tags/java.md>), [library](<https://devfeed.tech/tags/library.md>), [memory](<https://devfeed.tech/tags/memory.md>), [pinot](<https://devfeed.tech/tags/pinot.md>), [range](<https://devfeed.tech/tags/range.md>), [roaring](<https://devfeed.tech/tags/roaring.md>), [speed](<https://devfeed.tech/tags/speed.md>)

### AI overview

This article evaluates equality and inequality queries using RangeBitmap in the RoaringBitmap library. It explains how the enhancement can support equality filtering as a compact inverted-index alternative, including as a fallback for Apache Pinot range indexes, and reports faster selection than a stream-based scan in the described example while using less space than some inverted indexes.

### Source excerpt

I have just implemented support for (in)equality queries against a RangeBitmap, a succinct data structure in the RoaringBitmap library which supports range queries. RangeBitmap was designed to support range queries in Apache Pinot (more details here) but this enhancement would allow a range index to be used as a fallback for (in)equality queries in case nothing better is available. Supporting (in)equality queries allows a RangeBitmap to be used as a kind of compact inverted index, trading space for time, capable of supporting high cardinality gracefully. Since RangeBitmap supports memory mapping from files, I think that it could be used for data engineering beyond Apache Pinot.

## How a strategically designed data platform elevates product outcomes

DevFeed: [How a strategically designed data platform elevates product outcomes](<https://devfeed.tech/articles/how-a-strategically-designed-data-platform-elevates-product-outcomes-20030.md>)

Original publisher: [Read original article](<https://technology.doximity.com/articles/how-a-strategically-designed-data-platform-elevates-product-outcomes>)

Author: Doximity

Published: 2022-12-09T16:00:00Z

Content type: article

Language: en

Sources: [Doximity](<https://devfeed.tech/sources/doximity.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Domain-driven design (DDD)](<https://devfeed.tech/topics/domain-driven-design.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>)

Tags: [building](<https://devfeed.tech/tags/building.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [domain-driven-design](<https://devfeed.tech/tags/domain-driven-design.md>), [engineering](<https://devfeed.tech/tags/engineering.md>)

### AI overview

Doximity describes organizing its data function around cross-functional product teams that combine data analysts, data engineers, and product managers. The article contrasts this model with functionally separated data teams and argues that domain-focused ownership better supports data-driven products.

### Source excerpt

If this article piques your interest and you would like an opportunity to do your best data work in a unique and truly data-driven environment, you're in luck: ⚡ we're hiring! ⚡ Building a company where data is fundamental has resulted in many unique challenges--one of which is how to organize our data teams. But before diving into Doximity's data team structure, let's set the stage. Doximity is commonly mentioned as the "LinkedIn for doctors" and although partially accurate it only captures a portion of our story. Today, Doximity assists medical professionals with career navigation, secure peer-to-peer communication, telehealth (Dialer), medical news and research discovery, Locum Tenens opportunities (Curative), and, most recently, clinical scheduling (Amion). I like to say that we are more of a Swiss army knife for medical professionals. Already unique in terms of our offerings and network, which includes over 2 million U.S. healthcare professionals, including over 80% of U.S. physicians, we are also unique in that we are a profoundly data-driven company. Not only do we leverage data for business intelligence and decision support, but data directly drives the majority of our products. Through talking with industry peers, it is my experience that even in today's increasingly data-driven world, many organizations still leverage an older strategy for organizing their data teams. In short, most companies have a data analyst team and a data engineering team. The division is purely functional, and if there's a need for data infrastructure, you will often find it managed by members of the general operations team. Being big proponents of domain-driven design we do not believe it is possible to achieve the best performance for modern high-performance data organizations using this model. Instead, at Doximity, we have multiple cross-functional product teams where each team is responsible for a specific product (or part of a product). These product teams are generally composed

## Algo Hour - Large Scale Data & ML Monitoring with whylogs | Alessya Visnjic

DevFeed: [Algo Hour - Large Scale Data & ML Monitoring with whylogs | Alessya Visnjic](<https://devfeed.tech/articles/algo-hour-large-scale-data-ml-monitoring-with-whylogs-alessya-visnjic-29338.md>)

Original publisher: [Read original article](<https://multithreaded.stitchfix.com/blog/2022/09/29/alessya-algo-hour-announcement/>)

Published: 2022-09-29T09:00:00Z

Content type: article

Language: en

Sources: [Stitch Fix](<https://devfeed.tech/sources/stitch-fix.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-ml](<https://devfeed.tech/tags/data-ml.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

### AI overview

This talk explains how the open-source whylogs library supports end-to-end data quality and monitoring across machine learning pipelines. It covers whylogs' lightweight statistical data collection, language- and platform-agnostic approach, architecture, and application to existing data and ML pipelines.

### Source excerpt

Title: Large Scale Data & ML Monitoring with whylogs Talk Abstract: In the era of microservices, decentralized ML architectures and complex data pipelines, data quality has become a bigger challenge than ever. When data is involved in complex business processes and decisions, bad data can, and will, affect the bottom line. As a result, ensuring data quality across the entire ML pipeline is both costly, and cumbersome while data monitoring is often fragmented and performed ad hoc. An open source library called whylogs is built to address these challenges. It is a lightweight data profiling library that enables end-to-end data monitoring across the entire software stack. The library implements a language and platform agnostic approach to data quality and data monitoring. It's been deployed at massive-scale data environments, on structured and unstructured data modalities, and across a range of points in the ML lifecycle. In this talk, we will provide an overview of the whylogs architecture, including its lightweight statistical data collection approach and we will show how users can apply this library to existing data and ML pipelines. Date and Time: The talk will be held on Tuesday, October 11th at 1:00PM PDT. Recording Info: This talk was recorded live and is viewable below: Speaker Info: Alessya Visnjic is the CEO of WhyLabs, the AI Observability company building tools that power robust and responsible AI deployment. Prior to WhyLabs, Alessya was a CTO-in-residence at the Allen Institute for AI, where she evaluated commercial potential for the latest AI research. Earlier, Alessya spent 9 years at Amazon leading ML initiatives, including forecasting and data science platforms. Alessya is also the founder of Rsqrd AI, a global community of 1,000+ AI practitioners who are committed to making enterprise AI technology responsible.

## Announcing Presto Conference Tokyo 2020

DevFeed: [Announcing Presto Conference Tokyo 2020](<https://devfeed.tech/articles/announcing-presto-conference-tokyo-2020-8655.md>)

Original publisher: [Read original article](<https://trino.io/blog/2020/10/21/announcing-presto-conference-tokyo-2020.html>)

Author: Yuya Ebihara, LINE

Published: 2020-10-21T00:00:00Z

Content type: news

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [data observability](<https://devfeed.tech/topics/data-observability.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [amazon-web-services](<https://devfeed.tech/tags/amazon-web-services.md>), [conference](<https://devfeed.tech/tags/conference.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [developer](<https://devfeed.tech/tags/developer.md>), [event](<https://devfeed.tech/tags/event.md>), [japan](<https://devfeed.tech/tags/japan.md>)

### AI overview

Presto Conference Tokyo 2020 will be held online on November 20, featuring six sessions from Treasure Data, Amazon Web Services Japan, Repro, and LINE, plus open sessions with Presto community members. The event is aimed at people using Presto for data engineering and those interested in adopting it.

### Source excerpt

Last year, Presto Conference Tokyo 2019 was held in Japan with Martin Traverso, Dain Sundstrom and David Phillips, the founders of the Presto Software Foundation. This year, the event changes to be an online only event. Presto Conference Tokyo 2020 is happening on the 20th of November. You can find out details and register right now! The event includes six sessions from Treasure Data, Amazon Web Services Japan, Repro and LINE, as well as open sessions with Martin and Brian Olsen, a Developer Advocate at Starburst Data. This is a valuable opportunity to hear from engineers who are actually using Presto. It has something for those who are using Presto for data engineering and those who don't use Presto yet but are interested in it.

## Why Data Pipelines Should Avoid Monolithic Queue Processors

DevFeed: [Why Data Pipelines Should Avoid Monolithic Queue Processors](<https://devfeed.tech/articles/building-data-pipelines-learning-number-01-die-monolith-die-28152.md>)

Original publisher: [Read original article](<http://fuzzyblog.io/blog/data_pipeline/2020/07/20/building-data-pipelines-die-monolith-die.html>)

Author: Fuzzygroup

Published: 2020-07-20T00:00:00Z

Content type: article

Language: en

Sources: [Scott Johnson](<https://devfeed.tech/sources/scott-johnson.md>)

Topics: [data observability](<https://devfeed.tech/topics/data-observability.md>), [Amazon Simple Queue Service (SQS)](<https://devfeed.tech/topics/amazon-simple-queue-service-sqs.md>), [debugging](<https://devfeed.tech/topics/debugging.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [Tensorflow](<https://devfeed.tech/topics/tensorflow.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [data-pipeline](<https://devfeed.tech/tags/data-pipeline.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [process](<https://devfeed.tech/tags/process.md>), [sqs](<https://devfeed.tech/tags/sqs.md>), [tensorflow](<https://devfeed.tech/tags/tensorflow.md>)

### AI overview

The article describes lessons from building a high-performance, near-real-time data pipeline on AWS with SQS and machine learning components. It argues that a monolithic queue processor can increase cloud costs when GPU-backed processing is applied to work that does not require a GPU, make debugging harder, and limit independent scaling of different processing routines.

### Source excerpt

I am in the process of wrapping up a year long engagement where I: Built a high performance data pipeline Capable of processing all of Twitter in real time / near real time Applied multiple tools to the data at different stages of the pipeline Applied one or more Machine Learning models at different stages of the pipeline Operated on AWS using SQS as the queueing structure This blog post talks about one of the key things I learned in terms of the data pipeline, specifically: **Do Not Build Data Pipelines Around a Monolithic Queue Processor ** Note: By queue processor I mean the bit of software which pulls the data of the queue, operates on it and then puts it back. This lesson may be obvious to some but we took a meandering approach to this problem where we started with the idea of distributed pipeline components, moved to a monolithic approach and then ended up back at a distributed approach. As with a lot of research endeavors, the obvious conclusion wasn't quite so obvious in the throes of the research. Lesson 01: Monolithic Queue Processing Raises Your Costs At the heart of our queue processing were a number of Machine Learning components (python / tensorflow) that really needed a GPU for efficient data processing. The problem here is that when you have a monolithic queue processor, all your processing happens on a box with the GPU whether or not all that processing needs the GPU. When you are using cloud computing, you pay for the GPU whether not not it is being used for a given operation. And since GPU boxes generally cost at least 5x to 6x more than CPU only boxes, well, our monolithic queue processor proved to be an economic disaster. Lesson 02: Monolithic Queue Processing Is Harder to Debug After realizing Lesson 01, I took our monolithic queue processor apart and broke it down into 8 (ultimately 9) individual queue processors. One thing that I quickly found is that debugging the 8 individual queue processors was dramatically easier than debugging the singl