# streaming-data-processing

A data-processing approach that continuously processes data from unbounded streams as it arrives.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Redpanda named a Leader in G2's Fall 2026 Event Stream Processing Reports

DevFeed: [Redpanda named a Leader in G2's Fall 2026 Event Stream Processing Reports](<https://devfeed.tech/articles/redpanda-named-a-leader-in-g2-s-fall-2026-event-stream-processing-reports-12700.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/g2-fall-2026-results>)

Author: Redpanda

Published: 2026-09-01T00:00:00Z

Content type: article

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [stream-processing](<https://devfeed.tech/topics/stream-processing.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Usability](<https://devfeed.tech/topics/usability.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Software](<https://devfeed.tech/topics/software.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [event](<https://devfeed.tech/tags/event.md>), [performance](<https://devfeed.tech/tags/performance.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [reports](<https://devfeed.tech/tags/reports.md>), [software](<https://devfeed.tech/tags/software.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [usability](<https://devfeed.tech/tags/usability.md>)

### AI overview

Redpanda reports that it was recognized in 10 G2 Fall 2026 Event Stream Processing reports, earning first place in eight indexes, including implementation, usability, results, and relationship. The article also highlights Redpanda Shadowing for migration from Confluent and cites a 4.7 out of 5 average rating across 53 G2 reviews.

### Source excerpt

Redpanda ranked #1 in Implementation, Usability, Results, and Relationship Indexes in G2's Fall 2026 Event Stream Processing Reports.

## New in Confluent Intelligence and AI Tools: Making Agents Native to the Stream, Expanded Model Support, New Agent Skills, and Copilot

DevFeed: [New in Confluent Intelligence and AI Tools: Making Agents Native to the Stream, Expanded Model Support, New Agent Skills, and Copilot](<https://devfeed.tech/articles/new-in-confluent-intelligence-and-ai-tools-making-agents-native-to-the-stream-expanded-model-support-new-agent-skills-and-copilot-11547.md>)

Original publisher: [Read original article](<https://www.confluent.io/blog/2026-q3-confluent-intelligence-ai-update/>)

Author: Confluent Staff

Published: 2026-08-18T14:00:10Z

Content type: article

Language: en

Sources: [Confluent: Data in motion](<https://devfeed.tech/sources/confluent-data-in-motion.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [apache-flink](<https://devfeed.tech/topics/apache-flink.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Agent Skills](<https://devfeed.tech/topics/agent-skills.md>), [MCP](<https://devfeed.tech/topics/mcp.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Google](<https://devfeed.tech/topics/google.md>), [ibm](<https://devfeed.tech/topics/ibm.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-ready-data](<https://devfeed.tech/tags/ai-ready-data.md>), [ai-tools](<https://devfeed.tech/tags/ai-tools.md>), [apache-flink](<https://devfeed.tech/tags/apache-flink.md>), [confluent](<https://devfeed.tech/tags/confluent.md>), [confluent-cloud](<https://devfeed.tech/tags/confluent-cloud.md>), [google](<https://devfeed.tech/tags/google.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [skills](<https://devfeed.tech/tags/skills.md>), [streaming-data-processing](<https://devfeed.tech/tags/streaming-data-processing.md>), [support](<https://devfeed.tech/tags/support.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

Confluent announces updates to Confluent Intelligence and related AI tools for building production AI systems on Apache Kafka and Apache Flink. The release adds expanded time-series model support, a generally available Real-Time Context Engine, updates to the fully managed MCP Server, Agent Skills for AI coding assistants, and Confluent Copilot. The features are intended to provide agents and applications with fresh business context and governed access to live data and infrastructure.

### Source excerpt

Explore new AI features and AI tools: support for IBM Granite Time Series models and TimesFM models (EA), enhanced Real-Time Context Engine experience, new Agent Skills, Confluent Copilot

## Real-time streaming for the agentic era with NVIDIA

DevFeed: [Real-time streaming for the agentic era with NVIDIA](<https://devfeed.tech/articles/real-time-streaming-for-the-agentic-era-with-nvidia-12719.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/nvidia-ai-ecosystem>)

Author: Melissa Czapiga

Published: 2026-06-01T00:00:00Z

Content type: article

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Vera CPU](<https://devfeed.tech/topics/vera-cpu.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [agent observability](<https://devfeed.tech/topics/agent-observability.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [performance](<https://devfeed.tech/tags/performance.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

Redpanda describes its collaboration with NVIDIA to run streaming workloads on NVIDIA Vera CPUs, reporting 5.5x lower latency for AI agents in mission-critical environments. The article discusses testing on workloads processing billions of messages daily, real-time data infrastructure, and potential benefits for enterprise AI and agentic applications.

### Source excerpt

NVIDIA Vera launches today with Redpanda as part of the ecosystem, delivering 5.5x lower latencies for agents running in mission-critical environments.

## Event-Driven Architecture with Apache Kafka

DevFeed: [Event-Driven Architecture with Apache Kafka](<https://devfeed.tech/articles/mastering-event-driven-architecture-with-apache-kafka-39559.md>)

Original publisher: [Read original article](<https://ankit-rana.com/logs/07-kafka-event-driven-architecture/>)

Author: hello@ankit-rana.com

Published: 2026-03-16T00:00:00Z

Content type: tutorial

Language: en

Sources: [Ankit Rana | Mechanical Sympathy](<https://devfeed.tech/sources/ankit-rana-mechanical-sympathy.md>)

Topics: [event driven](<https://devfeed.tech/topics/event-driven.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Apache-Kafka](<https://devfeed.tech/topics/apache-kafka.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [apache-kafka](<https://devfeed.tech/tags/apache-kafka.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [asynchronous](<https://devfeed.tech/tags/asynchronous.md>), [consumer](<https://devfeed.tech/tags/consumer.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [event-driven-architecture](<https://devfeed.tech/tags/event-driven-architecture.md>), [event-sourcing](<https://devfeed.tech/tags/event-sourcing.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [partitions](<https://devfeed.tech/tags/partitions.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [schemas](<https://devfeed.tech/tags/schemas.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>)

### AI overview

This tutorial explains event-driven architecture and how Apache Kafka supports asynchronous, real-time data processing. It covers producers, consumers, immutable events, event sourcing, scalability, resilience, stream-processing pipelines, and challenges such as ordering, debugging, and eventual consistency.

### Source excerpt

Event-driven architecture replaces synchronous point-to-point calls with immutable events on a durable log, so producers and consumers scale and fail independently. Kafka provides that log: topics sharded into ordered append-only partitions, replicated across brokers, with consumer groups sharing partitions to scale read throughput.

## Why Flink may be unnecessarily complex for most streaming data processing users

DevFeed: [Why Flink may be unnecessarily complex for most streaming data processing users](<https://devfeed.tech/articles/flink-s-95-problem-18498.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/flink-is-95-problem>)

Author: Javi Santana

Published: 2025-10-21T00:00:00Z

Content type: opinion

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [flink](<https://devfeed.tech/tags/flink.md>), [scalable-analytics-architecture](<https://devfeed.tech/tags/scalable-analytics-architecture.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [streaming-data-processing](<https://devfeed.tech/tags/streaming-data-processing.md>)

### AI overview

The article argues that Flink's complexity may make it unnecessary for most people who need streaming data processing.

### Source excerpt

Flink might sound like the holy grail of streaming data processing, but for 95% of us, it's just a complex headache we don't need.

## Introducing multi-language dynamic plugins for Redpanda Connect

DevFeed: [Introducing multi-language dynamic plugins for Redpanda Connect](<https://devfeed.tech/articles/introducing-multi-language-dynamic-plugins-for-redpanda-connect-12717.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/multi-language-redpanda-connect-plugins>)

Author: James Kinley

Published: 2025-06-17T00:00:00Z

Content type: release

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [gRPC](<https://devfeed.tech/topics/grpc.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [Go Language](<https://devfeed.tech/topics/go-language.md>), [Python](<https://devfeed.tech/topics/python.md>), [Processes](<https://devfeed.tech/topics/processes.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Unix](<https://devfeed.tech/topics/unix.md>)

Tags: [ai-ml-capabilities-in-streaming-data](<https://devfeed.tech/tags/ai-ml-capabilities-in-streaming-data.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [creating-plugins-in-golang-and-python](<https://devfeed.tech/tags/creating-plugins-in-golang-and-python.md>), [dynamic-vs-compiled-plugins](<https://devfeed.tech/tags/dynamic-vs-compiled-plugins.md>), [framework](<https://devfeed.tech/tags/framework.md>), [go](<https://devfeed.tech/tags/go.md>), [grpc-plugin-system](<https://devfeed.tech/tags/grpc-plugin-system.md>), [integration](<https://devfeed.tech/tags/integration.md>), [interfaces](<https://devfeed.tech/tags/interfaces.md>), [ipc](<https://devfeed.tech/tags/ipc.md>), [language-agnostic-plugin-system](<https://devfeed.tech/tags/language-agnostic-plugin-system.md>), [multi-language-plugin-development](<https://devfeed.tech/tags/multi-language-plugin-development.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [plugin-development](<https://devfeed.tech/tags/plugin-development.md>), [plugins](<https://devfeed.tech/tags/plugins.md>), [process](<https://devfeed.tech/tags/process.md>), [processes](<https://devfeed.tech/tags/processes.md>), [product](<https://devfeed.tech/tags/product.md>), [protocol](<https://devfeed.tech/tags/protocol.md>), [python](<https://devfeed.tech/tags/python.md>), [python-sdk-for-streaming-data](<https://devfeed.tech/tags/python-sdk-for-streaming-data.md>), [redpanda-connect](<https://devfeed.tech/tags/redpanda-connect.md>), [redpanda-connect-dynamic-plugins](<https://devfeed.tech/tags/redpanda-connect-dynamic-plugins.md>), [redpanda-streaming-infrastructure](<https://devfeed.tech/tags/redpanda-streaming-infrastructure.md>), [runtime-loaded-plugins-grpc](<https://devfeed.tech/tags/runtime-loaded-plugins-grpc.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [sdks](<https://devfeed.tech/tags/sdks.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming-data-pipeline-plugins](<https://devfeed.tech/tags/streaming-data-pipeline-plugins.md>)

### AI overview

Redpanda Connect introduces Apache 2.0-licensed dynamic plugins in Beta with version 4.56.0. Plugins can be loaded at runtime as external executables that communicate with the main process through gRPC over Unix sockets, enabling plugin development in Go and Python. Official SDKs support both languages, while native Go plugins remain the preferred option for performance-critical workloads.

### Source excerpt

Redpanda Connect dynamic plugins framework allows you to create and load plugins at runtime, opening up a world of new integration possibilities beyond Go.

## Streaming and Batch Processing Are Complementary; Pull Versus Push Is the Key Distinction

DevFeed: [Streaming and Batch Processing Are Complementary; Pull Versus Push Is the Key Distinction](<https://devfeed.tech/articles/streaming-vs-batch-is-a-wrong-dichotomy-and-i-think-it-s-confusing-18872.md>)

Original publisher: [Read original article](<https://www.morling.dev/blog/streaming-vs-batch-wrong-dichotomy/>)

Published: 2025-05-14T08:10:00Z

Content type: opinion

Language: en

Sources: [Gunnar Morling](<https://devfeed.tech/sources/gunnar-morling.md>)

Topics: [Streaming](<https://devfeed.tech/topics/streaming.md>), [stream-processing](<https://devfeed.tech/topics/stream-processing.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [systems](<https://devfeed.tech/topics/systems.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [kafka](<https://devfeed.tech/tags/kafka.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [performance](<https://devfeed.tech/tags/performance.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

The article argues that streaming systems commonly use batching to improve throughput, so streaming and batch processing are not opposites. It proposes pull versus push semantics as the more meaningful distinction and explains that push-based streaming can provide timely updates, while adding complexity around state, joins, and out-of-order data.

### Source excerpt

Often times, "Stream vs. Batch" is discussed as if it's one or the other, but to me this does not make that much sense really.

## Snyk Helps Secure the Golang Bento Project

DevFeed: [Snyk Helps Secure the Golang Bento Project](<https://devfeed.tech/articles/snyk-helps-secure-the-golang-bento-project-8139.md>)

Original publisher: [Read original article](<https://snyk.io/blog/snyk-helps-secure-the-golang-bento-project/>)

Author: Phill Garrett

Published: 2025-03-12T04:00:00Z

Content type: article

Language: en

Sources: [Blog RSS Feed | Snyk](<https://devfeed.tech/sources/blog-rss-feed-snyk.md>)

Topics: [Vulnerabilities](<https://devfeed.tech/topics/vulnerabilities.md>), [open-source-security](<https://devfeed.tech/topics/open-source-security.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>)

Tags: [blog](<https://devfeed.tech/tags/blog.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [container-security](<https://devfeed.tech/tags/container-security.md>), [developer](<https://devfeed.tech/tags/developer.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [go](<https://devfeed.tech/tags/go.md>), [golang](<https://devfeed.tech/tags/golang.md>), [interest](<https://devfeed.tech/tags/interest.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [maintainers](<https://devfeed.tech/tags/maintainers.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-security](<https://devfeed.tech/tags/open-source-security.md>), [pull-request](<https://devfeed.tech/tags/pull-request.md>), [releases](<https://devfeed.tech/tags/releases.md>), [security](<https://devfeed.tech/tags/security.md>), [snyk-open-source](<https://devfeed.tech/tags/snyk-open-source.md>), [streaming-data-processing](<https://devfeed.tech/tags/streaming-data-processing.md>), [vulnerability](<https://devfeed.tech/tags/vulnerability.md>)

### AI overview

Snyk describes contributing fixes to the open-source Go project Bento after finding a denial-of-service vulnerability in its SSH dependency. The post also outlines Bento's streaming-data role and Snyk's program for open-source maintainers.

### Source excerpt

Discover how Snyk helps secure the open source Golang project Bento by contributing vulnerability fixes and leveraging AI-powered tools. Learn about our efforts to enhance Bento's security and support open source maintainers through the Snyk Secure Developer Program.

## Latest Chainguard Images: FIPS, Harbor stack, Apache, and more!

DevFeed: [Latest Chainguard Images: FIPS, Harbor stack, Apache, and more!](<https://devfeed.tech/articles/latest-chainguard-images-fips-harbor-stack-apache-and-more-13139.md>)

Original publisher: [Read original article](<https://www.chainguard.dev/unchained/latest-chainguard-images-fips-harbor-stack-apache-and-more>)

Published: 2024-06-20T00:00:00Z

Content type: article

Language: en

Sources: [Chainguard: Unchained](<https://devfeed.tech/sources/chainguard-unchained.md>)

Topics: [chainguard images](<https://devfeed.tech/topics/chainguard-images.md>), [container images](<https://devfeed.tech/topics/container-images.md>), [Security](<https://devfeed.tech/topics/security.md>), [open-source-security](<https://devfeed.tech/topics/open-source-security.md>), [supply-chain-security](<https://devfeed.tech/topics/supply-chain-security.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [apache](<https://devfeed.tech/tags/apache.md>), [apache-zookeeper](<https://devfeed.tech/tags/apache-zookeeper.md>), [chainguard](<https://devfeed.tech/tags/chainguard.md>), [chainguard-images](<https://devfeed.tech/tags/chainguard-images.md>), [container-image](<https://devfeed.tech/tags/container-image.md>), [container-images](<https://devfeed.tech/tags/container-images.md>), [fips](<https://devfeed.tech/tags/fips.md>), [harbor](<https://devfeed.tech/tags/harbor.md>), [jitsu](<https://devfeed.tech/tags/jitsu.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [security](<https://devfeed.tech/tags/security.md>), [software-supply-chain](<https://devfeed.tech/tags/software-supply-chain.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [vulnerabilities](<https://devfeed.tech/tags/vulnerabilities.md>)

### AI overview

Chainguard's monthly roundup announces 57 new hardened, minimal container images, including expanded Apache, Jitsu, and Harbor stacks. More than half of the releases are FIPS compliant, bringing the total to nearly 300 FIPS-validated Images. The article highlights security, reduced CVE exposure, software supply chain protection, and images for Apache Airflow, Superset, Zookeeper, and Nifi.

### Source excerpt

Upgrade your software security with the latest Chainguard Images, featuring FIPS, Harbor, Apache, and more. Streamline your development and enhance protection.

## When Change Data Capture Can Break Application Encapsulation

DevFeed: [When Change Data Capture Can Break Application Encapsulation](<https://devfeed.tech/articles/change-data-capture-breaks-encapsulation-does-it-though-18808.md>)

Original publisher: [Read original article](<https://www.morling.dev/blog/change-data-capture-breaks-encapsulation-does-it-though/>)

Published: 2023-11-21T00:00:00Z

Content type: article

Language: en

Sources: [Gunnar Morling](<https://devfeed.tech/sources/gunnar-morling.md>)

Topics: [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [apache-kafka](<https://devfeed.tech/tags/apache-kafka.md>), [data](<https://devfeed.tech/tags/data.md>), [debezium](<https://devfeed.tech/tags/debezium.md>), [schema](<https://devfeed.tech/tags/schema.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

This article examines whether exposing database change event feeds through Change Data Capture (CDC) breaks application encapsulation. It explains that CDC can expose internal table models as event APIs and create schema and downstream-consumer risks, then discusses ways to address those risks.

### Source excerpt

Table of Contents CDC--A Quick Primer Does CDC Break Encapsulation? Entering Data Contracts Implementation Approaches For Data Contracts The Outbox Pattern Stream Processing Streaming Data Contracts--Beyond the Basics Handling Schema Changes Summary This post originally appeared on the Decodable blog. All rights reserved. Having worked on Debezium--an open-source platform for Change Data Capture (CDC)--for several years, one concern I've heard repeatedly is this: aren't you breaking the encapsulation of your application when you expose change event feeds directly from your database? After all, CDC exposes your internal persistent data model to the outside world, which may have unintended consequences, e.g. in terms of data exposure but also when it comes to changes to the schema of your data, which may break downstream consumers.

## CDC Use Cases: 7 Ways to Put CDC to Work

DevFeed: [CDC Use Cases: 7 Ways to Put CDC to Work](<https://devfeed.tech/articles/cdc-use-cases-7-ways-to-put-cdc-to-work-18807.md>)

Original publisher: [Read original article](<https://www.morling.dev/blog/cdc-use-cases/>)

Published: 2023-11-02T00:00:00Z

Content type: article

Language: en

Sources: [Gunnar Morling](<https://devfeed.tech/sources/gunnar-morling.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [apache-flink](<https://devfeed.tech/topics/apache-flink.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [Microservice](<https://devfeed.tech/topics/microservice.md>), [migration](<https://devfeed.tech/topics/migration.md>)

Tags: [apache-flink](<https://devfeed.tech/tags/apache-flink.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [debezium](<https://devfeed.tech/tags/debezium.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [migration](<https://devfeed.tech/tags/migration.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

This article explains change data capture (CDC), focusing on log-based CDC and its low-latency, resource-efficient handling of database changes. It introduces Debezium and describes how CDC works with data streaming and stream-processing tools such as Apache Kafka and Apache Flink, before outlining seven common use cases.

### Source excerpt

Table of Contents What is CDC? CDC Tools Analytics Data Platforms Application Caches Full-Text Search Audit Logs Continuous Queries Microservices Data Exchange Monolith-to-Microservices Migration Summary This post originally appeared on the Decodable blog. All rights reserved. Change Data Capture (CDC) is a powerful tool in data engineering and has seen a tremendous uptake in organizations of all kinds over the last few years. This is because it enables the tight integration of transactional databases into many other systems in your business at a very low latency.

## Tinybird: A ksqlDB alternative for stateful stream processing

DevFeed: [Tinybird: A ksqlDB alternative for stateful stream processing](<https://devfeed.tech/articles/tinybird-a-ksqldb-alternative-for-stateful-stream-processing-18552.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/ksqldb-alternative>)

Author: Alasdair Brown

Published: 2023-07-27T00:00:00Z

Content type: comparison

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [stream-processing](<https://devfeed.tech/topics/stream-processing.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>)

Tags: [compare](<https://devfeed.tech/tags/compare.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [scalable-analytics-architecture](<https://devfeed.tech/tags/scalable-analytics-architecture.md>), [stateful](<https://devfeed.tech/tags/stateful.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

This comparison presents Tinybird as an alternative to ksqlDB for stateful stream processing and streaming SQL, emphasizing an approach intended to avoid the complexity associated with Kafka Streams.

### Source excerpt

Looking for a ksqlDB alternative? Tinybird gives you streaming SQL without the Kafka Streams complexity. Compare the approaches.

## Why I Joined Decodable

DevFeed: [Why I Joined Decodable](<https://devfeed.tech/articles/why-i-joined-decodable-18891.md>)

Original publisher: [Read original article](<https://www.morling.dev/blog/why-i-joined-decodable/>)

Published: 2022-11-03T14:00:00Z

Content type: opinion

Language: en

Sources: [Gunnar Morling](<https://devfeed.tech/sources/gunnar-morling.md>)

Topics: [stream-processing](<https://devfeed.tech/topics/stream-processing.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Apache Pulsar](<https://devfeed.tech/topics/pulsar.md>), [amazon](<https://devfeed.tech/topics/amazon.md>)

Tags: [amazon](<https://devfeed.tech/tags/amazon.md>), [apache](<https://devfeed.tech/tags/apache.md>), [apache-kafka](<https://devfeed.tech/tags/apache-kafka.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [news](<https://devfeed.tech/tags/news.md>), [platform](<https://devfeed.tech/tags/platform.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>)

### AI overview

A software engineer explains why he joined Decodable after several years at Red Hat, citing the field of real-time stream processing, the startup environment, and the team. He describes his experience with Debezium and his interest in helping users implement end-to-end streaming use cases.

### Source excerpt

Table of Contents The Space: Real-Time Stream Processing The Environment: A Start-up The Team: One of a Kind Outlook It's my first week as a software engineer at Decodable, a start-up building a serverless real-time data platform! When I shared this news on social media yesterday, folks were not only super supportive and excited for me (thank you so much for all the nice words and wishes!), but some also asked about the reasons behind my decision for switching jobs and going to a start-up, after having worked for Red Hat for the last few years.

## Scaling Shopify's BFCM Live Map: An Apache Flink Redesign

DevFeed: [Scaling Shopify's BFCM Live Map: An Apache Flink Redesign](<https://devfeed.tech/articles/scaling-shopify-s-bfcm-live-map-an-apache-flink-redesign-1306.md>)

Original publisher: [Read original article](<https://shopify.engineering/bfcm-live-map-2021-apache-flink-redesign>)

Author: Berkay Antmen

Published: 2021-12-10T19:00:00Z

Content type: article

Language: en

Sources: [Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering.md>), [Shopify Engineering - Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering-shopify-engineering.md>)

Topics: [Shopify](<https://devfeed.tech/topics/shopify.md>), [apache-flink](<https://devfeed.tech/topics/apache-flink.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>)

Tags: [apache-flink](<https://devfeed.tech/tags/apache-flink.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [shopify](<https://devfeed.tech/tags/shopify.md>), [stateful](<https://devfeed.tech/tags/stateful.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Shopify Data Platform Engineering redesigned the infrastructure behind the BFCM live map with Apache Flink to support more than 1.7 million merchants, richer insights, higher data volume, and higher uptime without manual intervention.

### Source excerpt

A deep dive into how Shopify Data revamped the data infrastructure powering our BFCM live map using Apache Flink.

## Emerging Kafka Trends in Connectors, Self-Service Environments, and Stream Processing

DevFeed: [Emerging Kafka Trends in Connectors, Self-Service Environments, and Stream Processing](<https://devfeed.tech/articles/three-plus-some-lovely-kafka-trends-18882.md>)

Original publisher: [Read original article](<https://www.morling.dev/blog/three-plus-some-lovely-kafka-trends/>)

Published: 2021-05-28T08:30:00Z

Content type: opinion

Language: en

Sources: [Gunnar Morling](<https://devfeed.tech/sources/gunnar-morling.md>)

Topics: [Kafka](<https://devfeed.tech/topics/kafka.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [kafka](<https://devfeed.tech/tags/kafka.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [review](<https://devfeed.tech/tags/review.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>)

### AI overview

The article identifies recurring themes in Kafka community activity: the expanding connector ecosystem and its operational maturity, self-service Kafka environments, and growing adoption of stream processing. It also mentions related concerns including access control, schema management, data lineage, quality, compliance, privacy, and observability.

### Source excerpt

Table of Contents Cambrian Explosion of Connectors Democratization of Data Pipelines Stream Processing for Everyone Honorable Mentions Over the course of the last few months, I've had the pleasure to serve on the Kafka Summit program committee and review several hundred session abstracts for the three Summits happening this year (Europe, APAC, Americas). That's not only a big honour, but also a unique opportunity to learn what excites people currently in the Kafka eco-system (and yes, it's a fair amount of work, too ;). While voting on the proposals, and also generally aspiring to stay informed of what's going on in the Kafka community at large, I noticed a few repeating themes and topics which I thought would be interesting to share (without touching on any specific talks of course). At first I meant to put this out via a Twitter thread, but then it became a bit too long for that, so I decided to write this quick blog post instead. Here it goes!

## From pipeline to beyond

DevFeed: [From pipeline to beyond](<https://devfeed.tech/articles/from-pipeline-to-beyond-19833.md>)

Original publisher: [Read original article](<https://tech.gc.com/from-pipeline-to-beyond/>)

Author: GameChanger

Published: 2021-05-05T09:00:18Z

Content type: article

Language: en

Sources: [GameChanger](<https://devfeed.tech/sources/gamechanger.md>)

Topics: [Kafka](<https://devfeed.tech/topics/kafka.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [data](<https://devfeed.tech/topics/data.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Back end](<https://devfeed.tech/topics/backend.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [backend](<https://devfeed.tech/tags/backend.md>), [consul](<https://devfeed.tech/tags/consul.md>), [data](<https://devfeed.tech/tags/data.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [redshift](<https://devfeed.tech/tags/redshift.md>), [s3](<https://devfeed.tech/tags/s3.md>), [terraform](<https://devfeed.tech/tags/terraform.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article examines options for moving data out of Kafka into a data warehouse and archive. It discusses off-the-shelf tools such as Kafka Connect, Secor, and Gobblin, the limitations encountered, and the development of a custom solution. The requirements include preserving Avro data and schemas, writing to S3, and partitioning data by processing time or event time.

### Source excerpt

An overview of off-the-shelf solutions for moving data out of Kafka, problems we had making those systems work, and how we wrote our own solution and stood it up for those in a similar situation. You cannot know everything a system will be used for when you start: it is only at the end of its life you can have such certainty. [1] Many moons ago I wrote about our design for upgrading our data pipeline, which lightly touched on how we'd move data out of our pipeline (Kafka) to downstream systems, namely our data warehouse and data archive. At that time we hadn't really been able to dive into focusing on getting the data out of Kafka, because getting data in to Kafka is often much more custom and complex, and we thought we'd be able to use an off the shelf solution like Kafka Connect to move data out, don't even worry about it. We were, uh -- we were wrong. Let me take you on our journey, in case you're on this journey too. The problem space Solution 1: Kafka Connect Solution 2: Secor or Gobblin Solution 3: we'll do this ourselves Tangent: naming things How you can do this yourselves, code edition How you can do this yourselves, infrastructure and metrics edition Takeaways The problem space Programmers are not to be measured by their ingenuity and their logic but by the completeness of their case analysis. [4] At a high level, the problem we needed a solution for was as follows: Data enters the data pipeline from numerous backend systems. This crossover point is producers into the pipeline, which we'd already implemented. Data from the data pipeline needs to move into the data warehouse and the data archive. This crossover point would be a consumer on the pipeline. Ideally we'd like the same consumer for both needs that we can simply configure differently. We want to preserve our data's Avro format along side its schemas. This would allow every system that interacts with the data to use the same language. We want to write our data to S3. data warehouse: This will be our

## Apache Beam for Search: Getting Started by Hacking Time

DevFeed: [Apache Beam for Search: Getting Started by Hacking Time](<https://devfeed.tech/articles/apache-beam-for-search-getting-started-by-hacking-time-1294.md>)

Original publisher: [Read original article](<https://shopify.engineering/apache-beam-for-search-getting-started-by-hacking-time>)

Author: Doug Turnbull

Published: 2021-01-08T15:00:01Z

Content type: tutorial

Language: en

Sources: [Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering.md>), [Shopify Engineering - Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering-shopify-engineering.md>)

Topics: [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [apache](<https://devfeed.tech/tags/apache.md>), [apis](<https://devfeed.tech/tags/apis.md>), [batch](<https://devfeed.tech/tags/batch.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [event](<https://devfeed.tech/tags/event.md>), [events](<https://devfeed.tech/tags/events.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [search](<https://devfeed.tech/tags/search.md>), [shopify](<https://devfeed.tech/tags/shopify.md>), [spark](<https://devfeed.tech/tags/spark.md>), [statistics](<https://devfeed.tech/tags/statistics.md>), [storage](<https://devfeed.tech/tags/storage.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming-data-processing](<https://devfeed.tech/tags/streaming-data-processing.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

An introduction to using Apache Beam for search-related streaming data processing. It explains how unified batch and streaming workflows can process clickstream data for real-time relevance tuning, and introduces event time, delayed events, and out-of-order data as core challenges.

### Source excerpt

To create relevant search, processing clickstream data is key: you frequently want to promote search results that are being clicked on and purchased, and demote those things users don't love. Typically search systems think of processing clickstream data as a batch job run over historical data, perhaps using a system like Spark.

## Bullet Updates - Windowing, Apache Pulsar PubSub, Configuration-based Data Ingestion, and More

DevFeed: [Bullet Updates - Windowing, Apache Pulsar PubSub, Configuration-based Data Ingestion, and More](<https://devfeed.tech/articles/bullet-updates-windowing-apache-pulsar-pubsub-configuration-based-data-ingestion-and-more-20497.md>)

Original publisher: [Read original article](<https://yahooeng.tumblr.com/post/183315480351>)

Author: rosaliebeevm-blog

Published: 2019-03-08T17:12:50Z

Content type: release

Language: en

Sources: [Yahoo](<https://devfeed.tech/sources/yahoo.md>)

Topics: [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [stream-processing](<https://devfeed.tech/topics/stream-processing.md>), [Query (disambiguation)](<https://devfeed.tech/topics/query.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>)

Tags: [announcements](<https://devfeed.tech/tags/announcements.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [code](<https://devfeed.tech/tags/code.md>), [java](<https://devfeed.tech/tags/java.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [releases](<https://devfeed.tech/tags/releases.md>), [spark](<https://devfeed.tech/tags/spark.md>), [stream](<https://devfeed.tech/tags/stream.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [updates](<https://devfeed.tech/tags/updates.md>), [yahoo-engineering](<https://devfeed.tech/tags/yahoo-engineering.md>)

### AI overview

This update describes new windowing support in Bullet, an open-source query system for data flowing through streaming systems. It explains time- and record-based windows, including tumbling, sliding, and hopping window patterns, for returning intermediate query results.

### Source excerpt

yahoodevelopers: By Akshay Sarma, Principal Engineer, Verizon Media & Brian Xiao, Software Engineer, Verizon Media This is the first of an ongoing series of blog posts sharing releases and announcements for Bullet, an open-sourced lightweight, scalable, pluggable, multi-tenant query system. Bullet allows you to query any data flowing through a streaming system without having to store it first through its UI or API. The queries are injected into the running system and have minimal overhead. Running hundreds of queries generally fit into the overhead of just reading the streaming data. Bullet requires running an instance of its backend on your data. This backend runs on common stream processing frameworks (Storm and Spark Streaming currently supported). The data on which Bullet sits determines what it is used for. For example, our team runs an instance of Bullet on user engagement data (~1M events/sec) to let developers find their own events to validate their code that produces this data. We also use this instance to interactively explore data, throw up quick dashboards to monitor live releases, count unique users, debug issues, and more. Since open sourcing Bullet in 2017, we've been hard at work adding many new features! We'll highlight some of these here and continue sharing update posts for future releases. Windowing Bullet used to operate in a request-response fashion - you would submit a query and wait for the query to meet its termination conditions (usually duration) before receiving results. For short-lived queries, say, a few seconds, this was fine. But as we started fielding more interactive and iterative queries, waiting even a minute for results became too cumbersome. Enter windowing! Bullet now supports time and record-based windowing. With time windowing, you can break up your query into chunks of time over its duration and retrieve results for each chunk. For example, you can calculate the average of a field, and stream back results every second: In th

## One-box stream processing with CSP

DevFeed: [One-box stream processing with CSP](<https://devfeed.tech/articles/one-box-stream-processing-with-csp-32160.md>)

Original publisher: [Read original article](<https://adambard.com/blog/stream-processing-core-async/>)

Published: 2018-02-18T00:00:00Z

Content type: tutorial

Language: en

Sources: [Adam Bard](<https://devfeed.tech/sources/adam-bard.md>)

Topics: [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Clojure](<https://devfeed.tech/topics/clojure.md>), [stream-processing](<https://devfeed.tech/topics/stream-processing.md>), [Concurrent Programming](<https://devfeed.tech/topics/concurrent-programming.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>)

Tags: [clojure](<https://devfeed.tech/tags/clojure.md>), [concurrent](<https://devfeed.tech/tags/concurrent.md>), [core-async](<https://devfeed.tech/tags/core-async.md>), [coroutine](<https://devfeed.tech/tags/coroutine.md>), [distributed-computing](<https://devfeed.tech/tags/distributed-computing.md>), [distributed-system](<https://devfeed.tech/tags/distributed-system.md>), [flink](<https://devfeed.tech/tags/flink.md>), [spark](<https://devfeed.tech/tags/spark.md>), [stream](<https://devfeed.tech/tags/stream.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>)

### AI overview

This article presents a small-scale stream-processing architecture in Clojure using core.async and component libraries. It describes modular components connected by asynchronous queues and explains how stream-processing design principles can support reusable, loosely coupled systems without requiring a distributed stream processor.

### Source excerpt

If you're like me (that is, employed by an ad tech company), stream processing is usually associated with frameworks like Storm, Flink, Spark Streaming, and other such solutions. However, a lot of real-life software can be described as stream processing - data comes in one end, is transformed or aggregated, and goes somewhere else. Many of these workloads don't justify the overhead of a stream processor, but that doesn't mean they can't benefit from some of the lessons of stream processing systems.

## Real-time Big Data at Target

DevFeed: [Real-time Big Data at Target](<https://devfeed.tech/articles/real-time-big-data-at-target-20394.md>)

Original publisher: [Read original article](<https://target.github.io/analytics/big-data-storm>)

Author: Target Brands, Inc

Published: 2015-11-11T06:00:00Z

Content type: article

Language: en

Sources: [Target](<https://devfeed.tech/sources/target.md>)

Topics: [big-data](<https://devfeed.tech/topics/big-data.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [stream-processing](<https://devfeed.tech/topics/stream-processing.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [apache-kafka](<https://devfeed.tech/tags/apache-kafka.md>), [batch](<https://devfeed.tech/tags/batch.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [data](<https://devfeed.tech/tags/data.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [storm](<https://devfeed.tech/tags/storm.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Target's Big Data platform team describes its effort to move data into Hadoop in real time using flexible open source components. The article outlines requirements for resilience, message-format support, usability, and low latency, then evaluates Apache Storm and identifies recovery and batching-related latency issues in testing.

### Source excerpt

An enterprise as large as Target generates a lot of data and on my Big Data platform team we want to make it as easy as possible for our users to get it into Hadoop in real-time. I want to discuss how we are starting to approach this problem, what we've done so far, and what is still to come. Our requirements We wanted to build a system with flexible open source components. Our experience with proprietary products on Hadoop is that they tend to be inflexible and only work with a narrow set of use cases. That can also be true of open source products, but we have found them to be easier to adapt to our needs. Indeed, that ended up being the case as we went through this journey and made several contributions to Apache projects. More specifically, we wanted to create a system that: was highly resilient to the failure of any individual component supported a variety of message formats delivered data to Hadoop with very low latency made the data immedaiately usable streamed data from Apache Kafka sources A Streaming Framework There are many excellent comparisons of streaming frameworks available, and I won't attempt to recreate them here. The first criteria we considered was what tooling was needed to monitor and administer the streaming framework, with a strong preference to use our primary Hadoop administration tool, Apache Ambari. Apache Storm fit that bill and was also a proven solution for stream processing. If Storm could meet our other requirements, it would be our first choice. To test its resiliency we ran a simple scenario: start a data stream into Hadoop, disable HDFS, and then reenable it. Streaming would obviously fail while HDFS was disabled, but we needed the system to recover gracefully when HDFS came back online. Unfortunately our first test of this scenario left our Storm topology in an unrecoverable state, which required a manual restart. That's not something we could live with. We also needed very fine control over the latency of arriving data. In gener