# Data Infrastructure

Published articles for Data Infrastructure.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Seagate and WD AI Storage Research Finds Enterprises Rank Storage Above Compute as the AI Bottleneck

DevFeed: [Seagate and WD AI Storage Research Finds Enterprises Rank Storage Above Compute as the AI Bottleneck](<https://devfeed.tech/articles/seagate-and-wd-ai-storage-research-finds-enterprises-rank-storage-above-compute-as-the-ai-bottleneck-26756.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/seagate-and-wd-ai-storage-research-finds-enterprises-rank-storage-above-compute-as-the-ai-bottleneck>)

Author: Lyle Smith

Published: 2026-09-15T17:23:54Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [idc](<https://devfeed.tech/topics/idc.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Retrieval-Augmented Generation](<https://devfeed.tech/topics/retrieval-augmented-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [genai](<https://devfeed.tech/tags/genai.md>), [hdd](<https://devfeed.tech/tags/hdd.md>), [idc](<https://devfeed.tech/tags/idc.md>), [inference](<https://devfeed.tech/tags/inference.md>), [reports](<https://devfeed.tech/tags/reports.md>), [research](<https://devfeed.tech/tags/research.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [storage](<https://devfeed.tech/tags/storage.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>)

### AI overview

Seagate and WD published separate studies indicating that AI is increasing enterprise storage requirements and extending data retention. Although their headline percentages differ because they asked different questions, both reports point to storage becoming a larger part of AI infrastructure planning alongside growing archive and retrieval needs.

### Source excerpt

Seagate and WD published separate AI storage studies within days of each other; the headline numbers: Seagate says 99% of enterprises expect AI to increase their storage requirements over the next three years, while WD's IDC research puts the comparable figure at 74%. Read the fine print, and both reports land in the same directional The post Seagate and WD AI Storage Research Finds Enterprises Rank Storage Above Compute as the AI Bottleneck appeared first on StorageReview.com.

## Streamhouse Working Group proposes a vendor-neutral architecture for real-time data infrastructure

DevFeed: [Streamhouse Working Group proposes a vendor-neutral architecture for real-time data infrastructure](<https://devfeed.tech/articles/why-streamhouse-mission-critical-data-and-ai-need-infrastructure-built-for-live-26772.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/streamhouse-working-group-ai-infrastructure>)

Author: Alexander Gallego

Published: 2026-09-15T00:00:00Z

Content type: opinion

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Redpanda-Connect](<https://devfeed.tech/topics/redpanda-connect.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [redpanda-connect](<https://devfeed.tech/tags/redpanda-connect.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Redpanda says it has joined Aiven, Confluent, StreamNative, and Ververica to form the Streamhouse Working Group, which proposes a vendor-neutral architecture for real-time operational data infrastructure. The article distinguishes Streamhouse from the lakehouse by focusing on continuously processing current data for production applications and AI agents.

### Source excerpt

Redpanda has joined Aiven, Confluent, StreamNative, and Ververica to form the Streamhouse Working Group. Here's what that means for the future of real-time data infrastructure.

## VDURA Deploys High-Performance Storage Platform for AI and HPC at New Mexico State University

DevFeed: [VDURA Deploys High-Performance Storage Platform for AI and HPC at New Mexico State University](<https://devfeed.tech/articles/vdura-deploys-high-performance-storage-platform-for-ai-and-hpc-at-new-mexico-state-university-12380.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/vdura-deploys-high-performance-storage-platform-for-ai-and-hpc-at-new-mexico-state-university>)

Author: Harold Fritts

Published: 2026-09-02T10:00:00Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [InfiniBand](<https://devfeed.tech/topics/infiniband.md>), [Post-quantum cryptography](<https://devfeed.tech/topics/post-quantum-cryptography.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [Quantum Computing](<https://devfeed.tech/topics/quantum-computing.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [enterprise-storage](<https://devfeed.tech/tags/enterprise-storage.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [infiniband](<https://devfeed.tech/tags/infiniband.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [post-quantum-cryptography-pqc](<https://devfeed.tech/tags/post-quantum-cryptography-pqc.md>), [quantum](<https://devfeed.tech/tags/quantum.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

VDURA has moved its storage platform at New Mexico State University into full production to support AI and high-performance computing research. The deployment combines NVMe flash and high-density HDD tiers through a global namespace over InfiniBand, allowing separate scaling of performance and capacity. It also supports NMSU's post-quantum cryptography research and large-scale data pipeline projects.

### Source excerpt

VDURA has completed the deployment of its data platform at New Mexico State University (NMSU), moving the system into full production. The infrastructure is designed to serve the university's research community with a high-durability, high-throughput storage environment tailored specifically for artificial intelligence and high-performance computing (HPC) workloads. NMSU, which holds Carnegie R1 status and manages The post VDURA Deploys High-Performance Storage Platform for AI and HPC at New Mexico State University appeared first on StorageReview.com.

## MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet

DevFeed: [MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet](<https://devfeed.tech/articles/metaroce-a-new-rdma-transport-built-for-ai-scale-ethernet-130.md>)

Original publisher: [Read original article](<https://engineering.fb.com/2026/08/24/networking-traffic/metaroce-rdma-transport-ai-ethernet/>)

Author: Arvind Srinivasan; Neil Spring; Omar Baldonado; Rajiv Krishnamurthy

Published: 2026-08-24T18:02:29Z

Content type: article

Language: en

Sources: [Engineering at Meta](<https://devfeed.tech/sources/engineering-at-meta.md>)

Topics: [Ethernet](<https://devfeed.tech/topics/ethernet.md>), [Networks](<https://devfeed.tech/topics/networks.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [data centers](<https://devfeed.tech/topics/data-centers.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [data-center-engineering](<https://devfeed.tech/tags/data-center-engineering.md>), [data-centers](<https://devfeed.tech/tags/data-centers.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [ethernet](<https://devfeed.tech/tags/ethernet.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [networking-traffic](<https://devfeed.tech/tags/networking-traffic.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Meta introduces MetaRoCE, an RDMA transport protocol designed for AI workloads on commodity Ethernet at million-GPU scale. The article describes its release through the Open Compute Project and explains how endpoint intelligence, packet spraying, fine-grained logical paths, and real-time telemetry aim to provide high throughput, low tail latency, and operational simplicity for distributed training and inference.

### Source excerpt

Training and serving frontier AI models depends on fast, reliable networks that move data between GPUs without wasting compute cycles. To meet this challenge at scale, Meta designed MetaRoCE - a clean-sheet RDMA transport protocol purpose-built for AI workloads on commodity Ethernet. We're releasing the MetaRoCE specification, a reference software implementation and a compliance test [...] Read More... The post MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet appeared first on Engineering at Meta.

## MTIA 300: Meta's First Training Chip with Built-in NICs and Communication-Offloading Engines

DevFeed: [MTIA 300: Meta's First Training Chip with Built-in NICs and Communication-Offloading Engines](<https://devfeed.tech/articles/mtia-300-meta-s-first-training-chip-with-built-in-nics-and-communication-offloading-engines-131.md>)

Original publisher: [Read original article](<https://engineering.fb.com/2026/08/24/networking-traffic/mtia-300-meta-training-chip-built-in-nics/>)

Author: Rajiv Krishnamurthy; Wes Bland

Published: 2026-08-24T17:45:52Z

Content type: article

Language: en

Sources: [Engineering at Meta](<https://devfeed.tech/sources/engineering-at-meta.md>)

Topics: [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [chip-design](<https://devfeed.tech/tags/chip-design.md>), [communication](<https://devfeed.tech/tags/communication.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [devinfra](<https://devfeed.tech/tags/devinfra.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [meta](<https://devfeed.tech/tags/meta.md>), [networking-traffic](<https://devfeed.tech/tags/networking-traffic.md>), [performance](<https://devfeed.tech/tags/performance.md>), [production-engineering](<https://devfeed.tech/tags/production-engineering.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Meta describes MTIA 300, an in-house accelerator for training ranking and recommendation models, with built-in network chiplets and a co-designed HCCL communication library. The design targets communication-heavy distributed training by integrating RDMA NICs into the chip package and offloading communication work.

### Source excerpt

MTIA 300 is the first of Meta's family of in-house training and inference accelerators optimized for training ranking and recommendation models. We're sharing how MTIA 300's built-in NIC chiplets allow it to meet the communication needs associated with training recommendation models with superior performance over general-purpose GPUs. By co-designing MTIA's communication library, HCCL, alongside the [...] Read More... The post MTIA 300: Meta's First Training Chip with Built-in NICs and Communication-Offloading Engines appeared first on Engineering at Meta.

## Core Banking Modernization with Temenos Core and CockroachDB

DevFeed: [Core Banking Modernization with Temenos Core and CockroachDB](<https://devfeed.tech/articles/core-banking-modernization-with-temenos-core-and-cockroachdb-23772.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/core-banking-modernization-temenos-cockroachdb>)

Author: Nanda Badrappan,Muruga Balakrishnan

Published: 2026-08-24T00:00:00Z

Content type: article

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [CockroachDB](<https://devfeed.tech/topics/cockroachdb.md>), [Cockroach Labs](<https://devfeed.tech/topics/cockroach-labs.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Transactions](<https://devfeed.tech/topics/transactions.md>), [legacy](<https://devfeed.tech/topics/legacy.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [systems](<https://devfeed.tech/topics/systems.md>), [on-prem](<https://devfeed.tech/topics/on-prem.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cockroach-labs](<https://devfeed.tech/tags/cockroach-labs.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [legacy](<https://devfeed.tech/tags/legacy.md>), [on-prem](<https://devfeed.tech/tags/on-prem.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [sql](<https://devfeed.tech/tags/sql.md>), [transactions](<https://devfeed.tech/tags/transactions.md>)

### AI overview

This article explains why banks are modernizing legacy core banking infrastructure. It describes how Temenos Core and CockroachDB are positioned to support real-time transaction processing, continuous availability, distributed operations, resilience, and regulatory requirements.

### Source excerpt

Banking has become an always-on business. Customers expect instant payments, accurate balances, and continuous access to financial services...

## From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta's Ads Ranking

DevFeed: [From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta's Ads Ranking](<https://devfeed.tech/articles/from-user-sequences-to-scaling-laws-a-multi-stage-architecture-for-meta-s-ads-ranking-128.md>)

Original publisher: [Read original article](<https://engineering.fb.com/2026/08/05/ml-applications/from-user-sequences-to-scaling-laws-a-multi-stage-architecture-for-metas-ads-ranking/>)

Author: Steven De Gryze; Parshva Doshi; Sean O'Byrne; Arnold Overwijk; Dinesh Ramasamy; Lee Xiong

Published: 2026-08-05T19:20:20Z

Content type: article

Language: en

Sources: [Engineering at Meta](<https://devfeed.tech/sources/engineering-at-meta.md>), [Meta ML Applications](<https://devfeed.tech/sources/meta-ml-applications.md>)

Topics: [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [Temporal data](<https://devfeed.tech/topics/temporal-data.md>)

Tags: [ads](<https://devfeed.tech/tags/ads.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [meta](<https://devfeed.tech/tags/meta.md>), [ml-applications](<https://devfeed.tech/tags/ml-applications.md>), [production-engineering](<https://devfeed.tech/tags/production-engineering.md>), [recommendation-systems](<https://devfeed.tech/tags/recommendation-systems.md>), [scaling-laws](<https://devfeed.tech/tags/scaling-laws.md>), [tokenization](<https://devfeed.tech/tags/tokenization.md>)

### AI overview

Meta describes a multi-stage sequence-model architecture for ads ranking that separates offline user modeling from lightweight online ranking. It also uses dense tokenization and target-aware attention to learn feature interactions, with reported conversion and ad-click lifts across Instagram and Facebook.

### Source excerpt

Every day, Meta's recommendation platforms handle billions of user interactions, generating rich temporal signals that capture individual preferences and intent across products, ads, and content. In our 2024 post on sequence learning for ads recommendations, we showed how modeling the order and timing of user actions (rather than relying on static, manually engineered sparse features) [...] Read More... The post From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta's Ads Ranking appeared first on Engineering at Meta.

## GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

DevFeed: [GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model](<https://devfeed.tech/articles/gem-training-how-meta-doubled-the-efficiency-of-its-llm-scale-ads-foundation-model-127.md>)

Original publisher: [Read original article](<https://engineering.fb.com/2026/08/03/ml-applications/training-gem-at-llm-scale-meta-ads-recommendation-foundation-model/>)

Author: Darren Liu; Huayu Li; Raghav Boinepalli; Yuzhen Huang; Jackie (Jiaqi) Xu; Richard Qiu; Chunzhi Yang; Rich Zhu; Dev (Devashish) Shankar; Huaqing Xiong

Published: 2026-08-03T18:00:17Z

Content type: article

Language: en

Sources: [Engineering at Meta](<https://devfeed.tech/sources/engineering-at-meta.md>), [Meta AI Research](<https://devfeed.tech/sources/meta-ai-research.md>), [Meta ML Applications](<https://devfeed.tech/sources/meta-ml-applications.md>)

Topics: [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [networking](<https://devfeed.tech/topics/networking.md>)

Tags: [ads](<https://devfeed.tech/tags/ads.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [kernels](<https://devfeed.tech/tags/kernels.md>), [llm](<https://devfeed.tech/tags/llm.md>), [meta](<https://devfeed.tech/tags/meta.md>), [ml-applications](<https://devfeed.tech/tags/ml-applications.md>), [networking](<https://devfeed.tech/tags/networking.md>), [recommendation-systems](<https://devfeed.tech/tags/recommendation-systems.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Meta describes training its GEM ads recommendation foundation model at LLM scale. The article covers recommendation-specific kernels, ultra-low-precision training, and topology-aware parallelism that doubled end-to-end training efficiency to 20-25% MFU while increasing training FLOPs fourfold.

### Source excerpt

Meta's Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20-25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in [...] Read More... The post GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model appeared first on Engineering at Meta.

## Data Engineering Weekly #280

DevFeed: [Data Engineering Weekly #280](<https://devfeed.tech/articles/data-engineering-weekly-280-18260.md>)

Original publisher: [Read original article](<https://www.dataengineeringweekly.com/p/data-engineering-weekly-280>)

Author: Ananth Packkildurai

Published: 2026-07-27T03:37:19Z

Content type: article

Language: en

Sources: [Data Engineering Weekly](<https://devfeed.tech/sources/data-engineering-weekly.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [DuckDB](<https://devfeed.tech/topics/duckdb.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [database](<https://devfeed.tech/tags/database.md>), [duckdb](<https://devfeed.tech/tags/duckdb.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [python](<https://devfeed.tech/tags/python.md>)

### AI overview

Data Engineering Weekly #280 is a newsletter roundup covering updates to leetdata.ai and aidataengineer.io, agent-oriented data infrastructure, open-source modern data stack tools, data-tool landscapes, metric certification, and data quality in the AI era.

### Source excerpt

The Weekly Data Engineering Newsletter

## The data platform 1Password needed didn't exist. So we built it.

DevFeed: [The data platform 1Password needed didn't exist. So we built it.](<https://devfeed.tech/articles/the-data-platform-1password-needed-didn-t-exist-so-we-built-it-1969.md>)

Original publisher: [Read original article](<https://1password.com/blog/we-built-the-data-platform-1password-needed>)

Author: info@1password.com (Wayne Duso; Mandy Gu; and Amie Bright)

Published: 2026-07-22T00:00:00Z

Content type: article

Language: en

Sources: [Blog on 1Password Blog](<https://devfeed.tech/sources/blog-on-1password-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [data-governance](<https://devfeed.tech/topics/data-governance.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Unified Access](<https://devfeed.tech/topics/unified-access.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [building-1password](<https://devfeed.tech/tags/building-1password.md>), [data](<https://devfeed.tech/tags/data.md>), [data-governance](<https://devfeed.tech/tags/data-governance.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [developers](<https://devfeed.tech/tags/developers.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [sql](<https://devfeed.tech/tags/sql.md>), [systems](<https://devfeed.tech/tags/systems.md>), [unified-access](<https://devfeed.tech/tags/unified-access.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

1Password describes building an internal data platform to make business data accessible, trustworthy, and available in real time across product, finance, engineering, analytics, SQL, APIs, dashboards, and AI workflows. The effort addresses bottlenecks caused by a centralized data lake, bespoke pipelines, and tightly coupled storage, governance, and compute.

### Source excerpt

Consider a few tasks that take place across every business, every day: A product team ships a feature and wants to know if customers are using it successfully. A finance team needs customer and account information for planning. An analyst needs definitions to create a report. An AI assistant needs operational context to answer a business question. Those sound like different workflows, but they all rely on the same underlying data. And each one of these actors, across each of these teams, needs that data to be both accessible and trustworthy. At 1Password, trust is at the center of everything we build. Millions of people and businesses rely on us to protect the credentials, secrets, and access workflows that power modern work. The same principle applies to our own internal data. As our products, systems, and use of AI evolved, data became a shared dependency across the business. It powers everything from Unified Access andsecure agentic access patterns for customers, to the workflows used by product, finance, and engineering teams. As those systems grew, so did the number of people and applications that depended on our internal data. But our data infrastructure did not respond well to this. We had built a centralized data lake supported by a growing collection of bespoke data pipelines and one-off solutions. Each new use case required another integration or transformation. Over time, the data platform became a bottleneck: teams turned to CSVs to move faster, and data engineers spent more time maintaining pipelines than enabling new capabilities. What we learned was that manually moving data was no longer enough. Data needs to be available in real time, accessible wherever it's needed, and trusted through the forms our customers need, whether that is through SQL, APIs, dashboards or AI workflows. Privacy and governance need to be built into data the moment it's created so every downstream use remains secure and unambiguous by design. This is the foundation we're build

## ClickHouse announced as Principal Partner and Front of Shirt Sponsor of Fulham Football Club

DevFeed: [ClickHouse announced as Principal Partner and Front of Shirt Sponsor of Fulham Football Club](<https://devfeed.tech/articles/clickhouse-announced-as-principal-partner-and-front-of-shirt-sponsor-of-fulham-football-club-5068.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/clickhouse-announced-principal-partner-sponsor-fulham-fc>)

Author: ClickHouse

Published: 2026-07-21T08:09:34Z

Content type: news

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [data](<https://devfeed.tech/topics/data.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [applications](<https://devfeed.tech/tags/applications.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [company](<https://devfeed.tech/tags/company.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [global](<https://devfeed.tech/tags/global.md>), [innovation](<https://devfeed.tech/tags/innovation.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [performance](<https://devfeed.tech/tags/performance.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

ClickHouse has become Fulham Football Club's Principal Partner and Front of Shirt Sponsor of the Men's First Team in a multi-year partnership. The collaboration combines sponsorship, supporter-facing experiences, and opportunities to develop skills in data and artificial intelligence, while highlighting ClickHouse's role in real-time analytics and open-source database technology.

### Source excerpt

Fulham Football Club is proud to announce ClickHouse, a global technology company specialising in real-time analytics, data infrastructure and AI applications, as its newest Principal Partner and Front of Shirt Sponsor of the Men's First Team.

## One Driver, One Format, Every Language: ADBC

DevFeed: [One Driver, One Format, Every Language: ADBC](<https://devfeed.tech/articles/one-driver-one-format-every-language-adbc-5341.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/introducing-the-clickhouse-adbc-driver>)

Author: Luke Gannon

Published: 2026-07-10T15:00:55Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [data](<https://devfeed.tech/topics/data.md>), [Data Science](<https://devfeed.tech/topics/data-science.md>), [R](<https://devfeed.tech/topics/r.md>), [Ruby](<https://devfeed.tech/topics/ruby.md>), [Go Language](<https://devfeed.tech/topics/go-language.md>), [JavaScript](<https://devfeed.tech/topics/javascript.md>), [Python](<https://devfeed.tech/topics/python.md>)

Tags: [apache](<https://devfeed.tech/tags/apache.md>), [api](<https://devfeed.tech/tags/api.md>), [c](<https://devfeed.tech/tags/c.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data](<https://devfeed.tech/tags/data.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [database](<https://devfeed.tech/tags/database.md>), [databases](<https://devfeed.tech/tags/databases.md>), [go](<https://devfeed.tech/tags/go.md>), [java](<https://devfeed.tech/tags/java.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [python](<https://devfeed.tech/tags/python.md>), [ruby](<https://devfeed.tech/tags/ruby.md>)

### AI overview

ClickHouse has introduced an official ADBC driver, providing Arrow-native, zero-conversion access to ClickHouse across multiple programming languages. Built with Rust and distributed through the ADBC Driver Foundry, it gives languages such as Ruby, R, and C standardized database access without separate ClickHouse drivers.

### Source excerpt

ClickHouse now has an official ADBC driver, giving Ruby, R, C, and every other ADBC-aware tool zero-conversion, Arrow-native access to ClickHouse without a dedicated client for each language.

## 10 Years of Meta's Commitment to Python

DevFeed: [10 Years of Meta's Commitment to Python](<https://devfeed.tech/articles/10-years-of-meta-s-commitment-to-python-22582.md>)

Original publisher: [Read original article](<https://engineering.fb.com/2026/06/30/open-source/10-years-of-metas-commitment-to-python/>)

Author: Chris Wiltz

Published: 2026-06-30T16:00:46Z

Content type: opinion

Language: en

Sources: [Meta AI Research](<https://devfeed.tech/sources/meta-ai-research.md>), [Meta ML Applications](<https://devfeed.tech/sources/meta-ml-applications.md>)

Topics: [Meta](<https://devfeed.tech/topics/meta.md>), [Python](<https://devfeed.tech/topics/python.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Programming language](<https://devfeed.tech/topics/programming-language.md>), [Software](<https://devfeed.tech/topics/software.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>)

Tags: [ai-research](<https://devfeed.tech/tags/ai-research.md>), [culture](<https://devfeed.tech/tags/culture.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [devinfra](<https://devfeed.tech/tags/devinfra.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [meta](<https://devfeed.tech/tags/meta.md>), [ml-applications](<https://devfeed.tech/tags/ml-applications.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [production-engineering](<https://devfeed.tech/tags/production-engineering.md>), [programming](<https://devfeed.tech/tags/programming.md>), [python](<https://devfeed.tech/tags/python.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [software](<https://devfeed.tech/tags/software.md>)

### AI overview

Meta reflects on its 10th consecutive year sponsoring the Python Software Foundation and explains Python's importance across its engineering stack, including products, infrastructure, and AI research.

### Source excerpt

This year marks Meta's 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it. Python is one of the world's most influential programming languages, and we use it across our engineering stack, from [...] Read More... The post 10 Years of Meta's Commitment to Python appeared first on Engineering at Meta.

## Core dump epidemiology: fixing an 18-year-old bug

DevFeed: [Core dump epidemiology: fixing an 18-year-old bug](<https://devfeed.tech/articles/core-dump-epidemiology-fixing-an-18-year-old-bug-6359.md>)

Original publisher: [Read original article](<https://openai.com/index/core-dump-epidemiology-data-infrastructure-bug>)

Published: 2026-06-30T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [bug](<https://devfeed.tech/topics/bug.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Rockset](<https://devfeed.tech/topics/rockset.md>), [C++](<https://devfeed.tech/topics/c-plus-plus.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [race-condition](<https://devfeed.tech/topics/race-condition.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [azure](<https://devfeed.tech/tags/azure.md>), [bug](<https://devfeed.tech/tags/bug.md>), [c-plus-plus](<https://devfeed.tech/tags/c-plus-plus.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [openai](<https://devfeed.tech/tags/openai.md>), [race-condition](<https://devfeed.tech/tags/race-condition.md>), [software](<https://devfeed.tech/tags/software.md>)

### AI overview

OpenAI engineers analyzed a large population of core dumps to investigate rare crashes in the Rockset service. The investigation uncovered two unrelated causes: silent CPU corruption on an Azure host and an 18-year-old race condition in GNU libunwind.

### Source excerpt

OpenAI engineers used large-scale core dump analysis to debug rare infrastructure crashes, uncovering both a hardware fault and a long-standing software bug.

## Customer Analysis for SaaS: A Framework for Real Decisions

DevFeed: [Customer Analysis for SaaS: A Framework for Real Decisions](<https://devfeed.tech/articles/customer-analysis-for-saas-a-framework-for-real-decisions-9787.md>)

Original publisher: [Read original article](<https://dodopayments.com/blogs/customer-analysis-framework-saas/>)

Author: Ayush Agarwal

Published: 2026-06-13T00:00:00Z

Content type: tutorial

Language: en

Sources: [Dodo Payments Blog](<https://devfeed.tech/sources/dodo-payments-blog.md>)

Topics: [Software as a service](<https://devfeed.tech/topics/saas.md>), [data](<https://devfeed.tech/topics/data.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Framework](<https://devfeed.tech/topics/framework.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [churn](<https://devfeed.tech/tags/churn.md>), [data](<https://devfeed.tech/tags/data.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [growth](<https://devfeed.tech/tags/growth.md>), [guide](<https://devfeed.tech/tags/guide.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [retention](<https://devfeed.tech/tags/retention.md>), [saas](<https://devfeed.tech/tags/saas.md>)

### AI overview

This guide presents a framework for SaaS customer analysis, covering segmentation, customer behavior, willingness to pay, retention drivers, and the data infrastructure needed to turn customer data into actionable business decisions.

### Source excerpt

Customer analysis framework for SaaS founders. Segmentation, behavior, willingness-to-pay, retention drivers, and the data infrastructure to make it actionable.

## Scaling beyond one: How Airbnb evolved its data architecture for a multi-product world

DevFeed: [Scaling beyond one: How Airbnb evolved its data architecture for a multi-product world](<https://devfeed.tech/articles/scaling-beyond-one-how-airbnb-evolved-its-data-architecture-for-a-multi-product-world-1222.md>)

Original publisher: [Read original article](<https://medium.com/airbnb-engineering/scaling-beyond-one-how-airbnb-evolved-its-data-architecture-for-a-multi-product-world-6125645d470c?source=rss----53c7c27702d5---4>)

Author: Patrick Lam

Published: 2026-06-09T17:01:02Z

Content type: article

Language: en

Sources: [The Airbnb Tech Blog - Medium](<https://devfeed.tech/sources/the-airbnb-tech-blog-medium.md>)

Topics: [data-architecture](<https://devfeed.tech/topics/data-architecture.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [data-modeling](<https://devfeed.tech/topics/data-modeling.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [data](<https://devfeed.tech/tags/data.md>), [data-architecture](<https://devfeed.tech/tags/data-architecture.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [data-modeling](<https://devfeed.tech/tags/data-modeling.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [offline](<https://devfeed.tech/tags/offline.md>), [post](<https://devfeed.tech/tags/post.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

Airbnb's data and analytics engineering teams evolved a decade-old offline data warehouse to support Homes, Experiences, and Services. The article examines the trade-offs between separate product-specific data models and a unified monolithic model while describing the need for a consistent, flexible, and scalable data foundation.

### Source excerpt

How Airbnb's data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services. By: Patrick Lam, Namrata Lamba, Jamie Stober With the May 2025 Summer Release, Airbnb redesigned its app, relaunched Experiences, and debuted Services, pushing us beyond our traditional Homes focus. For the data teams, this meant rapidly evolving a decade-old infrastructure to integrate two brand-new product pillars. Our data engineers and analytics engineers rose to the challenge by building a consistent and flexible framework to serve as a robust and scalable data foundation for the next decade of growth. But getting there wasn't straightforward. This fundamental shift surfaced a critical question for our data organization: How do you evolve your offline data architecture to support new product lines without introducing disorder in vital analytics services? We knew the approach we took would have long-lasting implications. A fragmented strategy risked creating data silos, inconsistent analytics, and a tangled web of technical debt that would likely slow down future innovation. In this post, we'll take you behind the scenes to share key decisions that we made, the framework that emerged, and the lessons that helped reshape our offline data warehouse for the future. Note that we focus specifically on our offline data warehouse (the analytics-oriented data infrastructure owned by our data engineers and analytics engineers) rather than the online data systems that serve the app directly, as the two domains have fundamentally different requirements, constraints, and design philosophies that warrant separate treatment. The core dilemma: separate vs. monolithic The first and most critical question was how to structure offline data for the new, three-product world, with Homes, a refreshed Experiences product, and the new Services offering. This involved a trade-off between two main approaches: Separate

## Real-time streaming for the agentic era with NVIDIA

DevFeed: [Real-time streaming for the agentic era with NVIDIA](<https://devfeed.tech/articles/real-time-streaming-for-the-agentic-era-with-nvidia-12719.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/nvidia-ai-ecosystem>)

Author: Melissa Czapiga

Published: 2026-06-01T00:00:00Z

Content type: article

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Vera CPU](<https://devfeed.tech/topics/vera-cpu.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [agent observability](<https://devfeed.tech/topics/agent-observability.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [performance](<https://devfeed.tech/tags/performance.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

Redpanda describes its collaboration with NVIDIA to run streaming workloads on NVIDIA Vera CPUs, reporting 5.5x lower latency for AI agents in mission-critical environments. The article discusses testing on workloads processing billions of messages daily, real-time data infrastructure, and potential benefits for enterprise AI and agentic applications.

### Source excerpt

NVIDIA Vera launches today with Redpanda as part of the ecosystem, delivering 5.5x lower latencies for agents running in mission-critical environments.

## AI Agents Need Context to Reason, Not Just Data

DevFeed: [AI Agents Need Context to Reason, Not Just Data](<https://devfeed.tech/articles/ai-agents-need-context-to-reason-not-just-data-23742.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/ai-agent-context-management>)

Author: Quentin Packard

Published: 2026-05-28T00:00:00Z

Content type: article

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [context](<https://devfeed.tech/topics/context.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [Agentic AI Architecture](<https://devfeed.tech/topics/agentic-ai-architecture.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-architecture](<https://devfeed.tech/tags/ai-architecture.md>), [context](<https://devfeed.tech/tags/context.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [database](<https://devfeed.tech/tags/database.md>)

### AI overview

The article argues that production failures in AI agents often stem from inadequate context management rather than the model itself. Reliable agent behavior requires current data, memory, permissions, observability, and awareness of system constraints, making context management a data infrastructure problem beyond basic retrieval or prompt engineering.

### Source excerpt

When your AI agent makes a bad decision in production, what do you blame?

## ClickHouse tops $250M ARR and 4,000 customers, launches Claude-powered agents at Open House 2026

DevFeed: [ClickHouse tops $250M ARR and 4,000 customers, launches Claude-powered agents at Open House 2026](<https://devfeed.tech/articles/clickhouse-tops-250m-arr-and-4-000-customers-launches-claude-powered-agents-at-open-house-2026-5156.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/clickhouse-tops-250m-arr-and-4000-customers>)

Author: ClickHouse

Published: 2026-05-27T17:57:17Z

Content type: release

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude](<https://devfeed.tech/tags/claude.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [skills](<https://devfeed.tech/tags/skills.md>)

### AI overview

ClickHouse announced that its serverless cloud offering exceeded $250 million in annual run-rate revenue and reached 4,000 customers. It also introduced Claude-powered ClickHouse Agents and published the CostBench cloud data warehouse benchmark.

### Source excerpt

ClickHouse today opened Open House 2026, its second annual user conference, with a set of announcements that mark one of the company's most active quarters since founding.

## Making User-Sequence Data More Cost-Efficient, Faster, and Easier to Use

DevFeed: [Making User-Sequence Data More Cost-Efficient, Faster, and Easier to Use](<https://devfeed.tech/articles/making-user-sequence-data-more-cost-efficient-faster-and-easier-to-use-1231.md>)

Original publisher: [Read original article](<https://medium.com/pinterest-engineering/making-user-sequence-data-more-cost-efficient-faster-and-easier-to-use-2a56a928cae1?source=rss----4c5a5f6279b6---4>)

Author: Pinterest Engineering

Published: 2026-05-21T16:01:00Z

Content type: article

Language: en

Sources: [Pinterest Engineering Blog - Medium](<https://devfeed.tech/sources/pinterest-engineering-blog-medium.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [Feature Engineering](<https://devfeed.tech/topics/feature-engineering.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [feature-engineering](<https://devfeed.tech/tags/feature-engineering.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [pinterest](<https://devfeed.tech/tags/pinterest.md>), [production](<https://devfeed.tech/tags/production.md>), [recommendation-system](<https://devfeed.tech/tags/recommendation-system.md>), [recommendation-systems](<https://devfeed.tech/tags/recommendation-systems.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [systems](<https://devfeed.tech/tags/systems.md>), [train](<https://devfeed.tech/tags/train.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

Pinterest describes a redesign of its user-sequence platform for ranking, retrieval, and recommendation workloads. The article explains how enriched event sequences support training datasets, offline analysis, online inference, and latency-sensitive production use cases, with goals of reducing cost, improving extensibility, and simplifying debugging.

### Source excerpt

Authors (listed alphabetically) Ads Feature Engineering Infra team: Ajay Venkatakrishnan, Le Zhang Core ML Infra team: Eric Shang, Pihui Wei ML Data team: Connor Votroubek, Yi He User Understanding team: Camilo Munoz, Simin Li If you work on ranking, retrieval, or recommendation systems, you've probably asked for some version of the same thing: "Give me the last N meaningful actions this user took, with the right enrichments, in a format that's easy to train and serve ML models." On paper, that sounds simple. In practice, "user sequences" often become one of the most expensive and fragile parts of the ML data stack. They end up powering everything from training datasets to offline analysis and online inference, so they need to be fresh and complete at the same time. They must remain consistent as you add new events and enrichments. And they have to do all of this while serving latency-sensitive production workloads. This article walks through how we redesigned our user-sequence platform to make these sequences cheaper to run, faster to extend, and easier to debug, while still supporting demanding production use cases. What We Mean by "User Sequence" In this context, a user sequence is an ordered list of recent, relevant events for a user, along with the enrichments (signals) attached to each event. Here, enrichments mean all the extra signals we attach to raw events, so they're useful for models: embeddings (for example, Pin or query representations), contextual features (such as surface, device, or country), and derived attributes or counters that describe how the user interacted with a piece of content over time. A concrete example helps. Imagine a sequence made up of the last 500 engagements a user had with Pinterest Pins. Each event in that sequence might carry a timestamp, an action type, the surface where the action occurred, and a handful of embedding features or categorical attributes. As a data primitive, user sequences are powerful. They capture temporal b

## How We Built a Zero-Downtime Database Migration Service at Wix

DevFeed: [How We Built a Zero-Downtime Database Migration Service at Wix](<https://devfeed.tech/articles/how-we-built-a-zero-downtime-database-migration-service-at-wix-22636.md>)

Original publisher: [Read original article](<https://www.wix.engineering/post/how-we-built-a-zero-downtime-database-migration-service-at-wix>)

Author: Wix Engineering

Published: 2026-05-14T07:19:17Z

Content type: article

Language: en

Sources: [Wix Engineering](<https://devfeed.tech/sources/wix-engineering.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [migration](<https://devfeed.tech/topics/migration.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>)

Tags: [amazon-msk](<https://devfeed.tech/tags/amazon-msk.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [database](<https://devfeed.tech/tags/database.md>), [database-migration](<https://devfeed.tech/tags/database-migration.md>), [databases](<https://devfeed.tech/tags/databases.md>), [db-topology](<https://devfeed.tech/tags/db-topology.md>), [debezium-connector](<https://devfeed.tech/tags/debezium-connector.md>), [good-tools-doesn-t-fit](<https://devfeed.tech/tags/good-tools-doesn-t-fit.md>), [how-db-mover-works](<https://devfeed.tech/tags/how-db-mover-works.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [migration](<https://devfeed.tech/tags/migration.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [summary-6-key-focal-points](<https://devfeed.tech/tags/summary-6-key-focal-points.md>), [the-challenge](<https://devfeed.tech/tags/the-challenge.md>), [the-python-service](<https://devfeed.tech/tags/the-python-service.md>), [the-result](<https://devfeed.tech/tags/the-result.md>)

### AI overview

Wix describes DB Mover, an internal service for transparent, zero-downtime database migrations between shared and dedicated MySQL clusters. The article explains the operational risks of shared clusters, migration requirements, and the topology managed by Wix's Data Infrastructure team.

### Source excerpt

The Challenge At Wix, multiple applications share the same DB cluster. The reasons vary: consolidating apps from the same domain, grouping several small apps that don't justify a dedicated cluster, or simply optimizing cost and operational overhead. However, this setup comes with a significant risk: every application on the cluster has the potential to impact all the others. One real example: we had a critical service related to user authentication sharing a DB cluster with several other...

## Powering self-driving vehicle analytics at Avride with ClickHouse Cloud

DevFeed: [Powering self-driving vehicle analytics at Avride with ClickHouse Cloud](<https://devfeed.tech/articles/powering-self-driving-vehicle-analytics-at-avride-with-clickhouse-cloud-4971.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/avride>)

Author: ClickHouse

Published: 2026-05-11T00:00:00Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Grafana](<https://devfeed.tech/topics/grafana.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [aws](<https://devfeed.tech/tags/aws.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [driving](<https://devfeed.tech/tags/driving.md>), [grafana](<https://devfeed.tech/tags/grafana.md>), [latency](<https://devfeed.tech/tags/latency.md>), [lidar](<https://devfeed.tech/tags/lidar.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [robots](<https://devfeed.tech/tags/robots.md>), [storage](<https://devfeed.tech/tags/storage.md>), [streams](<https://devfeed.tech/tags/streams.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [tooling](<https://devfeed.tech/tags/tooling.md>), [visualization](<https://devfeed.tech/tags/visualization.md>)

### AI overview

Avride uses ClickHouse Cloud as the data backbone for its autonomous vehicles and delivery robots, supporting ride-data indexing, metrics, analytics, and internal tooling. Its migration from Apache Iceberg reduced index lookup and ingestion latency.

### Source excerpt

Avride replaced Apache Iceberg with ClickHouse Cloud, cutting index lookup latency from 20 seconds to under 100ms and ingestion from hours to seconds.

## From SSH to REST: A Security-Driven Modernization of Slack's EMR Data Pipelines

DevFeed: [From SSH to REST: A Security-Driven Modernization of Slack's EMR Data Pipelines](<https://devfeed.tech/articles/from-ssh-to-rest-a-security-driven-modernization-of-slack-s-emr-data-pipelines-146.md>)

Original publisher: [Read original article](<https://slack.engineering/from-ssh-to-rest-a-security-driven-modernization-of-slacks-emr-data-pipelines/>)

Author: Mahendran Vasagam

Published: 2026-05-05T14:00:01Z

Content type: article

Language: en

Sources: [Engineering at Slack](<https://devfeed.tech/sources/engineering-at-slack.md>)

Topics: [Security](<https://devfeed.tech/topics/security.md>), [OpenSSH](<https://devfeed.tech/topics/openssh.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [DataOps](<https://devfeed.tech/topics/dataops.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [aws](<https://devfeed.tech/tags/aws.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [modernization](<https://devfeed.tech/tags/modernization.md>), [security](<https://devfeed.tech/tags/security.md>), [ssh](<https://devfeed.tech/tags/ssh.md>), [uncategorized](<https://devfeed.tech/tags/uncategorized.md>)

### AI overview

Slack describes migrating more than 700 SSH-based data pipeline jobs to a REST-based architecture across eight data regions, eliminating SSH access to production AWS EMR clusters without downtime. The article explains the security and operational problems that motivated the modernization, including attack-surface exposure, key-management overhead, resource contention, broken connections, zombie jobs, and unreliable job-status detection.

### Source excerpt

Excerpt By 2024, Slack's data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We're talking daily search indexing that processed terabytes of data, analytics jobs powering business intelligence, the whole shebang. Every single one of these jobs required direct SSH access to production AWS Elastic MapReduce (EMR) clusters. We had a massive security...

## Gala supercharges analytics performance with ClickHouse on AWS

DevFeed: [Gala supercharges analytics performance with ClickHouse on AWS](<https://devfeed.tech/articles/gala-supercharges-analytics-performance-with-clickhouse-on-aws-5258.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/gala>)

Author: ClickHouse

Published: 2026-05-04T00:00:00Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Amazon Web Services (AWS)](<https://devfeed.tech/topics/amazon-web-services-aws.md>), [data analytics](<https://devfeed.tech/topics/data-analytics.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [data-platforms](<https://devfeed.tech/topics/data-platforms.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [databricks](<https://devfeed.tech/topics/databricks.md>), [Blockchain](<https://devfeed.tech/topics/blockchain.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [blockchain](<https://devfeed.tech/tags/blockchain.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [data-platforms](<https://devfeed.tech/tags/data-platforms.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Gala migrated from Databricks to ClickHouse on AWS to address growing data volumes and improve analytics performance. The article reports a threefold increase in analytics capacity, a 30% cost reduction, and query times reduced from minutes to sub-second.

### Source excerpt

Learn how Gala migrated to the ClickHouse to improve analytics performance and cut costs

[Next page](<https://devfeed.tech/tags/data-infrastructure.md?cursor=WyIyMDI2LTA1LTA0VDAwOjAwOjAwKzAwOjAwIiwgImJiZWI1NzUyLWU1YWUtNDlkYy04NWI3LWU0NTFhZWJkNWFkZCJd>)