# clusters

Published articles for clusters.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## How Intuit and AWS systematically improved resiliency on ElastiCache using AWS Fault Injection Service

DevFeed: [How Intuit and AWS systematically improved resiliency on ElastiCache using AWS Fault Injection Service](<https://devfeed.tech/articles/how-intuit-and-aws-systematically-improved-resiliency-on-elasticache-using-aws-fault-injection-service-42097.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/database/how-intuit-and-aws-systematically-improved-resiliency-on-elasticache-using-aws-fault-injection-service/>)

Author: Ryan Sheahan

Published: 2026-09-17T15:59:44Z

Content type: article

Language: en

Sources: [AWS Database Blog](<https://devfeed.tech/sources/aws-database-blog.md>)

Topics: [Amazon ElastiCache](<https://devfeed.tech/topics/amazon-elasticache.md>), [AWS Fault Injection Service (FIS)](<https://devfeed.tech/topics/aws-fault-injection-service-fis.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [reliability](<https://devfeed.tech/topics/reliability.md>), [Availability](<https://devfeed.tech/topics/availability.md>)

Tags: [advanced-300](<https://devfeed.tech/tags/advanced-300.md>), [amazon-elasticache](<https://devfeed.tech/tags/amazon-elasticache.md>), [applications](<https://devfeed.tech/tags/applications.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-fault-injection-service-fis](<https://devfeed.tech/tags/aws-fault-injection-service-fis.md>), [caching](<https://devfeed.tech/tags/caching.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [customer-experience](<https://devfeed.tech/tags/customer-experience.md>), [customer-solutions](<https://devfeed.tech/tags/customer-solutions.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [resiliency](<https://devfeed.tech/tags/resiliency.md>)

### AI overview

AWS and Intuit describe validating Amazon ElastiCache resilience during an Availability Zone impairment with AWS Fault Injection Service. The post covers experiment setup, measurements, configuration issues, and reported recovery improvements under production-level load.

### Source excerpt

Learn how Intuit and AWS validated Amazon ElastiCache resilience under a real Availability Zone impairment using AWS Fault Injection Service, cutting recovery from over 50 minutes to under 2 minutes with no manual intervention and reducing customer impact to effectively zero.

## Kubernetes Multi-Cluster Project Karmada Reaches CNCF Graduation

DevFeed: [Kubernetes Multi-Cluster Project Karmada Reaches CNCF Graduation](<https://devfeed.tech/articles/kubernetes-multi-cluster-project-karmada-reaches-cncf-graduation-41297.md>)

Original publisher: [Read original article](<https://www.infoq.com/news/2026/09/karmada-kubernetes-cncf/>)

Author: Claudio Masolo

Published: 2026-09-17T10:00:00Z

Content type: news

Language: en

Sources: [InfoQ](<https://devfeed.tech/sources/infoq.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Computing](<https://devfeed.tech/topics/computing.md>)

Tags: [ai-architecture](<https://devfeed.tech/tags/ai-architecture.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cluster](<https://devfeed.tech/tags/cluster.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [cncf](<https://devfeed.tech/tags/cncf.md>), [devops](<https://devfeed.tech/tags/devops.md>), [karmada-kubernetes-cncf](<https://devfeed.tech/tags/karmada-kubernetes-cncf.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [multi-cloud](<https://devfeed.tech/tags/multi-cloud.md>), [news](<https://devfeed.tech/tags/news.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [project](<https://devfeed.tech/tags/project.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

The CNCF announced that Karmada, a multi-cluster and multi-cloud Kubernetes orchestration project, graduated to its highest maturity tier. The announcement coincided with Karmada v1.19, which improves multi-component scheduling for AI training jobs and makes priority-based scheduling available by default in Beta.

### Source excerpt

The Cloud Native Computing Foundation (CNCF) announced on September 2026 that Karmada, a multi-cluster and multi-cloud Kubernetes orchestration project, has graduated. This multi-cluster and multi-cloud Kubernetes orchestration project reached CNCF's highest maturity tier. By Claudio Masolo

## Kubernetes v1.37: Memory QoS Graduates to Beta

DevFeed: [Kubernetes v1.37: Memory QoS Graduates to Beta](<https://devfeed.tech/articles/kubernetes-v1-37-memory-qos-graduates-to-beta-20863.md>)

Original publisher: [Read original article](<https://kubernetes.io/blog/2026/09/14/kubernetes-v1-37-memory-qos-graduates-to-beta/>)

Author: Qi Wang; Sohan Kunkerkar

Published: 2026-09-14T18:30:00Z

Content type: article

Language: en

Sources: [Kubernetes Blog](<https://devfeed.tech/sources/kubernetes-blog.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [releases](<https://devfeed.tech/topics/releases.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [Kernel](<https://devfeed.tech/topics/kernel.md>)

Tags: [clusters](<https://devfeed.tech/tags/clusters.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [container](<https://devfeed.tech/tags/container.md>), [kernel](<https://devfeed.tech/tags/kernel.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [linux](<https://devfeed.tech/tags/linux.md>), [memory](<https://devfeed.tech/tags/memory.md>), [qos](<https://devfeed.tech/tags/qos.md>), [releases](<https://devfeed.tech/tags/releases.md>), [upgrade](<https://devfeed.tech/tags/upgrade.md>), [v1](<https://devfeed.tech/tags/v1.md>)

### AI overview

Kubernetes v1.37 promotes Memory QoS to Beta and enables it by default on Linux nodes using cgroup v2. The article explains the updated defaults, configuration options for memory throttling and tiered memory protection, and upgrade behavior intended to preserve existing runtime behavior.

### Source excerpt

Memory QoS has graduated to Beta in Kubernetes v1.37 and is now enabled by default. On Linux nodes running cgroup v2, the feature uses the memory controller to give the kernel better guidance on how to treat container memory. It was first introduced as Alpha in v1.22, and expanded in v1.36 with tiered memory reservation. This post covers what changed in v1.37, what the Beta promotion means for cluster operators, and how to configure the feature. What changed in v1.37Memory QoS is Beta and enabled by default The MemoryQoS feature gate is now Beta in v1.37. This means every v1.37 kubelet has the feature gate turned on without any configuration change. Turning on the feature by default is safe because the default kubelet configuration does not enable memory throttling or memory reservation. No memory.high, memory.min, or memory.low values are written to cgroups unless you explicitly configure them. You can opt into specific behaviors through kubelet configuration fields: Set memoryThrottlingFactor (for example, 0.9) to enable memory.high throttling on Burstable and BestEffort containers. The default is null, which means no throttling. Set memoryReservationPolicy to TieredReservation to enable tiered memory protection via memory.min and memory.low. The default is None, which means no memory reservation. Default memoryThrottlingFactor changed to null In earlier Alpha releases, memoryThrottlingFactor defaulted to 0.9, which meant enabling the feature gate caused the kubelet to set memory.high on containers. In v1.37, the default is null, so the kubelet does not set memory.high unless you configure a value. This change was made because, with the feature gate now on by default, an automatic memory.high could throttle workloads that were previously running without throttling. Making it null ensures that upgrading to v1.37 does not change runtime behavior for existing clusters. If your kubelet configuration file already contains an explicit memoryThrottlingFactor value, that

## Hot Chips 2026: XCENA and Samsung's Near-Memory Compute CXL Device

DevFeed: [Hot Chips 2026: XCENA and Samsung's Near-Memory Compute CXL Device](<https://devfeed.tech/articles/hot-chips-2026-xcena-and-samsung-s-near-memory-compute-cxl-device-13999.md>)

Original publisher: [Read original article](<https://chipsandcheese.com/p/hot-chips-2026-xcena-and-samsungs>)

Author: Chester Lam

Published: 2026-08-30T07:25:37Z

Content type: article

Language: en

Sources: [Chips and Cheese](<https://devfeed.tech/sources/chips-and-cheese.md>)

Topics: [samsung](<https://devfeed.tech/topics/samsung.md>), [ddr5](<https://devfeed.tech/topics/ddr5.md>), [RISC-V](<https://devfeed.tech/topics/riscv.md>), [data](<https://devfeed.tech/topics/data.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [Arm](<https://devfeed.tech/topics/arm.md>), [intel](<https://devfeed.tech/topics/intel.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [arm](<https://devfeed.tech/tags/arm.md>), [cache](<https://devfeed.tech/tags/cache.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [compute](<https://devfeed.tech/tags/compute.md>), [core](<https://devfeed.tech/tags/core.md>), [ddr5](<https://devfeed.tech/tags/ddr5.md>), [dram](<https://devfeed.tech/tags/dram.md>), [intel](<https://devfeed.tech/tags/intel.md>), [memory](<https://devfeed.tech/tags/memory.md>), [parallel](<https://devfeed.tech/tags/parallel.md>), [pcie](<https://devfeed.tech/tags/pcie.md>), [performance](<https://devfeed.tech/tags/performance.md>), [risc-v](<https://devfeed.tech/tags/risc-v.md>), [samsung](<https://devfeed.tech/tags/samsung.md>)

### AI overview

The article examines XCENA and Samsung's MX1, a CXL memory expansion device that can host up to 2 TB of DDR5 memory, connect SSDs, and provide onboard compute through 3,072 RISC-V cores. It describes the device's memory bandwidth, cache hierarchy, power use, and focus on data-parallel workloads.

### Source excerpt

CXL memory expansion, with a side of compute

## pgwatch v6: dashboards reimagined, and a reaper that doesn't choke

DevFeed: [pgwatch v6: dashboards reimagined, and a reaper that doesn't choke](<https://devfeed.tech/articles/pgwatch-v6-dashboards-reimagined-and-a-reaper-that-doesn-t-choke-14490.md>)

Original publisher: [Read original article](<https://www.cybertec-postgresql.com/en/pgwatch-v6-dashboards-reimagined-and-a-reaper-that-doesnt-choke/>)

Author: Pavlo Golub

Published: 2026-08-28T03:00:58Z

Content type: release

Language: en

Sources: [CYBERTEC PostgreSQL | Services & Support](<https://devfeed.tech/sources/cybertec-postgresql-services-support.md>)

Topics: [Grafana](<https://devfeed.tech/topics/grafana.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>), [Prometheus](<https://devfeed.tech/topics/prometheus.md>), [incident](<https://devfeed.tech/topics/incident.md>)

Tags: [blog](<https://devfeed.tech/tags/blog.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [features](<https://devfeed.tech/tags/features.md>), [grafana](<https://devfeed.tech/tags/grafana.md>), [incident](<https://devfeed.tech/tags/incident.md>), [linux](<https://devfeed.tech/tags/linux.md>), [network](<https://devfeed.tech/tags/network.md>), [news](<https://devfeed.tech/tags/news.md>), [patroni](<https://devfeed.tech/tags/patroni.md>), [pgwatch](<https://devfeed.tech/tags/pgwatch.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [production](<https://devfeed.tech/tags/production.md>), [prometheus](<https://devfeed.tech/tags/prometheus.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

This release article describes pgwatch v6.0.0-beta, including a Grafana dashboard overhaul, first-class dashboards for Prometheus sources, Patroni cluster views, four new metrics, and fixes for collection hangs caused by production infrastructure failures.

### Source excerpt

This is an extended blog, and talks about the features of pgwatch v6.0.0 beta in more details. Read to know and contribute. The post pgwatch v6: dashboards reimagined, and a reaper that doesn't choke appeared first on CYBERTEC PostgreSQL | Services & Support.

## Why fleet management matters for Kubernetes at the edge

DevFeed: [Why fleet management matters for Kubernetes at the edge](<https://devfeed.tech/articles/kubernetes-at-the-edge-has-hit-a-wall-fleet-management-is-the-way-through-17620.md>)

Original publisher: [Read original article](<https://thenewstack.io/edge-kubernetes-fleet-management/>)

Author: Arvind Bhoj

Published: 2026-08-20T13:00:00Z

Content type: opinion

Language: en

Sources: [Kubernetes Overview, News and Trends | The New Stack](<https://devfeed.tech/sources/kubernetes-overview-news-and-trends-the-new-stack.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Cloud Native Ecosystem](<https://devfeed.tech/topics/cloud-native-ecosystem.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [edge-computing](<https://devfeed.tech/tags/edge-computing.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [management](<https://devfeed.tech/tags/management.md>), [nutanix](<https://devfeed.tech/tags/nutanix.md>), [operations](<https://devfeed.tech/tags/operations.md>), [post](<https://devfeed.tech/tags/post.md>), [security](<https://devfeed.tech/tags/security.md>), [sponsor-nutanix](<https://devfeed.tech/tags/sponsor-nutanix.md>), [sponsored](<https://devfeed.tech/tags/sponsored.md>), [sponsored-post](<https://devfeed.tech/tags/sponsored-post.md>)

### AI overview

The article argues that edge computing has become a mainstream strategic concern and that Kubernetes provides a suitable platform for edge workloads, including generative AI. It identifies dispersed, customized clusters and piecemeal configuration as operational bottlenecks that fleet management can address.

### Source excerpt

Not too long ago, edge computing was seen as a niche use case, limited to telcos, manufacturing plants, and large The post Kubernetes at the edge has hit a wall. Fleet management is the way through. appeared first on The New Stack.

## Nirmata's Cloud Agents Audited a 40-Cluster Kubernetes Fleet, and recovered 40% of the Cost

DevFeed: [Nirmata's Cloud Agents Audited a 40-Cluster Kubernetes Fleet, and recovered 40% of the Cost](<https://devfeed.tech/articles/nirmata-s-cloud-agents-audited-a-40-cluster-kubernetes-fleet-and-recovered-40-of-the-cost-17654.md>)

Original publisher: [Read original article](<https://nirmata.com/2026/08/12/how-nirmata-saved-40-in-kuberbnetes-cloud-cost/>)

Author: Anubhav Sharma

Published: 2026-08-12T20:52:35Z

Content type: article

Language: en

Sources: [Nirmata](<https://devfeed.tech/sources/nirmata.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [cncf](<https://devfeed.tech/tags/cncf.md>), [compute](<https://devfeed.tech/tags/compute.md>), [cost-savings](<https://devfeed.tech/tags/cost-savings.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [devsecops](<https://devfeed.tech/tags/devsecops.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kyverno](<https://devfeed.tech/tags/kyverno.md>), [policy-management](<https://devfeed.tech/tags/policy-management.md>), [resource](<https://devfeed.tech/tags/resource.md>), [verification](<https://devfeed.tech/tags/verification.md>)

### AI overview

Nirmata describes applying its Cost Analyzer and Resource Hygiene Cloud Agents across an enterprise customer's 40-cluster production Kubernetes fleet. The scans identified a roughly $107,000 monthly compute baseline and about 50% recoverable through right-sizing before stale-resource cleanup, while revealing recurring sources of waste and governance gaps.

### Source excerpt

Nirmata's Cloud Agents Audited a 40-Cluster Kubernetes Fleet, and recovered 40% of the Cost Most Kubernetes Cost overruns don't come from one singularly bad decision. They come from dozens of reasonable ones -- made independently, by different teams, at different times -- that... The post Nirmata's Cloud Agents Audited a 40-Cluster Kubernetes Fleet, and recovered 40% of the Cost first appeared on Nirmata.

## Introducing the CYBERTEC PG Operator

DevFeed: [Introducing the CYBERTEC PG Operator](<https://devfeed.tech/articles/introducing-the-cybertec-pg-operator-14487.md>)

Original publisher: [Read original article](<https://www.cybertec-postgresql.com/en/introducing-the-cybertec-pg-operator/>)

Author: Hans-Jürgen Schönig

Published: 2026-08-12T09:05:41Z

Content type: release

Language: en

Sources: [CYBERTEC PostgreSQL | Services & Support](<https://devfeed.tech/sources/cybertec-postgresql-services-support.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [clusters](<https://devfeed.tech/tags/clusters.md>), [failover](<https://devfeed.tech/tags/failover.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [news](<https://devfeed.tech/tags/news.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [operator](<https://devfeed.tech/tags/operator.md>), [pg-operator](<https://devfeed.tech/tags/pg-operator.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [postgresql-on-cloud](<https://devfeed.tech/tags/postgresql-on-cloud.md>), [replication](<https://devfeed.tech/tags/replication.md>)

### AI overview

CYBERTEC introduces an open-source Kubernetes operator for PostgreSQL. The CYBERTEC PG Operator provides PostgreSQL lifecycle management and supports Multi-site Clusters, including cross-site replication and automated failover across multiple Kubernetes locations.

### Source excerpt

CYBERTEC PG Operator is an open-source Kubernetes operator for PostgreSQL, with Multi-site Clusters for cross-site replication and automated failover. The post Introducing the CYBERTEC PG Operator appeared first on CYBERTEC PostgreSQL | Services & Support.

## Multi-region high availability for Kafka workloads with a single Stretch Cluster

DevFeed: [Multi-region high availability for Kafka workloads with a single Stretch Cluster](<https://devfeed.tech/articles/multi-region-high-availability-for-kafka-workloads-with-a-single-stretch-cluster-12718.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/multi-region-high-availability-kafka-stretch-clusters>)

Author: David Yu

Published: 2026-08-11T00:00:00Z

Content type: article

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Kafka](<https://devfeed.tech/topics/kafka.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [Raft](<https://devfeed.tech/topics/raft.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [stream-processing](<https://devfeed.tech/topics/stream-processing.md>)

Tags: [clusters](<https://devfeed.tech/tags/clusters.md>), [k8s](<https://devfeed.tech/tags/k8s.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [raft](<https://devfeed.tech/tags/raft.md>), [release](<https://devfeed.tech/tags/release.md>), [replication](<https://devfeed.tech/tags/replication.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [stream-processing](<https://devfeed.tech/tags/stream-processing.md>)

### AI overview

Redpanda Operator 26.2 introduces generally available Stretch Clusters, allowing one logical Redpanda cluster to span multiple Kubernetes clusters and regions. The release provides synchronous replication with Raft-based automatic failover, safer broker rolling restarts, Redpanda Connect pipelines as Kubernetes resources, and Gateway API support for Redpanda Console.

### Source excerpt

Redpanda Operator 26.2 brings GA Stretch Clusters for multi-region replication, Redpanda Connect pipelines as K8s resources, Gateway API support, and safer rolling restarts.

## AI for Kubernetes: Complete Guide

DevFeed: [AI for Kubernetes: Complete Guide](<https://devfeed.tech/articles/ai-for-kubernetes-complete-guide-17478.md>)

Original publisher: [Read original article](<https://kodekloud.com/blog/ai-for-kubernetes-complete-guide/>)

Author: Nimesha Jinarajadasa

Published: 2026-07-31T03:12:00Z

Content type: tutorial

Language: en

Sources: [Kubernetes - KodeKloud Blog | DevOps, Cloud, Kubernetes, AI Tutorials & More](<https://devfeed.tech/sources/kubernetes-kodekloud-blog-devops-cloud-kubernetes-ai-tutorials-more.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Tooling](<https://devfeed.tech/topics/tooling.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [GitOps](<https://devfeed.tech/topics/gitops.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [diagnostics](<https://devfeed.tech/tags/diagnostics.md>), [gitops](<https://devfeed.tech/tags/gitops.md>), [guide](<https://devfeed.tech/tags/guide.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [mcp](<https://devfeed.tech/tags/mcp.md>)

### AI overview

A guide to AI for Kubernetes that distinguishes read-only diagnostic tools, write-capable agents, and command copilots. It explains the role of MCP, outlines the limits and risks of each category, and recommends starting with diagnosis before supervised agents and GitOps-managed workflows.

### Source excerpt

The complete map of AI for Kubernetes: diagnostics, agents, and copilots, what each does today, the honest limits, and where to start.

## Using AI to Troubleshoot Kubernetes Clusters

DevFeed: [Using AI to Troubleshoot Kubernetes Clusters](<https://devfeed.tech/articles/using-ai-to-troubleshoot-kubernetes-clusters-17490.md>)

Original publisher: [Read original article](<https://kodekloud.com/blog/using-ai-to-troubleshoot-kubernetes/>)

Author: Nimesha Jinarajadasa

Published: 2026-07-18T08:34:19Z

Content type: tutorial

Language: en

Sources: [Kubernetes - KodeKloud Blog | DevOps, Cloud, Kubernetes, AI Tutorials & More](<https://devfeed.tech/sources/kubernetes-kodekloud-blog-devops-cloud-kubernetes-ai-tutorials-more.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [errors](<https://devfeed.tech/tags/errors.md>), [guide](<https://devfeed.tech/tags/guide.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [troubleshooting](<https://devfeed.tech/tags/troubleshooting.md>), [verify](<https://devfeed.tech/tags/verify.md>)

### AI overview

A practical tutorial on using k8sgpt to troubleshoot deliberately broken Kubernetes clusters. It covers diagnosing ImagePullBackOff and CrashLoopBackOff errors, generating AI-assisted explanations and fix steps, applying a suggested fix, and verifying recovery while a human remains responsible for the decision.

### Source excerpt

Use AI to troubleshoot Kubernetes with k8sgpt: diagnose real ImagePullBackOff and CrashLoopBackOff errors, get fix steps, and verify recovery.

## From intent to enforcement: Lessons from operating Kubernetes controllers at scale

DevFeed: [From intent to enforcement: Lessons from operating Kubernetes controllers at scale](<https://devfeed.tech/articles/from-intent-to-enforcement-lessons-from-operating-kubernetes-controllers-at-scale-17627.md>)

Original publisher: [Read original article](<https://thenewstack.io/kubernetes-controllers-at-scale/>)

Author: Sri Saran Balaji Vellore Rajakumar

Published: 2026-07-17T12:00:00Z

Content type: article

Language: en

Sources: [Kubernetes Overview, News and Trends | The New Stack](<https://devfeed.tech/sources/kubernetes-overview-news-and-trends-the-new-stack.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Amazon Elastic Kubernetes Service](<https://devfeed.tech/topics/amazon-elastic-kubernetes-service.md>), [VPC](<https://devfeed.tech/topics/vpc.md>)

Tags: [amazon-eks](<https://devfeed.tech/tags/amazon-eks.md>), [aws-marketplace](<https://devfeed.tech/tags/aws-marketplace.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [controllers](<https://devfeed.tech/tags/controllers.md>), [declarative](<https://devfeed.tech/tags/declarative.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [networking](<https://devfeed.tech/tags/networking.md>), [platform-engineering](<https://devfeed.tech/tags/platform-engineering.md>), [policy-controller](<https://devfeed.tech/tags/policy-controller.md>), [post-contributed](<https://devfeed.tech/tags/post-contributed.md>), [reconciliation](<https://devfeed.tech/tags/reconciliation.md>), [scale](<https://devfeed.tech/tags/scale.md>), [security](<https://devfeed.tech/tags/security.md>), [sponsor-aws-marketplace](<https://devfeed.tech/tags/sponsor-aws-marketplace.md>), [sponsored-post-contributed](<https://devfeed.tech/tags/sponsored-post-contributed.md>), [vpc](<https://devfeed.tech/tags/vpc.md>)

### AI overview

This article explains lessons from operating Kubernetes controllers at scale in Amazon EKS. It focuses on how the Network Policy Controller and VPC Resource Controller reconcile declared intent with changing cluster state to enforce network traffic and AWS resource access policies.

### Source excerpt

Kubernetes controllers are what make the platform's declarative model real. They observe the state, reconcile toward the intent, and keep The post From intent to enforcement: Lessons from operating Kubernetes controllers at scale appeared first on The New Stack.

## Securing kubectl on Remote Kubernetes Clusters Without Static Credentials or VPNs

DevFeed: [Securing kubectl on Remote Kubernetes Clusters Without Static Credentials or VPNs](<https://devfeed.tech/articles/securing-kubectl-on-remote-kubernetes-clusters-without-static-credentials-or-vpns-29735.md>)

Original publisher: [Read original article](<https://goteleport.com/blog/kubectl-remote-clusters/>)

Author: info@goteleport.com (Steven Martin)

Published: 2026-07-17T00:00:00Z

Content type: tutorial

Language: en

Sources: [Teleport](<https://devfeed.tech/sources/teleport.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Kubernetes clusters](<https://devfeed.tech/topics/kubernetes-clusters.md>), [k3s](<https://devfeed.tech/topics/k3s.md>), [Network](<https://devfeed.tech/topics/network.md>), [Networks](<https://devfeed.tech/topics/networks.md>)

Tags: [best-practices](<https://devfeed.tech/tags/best-practices.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [credentials](<https://devfeed.tech/tags/credentials.md>), [firewalls](<https://devfeed.tech/tags/firewalls.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [k3s](<https://devfeed.tech/tags/k3s.md>), [kubectl](<https://devfeed.tech/tags/kubectl.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kubernetes-clusters](<https://devfeed.tech/tags/kubernetes-clusters.md>), [network](<https://devfeed.tech/tags/network.md>), [networks](<https://devfeed.tech/tags/networks.md>), [remote](<https://devfeed.tech/tags/remote.md>)

### AI overview

A guide to securing kubectl access to remote Kubernetes clusters running on distributed edge devices. It explains how NAT, firewalls, kubeconfig sprawl, and static credentials create access and security risks, and discusses avoiding publicly exposed API servers and VPN-related operational challenges.

### Source excerpt

Learn how to secure Kubernetes access across remote fleets without creating risk.

## Fat Rail vs Thin Rail in AI compute clusters

DevFeed: [Fat Rail vs Thin Rail in AI compute clusters](<https://devfeed.tech/articles/fat-rail-vs-thin-rail-in-ai-compute-clusters-40152.md>)

Original publisher: [Read original article](<https://blog.j2sw.com/inetarch/fat-rail-explained/>)

Author: j2sw

Published: 2026-07-01T07:30:18Z

Content type: tutorial

Language: en

Sources: [Justin Wilson (j2sw)](<https://devfeed.tech/sources/justin-wilson-j2sw.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Job](<https://devfeed.tech/topics/job.md>), [datasets](<https://devfeed.tech/topics/datasets.md>)

Tags: [400g](<https://devfeed.tech/tags/400g.md>), [ai](<https://devfeed.tech/tags/ai.md>), [cluster](<https://devfeed.tech/tags/cluster.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data](<https://devfeed.tech/tags/data.md>), [fabric](<https://devfeed.tech/tags/fabric.md>), [fat-rail](<https://devfeed.tech/tags/fat-rail.md>), [internet-architecture](<https://devfeed.tech/tags/internet-architecture.md>), [switching](<https://devfeed.tech/tags/switching.md>), [thin-rail](<https://devfeed.tech/tags/thin-rail.md>)

### AI overview

The article explains "fat" and "thin" rails in AI compute clusters. The terms describe whether a node's connection to the fabric provides sufficient bandwidth for the workload, with the practical distinction determined by node-to-node traffic rather than cable count alone.

### Source excerpt

We are starting to hear about "Thin" and "Fat" rails in AI fabrics. These are just another way to say a connection from compute into the Fabric. The link either has sufficient capacity (Fat) or insufficient capacity (Thin). A thin rail provides the node with a single path into the fabric. That may be one ... Read more The post Fat Rail vs Thin Rail in AI compute clusters appeared first on Justin Wilson (j2sw).

## How Redpanda Cloud Topics rethinks Kafka compaction

DevFeed: [How Redpanda Cloud Topics rethinks Kafka compaction](<https://devfeed.tech/articles/how-redpanda-cloud-topics-rethinks-kafka-compaction-12705.md>)

Original publisher: [Read original article](<https://www.redpanda.com/blog/how-redpanda-cloud-topics-rethinks-kafka-compaction>)

Author: Willem Kaufmann

Published: 2026-06-30T00:00:00Z

Content type: article

Language: en

Sources: [Redpanda](<https://devfeed.tech/sources/redpanda.md>)

Topics: [Kafka](<https://devfeed.tech/topics/kafka.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Algorithm](<https://devfeed.tech/topics/algorithm.md>)

Tags: [apache](<https://devfeed.tech/tags/apache.md>), [apache-kafka](<https://devfeed.tech/tags/apache-kafka.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [cloud-storage](<https://devfeed.tech/tags/cloud-storage.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [cycles](<https://devfeed.tech/tags/cycles.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [object-storage](<https://devfeed.tech/tags/object-storage.md>), [reduce](<https://devfeed.tech/tags/reduce.md>), [retention](<https://devfeed.tech/tags/retention.md>), [scale](<https://devfeed.tech/tags/scale.md>), [storage](<https://devfeed.tech/tags/storage.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

The article explains how Redpanda Cloud Topics redesigns Kafka log compaction for cloud-native streaming. It describes how the architecture reduces redundant processing and CPU use, lowers cloud storage costs, and preserves Kafka behavior while addressing scaling challenges such as limited memory, tombstone removal, and rewriting large volumes of object storage data.

### Source excerpt

Compaction can overwhelm poorly sized Kafka clusters, leading to full disks and maxed-out CPUs. Learn how Redpanda's Cloud Topics architecture redesigns compaction to cut redundant work, reduce cloud storage costs, and preserve the Kafka semantics you rely on.

## Tracking Down Orphaned Objects in ClickHouse Clusters and Improving Cleanup and Recovery

DevFeed: [Tracking Down Orphaned Objects in ClickHouse Clusters and Improving Cleanup and Recovery](<https://devfeed.tech/articles/hunting-orphan-objects-45-off-our-clickhouse-storage-bill-and-a-near-data-loss-incident-18527.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/how-we-deal-with-cloud-orphan-objects>)

Author: Irene Martínez

Published: 2026-05-19T00:00:00Z

Content type: article

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [data](<https://devfeed.tech/topics/data.md>), [incident](<https://devfeed.tech/topics/incident.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>)

Tags: [cleanup](<https://devfeed.tech/tags/cleanup.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [data](<https://devfeed.tech/tags/data.md>), [engineering-excellence](<https://devfeed.tech/tags/engineering-excellence.md>), [incident](<https://devfeed.tech/tags/incident.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

This article describes how the team tracked down orphaned objects in its ClickHouse clusters, reduced monthly storage costs, and improved cleanup and recovery procedures after a near data-loss incident.

### Source excerpt

How we tracked down petabytes of orphaned objects in our ClickHouse clusters, recovered tens of thousands of dollars in monthly storage costs, and hardened our cleanup and recovery procedures after almost losing data along the way.

## Zero downtime Upgrade: Yelp's Cassandra 4.x Upgrade Story

DevFeed: [Zero downtime Upgrade: Yelp's Cassandra 4.x Upgrade Story](<https://devfeed.tech/articles/zero-downtime-upgrade-yelp-s-cassandra-4-x-upgrade-story-27423.md>)

Original publisher: [Read original article](<https://engineeringblog.yelp.com/2026/04/zero-downtime-upgrade-yelp-cassandra-upgrade-story.html>)

Author: Mark Surnin and Muhammad Junaid Muzammil, Software Engineer

Published: 2026-04-07T00:00:00Z

Content type: article

Language: en

Sources: [Yelp](<https://devfeed.tech/sources/yelp.md>)

Topics: [Apache Cassandra](<https://devfeed.tech/topics/cassandra.md>), [upgrade](<https://devfeed.tech/topics/upgrade.md>), [NoSQL](<https://devfeed.tech/topics/nosql.md>), [Database](<https://devfeed.tech/topics/database.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [apache](<https://devfeed.tech/tags/apache.md>), [cassandra](<https://devfeed.tech/tags/cassandra.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [downtime](<https://devfeed.tech/tags/downtime.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [nosql](<https://devfeed.tech/tags/nosql.md>), [observability](<https://devfeed.tech/tags/observability.md>), [reliability-engineering](<https://devfeed.tech/tags/reliability-engineering.md>), [upgrade](<https://devfeed.tech/tags/upgrade.md>)

### AI overview

Yelp's Database Reliability Engineering team describes upgrading more than a thousand Cassandra nodes from 3.11 to 4.1 on Kubernetes without downtime. The article covers the motivation, expected reliability and performance improvements, operational guardrails, certificate handling, repairs, logging, and compatibility work for related components.

### Source excerpt

The Database Reliability Engineering team at Yelp seamlessly upgraded more than a thousand Cassandra nodes with zero downtime. This post takes you behind the scenes of our upgrade strategy, from planning sessions to flawless rollouts. Background Motivation Apache Cassandra is a distributed wide-column NoSQL datastore and is used widely at Yelp for storing both primary and derived data. Yelp orchestrates Cassandra clusters on Kubernetes with the help of operators, as explained in our operator overview post. Upgrading from Cassandra 3.11 to 4.1 offered several observability and reliability improvements, in addition to performance gains. Based on public benchmarks, we expected to...

## Cockroach Labs announces CockroachDB capabilities for AI agents

DevFeed: [Cockroach Labs announces CockroachDB capabilities for AI agents](<https://devfeed.tech/articles/cockroachdb-is-built-for-ai-agents-23759.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/cockroachdb-ai-agents-agent-ready-database>)

Author: Lakshmi Kannan

Published: 2026-03-25T00:00:00Z

Content type: release

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [CockroachDB](<https://devfeed.tech/topics/cockroachdb.md>), [Cockroach Labs](<https://devfeed.tech/topics/cockroach-labs.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [Agent Skills](<https://devfeed.tech/topics/agent-skills.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [agent-skills](<https://devfeed.tech/tags/agent-skills.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [auditability](<https://devfeed.tech/tags/auditability.md>), [backend](<https://devfeed.tech/tags/backend.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [cockroach-labs](<https://devfeed.tech/tags/cockroach-labs.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [database](<https://devfeed.tech/tags/database.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [operations](<https://devfeed.tech/tags/operations.md>), [provisioning](<https://devfeed.tech/tags/provisioning.md>), [scale](<https://devfeed.tech/tags/scale.md>), [upgrades](<https://devfeed.tech/tags/upgrades.md>)

### AI overview

Cockroach Labs announces capabilities intended to make CockroachDB agent-ready, including a fully managed MCP Server, a redesigned ccloud CLI, and a public library of CockroachDB Agent Skills. The article describes requirements such as structured interfaces, scoped permissions, and auditability for agents operating across the database lifecycle.

### Source excerpt

Today Cockroach Labs is announcing new capabilities that make CockroachDB agent-ready, giving AI agents a secure, structured way to work with your database.

## Operating Trino at Scale With Trino Gateway

DevFeed: [Operating Trino at Scale With Trino Gateway](<https://devfeed.tech/articles/operating-trino-at-scale-with-trino-gateway-19736.md>)

Original publisher: [Read original article](<https://medium.com/expedia-group-tech/operating-trino-at-scale-with-trino-gateway-41824af788de?source=rss----38998a53046f---4>)

Author: Prakhar Sapre

Published: 2026-03-24T12:01:00Z

Content type: article

Language: en

Sources: [Expedia](<https://devfeed.tech/sources/expedia.md>)

Topics: [gateway](<https://devfeed.tech/topics/gateway.md>), [Load Balancing](<https://devfeed.tech/topics/load-balancing.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>), [data analytics](<https://devfeed.tech/topics/data-analytics.md>), [Authentication](<https://devfeed.tech/topics/authentication.md>), [SQL](<https://devfeed.tech/topics/sql.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [authentication](<https://devfeed.tech/tags/authentication.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [concurrency](<https://devfeed.tech/tags/concurrency.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [downtime](<https://devfeed.tech/tags/downtime.md>), [gateway](<https://devfeed.tech/tags/gateway.md>), [load-balancer](<https://devfeed.tech/tags/load-balancer.md>), [load-balancing](<https://devfeed.tech/tags/load-balancing.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [sql](<https://devfeed.tech/tags/sql.md>), [trino](<https://devfeed.tech/tags/trino.md>), [trino-gateway](<https://devfeed.tech/tags/trino-gateway.md>), [trinos](<https://devfeed.tech/tags/trinos.md>)

### AI overview

This article explains how Trino Gateway routes queries across multiple Trino clusters and centralizes routing, authentication, load balancing, monitoring, and cluster management. It describes the project's origins as Presto Gateway at Lyft and its role in supporting larger analytics platforms with more complex workloads and higher concurrency.

### Source excerpt

Expedia Group Technology -- DataWorkload-aware routing for TrinoPhoto by Joseph Barrientos on Unsplash Trino -- a fork of PrestoSQL -- is a powerful tool in modern data analytics, enabling organizations to query large datasets quickly and efficiently. As a distributed SQL query engine, Trino provides fast, scalable insights without requiring data relocation. While Trino is robust on its own, its capabilities are further enhanced when paired with a Gateway, which introduces features such as query routing, strong security, and streamlined cluster management. A brief overview The Gateway project originated at Lyft as Presto Gateway, serving as a proxy and load balancer for PrestoDB. It was later forked and integrated into the Trino ecosystem, with contributions from various organizations and the open-source community. The Gateway serves as a central point for managing and routing queries, providing a unified interface for users and administrators. As organizations scale their analytics platforms, they often encounter challenges such as increased query complexity, higher concurrency, and the need for specialized cluster configurations. Directing users to specific cluster endpoints becomes impractical as the user base grows. A Gateway addresses these challenges by routing queries to the most appropriate clusters based on workload, improving efficiency and responsiveness. The Gateway acts as a vital intermediary between users and the Trino query engine. By abstracting the complexities of distributed query execution, it manages critical functions such as routing, authentication, and load balancing across diverse backend clusters. This ensures that queries are efficiently directed to the optimal processing cluster. With an intuitive user interface, the Gateway transforms what was once a convoluted process into a manageable and transparent experience empowering administrators with real-time insights and precise control over their backend cluster infrastructure. Whether it's mon

## Replica management in dedicated clusters is now available via API

DevFeed: [Replica management in dedicated clusters is now available via API](<https://devfeed.tech/articles/cluster-management-is-now-scriptable-18450.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/cluster-management-api>)

Author: Adrián Miralles

Published: 2026-03-19T00:00:00Z

Content type: release

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [API](<https://devfeed.tech/topics/api.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>)

### AI overview

Replica management for dedicated clusters is now available through an API, enabling users to rebalance weights and add or remove replicas programmatically.

### Source excerpt

Replica management in dedicated clusters is now available via API. Rebalance weights, add replicas, remove replicas. All scriptable.

## A redesigned Tinybird UI for developers and operators

DevFeed: [A redesigned Tinybird UI for developers and operators](<https://devfeed.tech/articles/a-redesigned-tinybird-ui-for-developers-and-operators-18580.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/new-ui>)

Author: Nuria Mediavilla, Julia Vallina

Published: 2026-03-11T12:00:00Z

Content type: release

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [ui](<https://devfeed.tech/topics/ui.md>)

Tags: [clusters](<https://devfeed.tech/tags/clusters.md>), [developers](<https://devfeed.tech/tags/developers.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [ui](<https://devfeed.tech/tags/ui.md>)

### AI overview

Tinybird has redesigned its UI to provide a faster, information-dense way for developers and operators to manage workspaces and clusters.

### Source excerpt

A faster, information-dense UI to operate your Tinybird workspaces and clusters.

## Worth Reading: Modern Forwarding Architectures

DevFeed: [Worth Reading: Modern Forwarding Architectures](<https://devfeed.tech/articles/worth-reading-modern-forwarding-architectures-11335.md>)

Original publisher: [Read original article](<https://blog.ipspace.net/2026/02/worth-reading-modern-forwarding-architectures/>)

Published: 2026-02-17T06:55:00Z

Content type: opinion

Language: en

Sources: [ipSpace.net blog](<https://devfeed.tech/sources/ipspace-net-blog.md>)

Topics: [networking](<https://devfeed.tech/topics/networking.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [article](<https://devfeed.tech/tags/article.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [networking](<https://devfeed.tech/tags/networking.md>), [worth-reading](<https://devfeed.tech/tags/worth-reading.md>)

### AI overview

The article recommends an introduction to modern forwarding architectures, including networking infrastructure for AI clusters and distributed cell-based fabrics. It also criticizes the referenced article's mention of OpenFlow and other concepts.

### Source excerpt

Ignoring the obligatory misguided mention of OpenFlow and a few other unicorns, I found this article to be a nice introduction to modern forwarding architectures, including networking infrastructure for AI clusters and distributed cell-based fabrics.

## Why these three fintech companies scaled with distributed SQL

DevFeed: [Why these three fintech companies scaled with distributed SQL](<https://devfeed.tech/articles/why-these-three-fintech-companies-scaled-with-distributed-sql-23783.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/fintech-companies-scaled-distributed-sql>)

Author: Becca Weng

Published: 2026-02-02T00:00:00Z

Content type: article

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [CockroachDB](<https://devfeed.tech/topics/cockroachdb.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Microservices](<https://devfeed.tech/topics/microservices.md>), [migration](<https://devfeed.tech/topics/migration.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [backend](<https://devfeed.tech/tags/backend.md>), [caching](<https://devfeed.tech/tags/caching.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [compatibility](<https://devfeed.tech/tags/compatibility.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [databases](<https://devfeed.tech/tags/databases.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [financial](<https://devfeed.tech/tags/financial.md>), [fintech](<https://devfeed.tech/tags/fintech.md>), [growth](<https://devfeed.tech/tags/growth.md>), [india](<https://devfeed.tech/tags/india.md>), [migration](<https://devfeed.tech/tags/migration.md>), [multi-cloud](<https://devfeed.tech/tags/multi-cloud.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [payments](<https://devfeed.tech/tags/payments.md>), [scale](<https://devfeed.tech/tags/scale.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

This article describes how fintech platforms face growing transaction volumes, tighter regulations, and demanding financial workloads. Leaders from Groww, Global Payments, and Yubi explain their experiences, while the detailed case study of Groww covers its move from MySQL to CockroachDB to support horizontal scalability, multi-region resiliency, transactional guarantees, compliance, and growth with minimal application changes.

### Source excerpt

From retail investing and global card payments to large-scale lending, fintech platforms are being pushed to their limits. User expectations are rising, regulations are tightening, and transaction volumes are growing faster and more unpredictably than ever before. Across these domains, one theme is emerging clearly: traditional databases struggle to keep up with modern financial workloads at scale.

## An AI Engineer's Guide To Choosing GPUs

DevFeed: [An AI Engineer's Guide To Choosing GPUs](<https://devfeed.tech/articles/an-ai-engineer-s-guide-to-choosing-gpus-35011.md>)

Original publisher: [Read original article](<https://read.theaimerge.com/p/an-ai-engineers-guide-to-choosing>)

Author: Alex Razvant

Published: 2025-12-07T14:02:40Z

Content type: tutorial

Language: en

Sources: [Neural Bits](<https://devfeed.tech/sources/neural-bits.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Blackwell](<https://devfeed.tech/topics/blackwell.md>), [Hopper](<https://devfeed.tech/topics/hopper.md>), [lora](<https://devfeed.tech/topics/lora.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>), [Kernel](<https://devfeed.tech/topics/kernel.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-engineer](<https://devfeed.tech/tags/ai-engineer.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [deep-dive](<https://devfeed.tech/tags/deep-dive.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [guide](<https://devfeed.tech/tags/guide.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [hopper](<https://devfeed.tech/tags/hopper.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [ml](<https://devfeed.tech/tags/ml.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [pcie](<https://devfeed.tech/tags/pcie.md>), [software](<https://devfeed.tech/tags/software.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

A technical guide to choosing NVIDIA GPUs for AI workloads. It explains how GPU microarchitecture, memory subsystems, form factors, and interconnects affect capabilities, scaling, training, and inference, and compares consumer and data-center GPUs.

### Source excerpt

A deep dive on technical Hardware and Software details of NVIDIA GPUs for AI Workloads.

[Next page](<https://devfeed.tech/tags/clusters.md?cursor=WyIyMDI1LTEyLTA3VDE0OjAyOjQwKzAwOjAwIiwgImZhMzRjNzAxLWFkZjItNDQ0NC1hMGQzLTE1MmUxZmU0YzM1NiJd>)