# AI Training

Published articles for AI Training.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers

DevFeed: [Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers](<https://devfeed.tech/articles/cloudflare-just-gave-ai-training-bots-the-middle-finger-31387.md>)

Original publisher: [Read original article](<https://webdesignerdepot.com/cloudflare-just-gave-ai-training-bots-the-middle-finger/>)

Author: Alex Harper

Published: 2026-09-16T17:18:57Z

Content type: news

Language: en

Sources: [Web Designer Depot](<https://devfeed.tech/sources/web-designer-depot.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Google Search](<https://devfeed.tech/topics/google-search.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-tech](<https://devfeed.tech/tags/ai-tech.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [content-protection](<https://devfeed.tech/tags/content-protection.md>), [future-of-the-web](<https://devfeed.tech/tags/future-of-the-web.md>), [google-extended](<https://devfeed.tech/tags/google-extended.md>), [google-search](<https://devfeed.tech/tags/google-search.md>), [googlebot](<https://devfeed.tech/tags/googlebot.md>), [openai](<https://devfeed.tech/tags/openai.md>), [publishers](<https://devfeed.tech/tags/publishers.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [search](<https://devfeed.tech/tags/search.md>), [search-engines](<https://devfeed.tech/tags/search-engines.md>), [web](<https://devfeed.tech/tags/web.md>), [web-design](<https://devfeed.tech/tags/web-design.md>), [web-development](<https://devfeed.tech/tags/web-development.md>), [web-publishing](<https://devfeed.tech/tags/web-publishing.md>), [web-scraping](<https://devfeed.tech/tags/web-scraping.md>), [website-traffic](<https://devfeed.tech/tags/website-traffic.md>)

### AI overview

Cloudflare launched a Disallow AI Training setting that lets website owners allow traditional search crawlers while blocking training-only crawlers from companies including Amazon, Anthropic, Meta, and OpenAI. The article notes that robots.txt depends on crawler compliance and that blocking Google-Extended does not remove content from Google Search features such as AI Overviews or AI Mode.

### Source excerpt

Cloudflare just gave website owners a new weapon against AI crawlers: keep the search traffic, block the AI training. After years of watching bots consume the web's content, publishers finally have an easier way to tell AI companies where to go.

## Building a reliable cloud native foundation for distributed AI training

DevFeed: [Building a reliable cloud native foundation for distributed AI training](<https://devfeed.tech/articles/building-a-reliable-cloud-native-foundation-for-distributed-ai-training-4603.md>)

Original publisher: [Read original article](<https://www.cncf.io/blog/2026/09/11/building-a-reliable-cloud-native-foundation-for-distributed-ai-training/>)

Author: Abhi Kulkarni and Shishir Jindal, Atlassian

Published: 2026-09-11T11:00:00Z

Content type: article

Language: en

Sources: [Cloud Native Computing Foundation](<https://devfeed.tech/sources/cloud-native-computing-foundation.md>)

Topics: [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Network design](<https://devfeed.tech/topics/network-design.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [ai-training](<https://devfeed.tech/tags/ai-training.md>), [blog](<https://devfeed.tech/tags/blog.md>), [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

The article explains how to make multi-node AI training reliable by treating inter-node communication, shared storage, hardware placement, network topology, and validation as platform concerns. It identifies RDMA for GPU-node communication and Lustre for concurrent training-data and checkpoint access.

### Source excerpt

AI workloads are changing what platform teams need from infrastructure. Provisioning GPUs and standing up a cluster no longer makes a platform "AI-ready." Once training spans more than one node, the bottlenecks show up in places...

## Physical AI Takes the Wheel: How the World's Robotaxi Leaders Are Building With NVIDIA Technologies

DevFeed: [Physical AI Takes the Wheel: How the World's Robotaxi Leaders Are Building With NVIDIA Technologies](<https://devfeed.tech/articles/physical-ai-takes-the-wheel-how-the-world-s-robotaxi-leaders-are-building-with-nvidia-technologies-6959.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/robotaxi-leaders-full-stack-open-platform/>)

Author: Ali Kani

Published: 2026-09-10T16:00:04Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [autonomous-vehicles](<https://devfeed.tech/tags/autonomous-vehicles.md>), [cosmos](<https://devfeed.tech/tags/cosmos.md>), [customer-stories](<https://devfeed.tech/tags/customer-stories.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [driving](<https://devfeed.tech/tags/driving.md>), [mobility](<https://devfeed.tech/tags/mobility.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [nvidia-dgx](<https://devfeed.tech/tags/nvidia-dgx.md>), [nvidia-drive](<https://devfeed.tech/tags/nvidia-drive.md>), [nvidia-halos](<https://devfeed.tech/tags/nvidia-halos.md>), [omniverse](<https://devfeed.tech/tags/omniverse.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [simulation-and-design](<https://devfeed.tech/tags/simulation-and-design.md>)

### AI overview

NVIDIA describes an open robotaxi platform for training AI driving models, simulation and safety validation, and real-time in-vehicle computing.

### Source excerpt

The global robotaxi market -- physical AI's first commercial breakthrough -- is projected to reach $400 billion by 2035, with over 6 million commercial vehicles in operation as driverless fleets are already moving people through some of the world's busiest and most complex streets. Deploying a driverless vehicle is one challenge. Scaling a fleet is [...]

## China Merchants Bank Wins CNCF End User Case Study Contest for Unifying AI Training and Inference on Kubernetes

DevFeed: [China Merchants Bank Wins CNCF End User Case Study Contest for Unifying AI Training and Inference on Kubernetes](<https://devfeed.tech/articles/china-merchants-bank-wins-cncf-end-user-case-study-contest-for-unifying-ai-training-and-inference-on-kubernetes-4594.md>)

Original publisher: [Read original article](<https://www.cncf.io/announcements/2026/09/07/china-merchants-bank-wins-cncf-end-user-case-study-contest-for-unifying-ai-training-and-inference-on-kubernetes/>)

Author: Haley White

Published: 2026-09-08T01:54:31Z

Content type: news

Language: en

Sources: [Cloud Native Computing Foundation](<https://devfeed.tech/sources/cloud-native-computing-foundation.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [kueue](<https://devfeed.tech/topics/kueue.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Cloud Native Ecosystem](<https://devfeed.tech/topics/cloud-native-ecosystem.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [case-study](<https://devfeed.tech/tags/case-study.md>), [china](<https://devfeed.tech/tags/china.md>), [compute](<https://devfeed.tech/tags/compute.md>), [cost](<https://devfeed.tech/tags/cost.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [inference](<https://devfeed.tech/tags/inference.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kueue](<https://devfeed.tech/tags/kueue.md>), [lora](<https://devfeed.tech/tags/lora.md>)

### AI overview

China Merchants Bank won a CNCF case-study contest for a Kubernetes-based AI platform that shares nearly 10,000 accelerator cards across training, fine-tuning, and online inference. The bank reports increased average accelerator utilization and lower inference costs.

### Source excerpt

New cloud native platform lifted average accelerator compute utilization from 35% to more than 60% and cut inference cost per 1 million tokens by more than 60% Key Highlights SHANGHAI, China - KubeCon + CloudNativeCon +...

## The Pulse: Meta wanted to reduce teams by 60% because of AI

DevFeed: [The Pulse: Meta wanted to reduce teams by 60% because of AI](<https://devfeed.tech/articles/the-pulse-meta-wanted-to-reduce-teams-by-60-because-of-ai-40924.md>)

Original publisher: [Read original article](<https://blog.pragmaticengineer.com/the-pulse-meta-wanted-to-reduce-teams-by-60-because-of-ai/>)

Author: Ivan Klaric

Published: 2026-09-03T17:01:49Z

Content type: opinion

Language: en

Sources: [The Pragmatic Engineer](<https://devfeed.tech/sources/the-pragmatic-engineer-2.md>)

Topics: [Meta](<https://devfeed.tech/topics/meta.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Strategy](<https://devfeed.tech/topics/ai-strategy.md>), [Job](<https://devfeed.tech/topics/job.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [company](<https://devfeed.tech/tags/company.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [facebook](<https://devfeed.tech/tags/facebook.md>), [instagram](<https://devfeed.tech/tags/instagram.md>), [job-cuts](<https://devfeed.tech/tags/job-cuts.md>), [leadership](<https://devfeed.tech/tags/leadership.md>), [meta](<https://devfeed.tech/tags/meta.md>), [organization](<https://devfeed.tech/tags/organization.md>), [reporting](<https://devfeed.tech/tags/reporting.md>), [teams](<https://devfeed.tech/tags/teams.md>)

### AI overview

The article examines Meta's proposed Project Organization Transformation, which envisioned reducing many existing teams by 60% as AI took over more daily work. It also discusses earlier layoffs, reassignment of engineers to AI training, low morale, and operational outages, while noting that the larger layoff plan did not proceed.

### Source excerpt

An in-depth report by Reuters details how Meta's leadership decided to slash team sizes by 60%. Zuckerberg changed his mind, and now the company is stuck with low morale and culture turned mercenary.

## Delivering Vera: NVIDIA's First CPU Built for Agents Is Shipping Now

DevFeed: [Delivering Vera: NVIDIA's First CPU Built for Agents Is Shipping Now](<https://devfeed.tech/articles/delivering-vera-nvidia-s-first-cpu-built-for-agents-is-shipping-now-6962.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/vera-cpu-delivery/>)

Author: Ian Finder

Published: 2026-08-27T13:00:17Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [Vera CPU](<https://devfeed.tech/topics/vera-cpu.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [NVIDIA Vera](<https://devfeed.tech/topics/nvidia-vera.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [amazon-ec2](<https://devfeed.tech/tags/amazon-ec2.md>), [aws](<https://devfeed.tech/tags/aws.md>), [corporate](<https://devfeed.tech/tags/corporate.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [rubin-gpu](<https://devfeed.tech/tags/rubin-gpu.md>), [vera-cpu](<https://devfeed.tech/tags/vera-cpu.md>)

### AI overview

NVIDIA is shipping Vera, a CPU designed for agentic AI workloads, to AWS and other organizations across the AI ecosystem. The article describes Vera's 88 custom Olympus cores, 1.2TB/s memory bandwidth, and up to 1.8x faster per-core performance on agentic AI workloads, alongside its role in supporting concurrent CPU-intensive tasks such as tool calls, orchestration, retrieval, software testing, and data analysis.

### Source excerpt

NVIDIA Vice President of Hyperscale and HPC Ian Buck hand-delivers Vera CPU systems across the AI ecosystem as Vera begins shipping at scale.

## How Adobe Reduced GPU Idle Time in Generative AI Training Through Faster Data Access and Checkpointing

DevFeed: [How Adobe Reduced GPU Idle Time in Generative AI Training Through Faster Data Access and Checkpointing](<https://devfeed.tech/articles/why-your-gpu-is-sitting-idle-the-data-pipeline-problem-no-one-talks-about-12325.md>)

Original publisher: [Read original article](<https://www.backblaze.com/blog/why-your-gpu-is-sitting-idle-the-data-pipeline-problem-no-one-talks-about/>)

Author: Maddie Presland

Published: 2026-08-25T15:22:30Z

Content type: article

Language: en

Sources: [Backblaze Blog | Cloud Storage & Cloud Backup](<https://devfeed.tech/sources/backblaze-blog-cloud-storage-cloud-backup.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [networking](<https://devfeed.tech/topics/networking.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [b2cloud](<https://devfeed.tech/tags/b2cloud.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-storage](<https://devfeed.tech/tags/cloud-storage.md>), [compute](<https://devfeed.tech/tags/compute.md>), [ethernet](<https://devfeed.tech/tags/ethernet.md>), [featured](<https://devfeed.tech/tags/featured.md>), [featured-cloud-storage](<https://devfeed.tech/tags/featured-cloud-storage.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [model-training](<https://devfeed.tech/tags/model-training.md>), [networking](<https://devfeed.tech/tags/networking.md>)

### AI overview

Adobe's generative AI training pipeline left roughly two-thirds of GPU time waiting for data. The article attributes the waste to storage and retrieval bottlenecks, networking limits, and checkpointing overhead, and describes Adobe's use of a high-performance networking fabric and fragmented checkpoint storage to reduce delays.

### Source excerpt

Adobe's experience reveals why GPUs sit idle during AI model training: slow storage, insufficient throughput, and uneven workloads. Learn how storage bottlenecks and uneven data loading leave expensive GPUs idle--and how always-hot, high-throughput object storage keeps AI training pipelines running efficiently at scale while reducing wasted compute costs. The post Why Your GPU Is Sitting Idle: The Data Pipeline Problem No One Talks About appeared first on Backblaze Blog | Cloud Storage & Cloud Backup

## Say it once: Introducing Bot Preference Sync

DevFeed: [Say it once: Introducing Bot Preference Sync](<https://devfeed.tech/articles/say-it-once-introducing-bot-preference-sync-107.md>)

Original publisher: [Read original article](<https://blog.cloudflare.com/bot-preference-sync/>)

Author: Jin-Hee Lee

Published: 2026-08-21T23:19:57Z

Content type: article

Language: en

Sources: [Cloudflare Blog](<https://devfeed.tech/sources/cloudflare-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-bots](<https://devfeed.tech/tags/ai-bots.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [bots](<https://devfeed.tech/tags/bots.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [network-services](<https://devfeed.tech/tags/network-services.md>), [product-news](<https://devfeed.tech/tags/product-news.md>), [robots](<https://devfeed.tech/tags/robots.md>), [search](<https://devfeed.tech/tags/search.md>), [sync](<https://devfeed.tech/tags/sync.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Cloudflare introduces Bot Preference Sync, which updates robots.txt preferences to match configured AI bot policies for Search, Agent, and Training traffic.

### Source excerpt

Cloudflare's new Bot Preference Sync automatically aligns your robots.txt file with your AI bot policies for Search, Agent, and Training. Easily manage which bots access your content without maintaining static files.

## Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

DevFeed: [Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72](<https://devfeed.tech/articles/setting-a-world-record-for-moe-pre-training-on-nvidia-gb300-nvl72-6939.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/setting-a-world-record-for-moe-pre-training-on-nvidia-gb300-nvl72/>)

Author: Kirthi Devleker

Published: 2026-07-21T18:30:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>), [networking](<https://devfeed.tech/topics/networking.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-platform](<https://devfeed.tech/tags/ai-platform.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [collective](<https://devfeed.tech/tags/collective.md>), [communication](<https://devfeed.tech/tags/communication.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [featured](<https://devfeed.tech/tags/featured.md>), [frontier-model](<https://devfeed.tech/tags/frontier-model.md>), [gb300-nvl72](<https://devfeed.tech/tags/gb300-nvl72.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm-techniques](<https://devfeed.tech/tags/llm-techniques.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [megatron](<https://devfeed.tech/tags/megatron.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [moe](<https://devfeed.tech/tags/moe.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [performance](<https://devfeed.tech/tags/performance.md>), [top-stories](<https://devfeed.tech/tags/top-stories.md>), [train](<https://devfeed.tech/tags/train.md>), [training-ai-models](<https://devfeed.tech/tags/training-ai-models.md>)

### AI overview

The article explains how NVIDIA GB300 NVL72 achieved a world record for DeepSeek-V3 671B mixture-of-experts pre-training. It focuses on the communication demands of MoE models, including all-to-all traffic between GPUs, and the need for tightly coupled scale-up and predictable scale-out networking to sustain delivered training performance.

### Source excerpt

Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token...

## Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism

DevFeed: [Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism](<https://devfeed.tech/articles/enhancing-goodput-in-large-scale-llm-training-with-nonuniform-tensor-parallelism-6815.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/enhancing-goodput-in-large-scale-llm-training-with-nonuniform-tensor-parallelism/>)

Author: Michelle Horton

Published: 2026-07-06T21:44:23Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [LLM Techniques](<https://devfeed.tech/topics/llm-techniques.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Network](<https://devfeed.tech/topics/network.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [availability](<https://devfeed.tech/tags/availability.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llm-techniques](<https://devfeed.tech/tags/llm-techniques.md>), [mlops](<https://devfeed.tech/tags/mlops.md>), [network](<https://devfeed.tech/tags/network.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [performance](<https://devfeed.tech/tags/performance.md>), [scale](<https://devfeed.tech/tags/scale.md>), [training](<https://devfeed.tech/tags/training.md>), [work](<https://devfeed.tech/tags/work.md>)

### AI overview

This article explains how Nonuniform Tensor Parallelism can improve Goodput in large-scale LLM training by adapting tensor parallelism to changing GPU availability and overlapping data resharding. The experimental approach aims to reduce interruptions, lost throughput, and computational waste in tightly interconnected GPU clusters.

### Source excerpt

Training LLMs at massive scale brings unique infrastructure challenges, especially as jobs span thousands of GPUs and run for extended periods. The longer these...

## Decoupled DiLoCo: A new frontier for resilient, distributed AI training

DevFeed: [Decoupled DiLoCo: A new frontier for resilient, distributed AI training](<https://devfeed.tech/articles/decoupled-diloco-a-new-frontier-for-resilient-distributed-ai-training-6145.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/decoupled-diloco/>)

Author: Arthur Douillard and the DiLoCo team

Published: 2026-04-22T10:20:03Z

Content type: article

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [data centers](<https://devfeed.tech/topics/data-centers.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [communication](<https://devfeed.tech/tags/communication.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-centers](<https://devfeed.tech/tags/data-centers.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [networking](<https://devfeed.tech/tags/networking.md>), [performance](<https://devfeed.tech/tags/performance.md>), [research](<https://devfeed.tech/tags/research.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [scale](<https://devfeed.tech/tags/scale.md>), [tpu](<https://devfeed.tech/tags/tpu.md>), [train](<https://devfeed.tech/tags/train.md>)

### AI overview

Google describes Decoupled DiLoCo, a resilient distributed training architecture that trained a 12-billion-parameter model across four U.S. regions using achievable wide-area connectivity. By overlapping communication with computation, it was reported to be more than 20 times faster than conventional synchronization and could continue operating despite failures.

### Source excerpt

Google's new distributed architecture keeps AI training runs on track across distant data centers, with exceptional efficiency - even when hardware fails.

## A few updates: the DGX Spark giveaway, new GPU inference slides, and my Vision AI Course

DevFeed: [A few updates: the DGX Spark giveaway, new GPU inference slides, and my Vision AI Course](<https://devfeed.tech/articles/a-few-updates-the-dgx-spark-giveaway-new-gpu-inference-slides-and-my-vision-ai-course-35009.md>)

Original publisher: [Read original article](<https://read.theaimerge.com/p/a-few-updates-the-dgx-spark-giveaway>)

Author: Alex Razvant

Published: 2026-04-11T13:03:07Z

Content type: opinion

Language: en

Sources: [Neural Bits](<https://devfeed.tech/sources/neural-bits.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [DGX Spark](<https://devfeed.tech/topics/dgx-spark.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [FIRST](<https://devfeed.tech/topics/first.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [models](<https://devfeed.tech/tags/models.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [presentation](<https://devfeed.tech/tags/presentation.md>)

### AI overview

The author shares updates about the completed DGX Spark giveaway organized with NVIDIA and a GPU and AI slide deck that is still in development. The presentation aims to explain how GPUs are used at scale for AI training and inference.

### Source excerpt

What I've been working on behind the scenes, and what I'm preparing next.

## Protecting cities with AI-driven flash flood forecasting

DevFeed: [Protecting cities with AI-driven flash flood forecasting](<https://devfeed.tech/articles/protecting-cities-with-ai-driven-flash-flood-forecasting-6850.md>)

Original publisher: [Read original article](<https://research.google/blog/protecting-cities-with-ai-driven-flash-flood-forecasting/>)

Published: 2026-03-12T13:03:15Z

Content type: news

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [Machine Learning & Artificial Intelligence](<https://devfeed.tech/topics/machine-learning-artificial-intelligence.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [climate-sustainability](<https://devfeed.tech/tags/climate-sustainability.md>), [data](<https://devfeed.tech/tags/data.md>), [earth-ai](<https://devfeed.tech/tags/earth-ai.md>), [flash](<https://devfeed.tech/tags/flash.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [news](<https://devfeed.tech/tags/news.md>), [open-source-models-datasets](<https://devfeed.tech/tags/open-source-models-datasets.md>), [research](<https://devfeed.tech/tags/research.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [sustainability](<https://devfeed.tech/tags/sustainability.md>)

### AI overview

Google Research announces urban flash flood forecasts that use an AI-powered methodology to provide up to 24 hours of advance warning. The article describes the forecasting challenge posed by rapidly developing floods, limited ground-truth data, and the need to expand early-warning coverage for vulnerable communities.

### Source excerpt

Climate & Sustainability

## How Poolside is using ClickHouse to build next-gen AI for software development

DevFeed: [How Poolside is using ClickHouse to build next-gen AI for software development](<https://devfeed.tech/articles/how-poolside-is-using-clickhouse-to-build-next-gen-ai-for-software-development-5502.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/poolside-using-clickhouse-to-build-next-gen-ai-for-software-development>)

Author: ClickHouse

Published: 2025-03-03T00:00:00Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Developer Tools](<https://devfeed.tech/topics/developer-tools.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-development](<https://devfeed.tech/tags/ai-development.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>)

### AI overview

Poolside uses ClickHouse Cloud to query billions of records for analytics supporting the training and evaluation of AI models for software development. The faster queries help its team iterate on datasets, experiments, and model performance.

### Source excerpt

Read how Poolside is building next-generation AI for software engineering with ClickHouse Cloud at the core of their analytics workflows--querying billions of records in real time to analyze AI model performance and enable faster iteration.