# NVLink

NVIDIA's high-bandwidth, low-latency GPU-to-GPU interconnect and scale-up networking fabric for multi-GPU systems.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Cornelis and Delos Data propose open alternatives to Nvidia's NVLink for AI scale-up networking

DevFeed: [Cornelis and Delos Data propose open alternatives to Nvidia's NVLink for AI scale-up networking](<https://devfeed.tech/articles/ai-networking-startups-race-to-replace-nvidia-s-nvlink-26965.md>)

Original publisher: [Read original article](<https://www.theregister.com/systems/2026/09/15/ai-networking-startups-race-to-replace-nvidias-nvlink/5296672>)

Author: Tobias Mann

Published: 2026-09-15T20:56:48Z

Content type: news

Language: en

Sources: [www.theregister.com - Articles](<https://devfeed.tech/sources/www-theregister-com-articles.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [AI Networking](<https://devfeed.tech/topics/ai-networking.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-networking](<https://devfeed.tech/tags/ai-networking.md>), [cornelis-networks](<https://devfeed.tech/tags/cornelis-networks.md>), [delos-data](<https://devfeed.tech/tags/delos-data.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [network](<https://devfeed.tech/tags/network.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [performance](<https://devfeed.tech/tags/performance.md>), [software](<https://devfeed.tech/tags/software.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

The article reports that Cornelis Networks and Delos Data have entered the AI scale-up networking market with proposed alternatives to Nvidia's NVLink. Cornelis introduced its open Active Compute Fabric architecture, which combines programmable in-network compute with scale-up and scale-out networking, while the broader market is developing alternatives using protocols such as Ultra Ethernet and UALink.

### Source excerpt

Intel spin-off Cornelis and newcomer Delos Data pitch open alternatives for scaling AI beyond the rack

## AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

DevFeed: [AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories](<https://devfeed.tech/articles/ai-infra-summit-nvidia-vera-rubin-and-dsx-platform-advancements-showcase-energy-efficiencies-of-optimizing-tokens-per-watt-for-ai-factories-26942.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/>)

Author: NVIDIA Writers

Published: 2026-09-15T16:55:40Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [DSX](<https://devfeed.tech/topics/dsx.md>), [Vera Rubin](<https://devfeed.tech/topics/vera-rubin.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Conversational AI](<https://devfeed.tech/topics/conversational-ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [dsx](<https://devfeed.tech/tags/dsx.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infra](<https://devfeed.tech/tags/infra.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency-inference](<https://devfeed.tech/tags/low-latency-inference.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [nvidia-dsx](<https://devfeed.tech/tags/nvidia-dsx.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [nvidia-vera-rubin](<https://devfeed.tech/tags/nvidia-vera-rubin.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>)

### AI overview

NVIDIA's AI Infra Summit coverage describes collaborations and platform updates focused on improving AI factory efficiency. The article highlights Vera Rubin systems, DSX MaxLPS, Dynamo inference software, NVLink and networking technologies, including claims of up to 1.4x more tokens per megawatt through factory-wide power optimization.

### Source excerpt

Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech. Before a packed audience -- with more than 8,000 attendees this year, up from 3,500 last year -- [...]

## How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories

DevFeed: [How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories](<https://devfeed.tech/articles/how-nvidia-nvlink-6-delivers-multi-layer-resiliency-for-ai-factories-26914.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-nvidia-nvlink-6-delivers-multi-layer-resiliency-for-ai-factories/>)

Author: Elizabeth Goodman

Published: 2026-09-15T16:55:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [industry](<https://devfeed.tech/tags/industry.md>), [networking](<https://devfeed.tech/tags/networking.md>), [networking-communications](<https://devfeed.tech/tags/networking-communications.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [resiliency](<https://devfeed.tech/tags/resiliency.md>), [vera-rubin](<https://devfeed.tech/tags/vera-rubin.md>)

### AI overview

The article describes how NVIDIA NVLink 6 supports resiliency in large-scale AI factories. It explains that Vera Rubin NVL72 connects 72 Rubin GPUs into a single scale-up domain and outlines a multilayer approach using lossless networking, error correction, retry, flow control, and error containment to support continuous training and inference operations.

### Source excerpt

For operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale AI training, every GPU in the cluster...

## Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November

DevFeed: [Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November](<https://devfeed.tech/articles/fujitsu-monaka-server-brings-2nm-144-core-cpus-to-air-cooled-ai-inference-on-sale-in-november-17435.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/fujitsu-monaka-server-brings-2nm-144-core-cpus-to-air-cooled-ai-inference-on-sale-in-november>)

Author: Lyle Smith

Published: 2026-09-14T18:03:44Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [data centers](<https://devfeed.tech/topics/data-centers.md>), [Confidential Computing](<https://devfeed.tech/topics/confidential-computing.md>), [Arm](<https://devfeed.tech/topics/arm.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [arm](<https://devfeed.tech/tags/arm.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [data-centers](<https://devfeed.tech/tags/data-centers.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [fujitsu](<https://devfeed.tech/tags/fujitsu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>)

### AI overview

Fujitsu is introducing MONAKA Servers built around its 2nm FUJITSU-MONAKA processor for AI inference in air-cooled data centers. The servers offer up to 144 CPU cores, matrix instructions, SVE2 vector processing, hardware-level confidential computing, and planned NVLink Fusion integration with NVIDIA GPUs. Fujitsu claims higher inference throughput and reduced cooling power consumption, but the article notes that supporting benchmark details are unavailable.

### Source excerpt

Fujitsu is bringing its 2nm FUJITSU-MONAKA processor to AI infrastructure with a new server family designed to run AI inference in air-cooled data centers without requiring specialized liquid cooling. The MONAKA Server is designed, developed, and manufactured in Japan, with component and manufacturing traceability for sovereign AI deployments. The first MONAKA Servers will come in The post Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November appeared first on StorageReview.com.

## d-Matrix Joins the NVIDIA NVLink Fusion Platform

DevFeed: [d-Matrix Joins the NVIDIA NVLink Fusion Platform](<https://devfeed.tech/articles/d-matrix-joins-the-nvidia-nvlink-fusion-platform-14008.md>)

Original publisher: [Read original article](<https://www.servethehome.com/d-matrix-joins-the-nvidia-nvlink-fusion-platform/>)

Author: Cliff Robinson

Published: 2026-09-12T21:42:59Z

Content type: news

Language: en

Sources: [ServeTheHome](<https://devfeed.tech/sources/servethehome.md>)

Topics: [d-matrix](<https://devfeed.tech/topics/d-matrix.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [xpu](<https://devfeed.tech/topics/xpu.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [networking](<https://devfeed.tech/topics/networking.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [Spectrum-X](<https://devfeed.tech/topics/spectrum-x.md>)

Tags: [accelerators](<https://devfeed.tech/tags/accelerators.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-accelerator](<https://devfeed.tech/tags/ai-accelerator.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [d-matrix](<https://devfeed.tech/tags/d-matrix.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [ethernet](<https://devfeed.tech/tags/ethernet.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [server](<https://devfeed.tech/tags/server.md>), [xpu](<https://devfeed.tech/tags/xpu.md>)

### AI overview

d-Matrix and NVIDIA announced that d-Matrix will bring its next-generation XPUs to the NVLink Fusion platform. The integration is intended to support scaling from individual Raptor XPUs to larger rack-scale and clustered deployments for AI inference, alongside NVIDIA networking and CPU technologies.

### Source excerpt

d-Matrix and NVIDIA announced that d-Matrix will use NVLink Fusion to scale up and out with its next-gen Raptor AI accelerators The post d-Matrix Joins the NVIDIA NVLink Fusion Platform appeared first on ServeTheHome.

## The aircraft might not be flying, but the certificate has gone on vacation

DevFeed: [The aircraft might not be flying, but the certificate has gone on vacation](<https://devfeed.tech/articles/the-aircraft-might-not-be-flying-but-the-certificate-has-gone-on-vacation-8547.md>)

Original publisher: [Read original article](<https://www.theregister.com/offbeat/2026/09/12/the-aircraft-might-not-be-flying-but-the-certificate-has-gone-on-vacation/5295622>)

Author: Richard Speed

Published: 2026-09-12T09:00:00Z

Content type: news

Language: en

Sources: [www.theregister.com - Articles](<https://devfeed.tech/sources/www-theregister-com-articles.md>)

Topics: [Cybersecurity](<https://devfeed.tech/topics/cybersecurity.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Vibe coding](<https://devfeed.tech/topics/vibe-coding.md>), [ARKTunnel](<https://devfeed.tech/topics/arktunnel.md>), [Cybercrime](<https://devfeed.tech/topics/cybercrime.md>), [.NET](<https://devfeed.tech/topics/net.md>), [how to create smooth CSS transitions](<https://devfeed.tech/topics/how-to-create-smooth-css-transitions.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [bork](<https://devfeed.tech/tags/bork.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [malware](<https://devfeed.tech/tags/malware.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [offbeat](<https://devfeed.tech/tags/offbeat.md>), [phishing](<https://devfeed.tech/tags/phishing.md>), [security](<https://devfeed.tech/tags/security.md>), [tailwind](<https://devfeed.tech/tags/tailwind.md>), [ubuntu](<https://devfeed.tech/tags/ubuntu.md>)

### AI overview

A news roundup covering security incidents, AI-related developments, semiconductor infrastructure, phishing, ransomware, open-source software, and web development. The supplied title concerns an aircraft certificate, while the body mainly contains unrelated headlines.

### Source excerpt

Information is not forthcoming from this screen

## d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

DevFeed: [d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment](<https://devfeed.tech/articles/d-matrix-adopts-nvidia-nvlink-fusion-for-rack-scale-xpu-deployment-6947.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/>)

Author: Jesse Clayton

Published: 2026-09-10T13:00:21Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-platform](<https://devfeed.tech/tags/ai-platform.md>), [corporate](<https://devfeed.tech/tags/corporate.md>), [d-matrix](<https://devfeed.tech/tags/d-matrix.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [latency](<https://devfeed.tech/tags/latency.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [nvidia-vera-rubin](<https://devfeed.tech/tags/nvidia-vera-rubin.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [spectrum-x](<https://devfeed.tech/tags/spectrum-x.md>), [xpu](<https://devfeed.tech/tags/xpu.md>)

### AI overview

d-Matrix announced plans to use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs with NVIDIA AI infrastructure. The article describes using NVLink, Spectrum-X networking and MGX rack designs to support rack-scale, low-latency inference deployments.

### Source excerpt

AI inference chipmaker d-Matrix today announced it will use NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure platform -- joining a growing roster of ecosystem partners. By connecting Raptor to NVIDIA NVLink scale-up and Spectrum-X scale-out networking, the NVIDIA MGX rack architecture and the broader NVIDIA AI platform, NVLink Fusion gives [...]

## d-Matrix drinks the Nvidia Kool-Aid with NVLink Fusion and MGX rack designs

DevFeed: [d-Matrix drinks the Nvidia Kool-Aid with NVLink Fusion and MGX rack designs](<https://devfeed.tech/articles/d-matrix-drinks-the-nvidia-kool-aid-with-nvlink-fusion-and-mgx-rack-designs-8573.md>)

Original publisher: [Read original article](<https://www.theregister.com/systems/2026/09/10/d-matrix-drinks-the-nvidia-kool-aid-with-nvlink-fusion-and-mgx-rack-designs/5295403>)

Author: Tobias Mann

Published: 2026-09-10T13:00:00Z

Content type: news

Language: en

Sources: [www.theregister.com - Articles](<https://devfeed.tech/sources/www-theregister-com-articles.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [d-matrix](<https://devfeed.tech/topics/d-matrix.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [amazon](<https://devfeed.tech/tags/amazon.md>), [d-matrix](<https://devfeed.tech/tags/d-matrix.md>), [datacenter](<https://devfeed.tech/tags/datacenter.md>), [fujitsu](<https://devfeed.tech/tags/fujitsu.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [startup](<https://devfeed.tech/tags/startup.md>), [systems](<https://devfeed.tech/tags/systems.md>), [xpu](<https://devfeed.tech/tags/xpu.md>)

### AI overview

d-Matrix, an AI infrastructure startup, is described as joining other companies as an NVLink supporter. The headline also references NVLink Fusion and MGX rack designs.

### Source excerpt

AI infrastructure startup joins Qualcomm, Arm, Marvell, Amazon, Fujitsu, and MediaTek as NVLink true believers

## CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs

DevFeed: [CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs](<https://devfeed.tech/articles/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus-6789.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/>)

Author: Jonathan Bentz

Published: 2026-09-09T20:24:12Z

Content type: release

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [CUDA](<https://devfeed.tech/topics/cuda.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>)

Tags: [cli](<https://devfeed.tech/tags/cli.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [cuda-tile](<https://devfeed.tech/tags/cuda-tile.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [nsight-tools-compute](<https://devfeed.tech/tags/nsight-tools-compute.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [release](<https://devfeed.tech/tags/release.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [top-stories](<https://devfeed.tech/tags/top-stories.md>), [windows](<https://devfeed.tech/tags/windows.md>)

### AI overview

CUDA Toolkit 13.4 adds Windows on Arm support, preview support for the NVIDIA Rubin GPU architecture, and new GPU-sharing controls through MPS V3. It also introduces CUDA Compute Fabric Transport for data movement across NVIDIA NVLink fabric.

### Source excerpt

Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software...

## NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure

DevFeed: [NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure](<https://devfeed.tech/articles/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure-6903.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/>)

Author: Farshad Ghodsian

Published: 2026-08-26T21:06:58Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [accelerate](<https://devfeed.tech/tags/accelerate.md>), [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [integration](<https://devfeed.tech/tags/integration.md>), [memory](<https://devfeed.tech/tags/memory.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [networking-communications](<https://devfeed.tech/tags/networking-communications.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [performance](<https://devfeed.tech/tags/performance.md>), [platform](<https://devfeed.tech/tags/platform.md>), [scale](<https://devfeed.tech/tags/scale.md>), [support](<https://devfeed.tech/tags/support.md>), [top-stories](<https://devfeed.tech/tags/top-stories.md>)

### AI overview

NVIDIA NVLink Fusion connects custom XPUs and CPUs to NVIDIA's AI infrastructure platform, while NVHBM provides validated HBM base-die technology intended to increase memory bandwidth, save package area, and reduce power consumption. The article describes benefits for training and large-scale inference, including up to 30% more memory bandwidth per stack than standard HBM4e.

### Source excerpt

AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,...

## NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

DevFeed: [NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory](<https://devfeed.tech/articles/nvidia-nvlink-fusion-expands-with-nvhbm-custom-high-bandwidth-memory-6957.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/nvlink-fusion-nvhbm-custom-high-bandwidth-memory/>)

Author: Jesse Clayton

Published: 2026-08-26T21:05:30Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [amazon](<https://devfeed.tech/topics/amazon.md>)

Tags: [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [amazon](<https://devfeed.tech/tags/amazon.md>), [aws](<https://devfeed.tech/tags/aws.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [performance](<https://devfeed.tech/tags/performance.md>), [xpu](<https://devfeed.tech/tags/xpu.md>)

### AI overview

NVIDIA expands NVLink Fusion with NVHBM, a high-bandwidth memory technology for custom AI infrastructure. By moving the memory controller into the HBM base die, NVHBM is designed to provide up to 30% greater memory bandwidth, 15% lower HBM power consumption, and up to 25% more XPU compute-die area than standard HBM4E. Amazon's Annapurna Labs will be the first memory partner to work with NVIDIA on the technology, alongside collaboration on NVLink scale-up architecture for future AWS Trainium systems.

### Source excerpt

The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system. To help hyperscalers and AI innovators build the next generation [...]

## How XPUs Meet a World-Class AI Factory

DevFeed: [How XPUs Meet a World-Class AI Factory](<https://devfeed.tech/articles/how-xpus-meet-a-world-class-ai-factory-6958.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/nvlink-fusion-xpu-ai-factory/>)

Author: Jesse Clayton

Published: 2026-08-24T15:00:54Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Network](<https://devfeed.tech/topics/network.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ecosystem](<https://devfeed.tech/tags/ecosystem.md>), [ethernet](<https://devfeed.tech/tags/ethernet.md>), [gb300-nvl72](<https://devfeed.tech/tags/gb300-nvl72.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia-dsx](<https://devfeed.tech/tags/nvidia-dsx.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [software](<https://devfeed.tech/tags/software.md>), [time](<https://devfeed.tech/tags/time.md>), [xpu](<https://devfeed.tech/tags/xpu.md>)

### AI overview

The article explains how NVLink Fusion combines custom XPUs with NVIDIA's established AI infrastructure to help build semi-custom AI factories. It focuses on scale-up networking, performance, resiliency, telemetry, platform maturity, and the economics of large-scale AI workloads.

### Source excerpt

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. Hyperscalers and AI-native companies building custom XPUs must consider [...]

## Run Local Agentic AI Workflows with Meta's Muse Glimmer on NVIDIA

DevFeed: [Run Local Agentic AI Workflows with Meta's Muse Glimmer on NVIDIA](<https://devfeed.tech/articles/run-local-agentic-ai-workflows-with-meta-s-muse-glimmer-on-nvidia-6932.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/run-local-agentic-ai-workflows-with-metas-muse-glimmer-on-nvidia/>)

Author: Michelle Horton

Published: 2026-08-10T13:27:19Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Local AI](<https://devfeed.tech/topics/local-ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [NVIDIA DGX](<https://devfeed.tech/topics/nvidia-dgx.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Jetson](<https://devfeed.tech/topics/jetson.md>), [Automation](<https://devfeed.tech/topics/automation.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [dgx-station](<https://devfeed.tech/tags/dgx-station.md>), [edge-computing](<https://devfeed.tech/tags/edge-computing.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [jetson](<https://devfeed.tech/tags/jetson.md>), [local-ai](<https://devfeed.tech/tags/local-ai.md>), [nemoclaw](<https://devfeed.tech/tags/nemoclaw.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-dgx](<https://devfeed.tech/tags/nvidia-dgx.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [top-stories](<https://devfeed.tech/tags/top-stories.md>)

### AI overview

Meta's Muse Glimmer is a 30B open-weight dense model designed for local agentic AI workflows. With a 120K+ context window and performance of up to 20K tokens per second on a single GPU, it supports sustained, multi-step tool use and local processing of sensitive data.

### Source excerpt

Meta returns to the open source ecosystem with the release of Muse Glimmer, a 30B open-weight dense model with a 120K+ context window built for local AI...

## How DigitalOcean Served Kimi K3 on Day Zero

DevFeed: [How DigitalOcean Served Kimi K3 on Day Zero](<https://devfeed.tech/articles/under-the-hood-serving-kimi-k3-19944.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/serving-kimi-k3-inference-engine>)

Author: Shree Murthy

Published: 2026-07-30T17:10:40Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

DigitalOcean describes how it served the Kimi K3 model on its Inference Engine from day zero, including GPU selection, distributed serving with llm-d, vLLM tuning, and verification against Moonshot AI's benchmarks.

### Source excerpt

DigitalOcean launched Kimi K3 on day 0. It's already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a model this size running well on day zero took real work across several teams. Thanks to Moonshot AI, Inferact, RadixArk, NVIDIA, and AMD for the help getting there. Standing up a new model, integrating it into DigitalOcean's Inference Engine, and showcasing its unique attributes on day 0 takes three things: the right hardware, a tuned serving stack, and rigorous verification against Moonshot's own benchmarks. Here are the lessons we learned along the way: Hardware selection and implementation We selected NVIDIA HGX™ B300 and AMD Instinct™ MI350x GPUs to run K3 because these instances provide the memory capacity, FLOPs, and interconnect horsepower necessary for a model of K3's size and architecture. We built our distributed inference stack with llm-d because it includes native support for GPU type heterogeneity. This let us quickly onboard K3 to both AMD and NVIDIA platforms. Kimi K3 has roughly 2.78 trillion total parameters, 896 routed experts, and an attention stack that interleaves 69 Kimi Delta Attention (KDA) layers with 24 Gated Multi-head Latent Attention (MLA) layers. Kimi-K3 weights are ~1.56 TB in total, which requires about 195 GiB per GPU. Given such a large memory footprint for the weights alone, and a need to keep enough headroom for KV cache and activations, the practical unit of deployment is an 8x NVIDIA HGX B300 or AMD Instinct MI350X server. Both have 288GB of VRAM capacity, and after loading the weights, there is still some amount of practical memory left for the KV cache. Entire weights cannot be loaded on a single GPU. That's where the high-speed scaled-up NVIDIA's NVLink or AMD's Infinity Fabric is critical to ensure there is enough interconnect horsepower for bandwidth intensive, latency sensitive attention and expert parallel computations. Model

## AMD Instinct MI455X GPU Detailed for Rack-Scale AI Deployments

DevFeed: [AMD Instinct MI455X GPU Detailed for Rack-Scale AI Deployments](<https://devfeed.tech/articles/amd-s-instinct-mi455x-aiming-for-the-sun-13987.md>)

Original publisher: [Read original article](<https://chipsandcheese.com/p/amds-instinct-mi455x-aiming-for-the>)

Author: George Cozma

Published: 2026-07-23T17:37:10Z

Content type: article

Language: en

Sources: [Chips and Cheese](<https://devfeed.tech/sources/chips-and-cheese.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [networking](<https://devfeed.tech/topics/networking.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Ethernet](<https://devfeed.tech/topics/ethernet.md>), [SOC](<https://devfeed.tech/topics/soc.md>), [tsmc](<https://devfeed.tech/topics/tsmc.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [ai](<https://devfeed.tech/tags/ai.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [article](<https://devfeed.tech/tags/article.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [communication](<https://devfeed.tech/tags/communication.md>), [compute](<https://devfeed.tech/tags/compute.md>), [ethernet](<https://devfeed.tech/tags/ethernet.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [memory](<https://devfeed.tech/tags/memory.md>), [networking](<https://devfeed.tech/tags/networking.md>), [performance](<https://devfeed.tech/tags/performance.md>), [processors](<https://devfeed.tech/tags/processors.md>), [scale](<https://devfeed.tech/tags/scale.md>), [soc](<https://devfeed.tech/tags/soc.md>), [tsmc](<https://devfeed.tech/tags/tsmc.md>)

### AI overview

The article examines AMD's Instinct MI455X, a CDNA5 GPU designed for rack-scale AI deployments. It describes the Helios rack-scale system, UALink over Ethernet networking, compute and memory specifications, and changes to the WGP, register-file, and matrix-unit designs.

### Source excerpt

Editor's Note (7/25/2026): The article has been edited with more information about the L2 behavior along with the bandwidth of the die to die interface.

## NVIDIA NVLink: The Scale-Up Network for AI Factories

DevFeed: [NVIDIA NVLink: The Scale-Up Network for AI Factories](<https://devfeed.tech/articles/nvidia-nvlink-the-scale-up-network-for-ai-factories-6905.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/nvidia-nvlink-the-scale-up-network-for-ai-factories/>)

Author: Elizabeth Goodman

Published: 2026-07-20T15:46:28Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [collective](<https://devfeed.tech/tags/collective.md>), [communication](<https://devfeed.tech/tags/communication.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infiniband](<https://devfeed.tech/tags/infiniband.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [production](<https://devfeed.tech/tags/production.md>), [scale](<https://devfeed.tech/tags/scale.md>), [spectrum-ethernet](<https://devfeed.tech/tags/spectrum-ethernet.md>), [spectrum-x](<https://devfeed.tech/tags/spectrum-x.md>), [speed](<https://devfeed.tech/tags/speed.md>), [systems](<https://devfeed.tech/tags/systems.md>), [vera-rubin](<https://devfeed.tech/tags/vera-rubin.md>)

### AI overview

NVIDIA NVLink is presented as a scale-up networking fabric for AI factories. It provides high-bandwidth, low-latency GPU-to-GPU communication for large AI inference, training, and parallel-computing workloads, with collective-operation acceleration and rack-level resiliency.

### Source excerpt

The demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute...

## Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading

DevFeed: [Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading](<https://devfeed.tech/articles/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading-6925.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/>)

Author: Tanya Lenz

Published: 2026-07-10T18:17:40Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [LLM Techniques](<https://devfeed.tech/topics/llm-techniques.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [compute](<https://devfeed.tech/tags/compute.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gb200](<https://devfeed.tech/tags/gb200.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [grace-cpu](<https://devfeed.tech/tags/grace-cpu.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llm-techniques](<https://devfeed.tech/tags/llm-techniques.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [training-ai-models](<https://devfeed.tech/tags/training-ai-models.md>), [vera-rubin](<https://devfeed.tech/tags/vera-rubin.md>), [work](<https://devfeed.tech/tags/work.md>)

### AI overview

This article explains how host offloading in JAX-based large language model training reduces GPU high-bandwidth memory pressure by moving selected activations to pinned host memory and streaming them back during the backward pass. It discusses activation-transfer overlap, NVIDIA Grace Blackwell and GB200 NVL72 systems, and experiments involving Llama 3.1 405B and DeepSeek-V3 671B.

### Source excerpt

Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states,...

## Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72

DevFeed: [Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72](<https://devfeed.tech/articles/running-low-latency-analytical-workloads-with-gpu-accelerated-presto-on-nvidia-gb200-nvl72-6935.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/running-low-latency-analytical-workloads-with-gpu-accelerated-presto-on-nvidia-gb200-nvl72/>)

Author: Tanya Lenz

Published: 2026-07-08T16:05:25Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [GPUDirect](<https://devfeed.tech/topics/gpudirect.md>), [NVIDIA DGX](<https://devfeed.tech/topics/nvidia-dgx.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [C++](<https://devfeed.tech/topics/c-plus-plus.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>)

Tags: [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [c-plus-plus](<https://devfeed.tech/tags/c-plus-plus.md>), [cache](<https://devfeed.tech/tags/cache.md>), [communication](<https://devfeed.tech/tags/communication.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [data](<https://devfeed.tech/tags/data.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gb200](<https://devfeed.tech/tags/gb200.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpudirect](<https://devfeed.tech/tags/gpudirect.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [multi-gpu](<https://devfeed.tech/tags/multi-gpu.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-dgx](<https://devfeed.tech/tags/nvidia-dgx.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>)

### AI overview

This article presents GPU-accelerated Presto for low-latency analytical SQL workloads on large datasets. It compares single-node multi-GPU execution on NVIDIA DGX B200 and multinode NVIDIA GB200 NVL72 systems with CPU-based Presto, highlighting NVLink communication, GPUDirect Storage, cuDF algorithms, Parquet data, and benchmark results.

### Source excerpt

Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance...

## Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism

DevFeed: [Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism](<https://devfeed.tech/articles/enhancing-goodput-in-large-scale-llm-training-with-nonuniform-tensor-parallelism-6815.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/enhancing-goodput-in-large-scale-llm-training-with-nonuniform-tensor-parallelism/>)

Author: Michelle Horton

Published: 2026-07-06T21:44:23Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [LLM Techniques](<https://devfeed.tech/topics/llm-techniques.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Network](<https://devfeed.tech/topics/network.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [availability](<https://devfeed.tech/tags/availability.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llm-techniques](<https://devfeed.tech/tags/llm-techniques.md>), [mlops](<https://devfeed.tech/tags/mlops.md>), [network](<https://devfeed.tech/tags/network.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [performance](<https://devfeed.tech/tags/performance.md>), [scale](<https://devfeed.tech/tags/scale.md>), [training](<https://devfeed.tech/tags/training.md>), [work](<https://devfeed.tech/tags/work.md>)

### AI overview

This article explains how Nonuniform Tensor Parallelism can improve Goodput in large-scale LLM training by adapting tensor parallelism to changing GPU availability and overlapping data resharding. The experimental approach aims to reduce interruptions, lost throughput, and computational waste in tightly interconnected GPU clusters.

### Source excerpt

Training LLMs at massive scale brings unique infrastructure challenges, especially as jobs span thousands of GPUs and run for extended periods. The longer these...

## Нейро сети для самых маленьких. Часть первая (которая после нулевой). Удобство в прокрустовом ложе оптимизации

DevFeed: [Нейро сети для самых маленьких. Часть первая (которая после нулевой). Удобство в прокрустовом ложе оптимизации](<https://devfeed.tech/articles/article-24859.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1047072/>)

Author: eucariot (Яндекс, Yandex Cloud & Yandex Infrastructure)

Published: 2026-07-01T07:00:06Z

Content type: article

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [InfiniBand](<https://devfeed.tech/topics/infiniband.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Linux](<https://devfeed.tech/topics/linux.md>)

Tags: [ethernet](<https://devfeed.tech/tags/ethernet.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpudirect-rdma](<https://devfeed.tech/tags/gpudirect-rdma.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [infiniband](<https://devfeed.tech/tags/infiniband.md>), [linux](<https://devfeed.tech/tags/linux.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [performance](<https://devfeed.tech/tags/performance.md>), [rdma](<https://devfeed.tech/tags/rdma.md>), [roce](<https://devfeed.tech/tags/roce.md>), [tcp](<https://devfeed.tech/tags/tcp.md>), [zero-copy](<https://devfeed.tech/tags/zero-copy.md>)

### AI overview

This introductory article in a series explains the infrastructure used to train and run neural networks and for high-performance computing. It surveys specialized technologies including GPUs and TPUs, RDMA, kernel bypass, NVLink, InfiniBand, and RoCE, arguing that specialized solutions can outperform and cost less than a generic Linux and Ethernet/IP stack at scale.

### Source excerpt

Это первая (после нулевой) статья из серии Нейро сети для самых маленьких, в которой мы разбираем инфраструктуру для запуска нейронных сетей. Для обучения и инференса нейросетей и для любых видов High Performance Computing используются специализированные технологии: GPU/TPU, RDMA, Kernel bypass, NVLink, InfiniBand, RoCE и другие. Про некоторые из них большинство только что-то слышали, но сталкиваться с ними не приходилось. Нельзя просто взять ванильный стек Linux, воткнуть в него 400Gb Ethernet+IP и получить рабочее решение. Почему? Потому что общее решение на масштабе в большинстве случаев проигрывает специализированным как в скорости, так и в стоимости. Как бы странно последнее ни звучало. Читать далее

## Designing GPU-Accelerated Query Engines with NVIDIA GQE

DevFeed: [Designing GPU-Accelerated Query Engines with NVIDIA GQE](<https://devfeed.tech/articles/designing-gpu-accelerated-query-engines-with-nvidia-gqe-6799.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/designing-gpu-accelerated-query-engines-with-nvidia-gqe/>)

Author: Michelle Horton

Published: 2026-06-30T17:36:43Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>), [CUDA](<https://devfeed.tech/topics/cuda.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [IO](<https://devfeed.tech/topics/io.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Parser](<https://devfeed.tech/topics/parser.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [accelerate](<https://devfeed.tech/tags/accelerate.md>), [compression](<https://devfeed.tech/tags/compression.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [cuda-x](<https://devfeed.tech/tags/cuda-x.md>), [data](<https://devfeed.tech/tags/data.md>), [data-analytics-processing](<https://devfeed.tech/tags/data-analytics-processing.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [databases](<https://devfeed.tech/tags/databases.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [memory](<https://devfeed.tech/tags/memory.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [performance](<https://devfeed.tech/tags/performance.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

This article presents GQE, a reference architecture for executing SQL queries on GPUs. It explains how NVIDIA hardware and CUDA-X libraries address memory, I/O, data movement, decompression, and end-to-end performance challenges for large datasets.

### Source excerpt

GPU-accelerated query engines are often constrained by memory and I/O bandwidth. NVIDIA hardware advances--including high bandwidth memory (HBM), NVIDIA...

## Ubuntu Server on the NVIDIA DGX Spark (Without the Desktop)

DevFeed: [Ubuntu Server on the NVIDIA DGX Spark (Without the Desktop)](<https://devfeed.tech/articles/ubuntu-server-on-the-nvidia-dgx-spark-without-the-desktop-10683.md>)

Original publisher: [Read original article](<https://technotim.com/posts/ubuntu-gb10/>)

Author: Techno Tim

Published: 2026-06-22T13:00:00Z

Content type: tutorial

Language: en

Sources: [Techno Tim](<https://devfeed.tech/sources/techno-tim.md>)

Topics: [DGX Spark](<https://devfeed.tech/topics/dgx-spark.md>), [Ubuntu](<https://devfeed.tech/topics/ubuntu.md>), [NVIDIA DGX](<https://devfeed.tech/topics/nvidia-dgx.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [CUDA](<https://devfeed.tech/topics/cuda.md>), [Docker](<https://devfeed.tech/topics/docker.md>), [Ansible](<https://devfeed.tech/topics/ansible.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [Blackwell](<https://devfeed.tech/topics/blackwell.md>), [networking](<https://devfeed.tech/topics/networking.md>), [GitHub](<https://devfeed.tech/topics/github.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ansible](<https://devfeed.tech/tags/ansible.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [dell](<https://devfeed.tech/tags/dell.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [docker](<https://devfeed.tech/tags/docker.md>), [github](<https://devfeed.tech/tags/github.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [homelab](<https://devfeed.tech/tags/homelab.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia-dgx](<https://devfeed.tech/tags/nvidia-dgx.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [ubuntu](<https://devfeed.tech/tags/ubuntu.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

A guide to replacing DGX OS with a minimized Ubuntu 24.04 Server installation on GB10 systems such as the NVIDIA DGX Spark and ASUS Ascent GX10. It explains the memory and power benefits of removing GNOME while retaining the NVIDIA drivers, CUDA, Docker, and NVIDIA Container Toolkit, and covers ConnectX-7 networking, dual-node setup, and Ansible automation.

### Source excerpt

When you buy an NVIDIA DGX Spark or an ASUS Ascent GX10, it ships with DGX OS. DGX OS is NVIDIA's managed Ubuntu image, and it is fine - if you want a full GNOME desktop on an AI box. I did not want that. The GB10 has 128 GB of unified memory shared between the CPU and GPU over NVLink-C2C. Every gigabyte the OS and desktop environment consume is a gigabyte not available to your model. On the ...