# Multi-GPU

Multi-GPU programming enables applications to use multiple GPUs concurrently for greater aggregate performance, memory capacity, and bandwidth.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## KDE Plasma 6.8 Beta Released With Many Great Improvements, Kup Backup Scheduler

DevFeed: [KDE Plasma 6.8 Beta Released With Many Great Improvements, Kup Backup Scheduler](<https://devfeed.tech/articles/kde-plasma-6-8-beta-released-with-many-great-improvements-kup-backup-scheduler-12413.md>)

Original publisher: [Read original article](<https://www.phoronix.com/news/KDE-Plasma-6.8-Beta>)

Author: Michael Larabel

Published: 2026-09-10T16:39:24Z

Content type: release

Language: en

Sources: [Phoronix](<https://devfeed.tech/sources/phoronix.md>)

Topics: [Wayland](<https://devfeed.tech/topics/wayland.md>), [Multi-GPU](<https://devfeed.tech/topics/multi-gpu.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [monitor](<https://devfeed.tech/topics/monitor.md>)

Tags: [backup](<https://devfeed.tech/tags/backup.md>), [desktop](<https://devfeed.tech/tags/desktop.md>), [desktop-linux](<https://devfeed.tech/tags/desktop-linux.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [linux-benchmarking](<https://devfeed.tech/tags/linux-benchmarking.md>), [linux-hardware-benchmarks](<https://devfeed.tech/tags/linux-hardware-benchmarks.md>), [linux-hardware-reviews](<https://devfeed.tech/tags/linux-hardware-reviews.md>), [linux-how-to](<https://devfeed.tech/tags/linux-how-to.md>), [linux-performance](<https://devfeed.tech/tags/linux-performance.md>), [linux-server-benchmarks](<https://devfeed.tech/tags/linux-server-benchmarks.md>), [monitor](<https://devfeed.tech/tags/monitor.md>), [multi-gpu](<https://devfeed.tech/tags/multi-gpu.md>), [open-source-graphics](<https://devfeed.tech/tags/open-source-graphics.md>), [phoronix](<https://devfeed.tech/tags/phoronix.md>), [phoronix-test-suite](<https://devfeed.tech/tags/phoronix-test-suite.md>), [release](<https://devfeed.tech/tags/release.md>), [remote](<https://devfeed.tech/tags/remote.md>), [server](<https://devfeed.tech/tags/server.md>), [tool](<https://devfeed.tech/tags/tool.md>), [ubuntu-benchmarks](<https://devfeed.tech/tags/ubuntu-benchmarks.md>), [ubuntu-hardware](<https://devfeed.tech/tags/ubuntu-hardware.md>)

### AI overview

KDE Plasma 6.8 Beta introduces improvements to Wayland support, vRAM usage, remote desktops, multi-GPU and multi-monitor handling, theming, and other desktop functionality. It also includes Kup as the Plasma desktop backup scheduler.

### Source excerpt

Ahead of the stable release of Plasma 6.8 due out on 14 October that also marks the 30th anniversary of the KDE project, out today is the much anticipated beta release...

## How Replicate Handles Billing: A Complete Breakdown

DevFeed: [How Replicate Handles Billing: A Complete Breakdown](<https://devfeed.tech/articles/how-replicate-handles-billing-a-complete-breakdown-10310.md>)

Original publisher: [Read original article](<https://dodopayments.com/blogs/replicate-billing-model/>)

Author: Ayush Agarwal

Published: 2026-04-09T00:00:00Z

Content type: article

Language: en

Sources: [Dodo Payments Blog](<https://devfeed.tech/sources/dodo-payments-blog.md>)

Topics: [AI Platforms/Deployment](<https://devfeed.tech/topics/ai-platforms-deployment.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Multi-GPU](<https://devfeed.tech/topics/multi-gpu.md>), [llama](<https://devfeed.tech/topics/llama.md>), [stable-diffusion](<https://devfeed.tech/topics/stable-diffusion.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-platform](<https://devfeed.tech/tags/ai-platform.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [billing](<https://devfeed.tech/tags/billing.md>), [compute](<https://devfeed.tech/tags/compute.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [llama](<https://devfeed.tech/tags/llama.md>), [models](<https://devfeed.tech/tags/models.md>), [multi-gpu](<https://devfeed.tech/tags/multi-gpu.md>), [stable-diffusion](<https://devfeed.tech/tags/stable-diffusion.md>), [usage-based-billing](<https://devfeed.tech/tags/usage-based-billing.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The article analyzes Replicate's usage-based billing model, which charges for compute time by hardware type rather than by subscription, model, or token package. It explains hardware-tier pricing, multi-GPU committed-spend requirements, and model-agnostic billing, and discusses how to implement similar per-second billing for an AI platform.

### Source excerpt

A detailed analysis of Replicate's pure usage-based billing model - per-second compute pricing across hardware tiers, cold start costs, and how to build the same pay-per-second infrastructure billing for your own AI platform.

## DigitalOcean Announces GPU Droplets Accelerated by NVIDIA HGX B300

DevFeed: [DigitalOcean Announces GPU Droplets Accelerated by NVIDIA HGX B300](<https://devfeed.tech/articles/powering-the-next-leap-in-ai-gpu-droplets-accelerated-by-nvidia-hgxtm-b300-are-now-available-on-digitalocean-19865.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/coming-soon-gpu-droplets-nvidia-b300s>)

Author: Waverly Swinton

Published: 2025-12-15T17:51:51Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Blackwell](<https://devfeed.tech/topics/blackwell.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [virtual machines](<https://devfeed.tech/topics/virtual-machines.md>), [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [data analytics](<https://devfeed.tech/topics/data-analytics.md>), [Multi-GPU](<https://devfeed.tech/topics/multi-gpu.md>), [long-context](<https://devfeed.tech/topics/long-context.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [compute](<https://devfeed.tech/tags/compute.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [multi-gpu](<https://devfeed.tech/tags/multi-gpu.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

DigitalOcean announces GPU Droplets accelerated by NVIDIA HGX B300, describing the platform's intended benefits for AI training, inference, generative AI, data analytics, and high-performance computing workloads.

### Source excerpt

AI continues to evolve at an unprecedented pace, with new models and demanding workloads pushing the boundaries of what's possible. From complex large language models (LLMs) to intricate scientific simulations, developers and businesses need access to the most powerful and efficient computing infrastructure. At DigitalOcean, we're committed to providing the cutting-edge tools you need to build, deploy, and scale your AI initiatives with simplicity and affordability. That's why we're excited to announce that GPU Droplets accelerated by NVIDIA HGX™ B300 are coming soon to DigitalOcean, marking a significant upgrade to our GPU offerings. Why NVIDIA HGX™ B300? The NVIDIA Blackwell Ultra accelerated computing platform represents a leap forward in AI reasoning. Designed for both training and inference, the NVIDIA HGX B300 offers substantial improvements in computational power, memory bandwidth, and energy efficiency compared to previous generations. The NVIDIA Blackwell architecture at the heart of the HGX B300 is not just about raw power; it's also about efficiency and innovation. With 1.5X more dense Tensor Core FLOPS, enhanced attention performance, and significantly expanded memory, the HGX B300 is optimized for the most demanding AI workloads including generative AI, data analytics, and high-performance computing (HPC). Featuring 7X more AI compute than NVIDIA Hopper platforms, 2.1TB of HBM3e memory, and high-performance networking integration with NVIDIA ConnectX-8 SuperNICs, Blackwell Ultra delivers breakthrough performance on the most complex workloads from agentic systems and reasoning, to real-time video generation. For AI-native enterprises running large reasoning models and long-context workloads, this enables: -Reduced model offloading and improved time-to-first-token -Higher sustained throughput under concurrency -More efficient multi-GPU scaling -Improved tokens-per-second per dollar Unlike GPU capacity providers, DigitalOcean integrates inference-optimized

## 20x Faster TRL Fine-tuning with RapidFire AI

DevFeed: [20x Faster TRL Fine-tuning with RapidFire AI](<https://devfeed.tech/articles/20x-faster-trl-fine-tuning-with-rapidfire-ai-7452.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/rapidfireai>)

Author: Kamran Bigdely; Arun Kumar; Quentin Gallouédec

Published: 2025-11-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [trl](<https://devfeed.tech/topics/trl.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [Multi-GPU](<https://devfeed.tech/topics/multi-gpu.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [dashboards](<https://devfeed.tech/tags/dashboards.md>), [data](<https://devfeed.tech/tags/data.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llms](<https://devfeed.tech/tags/llms.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [multi-gpu](<https://devfeed.tech/tags/multi-gpu.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [rapidfireai](<https://devfeed.tech/tags/rapidfireai.md>), [trl](<https://devfeed.tech/tags/trl.md>)

### AI overview

RapidFire AI accelerates LLM fine-tuning and post-training experimentation by running multiple TRL configurations concurrently, including on a single GPU. Its adaptive chunk-based scheduling, live metrics dashboard, multi-GPU orchestration, and interactive controls help teams compare configurations sooner, stop weak runs, and clone promising ones. The article cites internal benchmarks reporting approximately 16-24x higher experimentation throughput than sequential comparison.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## From Single-Node to Multi-GPU Clusters: How Discord Made Distributed Compute Easy for ML Engineers

DevFeed: [From Single-Node to Multi-GPU Clusters: How Discord Made Distributed Compute Easy for ML Engineers](<https://devfeed.tech/articles/from-single-node-to-multi-gpu-clusters-how-discord-made-distributed-compute-easy-for-ml-engineers-239.md>)

Original publisher: [Read original article](<https://discord.com/blog/from-single-node-to-multi-gpu-clusters-how-discord-made-distributed-compute-easy-for-ml-engineers>)

Author: Serrana Aguirregaray; Nathaniel Jenkins

Published: 2025-10-09T00:00:00Z

Content type: article

Language: en

Sources: [Discord Blog](<https://devfeed.tech/sources/discord-blog.md>)

Topics: [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Developer experience](<https://devfeed.tech/topics/developer-experience.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Multi-GPU](<https://devfeed.tech/topics/multi-gpu.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Tooling](<https://devfeed.tech/topics/tooling.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>)

Tags: [clusters](<https://devfeed.tech/tags/clusters.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [developer-experience](<https://devfeed.tech/tags/developer-experience.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [multi-gpu](<https://devfeed.tech/tags/multi-gpu.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [tooling](<https://devfeed.tech/tags/tooling.md>)

### AI overview

Discord describes building a developer-friendly distributed machine learning platform on Ray to support multi-GPU training, larger datasets, and production workloads. The platform combines custom CLI tooling, Dagster and KubeRay orchestration, and an observability layer called X-Ray; the article says this work enabled an Ads Ranking model that improved business metrics by 200%.

### Source excerpt

From manual GPU configs to one-command clusters: Join Serrana Aguirregaray and Nathaniel Jenkins as they tell the story of Discord's journey to build a developer-friendly ML infrastructure on Ray for 200M+ monthly active users.

## Implementing Multi-GPU Distributed Training for Stitch Fix's Personalized Recommendations

DevFeed: [Implementing Multi-GPU Distributed Training for Stitch Fix's Personalized Recommendations](<https://devfeed.tech/articles/accelerating-ai-implementing-multi-gpu-distributed-training-for-personalized-recommendations-29344.md>)

Original publisher: [Read original article](<https://multithreaded.stitchfix.com/blog/2023/06/08/distributed-model-training/>)

Published: 2023-06-08T09:00:00Z

Content type: article

Language: en

Sources: [Stitch Fix](<https://devfeed.tech/sources/stitch-fix.md>)

Topics: [distributed-training](<https://devfeed.tech/topics/distributed-training.md>), [recommendations](<https://devfeed.tech/topics/recommendations.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [sharding](<https://devfeed.tech/topics/sharding.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Multi-GPU](<https://devfeed.tech/topics/multi-gpu.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>)

Tags: [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [model-training](<https://devfeed.tech/tags/model-training.md>), [multi-gpu](<https://devfeed.tech/tags/multi-gpu.md>), [parallel](<https://devfeed.tech/tags/parallel.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [recommendations](<https://devfeed.tech/tags/recommendations.md>), [sharding](<https://devfeed.tech/tags/sharding.md>)

### AI overview

This Stitch Fix engineering article explains how the company implemented multi-GPU distributed training for its Client Time Series Model (CTSM), a PyTorch-based model used in personalized recommendations. It describes sharding training data across GPUs and training mini-batches in parallel to reduce training time, along with the surrounding retraining and deployment workflow.

### Source excerpt

Stitch Fix uses a cutting-edge multi-tiered recommender system stack to personalize styling recommendations at scale. This stack comprises several critical components, including feature generation, scoring, ranking, and inventory optimization techniques. Our scoring module is based on the Client Time Series Model (CTSM) which is an award winning novel sequence based model that uses temporally masked encoders. CTSM is built using PyTorch, and was initially trained on a single Graphics Processing Unit (GPU) instance. Since we first put this model into production last year, we have launched several updates to the model that improved its performance. Many of these improvements involved adding new features or increasing the time window of our training data. As a result, the model training time increased significantly, making it harder for us to iterate quickly and get feedback on new ideas we want to try for improving the model. We needed a way to reduce the model training time. This blog delves into the steps we followed to overcome this challenge and our journey to implement multi-GPU distributed model training for CTSM. By sharding the training data across multiple GPUs and training multiple mini-batches in parallel, we aimed to achieve significant reductions in training time. We present empirical results showcasing the observed reduction in training time when we scaled up resources from 1 to N GPUs, and share some future directions we are considering in our continued effort to speed up model training. Model Training Workflow The scores generated by CTSM are leveraged by multiple downstream services to get insight into what items a client is likely to purchase. The model is retrained at a regular cadence to ensure that it is using the most updated information about each client when making predictions and does not degrade in its performance. We leverage configuration driven machine learning pipelines to set up a Directed Acyclic Graph (DAG) that automatically retrains