# Scaling AI

Published articles for Scaling AI.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## IBM Research and Red Hat benchmark llm-d serving GLM-5.2 on H100 GPUs

DevFeed: [IBM Research and Red Hat benchmark llm-d serving GLM-5.2 on H100 GPUs](<https://devfeed.tech/articles/how-llm-d-makes-the-most-of-the-hardware-you-already-have-17349.md>)

Original publisher: [Read original article](<https://research.ibm.com/blog/running-open-models-on-h100-gpus-with-llmd>)

Author: Peter Hess

Published: 2026-09-08T12:00:00Z

Content type: article

Language: en

Sources: [IBM Research](<https://devfeed.tech/sources/ibm-research.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Self-hosted](<https://devfeed.tech/topics/self-hosted.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [ibm](<https://devfeed.tech/topics/ibm.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-for-code](<https://devfeed.tech/tags/ai-for-code.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hybrid-cloud](<https://devfeed.tech/tags/hybrid-cloud.md>), [hybrid-cloud-platform](<https://devfeed.tech/tags/hybrid-cloud-platform.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [news](<https://devfeed.tech/tags/news.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [scaling-ai](<https://devfeed.tech/tags/scaling-ai.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>)

### AI overview

IBM Research, Red Hat, and collaborators report a benchmark deployment of the llm-d open-source inference framework serving the open-weight GLM-5.2 model on 544 NVIDIA H100 GPUs. The reported workload reached more than 6.6 million output tokens per minute and up to 3,000 concurrent coding agents without preemptions.

### Source excerpt

IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.

## Scaling AI in CPG: How Adaptive Teams Can Unlock Consumer Goods Growth

DevFeed: [Scaling AI in CPG: How Adaptive Teams Can Unlock Consumer Goods Growth](<https://devfeed.tech/articles/scaling-ai-in-cpg-how-adaptive-teams-can-unlock-consumer-goods-growth-4464.md>)

Original publisher: [Read original article](<https://www.toptal.com/executive-guidance/consumer-products-services/ai-in-cpg>)

Author: CHRIS DANIEL, GM, CONSUMER PRODUCTS & SERVICES @ TOPTAL

Published: 2026-06-22T22:00:00Z

Content type: opinion

Language: en

Sources: [Toptal Blog](<https://devfeed.tech/sources/toptal-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Strategy](<https://devfeed.tech/topics/ai-strategy.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [business](<https://devfeed.tech/tags/business.md>), [digital-transformation](<https://devfeed.tech/tags/digital-transformation.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [industry](<https://devfeed.tech/tags/industry.md>), [scaling-ai](<https://devfeed.tech/tags/scaling-ai.md>), [strategy](<https://devfeed.tech/tags/strategy.md>)

### AI overview

This article explains why AI initiatives in the consumer packaged goods industry often remain stuck in pilot mode. It argues that the main barrier is not access to technology but legacy operating models, fragmented data, rigid workflows, and insufficient governance. The article presents adaptive teams, aligned outcomes, and flexible operating structures as essential to embedding AI into everyday planning, innovation, and execution.

### Source excerpt

AI initiatives across the consumer goods industry often stall in pilot mode, constrained not by technology but by legacy operating models. Learn how successful CPG companies scale artificial intelligence through adaptive teams, outcome alignment, and smarter execution.

## How Scania accelerates work with AI across its global workforce

DevFeed: [How Scania accelerates work with AI across its global workforce](<https://devfeed.tech/articles/how-scania-accelerates-work-with-ai-across-its-global-workforce-6644.md>)

Original publisher: [Read original article](<https://openai.com/index/scania>)

Published: 2025-11-19T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [ChatGPT](<https://devfeed.tech/topics/chatgpt.md>), [AI Strategy](<https://devfeed.tech/topics/ai-strategy.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-adoption](<https://devfeed.tech/tags/ai-adoption.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [company](<https://devfeed.tech/tags/company.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [global](<https://devfeed.tech/tags/global.md>), [governance](<https://devfeed.tech/tags/governance.md>), [industrial](<https://devfeed.tech/tags/industrial.md>), [innovation](<https://devfeed.tech/tags/innovation.md>), [onboarding](<https://devfeed.tech/tags/onboarding.md>), [openai](<https://devfeed.tech/tags/openai.md>), [operations](<https://devfeed.tech/tags/operations.md>), [productivity](<https://devfeed.tech/tags/productivity.md>), [quality](<https://devfeed.tech/tags/quality.md>), [scaling-ai](<https://devfeed.tech/tags/scaling-ai.md>), [security](<https://devfeed.tech/tags/security.md>), [team](<https://devfeed.tech/tags/team.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Scania is scaling ChatGPT Enterprise across its global workforce, using team-based onboarding, broad experimentation, and governance guardrails to improve productivity, quality, innovation, and operational workflows.

### Source excerpt

Global manufacturer Scania is scaling AI with ChatGPT Enterprise. With team-based onboarding and strong guardrails, AI is boosting productivity, quality, and innovation.

## Ensuring your AI systems can scale to meet demand

DevFeed: [Ensuring your AI systems can scale to meet demand](<https://devfeed.tech/articles/ensuring-your-ai-systems-can-scale-to-meet-demand-11566.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/ensuring-your-ai-systems-can-scale-to-meet-demand>)

Author: Andre Newman

Published: 2025-04-01T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Terraform](<https://devfeed.tech/topics/terraform.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [industry](<https://devfeed.tech/tags/industry.md>), [inference](<https://devfeed.tech/tags/inference.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [openai](<https://devfeed.tech/tags/openai.md>), [s3](<https://devfeed.tech/tags/s3.md>), [scaling-ai](<https://devfeed.tech/tags/scaling-ai.md>), [terraform](<https://devfeed.tech/tags/terraform.md>)

### AI overview

This tutorial explains why AI workloads are harder to scale than traditional services, highlighting large model sizes, network and memory transfer costs, and unpredictable demand. It introduces scalable infrastructure patterns used by OpenAI and Anthropic and outlines a three-step approach: choose scaling metrics, set thresholds, and test them with simulated load.

### Source excerpt

Demand for AI services is ever-increasing. Are your systems prepared? This blog teaches you how to prepare for sudden demand surges.