# IBM Research and Red Hat benchmark llm-d serving GLM-5.2 on H100 GPUs

DevFeed: [IBM Research and Red Hat benchmark llm-d serving GLM-5.2 on H100 GPUs](<https://devfeed.tech/articles/how-llm-d-makes-the-most-of-the-hardware-you-already-have-17349.md>)

Original publisher: [Read original article](<https://research.ibm.com/blog/running-open-models-on-h100-gpus-with-llmd>)

Author: Peter Hess

Published: 2026-09-08T12:00:00Z

Content type: article

Language: en

Sources: [IBM Research](<https://devfeed.tech/sources/ibm-research.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Self-hosted](<https://devfeed.tech/topics/self-hosted.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [ibm](<https://devfeed.tech/topics/ibm.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-for-code](<https://devfeed.tech/tags/ai-for-code.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hybrid-cloud](<https://devfeed.tech/tags/hybrid-cloud.md>), [hybrid-cloud-platform](<https://devfeed.tech/tags/hybrid-cloud-platform.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [news](<https://devfeed.tech/tags/news.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [scaling-ai](<https://devfeed.tech/tags/scaling-ai.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>)

## AI overview

IBM Research, Red Hat, and collaborators report a benchmark deployment of the llm-d open-source inference framework serving the open-weight GLM-5.2 model on 544 NVIDIA H100 GPUs. The reported workload reached more than 6.6 million output tokens per minute and up to 3,000 concurrent coding agents without preemptions.

## Source excerpt

IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.