# Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1

DevFeed: [Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1](<https://devfeed.tech/articles/red-hat-delivers-peak-performance-on-kubernetes-and-cpus-in-mlperf-inference-v6-1-61669.md>)

Original publisher: [Read original article](<https://www.redhat.com/en/blog/red-hat-delivers-peak-performance-kubernetes-cpus-mlperf-inference-v61>)

Author: Alberto Perdomo; Ashish Kamra; Diane Feddema; Harika Pothina; Michael Goin; Michey Mehta; Naveen Miriyalu; Nikhil Palaskar; Thameem Abbas Ibrahim Bathusha

Published: 2026-09-29T00:00:00Z

Content type: article

Language: en

Sources: [Red Hat Blog](<https://devfeed.tech/sources/red-hat-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [GB200](<https://devfeed.tech/topics/gb200.md>), [performance-engineering](<https://devfeed.tech/topics/performance-engineering.md>), [rhel](<https://devfeed.tech/topics/rhel.md>), [armv9](<https://devfeed.tech/topics/armv9.md>), [Load Balancing](<https://devfeed.tech/topics/load-balancing.md>), [Arm](<https://devfeed.tech/topics/arm.md>)

Tags: [cpu](<https://devfeed.tech/tags/cpu.md>), [gb200](<https://devfeed.tech/tags/gb200.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [instinct](<https://devfeed.tech/tags/instinct.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [mlperf](<https://devfeed.tech/tags/mlperf.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

## AI overview

Red Hat reports its MLPerf Inference v6.1 results, covering Kubernetes based inference with vLLM across GPU and CPU systems. Highlights include GPT-OSS-120B throughput on GB200 systems, Qwen3-VL results on GB200 NVL4, and CPU only Llama-3.1-8B and Whisper inference results on Intel Xeon 6 servers.

## Source excerpt

Red Hat is proud to announce our results from the industry-standard MLPerf Inference v6.1 benchmark. This submission builds on our track record across recent rounds: In v5.1, we demonstrated cost-effective Llama-3.1-8B-FP8 inference with vLLM on NVIDIA H100 and L40S GPUs, and in v6.0 we delivered results across Qwen3-VL, Whisper, and gpt-oss-120b on NVIDIA H200 and B200 and AMD Instinct MI350 GPUs, including the first Kubernetes-based submission.Our v6.1 results highlight 3 things: peak performance on Kubernetes, 1 inference engine (vLLM) spanning graphics processing units (GPUs) and central p