# Runpod RTX 5090 Benchmarks Report Up to 65,000 Tokens per Second on Qwen2-0.5B

DevFeed: [Runpod RTX 5090 Benchmarks Report Up to 65,000 Tokens per Second on Qwen2-0.5B](<https://devfeed.tech/articles/the-rtx-5090-is-here-serve-65-000-tokens-per-second-on-runpod-78129.md>)

Original publisher: [Read original article](<https://www.runpod.io/blog/rtx-5090-launch-runpod>)

Author: Alyssa Mazzina

Published: 2026-09-13T15:06:30Z

Content type: release

Language: en

Sources: [Runpod Blog.](<https://devfeed.tech/sources/runpod-blog.md>)

Topics: [NVIDIA RTX](<https://devfeed.tech/topics/nvidia-rtx.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [gpu](<https://devfeed.tech/tags/gpu.md>), [gpu-memory](<https://devfeed.tech/tags/gpu-memory.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-engine](<https://devfeed.tech/tags/inference-engine.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [rtx-5090](<https://devfeed.tech/tags/rtx-5090.md>)

## AI overview

Runpod announces availability of the NVIDIA RTX 5090 for on-demand and containerized workloads and reports internal benchmarks using vLLM. Qwen2-0.5B exceeded 65,000 tokens per second and 250 requests per second at 1,024 concurrent prompts, while Phi-3-mini-4k-instruct reached 6,400 tokens per second and about 25 requests per second under the same conditions. The article notes that results reflect high-concurrency performance and that single-prompt throughput is lower.

## Source excerpt

The new NVIDIA RTX 5090 is now live on Runpod. With blazing-fast inference speeds and large memory capacity, it's ideal for real-time LLM workloads and AI.