# Scality AI Inference Factory Serves KV Cache From Object Storage Over RDMA, 14x Faster Than Recompute

DevFeed: [Scality AI Inference Factory Serves KV Cache From Object Storage Over RDMA, 14x Faster Than Recompute](<https://devfeed.tech/articles/scality-ai-inference-factory-serves-kv-cache-from-object-storage-over-rdma-14x-faster-than-recompute-77324.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/scality-ai-inference-factory-kv-cache-object-storage-rdma>)

Author: Harold Fritts

Published: 2026-10-08T16:32:44Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [rdma](<https://devfeed.tech/topics/rdma.md>), [gpudirect rdma](<https://devfeed.tech/topics/gpudirect-rdma.md>), [Dynamo-Triton](<https://devfeed.tech/topics/dynamo-triton.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [versioning](<https://devfeed.tech/topics/versioning.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-data-infrastructure](<https://devfeed.tech/tags/ai-data-infrastructure.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [bandwidth](<https://devfeed.tech/tags/bandwidth.md>), [code](<https://devfeed.tech/tags/code.md>), [concurrent](<https://devfeed.tech/tags/concurrent.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [object-storage](<https://devfeed.tech/tags/object-storage.md>), [stack](<https://devfeed.tech/tags/stack.md>)

## AI overview

Scality launched AI Inference Factory, an open-code stack for running open-weight models on customer-owned infrastructure. It uses ADI object storage as a shared KV cache, with vLLM and Dynamo separating prefill and decode, and an RDMA path for moving cache data to GPUs. In a small lab, Scality reported that restoring a 14K-token context took 166 ms, compared with 2.3 seconds to recompute it; it also reported serving 8 TB of KV cache with first-token latency between 355 and 373 ms across the tested cache sizes.

## Source excerpt

Scality today launched Scality AI Inference Factory, an open-code software stack for running open-weight models on infrastructure the customer owns. Scality validates, ships, and maintains the whole stack, which places its AI Data Infrastructure (ADI) object storage under the inference layer as a shared key-value (KV) cache, and it is aimed at enterprises, government agencies, The post Scality AI Inference Factory Serves KV Cache From Object Storage Over RDMA, 14x Faster Than Recompute appeared first on StorageReview.com.