# Introducing Qdrant Cloud Inference

DevFeed: [Introducing Qdrant Cloud Inference](<https://devfeed.tech/articles/introducing-qdrant-cloud-inference-46708.md>)

Original publisher: [Read original article](<https://qdrant.tech/blog/qdrant-cloud-inference-launch/>)

Author: info@qdrant.tech (Andrey Vasnetsov)

Published: 2025-07-15T00:00:00Z

Content type: release

Language: en

Sources: [Qdrant Blog on Qdrant - Vector Search Engine](<https://devfeed.tech/sources/qdrant-blog-on-qdrant-vector-search-engine.md>)

Topics: [Qdrant](<https://devfeed.tech/topics/qdrant.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [hybrid-search](<https://devfeed.tech/topics/hybrid-search.md>), [multimodal-ai](<https://devfeed.tech/topics/multimodal-ai.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [Application Development](<https://devfeed.tech/topics/application-development.md>), [API](<https://devfeed.tech/topics/api.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [application-development](<https://devfeed.tech/tags/application-development.md>), [approximate-nearest-neighbor-search](<https://devfeed.tech/tags/approximate-nearest-neighbor-search.md>), [bert](<https://devfeed.tech/tags/bert.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [fasttext](<https://devfeed.tech/tags/fasttext.md>), [hnsw](<https://devfeed.tech/tags/hnsw.md>), [hybrid-search](<https://devfeed.tech/tags/hybrid-search.md>), [image-search](<https://devfeed.tech/tags/image-search.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [knn-algorithm](<https://devfeed.tech/tags/knn-algorithm.md>), [matching](<https://devfeed.tech/tags/matching.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [neural-network](<https://devfeed.tech/tags/neural-network.md>), [qdrant](<https://devfeed.tech/tags/qdrant.md>), [rag](<https://devfeed.tech/tags/rag.md>), [recommender-system](<https://devfeed.tech/tags/recommender-system.md>), [saas](<https://devfeed.tech/tags/saas.md>), [simaes-networks](<https://devfeed.tech/tags/simaes-networks.md>), [similarity](<https://devfeed.tech/tags/similarity.md>), [transformer](<https://devfeed.tech/tags/transformer.md>), [vector-search](<https://devfeed.tech/tags/vector-search.md>), [vector-search-engine](<https://devfeed.tech/tags/vector-search-engine.md>), [vectors](<https://devfeed.tech/tags/vectors.md>), [word2vec](<https://devfeed.tech/tags/word2vec.md>)

## AI overview

Qdrant announces Qdrant Cloud Inference, which generates, stores, and indexes embeddings in a single API call. The service integrates model inference into Qdrant Cloud to reduce separate infrastructure, manual pipelines, data transfers, and network overhead for applications using RAG, multimodal, and hybrid search.

## Source excerpt

Introducing Qdrant Cloud Inference Today, we're announcing the launch of Qdrant Cloud Inference (get started in your cluster). With Qdrant Cloud Inference, users can generate, store and index embeddings in a single API call, turning unstructured text and images into search-ready vectors in a single environment. Directly integrating model inference into Qdrant Cloud removes the need for separate inference infrastructure, manual pipelines, and redundant data transfers. This simplifies workflows, accelerates development cycles, and eliminates unnecessary network hops for developers. With a single API call, you can now embed, store, and index your data more quickly and more simply. This speeds up application development for RAG, Multimodal, Hybrid search, and more.