# Announcing vLLM AFD Plugin: Disaggregating Attention and FFN for Flexible MoE Serving

DevFeed: [Announcing vLLM AFD Plugin: Disaggregating Attention and FFN for Flexible MoE Serving](<https://devfeed.tech/articles/announcing-vllm-afd-plugin-disaggregating-attention-and-ffn-for-flexible-moe-serving-79465.md>)

Original publisher: [Read original article](<https://vllm.ai/blog/2026-07-23-vllm-afd-plugin>)

Author: AFD Plugin Contributors

Published: 2026-07-23T00:00:00Z

Content type: release

Language: en

Sources: [vLLM Blog](<https://devfeed.tech/sources/vllm-blog.md>)

Topics: [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [tgi](<https://devfeed.tech/topics/tgi.md>)

Tags: [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [attention](<https://devfeed.tech/tags/attention.md>), [disaggregation](<https://devfeed.tech/tags/disaggregation.md>), [ecosystem](<https://devfeed.tech/tags/ecosystem.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm-inference](<https://devfeed.tech/tags/llm-inference.md>), [model-serving](<https://devfeed.tech/tags/model-serving.md>), [moe](<https://devfeed.tech/tags/moe.md>), [npu](<https://devfeed.tech/tags/npu.md>), [open-source-ai](<https://devfeed.tech/tags/open-source-ai.md>), [vllm](<https://devfeed.tech/tags/vllm.md>), [vllm-blog](<https://devfeed.tech/tags/vllm-blog.md>)

## AI overview

The vLLM AFD Plugin is an experimental plugin that separates Attention and FFN into independently deployed services for MoE inference. It retains vLLM's request lifecycle and OpenAI-compatible serving interface, with GPU and Ascend NPU backends and synchronous and asynchronous connectors. Benchmarks show throughput varies by Attention-to-FFN allocation: one tested configuration exceeded the EP64 baseline, while another fell below it. The reported results are limited experiments, not general production performance claims.

## Source excerpt

What vLLM AFD Plugin adds to the vLLM ecosystem: Attention-FFN disaggregation for MoE serving, GPU and Ascend NPU backends, connector-based execution, and graph and ubatching support.