# GLM 5.2 Fast via Wafer now available on AI Gateway

DevFeed: [GLM 5.2 Fast via Wafer now available on AI Gateway](<https://devfeed.tech/articles/glm-5-2-fast-via-wafer-now-available-on-ai-gateway-954.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/glm-5-2-fast-via-wafer-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-06-24T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [API](<https://devfeed.tech/topics/api.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [api](<https://devfeed.tech/tags/api.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [inference](<https://devfeed.tech/tags/inference.md>), [performance](<https://devfeed.tech/tags/performance.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [speed](<https://devfeed.tech/tags/speed.md>)

## AI overview

GLM 5.2 Fast is now available through AI Gateway via Wafer. Vercel reports that Wafer provides twice the throughput of other providers serving GLM-5.2 on serverless, with measured speeds above 170 tokens per second for small contexts and 200 tokens per second for large contexts.

## Source excerpt

GLM 5.2 Fast via Wafer is now available on AI Gateway. Based on our own benchmarking across small-context, large-context, and tool-call scenarios, Wafer delivers a 2x higher throughput than other providers serving GLM-5.2 on serverless, leading on decode and end-to-end speed for sustained generation in the small- and large-context cases. In our testing, GLM 5.2 Fast on Wafer measured: Small context: 170+ tok/s Large context: 200+ tok/s To use GLM 5.2 Fast, set model to zai/glm-5.2-fast in the AI SDK: AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try GLM 5.2 Fast in the model playground. Read more