# qwen

Qwen is a series of large language and large multimodal models from the Qwen Team at Alibaba Group.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Qwen выпустила Qwen3.8-Omni-Flash

DevFeed: [Qwen выпустила Qwen3.8-Omni-Flash](<https://devfeed.tech/articles/qwen-qwen3-8-omni-flash-42727.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/bothub/news/1083716/>)

Author: DashaPasha (BotHub)

Published: 2026-09-18T08:40:54Z

Content type: release

Language: ru

Sources: [Tagir Valeev](<https://devfeed.tech/sources/tagir-valeev.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [qwen3](<https://devfeed.tech/topics/qwen3.md>), [API](<https://devfeed.tech/topics/api.md>), [function calling](<https://devfeed.tech/topics/function-calling.md>)

Tags: [alibaba](<https://devfeed.tech/tags/alibaba.md>), [api](<https://devfeed.tech/tags/api.md>), [function-calling](<https://devfeed.tech/tags/function-calling.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [qwen-3-8](<https://devfeed.tech/tags/qwen-3-8.md>), [qwen3](<https://devfeed.tech/tags/qwen3.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [responses](<https://devfeed.tech/tags/responses.md>), [tag-61cd5a476b1d](<https://devfeed.tech/tags/tag-61cd5a476b1d.md>), [tag-67847f6cdd99](<https://devfeed.tech/tags/tag-67847f6cdd99.md>), [tag-6e9d47cd6240](<https://devfeed.tech/tags/tag-6e9d47cd6240.md>), [tag-dc29f2a3d816](<https://devfeed.tech/tags/tag-dc29f2a3d816.md>)

### AI overview

Qwen introduced Qwen3.8-Omni-Flash, a multimodal model focused on long-video understanding, detailed reporting, content description, and more precise timestamps. It accepts text, images, audio, and video, supports function calling, web search, reasoning controls, and API access through Chat Completions and Responses.

### Source excerpt

Команда Qwen представила Qwen3.8-Omni-Flash -- новую мультимодальную модель. Главный акцент в новой версии сделали на понимании видео. Qwen3.8-Omni-Flash умеет анализировать длинные ролики, составлять по ним подробные отчёты и генерировать описания содержимого. В Qwen отдельно отмечают улучшение работы с временными метками: модель должна точнее связывать события с конкретными моментами видео. При этом текстовые возможности модели сохранили на уровне Qwen3.8-Flash. В API также доступны function calling и веб-поиск, а режим рассуждений включён по умолчанию и может настраиваться через reasoning_effort. Qwen3.8-Omni-Flash -- промежуточное звено для более сложных мультимодальных пайплайнов. Среди возможных сценариев разрабы называют монтаж видео, перевод, создание кинокомментариев, анализ и генерацию контента. Отдельный фокус релиза -- стоимость. По сравнению с предыдущей Qwen3.5-Omni-Flash расход токенов при работе с мультимодальными данными заметно снизился. В частности, Qwen заявляет о снижении стоимости обработки аудио примерно на 98%, а аудио-видео -- более чем на 93%. При этом заявленная производительность сопоставима с Gemini 3.8 Flash. Независимые данные также показывают снижение расхода токенов на одном из видеобенчмарков: с 145 736 до 79 117 токенов при росте результата OmniVideoBench с 63,4 до 67,8. Контекстное окно новой модели -- 1 млн токенов, максимальная длина генерируемого ответа -- 131 072 токена. Qwen3.8-Omni-Flash принимает текст, изображения, аудио и видео, а на выходе генерирует текст. Читать далее

## How and Why We Bought 4x DGX Sparks

DevFeed: [How and Why We Bought 4x DGX Sparks](<https://devfeed.tech/articles/how-and-why-we-bought-4x-dgx-sparks-26641.md>)

Original publisher: [Read original article](<https://blog.alexellis.io/how-and-why-we-bought-4-dgx-sparks/>)

Author: Alex Ellis

Published: 2026-09-15T00:00:00Z

Content type: opinion

Language: en

Sources: [Alex Ellis' Blog](<https://devfeed.tech/sources/alex-ellis-blog.md>)

Topics: [Local AI](<https://devfeed.tech/topics/local-ai.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [llm](<https://devfeed.tech/tags/llm.md>), [local-ai](<https://devfeed.tech/tags/local-ai.md>), [localai](<https://devfeed.tech/tags/localai.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [red-teaming](<https://devfeed.tech/tags/red-teaming.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

The author explains why OpenFaaS Ltd bought four DGX Sparks and what the team learned from deploying local AI. The article argues that local infrastructure can provide tangible privacy and risk-reduction benefits for business use cases, even though it is not primarily justified by cost per token.

### Source excerpt

In June we deployed an RTX 6000 Pro into production, a few weeks later, we're now operating DGX Sparks for the team. Learn how and why.

## Perplexity's new agent runs entirely on your GPU -- with one expensive catch

DevFeed: [Perplexity's new agent runs entirely on your GPU -- with one expensive catch](<https://devfeed.tech/articles/perplexity-s-new-agent-runs-entirely-on-your-gpu-with-one-expensive-catch-21600.md>)

Original publisher: [Read original article](<https://thenewstack.io/perplexity-portable-computer-windows/>)

Author: Amanda Caswell

Published: 2026-09-14T18:21:44Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Windows](<https://devfeed.tech/topics/windows.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [browser](<https://devfeed.tech/topics/browser.md>), [Ollama](<https://devfeed.tech/topics/ollama.md>)

Tags: [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [news](<https://devfeed.tech/tags/news.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [ollama](<https://devfeed.tech/tags/ollama.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [windows](<https://devfeed.tech/tags/windows.md>)

### AI overview

The article reports that Perplexity's Portable Computer, a local version of its Computer agent, is available in the Perplexity app for Windows on compatible Nvidia GeForce RTX and RTX PRO GPUs. It requires at least 24GB of VRAM and combines local models, orchestration, a browser, tool calling, and a proprietary SPACE sandbox. The article also discusses platform-specific engineering, external service connectors, and the boundary between local and cloud computing.

### Source excerpt

Running an LLM on your PC is easy enough, but putting an agent to work there is a different story. The post Perplexity's new agent runs entirely on your GPU -- with one expensive catch appeared first on The New Stack.

## Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX

DevFeed: [Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX](<https://devfeed.tech/articles/perplexity-portable-computer-is-now-available-on-windows-powered-by-nvidia-rtx-21586.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/local-ai-perplexity-windows-pcs/>)

Author: Gerardo Delgado

Published: 2026-09-14T15:00:52Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [Local AI](<https://devfeed.tech/topics/local-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [NVIDIA RTX](<https://devfeed.tech/topics/nvidia-rtx.md>), [Windows](<https://devfeed.tech/topics/windows.md>), [GeForce](<https://devfeed.tech/topics/geforce.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [NVIDIA DGX](<https://devfeed.tech/topics/nvidia-dgx.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [Slack](<https://devfeed.tech/topics/slack.md>), [Google](<https://devfeed.tech/topics/google.md>), [Microsoft](<https://devfeed.tech/topics/microsoft.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [drive](<https://devfeed.tech/tags/drive.md>), [geforce](<https://devfeed.tech/tags/geforce.md>), [github](<https://devfeed.tech/tags/github.md>), [google](<https://devfeed.tech/tags/google.md>), [local-ai](<https://devfeed.tech/tags/local-ai.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [nvidia-dgx](<https://devfeed.tech/tags/nvidia-dgx.md>), [nvidia-rtx](<https://devfeed.tech/tags/nvidia-rtx.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [rtx-pro](<https://devfeed.tech/tags/rtx-pro.md>), [rtx-spark](<https://devfeed.tech/tags/rtx-spark.md>), [slack](<https://devfeed.tech/tags/slack.md>), [windows](<https://devfeed.tech/tags/windows.md>)

### AI overview

Perplexity is adding Portable Computer to its Windows app for compatible NVIDIA GeForce RTX PCs and NVIDIA RTX PRO Workstations. The local agent uses NVIDIA-accelerated models to plan multistep tasks, analyze files, and keep sensitive information on the device, while users can authorize cloud support for more advanced research and reasoning.

### Source excerpt

As local models become more capable, AI agents can handle more work directly on a PC while keeping sensitive information on the device. Portable Computer is a local version of the agent Perplexity Computer that plans and carries out multistep tasks. Accelerated by NVIDIA GPUs, it uses local models to analyze data, bring together information [...]

## Optimize vLLM speculative decoding with FastMTP heads

DevFeed: [Optimize vLLM speculative decoding with FastMTP heads](<https://devfeed.tech/articles/optimize-vllm-speculative-decoding-with-fastmtp-heads-12348.md>)

Original publisher: [Read original article](<https://developers.redhat.com/articles/2026/09/08/optimize-vllm-speculative-decoding-fastmtp-heads>)

Author: Rahul Tuli

Published: 2026-09-08T14:20:16Z

Content type: article

Language: en

Sources: [Red Hat](<https://devfeed.tech/sources/red-hat.md>), [Red Hat Developer](<https://devfeed.tech/sources/red-hat-developer.md>)

Topics: [vllm](<https://devfeed.tech/topics/vllm.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>), [qwen](<https://devfeed.tech/topics/qwen.md>)

Tags: [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [data](<https://devfeed.tech/tags/data.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [developer-tools](<https://devfeed.tech/tags/developer-tools.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [production](<https://devfeed.tech/tags/production.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This article explains how FastMTP-style fine-tuning improves vLLM speculative decoding. It describes using native multi-token prediction heads as speculators, adapting a single head for recursive multi-step drafting, extracting weights from verifier checkpoints, and producing vLLM-ready checkpoints without training from scratch.

### Source excerpt

Autoregressive decoding makes large language model (LLM) inference memory-bandwidth bound: every token needs 1 full forward pass over billions of parameters, so the hardware spends most of its time moving weights rather than computing. MTP is a training objective: models like the DeepSeek and Qwen families learn to predict several future tokens at each position, which improves their data efficiency and quality. The post Optimize vLLM speculative decoding with FastMTP heads appeared first on Red Hat Developer.

## Qwen 3.8 Max 0902 now available on AI Gateway

DevFeed: [Qwen 3.8 Max 0902 now available on AI Gateway](<https://devfeed.tech/articles/qwen-3-8-max-0902-now-available-on-ai-gateway-1064.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/qwen-3-8-max-0902-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-09-01T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [models](<https://devfeed.tech/tags/models.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [release](<https://devfeed.tech/tags/release.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

Qwen 3.8 Max 0902, a new Alibaba snapshot, is now available through AI Gateway. The release focuses on coding for larger projects, unsupervised long-horizon tasks, agent runs, and improved vision handling for charts and dense documents. It can be selected directly, pinned by its dated model ID, or reached through a rewrite routing rule. It is also supported in several coding-agent integrations and the model playground.

### Source excerpt

Qwen 3.8 Max 0902 from Alibaba is now available on AI Gateway. This is a new snapshot of Qwen 3.8 Max, with the gains concentrated in coding on larger projects, long-horizon work that runs without supervision, and agent runs. Vision handling is more accurate on charts and dense documents. To use Qwen 3.8 Max 0902, set model to alibaba/qwen3.8-max-0902: The dated ID pins this snapshot, so a later release will not change what your requests run against. To move existing traffic onto it without a code change, add a rewrite routing rule. The gateway substitutes the destination transparently, so an application that still asks for alibaba/qwen3.8-max runs on the new snapshot: To use it in a coding agent, see the coding agents guide, then run vercel ai-gateway coding-agents setup to connect agents like Claude Code, Codex, OpenCode, Cursor, Pi, and more and select alibaba/qwen3.8-max-0902. Try Qwen3.8-Max-0902 in the model playground. You can view all language models available on AI Gateway. Read more

## Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding

DevFeed: [Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding](<https://devfeed.tech/articles/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding-6819.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/>)

Author: Michelle Horton

Published: 2026-08-26T17:07:12Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [agentic-coding](<https://devfeed.tech/topics/agentic-coding.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [NeMo](<https://devfeed.tech/topics/nemo.md>), [sglang](<https://devfeed.tech/topics/sglang.md>), [TensorRT-LLM](<https://devfeed.tech/topics/tensorrt-llm.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [rust-ai](<https://devfeed.tech/topics/rust-ai.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [gb300-nvl72](<https://devfeed.tech/tags/gb300-nvl72.md>), [inference](<https://devfeed.tech/tags/inference.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [tensorrt-llm](<https://devfeed.tech/tags/tensorrt-llm.md>), [top-stories](<https://devfeed.tech/tags/top-stories.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This NVIDIA developer article introduces Qwen3.8-Flash-Next, a multimodal mixture-of-experts model released by Alibaba for experimentation and evaluation. It explains the model's long-context hybrid architecture, including Gated DeltaNet and Qwen Sparse Attention, and discusses reported efficiency improvements for million-token workloads. The article also covers inference support through SGLang, vLLM, TensorRT-LLM, and NVIDIA NeMo, plus performance on the NVIDIA GB300 NVL72 platform.

### Source excerpt

Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It's...

## Qwen 3.8 Flash now available on AI Gateway

DevFeed: [Qwen 3.8 Flash now available on AI Gateway](<https://devfeed.tech/articles/qwen-3-8-flash-now-available-on-ai-gateway-1063.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/qwen-3-8-flash-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-08-26T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [codex](<https://devfeed.tech/topics/codex.md>), [cursor](<https://devfeed.tech/topics/cursor.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [api-keys](<https://devfeed.tech/tags/api-keys.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [cost](<https://devfeed.tech/tags/cost.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [opencode](<https://devfeed.tech/tags/opencode.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [routing](<https://devfeed.tech/tags/routing.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Qwen 3.8 Flash from Alibaba is now available through Vercel AI Gateway. The model accepts text and images, supports a 1-million-token context window, and can generate responses of up to 65,000 tokens. It is recommended for coding, tool use, and multi-step agent workflows, and can be used through the AI SDK, coding agents, and the model playground. AI Gateway offers unified model access, usage and cost tracking, reliability features, reporting, retention controls, API key budgets, routing, and provider pricing without markup or inference platform fees.

### Source excerpt

Qwen 3.8 Flash from Alibaba is now available on AI Gateway. It takes text and images as input, serves a context window of 1 million tokens, and can return up to 65k tokens in a response. Alibaba recommends it for coding, tool use, and multi-step agent workflows. To use Qwen3.8-Flash, set model to alibaba/qwen3.8-flash in the AI SDK: To use it in a coding agent, see the coding agents guide, then run vercel ai-gateway coding-agents setup to connect agents like Claude Code, Codex, OpenCode, Cursor, Pi, and more and select alibaba/qwen3.8-flash inside the agent. Try Qwen3.8-Flash in the model playground. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Read more

## Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

DevFeed: [Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things](<https://devfeed.tech/articles/qwen-3-8-27b-is-excellent-but-it-defaults-to-wildly-overthinking-things-30498.md>)

Original publisher: [Read original article](<https://simonwillison.net/2026/Aug/16/qwen-38-27b/>)

Author: Simon Willison

Published: 2026-08-16T22:00:39Z

Content type: opinion

Language: en

Sources: [Simon Willison](<https://devfeed.tech/sources/simon-willison.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [DGX Spark](<https://devfeed.tech/topics/dgx-spark.md>), [NVIDIA DGX](<https://devfeed.tech/topics/nvidia-dgx.md>), [SVG](<https://devfeed.tech/topics/svg.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-2-235](<https://devfeed.tech/tags/ai-2-235.md>), [ai-in-china](<https://devfeed.tech/tags/ai-in-china.md>), [ai-in-china-108](<https://devfeed.tech/tags/ai-in-china-108.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [coding-agents-248](<https://devfeed.tech/tags/coding-agents-248.md>), [cost](<https://devfeed.tech/tags/cost.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [generative-ai-1-981](<https://devfeed.tech/tags/generative-ai-1-981.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llama-cpp-29](<https://devfeed.tech/tags/llama-cpp-29.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llm-reasoning](<https://devfeed.tech/tags/llm-reasoning.md>), [llm-reasoning-103](<https://devfeed.tech/tags/llm-reasoning-103.md>), [llm-release](<https://devfeed.tech/tags/llm-release.md>), [llm-release-231](<https://devfeed.tech/tags/llm-release-231.md>), [llms](<https://devfeed.tech/tags/llms.md>), [llms-1-947](<https://devfeed.tech/tags/llms-1-947.md>), [lm-studio](<https://devfeed.tech/tags/lm-studio.md>), [lm-studio-23](<https://devfeed.tech/tags/lm-studio-23.md>), [local-llms](<https://devfeed.tech/tags/local-llms.md>), [local-llms-164](<https://devfeed.tech/tags/local-llms-164.md>), [nvidia-dgx](<https://devfeed.tech/tags/nvidia-dgx.md>), [nvidia-spark](<https://devfeed.tech/tags/nvidia-spark.md>), [nvidia-spark-6](<https://devfeed.tech/tags/nvidia-spark-6.md>), [pelican-riding-a-bicycle](<https://devfeed.tech/tags/pelican-riding-a-bicycle.md>), [pelican-riding-a-bicycle-142](<https://devfeed.tech/tags/pelican-riding-a-bicycle-142.md>), [pi](<https://devfeed.tech/tags/pi.md>), [pi-6](<https://devfeed.tech/tags/pi-6.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [qwen-61](<https://devfeed.tech/tags/qwen-61.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [speed](<https://devfeed.tech/tags/speed.md>), [svg](<https://devfeed.tech/tags/svg.md>)

### AI overview

Simon Willison evaluates Qwen 3.8 27B, a vision-capable 27-billion-parameter LLM that can run locally on suitable hardware. He finds that its default xhigh reasoning setting consumes substantial context and time, while adjusting the reasoning effort and increasing the context limit improves practicality. He also reports strong results generating an SVG locally.

### Source excerpt

Friday's big release was Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor Qwen 3.6 27B was impressive. Qwen's self-reported benchmarks for this model are eye-opening. They show a boost from both Qwen 3.6 27B and the closed-weight Qwen 3.7-Plus, which was one of Qwen's strongest models of any size as recently as May this year. It will be interesting to hear what independent benchmarks have to say about the model. I've been running the model on two different machines: my 128GB M5 Max MacBook Pro, and an NVIDIA DGX Spark. On both machines I'm running LM Studio and their 17GB Q4_K_M quantized build. I also tried using llama-server directly on the Spark. The default of extra high results in spectacular over-thinking Qwen's documentation describes the model as defaulting to xhigh for the reasoning effort, and the LM Studio GGUF I've been trying preserves that default: Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost: xhigh (default): for complex tasks demanding thorough analysis medium: balancing accuracy and speed low: efficient reasoning optimizing for speed and cost This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware. I've been finding the results extremely entertaining. I quickly ran into problems with LM Studio's default context limit of 8,192 tokens - Qwen was using them all up thinking about even the most mundane of problems. I loaded the model with the full 262,144 maximum context length and that problem went away. Here's the pelican riding a bicycle SVG I got from my first attempt with that increased context length. It took 21 minutes to generate, using 22,276 reasoning tokens to produce 3,223 tokens of output. You can read the reasoning trace here. Th

## State of Open Models: Summer 2026 Observations

DevFeed: [State of Open Models: Summer 2026 Observations](<https://devfeed.tech/articles/state-of-open-models-summer-2026-observations-7490.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/state-of-open-models-summer-2026>)

Author: Adina Yakefu; Apolinário from multimodal AI art; Irene Solaiman

Published: 2026-08-14T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [spaces](<https://devfeed.tech/topics/spaces.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [hub](<https://devfeed.tech/tags/hub.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [research](<https://devfeed.tech/tags/research.md>), [spaces](<https://devfeed.tech/tags/spaces.md>)

### AI overview

This article examines the summer 2026 state of open models, highlighting rapid growth in public model repositories, datasets, and Spaces; the dominance of a small number of repositories in downloads; the rising scale of Chinese open models; differing model portfolio strategies; and the strong role of AMD, NVIDIA, and community quantization in making large models accessible.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Qwen 3.8 Max now available on Vercel AI Gateway

DevFeed: [Qwen 3.8 Max now available on Vercel AI Gateway](<https://devfeed.tech/articles/qwen-3-8-max-now-available-on-vercel-ai-gateway-1065.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/qwen-3-8-max-now-available-on-vercel-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-08-02T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [code](<https://devfeed.tech/tags/code.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Qwen 3.8 Max is now available through Vercel AI Gateway. The 2.4-trillion-parameter model supports text-only and vision-language tasks, with a context window of up to 1 million tokens. It is positioned for software engineering, office productivity, and visual workflows such as converting screenshots or design files into working pages. AI Gateway also supports model playground access, coding-agent integrations, unified API calls, usage and cost tracking, routing, retries, failover, and provider pricing without a platform fee on inference.

### Source excerpt

Qwen 3.8 Max is now available on AI Gateway. Qwen 3.8 Max handles text-only and vision-language work in one model, with 2.4 trillion parameters and a context window of up to 1 million tokens. The model is suited for software engineering and office productivity, along with visual work like turning screenshots or design files into working pages, captioning video, and answering questions grounded in an image. To use Qwen 3.8 Max, set model to alibaba/qwen3.8-max. Try Qwen 3.8 Max in the model playground. To use it in a coding agent, run vercel ai-gateway coding-agents setup to connect Claude Code, Codex, OpenCode, or Pi, then select alibaba/qwen3.8-max inside the agent. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Read more

## 🗓 This Week In AI Research (1-8 July 26)

DevFeed: [🗓 This Week In AI Research (1-8 July 26)](<https://devfeed.tech/articles/this-week-in-ai-research-1-8-july-26-18283.md>)

Original publisher: [Read original article](<https://www.intoai.pub/p/this-week-in-ai-research-1-8-july>)

Author: Dr. Ashish Bamania

Published: 2026-07-12T11:25:32Z

Content type: article

Language: en

Sources: [Into AI](<https://devfeed.tech/sources/into-ai.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [Algorithms](<https://devfeed.tech/topics/algorithms.md>), [releases](<https://devfeed.tech/topics/releases.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [algorithm](<https://devfeed.tech/tags/algorithm.md>), [algorithms](<https://devfeed.tech/tags/algorithms.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [releases](<https://devfeed.tech/tags/releases.md>), [research](<https://devfeed.tech/tags/research.md>), [rl](<https://devfeed.tech/tags/rl.md>), [training](<https://devfeed.tech/tags/training.md>), [update](<https://devfeed.tech/tags/update.md>)

### AI overview

A weekly roundup of AI research papers and releases highlights findings that reinforcement-learning gains can be concentrated in a single transformer layer and presents LLM-as-a-Verifier, a framework for continuous scoring and ranking of agentic-task solutions.

### Source excerpt

The top 10 research papers and AI releases this week (SpaceXAI's Grok 4.5, OpenAI's GPT-Live voice models, Cognition's SWE-1.7, Meta's Muse Spark 1.1, and many more)

## Using local Gemma and Qwen models to triage OpenClaw issues and pull requests

DevFeed: [Using local Gemma and Qwen models to triage OpenClaw issues and pull requests](<https://devfeed.tech/articles/we-got-local-models-to-triage-the-openclaw-repo-for-free-7341.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/local-models-pr-triage>)

Author: Onur Solmaz; ben burtenshaw; shaun smith

Published: 2026-06-22T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [OpenClaw](<https://devfeed.tech/topics/openclaw.md>), [Agent Harness](<https://devfeed.tech/topics/agent-harness.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [gemma](<https://devfeed.tech/topics/gemma.md>), [qwen](<https://devfeed.tech/topics/qwen.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agent-harness](<https://devfeed.tech/tags/agent-harness.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [free](<https://devfeed.tech/tags/free.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [guide](<https://devfeed.tech/tags/guide.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-collab](<https://devfeed.tech/tags/open-source-collab.md>), [openclaw](<https://devfeed.tech/tags/openclaw.md>), [pull-requests](<https://devfeed.tech/tags/pull-requests.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

The article describes using local Gemma and Qwen models in an agent harness to classify and triage issues and pull requests in the OpenClaw repository. It presents local execution as a way to support near-real-time notifications without relying on a paid hosted-model quota, using structured outputs and a finite label set.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Local Qwen isn't a worse Opus, it's a different tool

DevFeed: [Local Qwen isn't a worse Opus, it's a different tool](<https://devfeed.tech/articles/local-qwen-isn-t-a-worse-opus-it-s-a-different-tool-26644.md>)

Original publisher: [Read original article](<https://blog.alexellis.io/local-ai-is-not-opus/>)

Author: Alex Ellis

Published: 2026-06-17T00:00:00Z

Content type: opinion

Language: en

Sources: [Alex Ellis' Blog](<https://devfeed.tech/sources/alex-ellis-blog.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-tools](<https://devfeed.tech/tags/ai-tools.md>), [claude](<https://devfeed.tech/tags/claude.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [llm](<https://devfeed.tech/tags/llm.md>), [localai](<https://devfeed.tech/tags/localai.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [source](<https://devfeed.tech/tags/source.md>)

### AI overview

The author shares a founder's experience using local Qwen models in a small software business and open source projects. The models have provided useful, specific value, but the author still does not trust them unsupervised because of infinite loops and hallucination risks, especially after quantization for a consumer GPU.

### Source excerpt

We've all heard people say that Qwen is near-Opus level, but I have receipts and am here to be transparent with you.

## Holo3.1: Fast & Local Computer Use Agents

DevFeed: [Holo3.1: Fast & Local Computer Use Agents](<https://devfeed.tech/articles/holo3-1-fast-local-computer-use-agents-7004.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/Hcompany/holo31>)

Author: Maxime Langevin; Hamza Benchekroun; Axel Moyal; Emrick Sinitambirivoutin; Antonio Loison; Avshalom Manevich; Tony Wu; Pierre-Louis Cedoz; Aurélien Lac; Ronan Riochet

Published: 2026-06-02T14:13:23Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [computer-use](<https://devfeed.tech/topics/computer-use.md>), [On-device AI](<https://devfeed.tech/topics/on-device-ai.md>), [Local AI](<https://devfeed.tech/topics/local-ai.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [NVFP4](<https://devfeed.tech/topics/nvfp4.md>), [browser](<https://devfeed.tech/topics/browser.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [browser](<https://devfeed.tech/tags/browser.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [computer-use](<https://devfeed.tech/tags/computer-use.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [devices](<https://devfeed.tech/tags/devices.md>), [inference](<https://devfeed.tech/tags/inference.md>), [json](<https://devfeed.tech/tags/json.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nvfp4](<https://devfeed.tech/tags/nvfp4.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [performance](<https://devfeed.tech/tags/performance.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [software](<https://devfeed.tech/tags/software.md>)

### AI overview

Holo3.1 is a family of computer-use models designed to operate across web, desktop, and mobile environments and integrate with different agent frameworks. The release adds quantized checkpoints for local inference, native function-calling support, and model sizes ranging from 0.8B to 35B-A3B, targeting private, cost-effective, and high-performance deployments.

### Source excerpt

Users want to run the same computer-use capabilities across desktop and mobile environments, with seamless integration with different agent frameworks. They want deployment flexibility, from cloud inference to fully local execution on end-user devices. This is why we are releasing the Holo3.1 family. Holo3.1 improves robustness across the three dimensions that matter most in production: environments (web, desktop, mobile), agent frameworks, and deployment targets.

## Qwen 3.7 Max now available on Vercel AI Gateway

DevFeed: [Qwen 3.7 Max now available on Vercel AI Gateway](<https://devfeed.tech/articles/qwen-3-7-max-now-available-on-vercel-ai-gateway-1061.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/qwen-3-7-max-now-available-on-vercel-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-05-21T07:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Front end](<https://devfeed.tech/topics/frontend.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [API](<https://devfeed.tech/topics/api.md>), [observability](<https://devfeed.tech/topics/observability.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [automation](<https://devfeed.tech/tags/automation.md>), [coding](<https://devfeed.tech/tags/coding.md>), [cost](<https://devfeed.tech/tags/cost.md>), [frontend](<https://devfeed.tech/tags/frontend.md>), [observability](<https://devfeed.tech/tags/observability.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Qwen 3.7 Max from Alibaba is now available through Vercel AI Gateway. The model targets coding, office workflow automation, frontend prototyping, complex multi-file engineering, multi-agent orchestration, and long-horizon tool-calling. AI Gateway provides model access through the AI SDK along with usage and cost tracking, retries, failover, provider routing, observability, and performance features.

### Source excerpt

Qwen 3.7 Max from Alibaba is now available on Vercel AI Gateway. The model is designed as an agent foundation, with capabilities spanning coding, office workflow automation, and long-horizon autonomous execution. Qwen 3.7 Max shows improvements in frontend prototyping and complex multi-file engineering. The model supports office and productivity tasks through multi-agent orchestration and sustains coherent reasoning across long-horizon tool-calling sessions. To use Qwen 3.7 Max, set model to alibaba/qwen-3.7-max in the AI SDK. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, observability, Bring Your Own Key support, and intelligent provider routing with automatic retries. Learn more about AI Gateway, view the AI Gateway model leaderboard or try it in our model playground. Read more

## How DigitalOcean optimized DeepSeek V3.2, MiniMax-M2.5, and Qwen 3.5 397B for Serverless Inference

DevFeed: [How DigitalOcean optimized DeepSeek V3.2, MiniMax-M2.5, and Qwen 3.5 397B for Serverless Inference](<https://devfeed.tech/articles/how-we-built-the-most-performant-deepseek-v3-2-minimax-m2-5-and-qwen-3-5-397b-on-digitalocean-serverless-inference-19888.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/how-we-built-fastest-deepseek-minimax-qwen-on-blackwell-ultra>)

Author: Bhaskar Dutt

Published: 2026-04-28T09:00:00Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>)

Tags: [deepseek](<https://devfeed.tech/tags/deepseek.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [low-latency-inference](<https://devfeed.tech/tags/low-latency-inference.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [serverless](<https://devfeed.tech/tags/serverless.md>)

### AI overview

DigitalOcean announces general availability of DeepSeek V3.2, MiniMax-M2.5, and Qwen 3.5 397B on DigitalOcean Serverless Inference. The article describes GPU-level optimization and serving-stack tuning, reporting 230 output tokens per second and sub-one-second time to first token for DeepSeek V3.2, plus top output-speed results in Artificial Analysis testing for DeepSeek V3.2 and Qwen 3.5 397B.

### Source excerpt

Today at Deploy, we are announcing the general availability of DeepSeek V3.2, MiniMax-M2.5, and Qwen 3.5 397B on DigitalOcean Serverless Inference. On DeepSeek V3.2 and Qwen 3.5 397B, we deliver #1 output speed across all providers Artificial Analysis tested. On DeepSeek V3.2 specifically, that translates to 230 output tokens per second and sub-1-second Time-to-First-Token (TTFT) for 10,000 input tokens. This post covers how we got there: the GPU-level work, the serving stack tuning, and the specific technical tradeoffs we made along the way. Why fast inference matters The focus in AI development has fundamentally shifted from the training of models to the efficiency of inference. This shift is driven by the proliferation of agentic workloads, copilots, and real-time systems that form the core of next-generation AI applications. For these applications, speed is no longer just a performance metric; it is the critical differentiator between an engaging product and one that users abandon. Specifically, low-latency inference is essential for a seamless end-user experience. For highly interactive applications like conversational agents and voice interfaces, any delay beyond a sub-1-second TTFT is perceived as sluggish. The importance of fast inference is compounded by the complexity of modern AI workflows. An agentic task, for instance, often involves dozens of sequential model calls, where even minute Time-Per-Output-Token (TPOT) delays can accumulate into several seconds of user-visible latency. Quick inference also helps businesses by providing reliable performance and lower costs. Optimization in this area, such as that provided by DigitalOcean's inference engine, allows enterprises to achieve superior token economics, sustained throughput, and predictable latency, which are essential for scaling their AI-native applications reliably and affordably. Leading the Artificial Analysis benchmarks on speed The benchmarks we're publishing today reflect this. On DeepSeek V3.

## Ecom-RLVE: Adaptive Verifiable Environments for E-Commerce Conversational Agents

DevFeed: [Ecom-RLVE: Adaptive Verifiable Environments for E-Commerce Conversational Agents](<https://devfeed.tech/articles/ecom-rlve-adaptive-verifiable-environments-for-e-commerce-conversational-agents-7178.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ecom-rlve>)

Author: Rahul Bajaj; Jaya Nupur; Anuj Garg; ben burtenshaw

Published: 2026-04-16T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [rlvr](<https://devfeed.tech/topics/rlvr.md>), [openenv](<https://devfeed.tech/topics/openenv.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [customer-service](<https://devfeed.tech/tags/customer-service.md>), [e-commerce](<https://devfeed.tech/tags/e-commerce.md>), [llm](<https://devfeed.tech/tags/llm.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [openenv](<https://devfeed.tech/tags/openenv.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [rlvr](<https://devfeed.tech/tags/rlvr.md>), [tools](<https://devfeed.tech/tags/tools.md>), [training](<https://devfeed.tech/tags/training.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Ecom-RLVE introduces EcomRLVE-GYM, a suite of verifiable, multi-turn, tool-augmented e-commerce environments for training conversational agents. It uses procedural task generation, adaptive difficulty, and algorithmically verifiable rewards across shopping and customer-service workflows, with early results from training a Qwen 3 8B model using DAPO.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Qwen 3.6 Plus on AI Gateway

DevFeed: [Qwen 3.6 Plus on AI Gateway](<https://devfeed.tech/articles/qwen-3-6-plus-on-ai-gateway-1066.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/qwen-3.6-plus-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-04-02T07:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [agentic-coding](<https://devfeed.tech/topics/agentic-coding.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [coding](<https://devfeed.tech/topics/coding.md>), [context window](<https://devfeed.tech/topics/context-window.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [API](<https://devfeed.tech/topics/api.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>)

Tags: [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [api](<https://devfeed.tech/tags/api.md>), [coding](<https://devfeed.tech/tags/coding.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [frontend](<https://devfeed.tech/tags/frontend.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [observability](<https://devfeed.tech/tags/observability.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [routing](<https://devfeed.tech/tags/routing.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Qwen 3.6 Plus from Alibaba is now available through Vercel AI Gateway. The model improves agentic coding, multimodal perception, reasoning, tool calling, long-horizon planning, and multilingual tasks, and provides a 1M context window.

### Source excerpt

Qwen 3.6 Plus from Alibaba is now available on Vercel AI Gateway. Compared to Qwen 3.5 Plus, this model adds stronger agentic coding capabilities, from frontend development to repository-level problem solving, along with improved multimodal perception and reasoning. It features a 1M context window and improved performance on tool-calling, long-horizon planning, and multilingual tasks. To use Qwen 3.6 Plus, set model to qwen/qwen3.6-plus in the AI SDK. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, observability, Bring Your Own Key support, and intelligent provider routing with automatic retries. Learn more about AI Gateway, view the AI Gateway model leaderboard or try it in our model playground. Read more

## Ask AIMee: An accessible accessibility-focused AI chatbot

DevFeed: [Ask AIMee: An accessible accessibility-focused AI chatbot](<https://devfeed.tech/articles/ask-aimee-an-accessible-accessibility-focused-ai-chatbot-9441.md>)

Original publisher: [Read original article](<https://webaim.org/blog/ask-aimee/>)

Author: Jared Smith

Published: 2026-03-31T16:49:07Z

Content type: release

Language: en

Sources: [WebAIM Blog](<https://devfeed.tech/sources/webaim-blog.md>)

Topics: [Accessibility](<https://devfeed.tech/topics/accessibility.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Chat Bot](<https://devfeed.tech/topics/chatbot.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [accessibility](<https://devfeed.tech/tags/accessibility.md>), [ai](<https://devfeed.tech/tags/ai.md>), [chat](<https://devfeed.tech/tags/chat.md>), [chatbots](<https://devfeed.tech/tags/chatbots.md>), [hallucinations](<https://devfeed.tech/tags/hallucinations.md>), [llm](<https://devfeed.tech/tags/llm.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [uncategorized](<https://devfeed.tech/tags/uncategorized.md>)

### AI overview

WebAIM introduces AIMee, an AI-powered conversational chatbot focused on accessibility. It is designed to support users with disabilities by answering accessibility questions, reviewing code, drafting policies, explaining concepts, and generating resources, checklists, and quizzes. AIMee primarily uses the Qwen 3 Coder large language model with additional safeguards, but its answers can still contain hallucinations and should be verified.

### Source excerpt

We're happy to introduce AIMee - an easy-to-use, AI-powered conversational chatbot focused on accessibility. AIMee has been designed to be highly accessible to users with disabilities. Ask her accessibility questions to get quick answers and guidance. The name "AIMee" plays off of the "AIM" (Accessibility In Mind) from "WebAIM" and also "AI". Here are some [...]

## The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+

DevFeed: [The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+](<https://devfeed.tech/articles/the-future-of-the-global-open-source-ai-ecosystem-from-deepseek-to-ai-7252.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/huggingface/one-year-since-the-deepseek-moment-blog-3>)

Author: Adina Yakefu; Irene Solaiman

Published: 2026-02-03T15:03:19Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [article](<https://devfeed.tech/tags/article.md>), [blog](<https://devfeed.tech/tags/blog.md>), [china](<https://devfeed.tech/tags/china.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [ecosystem](<https://devfeed.tech/tags/ecosystem.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [integration](<https://devfeed.tech/tags/integration.md>), [meta](<https://devfeed.tech/tags/meta.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [updates](<https://devfeed.tech/tags/updates.md>)

### AI overview

This third and final article in a series examines the trajectories of prominent Chinese AI organizations after the January 2025 "DeepSeek Moment." It argues that open source is becoming the dominant approach for Chinese AI organizations, supported by the sharing of models, papers, techniques, and deployment infrastructure. The article highlights DeepSeek, Qwen, and other organizations in the emergence of a collaborative Chinese and global open-source AI ecosystem.

### Source excerpt

This is the third and final blog in a three-part series on China's open source community's historical advancements since January 2025's "DeepSeek Moment." The first blog on strategic changes and open artifact growth is available here, and the second blog on architectural and hardware shifts is available here. In this third article, we examine paths and trajectories of prominent Chinese AI organizations, and posit future directions for open source.

## Technical Deep Dive: How DigitalOcean and AMD Delivered a 2x Production Inference Performance Increase for Character.ai

DevFeed: [Technical Deep Dive: How DigitalOcean and AMD Delivered a 2x Production Inference Performance Increase for Character.ai](<https://devfeed.tech/articles/technical-deep-dive-how-digitalocean-and-amd-delivered-a-2x-production-inference-performance-increase-for-character-ai-19948.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/technical-deep-dive-character-ai-amd>)

Author: Karnik Modi

Published: 2026-01-13T12:30:00Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [GPU optimization](<https://devfeed.tech/topics/gpu-optimization.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>), [qwen](<https://devfeed.tech/topics/qwen.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [news](<https://devfeed.tech/tags/news.md>), [performance](<https://devfeed.tech/tags/performance.md>), [platforms](<https://devfeed.tech/tags/platforms.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

This technical deep dive describes how Character.ai, AMD, and DigitalOcean optimized inference for the Qwen3-235B Instruct FP8 model on AMD Instinct MI300X and MI325X GPUs. The reported optimizations, including parallelization, FP8 execution paths, optimized kernels, topology-aware allocation, and Kubernetes orchestration, produced up to a 2x improvement in production request throughput under stated latency and concurrency constraints.

### Source excerpt

Background: How Character.ai worked with DigitalOcean and AMD to optimize performance Character.ai, a leading AI entertainment platform with about 20 million worldwide users, wanted to optimize GPU performance and achieve lower inference costs for its application, which requires low-latency performance at large scale. They approached DigitalOcean and AMD in order to achieve this goal. Working closely together, the Character.ai, AMD, and DigitalOcean teams optimized AMD Instinct™ MI300X and MI325X GPU platforms, resulting in a 2x production inference throughput. In optimized configurations, DigitalOcean delivered high request density per node while maintaining exceptional p90 responsiveness for initial token and sustained token generation throughput, outperforming prior deployments on generic, non-optimized GPU infrastructure. These gains were achieved through platform-level optimizations, including clever parallelization strategies for large Mixture-of-Experts models, efficient FP8 execution paths, optimized kernels with AITER, topology-aware GPU allocation, and production-ready Kubernetes orchestration through DigitalOcean Kubernetes (DOKS). Together, these capabilities allowed Character.ai to scale inference predictably without increasing operational burden. In this post, we will explore the specific orchestration and tuning strategies that made these gains possible. Technical deep dive overview Character.ai leverages multiple models like Qwen, Mistral and more to power their applications. This document is focused on how we optimized the Qwen3-235B Instruct FP8 model on a cluster of DigitalOcean featuring AMD Instinct GPUs. This workload was migrated from a generic, non-optimized setup on other providers to AMD Instinct™ MI325X platform on DigitalOcean, and following the outlined optimizations we were able to achieve up to a 2x improvement in request throughput (QPS) under strict latency and concurrency constraints. The Character.ai team has a demanding workload,

## LLM Duel: Gemini 3 Flash vs Opus 4.5 vs GPT-5.2-Codex vs GLM-4.7

DevFeed: [LLM Duel: Gemini 3 Flash vs Opus 4.5 vs GPT-5.2-Codex vs GLM-4.7](<https://devfeed.tech/articles/llm-duel-gemini-3-flash-vs-opus-4-5-vs-gpt-5-2-codex-vs-glm-4-7-27199.md>)

Original publisher: [Read original article](<https://antonioleiva.com/llm-duel-gemini-opus-gpt-glm>)

Published: 2025-12-23T00:00:00Z

Content type: comparison

Language: en

Sources: [Antonio Leiva](<https://devfeed.tech/sources/antonio-leiva.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [codex](<https://devfeed.tech/topics/codex.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [codex](<https://devfeed.tech/tags/codex.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [llm](<https://devfeed.tech/tags/llm.md>), [qwen](<https://devfeed.tech/tags/qwen.md>)

### AI overview

The article compares Gemini 3 Flash, Opus 4.5, GPT-5.2-Codex, and GLM-4.7 through a practical coding project: building an image editor for YouTube thumbnails. It evaluates not only code generation but also task planning and preliminary research, while questioning how well benchmarks reflect real-world use.

### Source excerpt

Everything Android, Kotlin and other random topics

## Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models

DevFeed: [Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models](<https://devfeed.tech/articles/accelerating-qwen3-8b-agent-on-intel-coretm-ultra-with-depth-pruned-draft-models-7292.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/intel-qwen3-agent>)

Author: Igor Margulis; Ofir Zafrir; Shira Guskin; Guy Boudoukh; Pedro Cuenca

Published: 2025-09-29T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [intel](<https://devfeed.tech/topics/intel.md>), [smolagents](<https://devfeed.tech/topics/smolagents.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [guide](<https://devfeed.tech/tags/guide.md>), [hub](<https://devfeed.tech/tags/hub.md>), [inference](<https://devfeed.tech/tags/inference.md>), [intel](<https://devfeed.tech/tags/intel.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [smolagents](<https://devfeed.tech/tags/smolagents.md>)

### AI overview

The article explains how to accelerate the Qwen3-8B agent on an Intel Lunar Lake integrated GPU using OpenVINO.GenAI. Speculative decoding with Qwen3-0.6B as a draft model achieves about a 1.3x speedup, while pruning the draft model increases the speedup to about 1.4x. It also demonstrates running a fast, local AI agent with smolagents.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

[Next page](<https://devfeed.tech/topics/qwen.md?cursor=WyIyMDI1LTA5LTI5VDAwOjAwOjAwKzAwOjAwIiwgIjRhYWJmODIzLTA4ZDYtNDQxZC1hNWM2LTRmZTE2MTk3NGI4MCJd>)