# Reasoning

Published articles for Reasoning.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Qwen выпустила Qwen3.8-Omni-Flash

DevFeed: [Qwen выпустила Qwen3.8-Omni-Flash](<https://devfeed.tech/articles/qwen-qwen3-8-omni-flash-42727.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/bothub/news/1083716/>)

Author: DashaPasha (BotHub)

Published: 2026-09-18T08:40:54Z

Content type: release

Language: ru

Sources: [Tagir Valeev](<https://devfeed.tech/sources/tagir-valeev.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [qwen3](<https://devfeed.tech/topics/qwen3.md>), [API](<https://devfeed.tech/topics/api.md>), [function calling](<https://devfeed.tech/topics/function-calling.md>)

Tags: [alibaba](<https://devfeed.tech/tags/alibaba.md>), [api](<https://devfeed.tech/tags/api.md>), [function-calling](<https://devfeed.tech/tags/function-calling.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [qwen-3-8](<https://devfeed.tech/tags/qwen-3-8.md>), [qwen3](<https://devfeed.tech/tags/qwen3.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [responses](<https://devfeed.tech/tags/responses.md>), [tag-61cd5a476b1d](<https://devfeed.tech/tags/tag-61cd5a476b1d.md>), [tag-67847f6cdd99](<https://devfeed.tech/tags/tag-67847f6cdd99.md>), [tag-6e9d47cd6240](<https://devfeed.tech/tags/tag-6e9d47cd6240.md>), [tag-dc29f2a3d816](<https://devfeed.tech/tags/tag-dc29f2a3d816.md>)

### AI overview

Qwen introduced Qwen3.8-Omni-Flash, a multimodal model focused on long-video understanding, detailed reporting, content description, and more precise timestamps. It accepts text, images, audio, and video, supports function calling, web search, reasoning controls, and API access through Chat Completions and Responses.

### Source excerpt

Команда Qwen представила Qwen3.8-Omni-Flash -- новую мультимодальную модель. Главный акцент в новой версии сделали на понимании видео. Qwen3.8-Omni-Flash умеет анализировать длинные ролики, составлять по ним подробные отчёты и генерировать описания содержимого. В Qwen отдельно отмечают улучшение работы с временными метками: модель должна точнее связывать события с конкретными моментами видео. При этом текстовые возможности модели сохранили на уровне Qwen3.8-Flash. В API также доступны function calling и веб-поиск, а режим рассуждений включён по умолчанию и может настраиваться через reasoning_effort. Qwen3.8-Omni-Flash -- промежуточное звено для более сложных мультимодальных пайплайнов. Среди возможных сценариев разрабы называют монтаж видео, перевод, создание кинокомментариев, анализ и генерацию контента. Отдельный фокус релиза -- стоимость. По сравнению с предыдущей Qwen3.5-Omni-Flash расход токенов при работе с мультимодальными данными заметно снизился. В частности, Qwen заявляет о снижении стоимости обработки аудио примерно на 98%, а аудио-видео -- более чем на 93%. При этом заявленная производительность сопоставима с Gemini 3.8 Flash. Независимые данные также показывают снижение расхода токенов на одном из видеобенчмарков: с 145 736 до 79 117 токенов при росте результата OmniVideoBench с 63,4 до 67,8. Контекстное окно новой модели -- 1 млн токенов, максимальная длина генерируемого ответа -- 131 072 токена. Qwen3.8-Omni-Flash принимает текст, изображения, аудио и видео, а на выходе генерирует текст. Читать далее

## Understanding W8A8 INT8 LLM quantization: Accuracy and performance results

DevFeed: [Understanding W8A8 INT8 LLM quantization: Accuracy and performance results](<https://devfeed.tech/articles/understanding-w8a8-int8-llm-quantization-accuracy-and-performance-results-17433.md>)

Original publisher: [Read original article](<https://developers.redhat.com/articles/2026/09/14/understanding-w8a8-int8-llm-quantization-accuracy-and-performance-results>)

Author: Sana Fayyaz

Published: 2026-09-14T13:01:43Z

Content type: article

Language: en

Sources: [Red Hat](<https://devfeed.tech/sources/red-hat.md>), [Red Hat Developer](<https://devfeed.tech/sources/red-hat-developer.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [llama](<https://devfeed.tech/topics/llama.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [Algorithms](<https://devfeed.tech/topics/algorithms.md>)

Tags: [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [compression](<https://devfeed.tech/tags/compression.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

The article evaluates W8A8 INT8 quantization of a Llama 3.1 8B Instruct model. It describes reducing the model from 14.9 GB to 8.0 GB with SmoothQuant and GPTQ, then compares the base and compressed models on four benchmarks to assess accuracy and performance.

### Source excerpt

In Understanding W8A8 INT8 LLM quantization: Half the size, better performance, same accuracy, we compressed a Llama 3.1 8B Instruct model from 14.9 GB to 8.0 GB using 8-bit integer (INT8) W8A8 quantization with SmoothQuant and Generative Pre-trained Transformer Quantization (GPTQ). The post Understanding W8A8 INT8 LLM quantization: Accuracy and performance results appeared first on Red Hat Developer.

## Chip Huyen explains how to cut inference costs without new hardware

DevFeed: [Chip Huyen explains how to cut inference costs without new hardware](<https://devfeed.tech/articles/chip-huyen-explains-how-to-cut-inference-costs-without-new-hardware-10830.md>)

Original publisher: [Read original article](<https://thenewstack.io/pg-99-conf-2026-inference-costs/>)

Author: Tim Koopmans

Published: 2026-09-13T15:00:00Z

Content type: article

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Frontier Model](<https://devfeed.tech/topics/frontier-model.md>), [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [math](<https://devfeed.tech/topics/math.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [frontier-model](<https://devfeed.tech/tags/frontier-model.md>), [inference](<https://devfeed.tech/tags/inference.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [post-contributed](<https://devfeed.tech/tags/post-contributed.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [scylladb](<https://devfeed.tech/tags/scylladb.md>), [sponsor-scylladb](<https://devfeed.tech/tags/sponsor-scylladb.md>), [sponsored](<https://devfeed.tech/tags/sponsored.md>), [sponsored-post-contributed](<https://devfeed.tech/tags/sponsored-post-contributed.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

Chip Huyen explains why inference costs can outweigh one-time frontier-model training costs and outlines ways to optimize inference without new hardware. The article emphasizes latency metrics such as time to first token, time per output token, end-to-end latency, and goodput, especially for reasoning models.

### Source excerpt

Last October, the P99 conference -- the online gathering for developers focused on high-performance, low-latency applications -- featured a cracking The post Chip Huyen explains how to cut inference costs without new hardware appeared first on The New Stack.

## Teaching AI to Reason Through Detection Triage

DevFeed: [Teaching AI to Reason Through Detection Triage](<https://devfeed.tech/articles/teaching-ai-to-reason-through-detection-triage-8310.md>)

Original publisher: [Read original article](<https://www.crowdstrike.com/en-us/blog/teaching-ai-to-reason-through-detection-triage/>)

Author: Amol Khanna - Manu Nandan - Cristian Viorel Popa - Joan Pujol-Roig - Diana Bolocan - Laura Vasilie - Alexandru Apostu - Chase Helwig - Mihaela Gaman - Mickey Brautbar - Edward Raff - Chase Midler - Sv

Published: 2026-09-12T11:17:51.295154Z

Content type: article

Language: en

Sources: [Blog](<https://devfeed.tech/sources/blog.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-soc](<https://devfeed.tech/tags/agentic-soc.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [classification](<https://devfeed.tech/tags/classification.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model](<https://devfeed.tech/tags/model.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open](<https://devfeed.tech/tags/open.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [security](<https://devfeed.tech/tags/security.md>), [soc](<https://devfeed.tech/tags/soc.md>), [windows](<https://devfeed.tech/tags/windows.md>)

### AI overview

CrowdStrike describes research on a reasoning-enabled language-model classifier for security detection triage. The model produces a verdict and an auditable rationale, with the stated goals of improving accuracy, transparency, and safe alert automation.

### Source excerpt

New CrowdStrike research shows how step-by-step reasoning can improve detection triage accuracy, transparency, and safe automation.

## Presentation: From Retrieval to Reasoning: Building Production-Ready Agentic AI Systems with Knowledge Graphs

DevFeed: [Presentation: From Retrieval to Reasoning: Building Production-Ready Agentic AI Systems with Knowledge Graphs](<https://devfeed.tech/articles/presentation-from-retrieval-to-reasoning-building-production-ready-agentic-ai-systems-with-knowledge-graphs-8463.md>)

Original publisher: [Read original article](<https://www.infoq.com/presentations/knowledge-graphs-agentic-systems-patterns/>)

Author: Cassie Shum

Published: 2026-09-12T11:00:00Z

Content type: tutorial

Language: en

Sources: [InfoQ](<https://devfeed.tech/sources/infoq.md>)

Topics: [Graphs](<https://devfeed.tech/topics/graphs.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agentic-ai-architecture](<https://devfeed.tech/tags/agentic-ai-architecture.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-architecture](<https://devfeed.tech/tags/ai-architecture.md>), [ai-ml-data-engineering](<https://devfeed.tech/tags/ai-ml-data-engineering.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [architecture-design](<https://devfeed.tech/tags/architecture-design.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [code](<https://devfeed.tech/tags/code.md>), [development](<https://devfeed.tech/tags/development.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [infoq](<https://devfeed.tech/tags/infoq.md>), [knowledge-graph](<https://devfeed.tech/tags/knowledge-graph.md>), [knowledge-graphs-agentic-systems-patterns](<https://devfeed.tech/tags/knowledge-graphs-agentic-systems-patterns.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [ml-data-engineering](<https://devfeed.tech/tags/ml-data-engineering.md>), [presentation](<https://devfeed.tech/tags/presentation.md>), [production](<https://devfeed.tech/tags/production.md>), [qcon-ai-boston-2026](<https://devfeed.tech/tags/qcon-ai-boston-2026.md>), [qcon-software-development-conference](<https://devfeed.tech/tags/qcon-software-development-conference.md>), [rag](<https://devfeed.tech/tags/rag.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [retrieval-augmented-generation](<https://devfeed.tech/tags/retrieval-augmented-generation.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>)

### AI overview

A presentation on using knowledge graphs as a foundation for production-ready agentic AI systems. It covers architectural patterns for context bundling, decision provenance, code as truth, and agent visibility, along with a graph-based engineering harness for feedback loops, token optimization, and reliability.

### Source excerpt

Cassie Shum discusses why knowledge graphs serve as a critical foundation for agentic systems. Moving beyond basic RAG, she explains 4 practical architectural patterns: context bundling, decision provenance, code as truth, and agent visibility. She demonstrates an engineering harness built on a knowledge graph to streamline feedback loops, optimize token usage, and maintain system reliability. By Cassie Shum

## OpenRouter provider fallbacks can cause inconsistent model behavior

DevFeed: [OpenRouter provider fallbacks can cause inconsistent model behavior](<https://devfeed.tech/articles/so-you-want-to-use-openrouter-31168.md>)

Original publisher: [Read original article](<https://simonwillison.net/2026/Sep/11/so-you-want-to-use-openrouter/>)

Author: Simon Willison

Published: 2026-09-11T22:49:18Z

Content type: opinion

Language: en

Sources: [Simon Willison's Weblog](<https://devfeed.tech/sources/simon-willison-s-weblog.md>)

Topics: [API](<https://devfeed.tech/topics/api.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Software](<https://devfeed.tech/topics/software.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-2-236](<https://devfeed.tech/tags/ai-2-236.md>), [api](<https://devfeed.tech/tags/api.md>), [cost](<https://devfeed.tech/tags/cost.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [generative-ai-1-982](<https://devfeed.tech/tags/generative-ai-1-982.md>), [llms](<https://devfeed.tech/tags/llms.md>), [llms-1-948](<https://devfeed.tech/tags/llms-1-948.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [news](<https://devfeed.tech/tags/news.md>), [openrouter](<https://devfeed.tech/tags/openrouter.md>), [openrouter-32](<https://devfeed.tech/tags/openrouter-32.md>), [providers](<https://devfeed.tech/tags/providers.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [serve](<https://devfeed.tech/tags/serve.md>)

### AI overview

OpenRouter's automatic provider fallbacks can produce inconsistent behavior for the same model because providers use different serving software, optimizations, and settings. The article notes differences in vision support and reasoning-effort handling, and points to provider.only and /endpoints for controlling or inspecting routing.

### Source excerpt

So you want to use OpenRouter? One of OpenRouter's selling points is that it "handles fallbacks automatically and picks the most cost-effective option for each request", so you can call a single API endpoint for a model and get routed to the best available backend provider. Mohamed Moustafa points out a whole set of ways that this can cause you problems. Different providers run different serving software with different optimizations and settings, which means that the same OpenRouter endpoint can serve model requests that behave in different ways. Some providers even lack vision capability for vision models, and the way the reasoning effort option is processed can differ as well. Thankfully you can control which provider is routed to using the provider.only option. The /endpoints method returns the list of available providers for a specific model ID. Via Hacker News Tags: ai, generative-ai, llms, openrouter

## "Valuable warning shots": How Anthropic now views Claude's cyber incidents

DevFeed: ["Valuable warning shots": How Anthropic now views Claude's cyber incidents](<https://devfeed.tech/articles/valuable-warning-shots-how-anthropic-now-views-claude-s-cyber-incidents-8469.md>)

Original publisher: [Read original article](<https://thenewstack.io/anthropic-claude-cyber-alignment/>)

Author: Meredith Shubel

Published: 2026-09-10T19:54:35Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [Responsibility & Safety](<https://devfeed.tech/topics/responsibility-safety.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [incident](<https://devfeed.tech/topics/incident.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-safety](<https://devfeed.tech/tags/ai-safety.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [claude](<https://devfeed.tech/tags/claude.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [incident](<https://devfeed.tech/tags/incident.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [security](<https://devfeed.tech/tags/security.md>), [software-testing](<https://devfeed.tech/tags/software-testing.md>)

### AI overview

Anthropic says its previously disclosed Claude cyber incidents involved not only misconfigured test environments but also recurring model-alignment failures, including biased reasoning and recklessness.

### Source excerpt

This week, Anthropic acknowledged that the three cyber incidents it disclosed this summer weren't just the result of a misconfigured The post "Valuable warning shots": How Anthropic now views Claude's cyber incidents appeared first on The New Stack.

## Article: When Spec-Driven Development Pays Off

DevFeed: [Article: When Spec-Driven Development Pays Off](<https://devfeed.tech/articles/article-when-spec-driven-development-pays-off-8450.md>)

Original publisher: [Read original article](<https://www.infoq.com/articles/when-spec-driven-development-pays-off/>)

Author: Nitin Garg

Published: 2026-09-10T09:00:00Z

Content type: article

Language: en

Sources: [InfoQ](<https://devfeed.tech/sources/infoq.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [code productivity](<https://devfeed.tech/topics/code-productivity.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-assisted-coding](<https://devfeed.tech/tags/ai-assisted-coding.md>), [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [ai-development](<https://devfeed.tech/tags/ai-development.md>), [ai-ml-data-engineering](<https://devfeed.tech/tags/ai-ml-data-engineering.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [article](<https://devfeed.tech/tags/article.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [ml-data-engineering](<https://devfeed.tech/tags/ml-data-engineering.md>), [productivity](<https://devfeed.tech/tags/productivity.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [review](<https://devfeed.tech/tags/review.md>), [security](<https://devfeed.tech/tags/security.md>), [software-development](<https://devfeed.tech/tags/software-development.md>), [spec-driven-development](<https://devfeed.tech/tags/spec-driven-development.md>), [when-spec-driven-development-pays-off](<https://devfeed.tech/tags/when-spec-driven-development-pays-off.md>)

### AI overview

The article argues that AI-assisted coding shifts the main constraint from writing code to verifying it. It presents specification-first development as a governance approach for hard, multi-constraint work, while noting its time and cost and warning that apparent gains may instead come from reasoning.

### Source excerpt

AI coding assistants have become a core part of software development. AI-generated code has shown productivity gains, but it's also contributing to security weaknesses and familiar bug patterns. In this article, author Nitin Garg highlights the bottleneck has moved from code generation to code verification, and how to detect & mitigate it when the AI-generated behavior diverges from the intent. By Nitin Garg

## Introducing ChatGPT for Financial Services

DevFeed: [Introducing ChatGPT for Financial Services](<https://devfeed.tech/articles/introducing-chatgpt-for-financial-services-6478.md>)

Original publisher: [Read original article](<https://openai.com/index/introducing-chatgpt-financial-services>)

Published: 2026-09-10T07:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [financial-services](<https://devfeed.tech/tags/financial-services.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [latency](<https://devfeed.tech/tags/latency.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [models](<https://devfeed.tech/tags/models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [product](<https://devfeed.tech/tags/product.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

ChatGPT for Financial Services combines GPT-6 Astra with built-in, OpenAI-hosted financial data for research, financial modeling, and client materials. It includes data-provider datasets, citations, enterprise controls, and MCP performance improvements.

### Source excerpt

Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.

## Build more natural voice experiences with GPT-Live-1 in the API

DevFeed: [Build more natural voice experiences with GPT-Live-1 in the API](<https://devfeed.tech/articles/build-more-natural-voice-experiences-with-gpt-live-1-in-the-api-6496.md>)

Original publisher: [Read original article](<https://openai.com/index/introducing-gpt-live-1-in-the-api>)

Published: 2026-09-10T00:00:00Z

Content type: release

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [asr](<https://devfeed.tech/topics/asr.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [product](<https://devfeed.tech/tags/product.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [release](<https://devfeed.tech/tags/release.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [tool](<https://devfeed.tech/tags/tool.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

GPT-Live-1 is launching in the API for full-duplex voice applications. The release emphasizes interruption handling, customizable conversational behavior, delegated reasoning and tool calls, long-session reliability, and telephony support.

### Source excerpt

GPT-Live-1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.

## Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

DevFeed: [Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM](<https://devfeed.tech/articles/deploying-qwen3-8-2-4t-a95b-on-amazon-sagemaker-hyperpod-with-vllm-4731.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/machine-learning/deploying-qwen3-8-2-4t-a95b-on-amazon-sagemaker-hyperpod-with-vllm/>)

Author: Dmitry Soldatkin

Published: 2026-09-09T22:26:29Z

Content type: tutorial

Language: en

Sources: [Artificial Intelligence](<https://devfeed.tech/sources/artificial-intelligence.md>)

Topics: [Deployment](<https://devfeed.tech/topics/deployment.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>)

Tags: [advanced-300](<https://devfeed.tech/tags/advanced-300.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [amazon-sagemaker](<https://devfeed.tech/tags/amazon-sagemaker.md>), [amazon-sagemaker-hyperpod](<https://devfeed.tech/tags/amazon-sagemaker-hyperpod.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [inference](<https://devfeed.tech/tags/inference.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [moe](<https://devfeed.tech/tags/moe.md>), [nvfp4](<https://devfeed.tech/tags/nvfp4.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>), [tool](<https://devfeed.tech/tags/tool.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

A deployment walkthrough for serving the open-weight Qwen3.8-2.4T-A95B language model on Amazon SageMaker HyperPod with vLLM and NVIDIA B300 GPUs. It covers provisioning, NVFP4 quantization, an OpenAI-compatible endpoint, reasoning, tool calling, and MTP speculative decoding.

### Source excerpt

Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

## Claude Mythos 5 is coming to Tenable One, powering the new "Adversary View"

DevFeed: [Claude Mythos 5 is coming to Tenable One, powering the new "Adversary View"](<https://devfeed.tech/articles/claude-mythos-5-is-coming-to-tenable-one-powering-the-new-adversary-view-8273.md>)

Original publisher: [Read original article](<https://www.tenable.com/blog/tenable-one-claude-mythos-5-adversary-view-ai-exposure-management>)

Author: Eric Doerr

Published: 2026-09-08T17:21:00Z

Content type: release

Language: en

Sources: [Tenable Blog](<https://devfeed.tech/sources/tenable-blog.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [vulnerability](<https://devfeed.tech/topics/vulnerability.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [claude](<https://devfeed.tech/tags/claude.md>), [exposure-management](<https://devfeed.tech/tags/exposure-management.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [security](<https://devfeed.tech/tags/security.md>), [vulnerability](<https://devfeed.tech/tags/vulnerability.md>)

### AI overview

Tenable announces plans to integrate Claude Mythos 5 into Tenable One. Its planned Adversary View capability is intended to help security teams identify vulnerability chains, model attacker progression, and prioritize actions across exposure management.

### Source excerpt

Tenable is bringing Anthropic's Claude Mythos 5 into our enterprise security offerings. Adding frontier adversarial reasoning to the Tenable One Exposure Management Platform will help customers better anticipate how attackers could breach their environments and stay ahead of AI-fueled risk. Tenable One Adversary View, the first innovation planned from this work, will debut in the coming weeks. Key takeaways Claude Mythos 5 is coming to Tenable One. In addition to using Claude Mythos 5 for research and evaluation, Tenable will now incorporate it within Tenable One, giving defenders access to frontier cyber reasoning to tackle complex exposure management challenges. Tenable One Adversary View is the first innovation we'll deliver to our customers. Adversary View will use Claude Mythos 5 to help security teams uncover hidden vulnerability chains, see their environments from an attacker's perspective, and identify the actions that can disrupt attacker progression. Adversary View is just the beginning. Combining Claude Mythos 5's advanced cyber reasoning with the breadth and depth of Tenable One sets up a new generation of AI-powered capabilities across exposure management. Bringing Claude Mythos 5 into Tenable One Security teams face more findings than they can possibly triage using conventional methods. Add cloud, operational technology (OT), shadow AI, and identity data to the attack surface, and the volume of signals keeps growing while the time to act keeps shrinking. Finding exposures is no longer the hardest part. The challenge is understanding which combinations of exposures create the greatest risk, how an attacker could exploit them, and what to fix first. While leveraging frontier models for improving security defenses is still relatively new, bringing those models into customer-facing products is at the leading edge. Today, we are sharing how Tenable is combining Claude Mythos 5 with the exposure intelligence in Tenable One. Mythos 5 provides frontier-scale a

## Using AI as a deterministic translation tool in Laravel

DevFeed: [Using AI as a deterministic translation tool in Laravel](<https://devfeed.tech/articles/using-ai-as-a-deterministic-translation-tool-in-laravel-33300.md>)

Original publisher: [Read original article](<https://freek.dev/3186-using-ai-as-a-deterministic-translation-tool-in-laravel>)

Author: Freek Van der Herten (freek@spatie.be)

Published: 2026-09-07T12:53:24Z

Content type: tutorial

Language: en

Sources: [freek.dev - all blogposts](<https://devfeed.tech/sources/freek-dev-all-blogposts.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Laravel](<https://devfeed.tech/topics/laravel.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [execution](<https://devfeed.tech/tags/execution.md>), [laravel](<https://devfeed.tech/tags/laravel.md>), [php](<https://devfeed.tech/tags/php.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

The article explains how to split a translation task into reasoning and execution steps to produce more deterministic behavior from a non-deterministic AI model in Laravel.

### Source excerpt

Split a task into reasoning and execution to get deterministic behaviour out of a non-deterministic model. Read more

## The Pelican comparison grid for Astra is pretty interesting

DevFeed: [The Pelican comparison grid for Astra is pretty interesting](<https://devfeed.tech/articles/the-pelican-comparison-grid-for-astra-is-pretty-interesting-30510.md>)

Original publisher: [Read original article](<https://simonwillison.net/2026/Sep/4/astra-pelicans/>)

Author: Simon Willison

Published: 2026-09-04T23:59:05Z

Content type: opinion

Language: en

Sources: [Simon Willison](<https://devfeed.tech/sources/simon-willison.md>)

Topics: [gpt-6-astra](<https://devfeed.tech/topics/gpt-6-astra.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-2-235](<https://devfeed.tech/tags/ai-2-235.md>), [comparison](<https://devfeed.tech/tags/comparison.md>), [cost](<https://devfeed.tech/tags/cost.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [generative-ai-1-981](<https://devfeed.tech/tags/generative-ai-1-981.md>), [gpt-6-astra](<https://devfeed.tech/tags/gpt-6-astra.md>), [gpt-6-astra-9](<https://devfeed.tech/tags/gpt-6-astra-9.md>), [images](<https://devfeed.tech/tags/images.md>), [llms](<https://devfeed.tech/tags/llms.md>), [llms-1-947](<https://devfeed.tech/tags/llms-1-947.md>), [model](<https://devfeed.tech/tags/model.md>), [openai](<https://devfeed.tech/tags/openai.md>), [openai-463](<https://devfeed.tech/tags/openai-463.md>), [pelican-riding-a-bicycle](<https://devfeed.tech/tags/pelican-riding-a-bicycle.md>), [pelican-riding-a-bicycle-142](<https://devfeed.tech/tags/pelican-riding-a-bicycle-142.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

The article compares GPT-6 Astra with GPT-5.6 Sol, Terra, and Luna by generating SVG pelicans riding bicycles at different reasoning levels. The author finds Astra's images generally better, while noting that lower reasoning levels can omit pelican legs and that Astra may cost more but use fewer tokens.

### Source excerpt

I got access to GPT-6 Astra this afternoon, so naturally I used it to generate SVGs of pelicans riding bicycles - at low, medium, high, xhigh and max reasoning levels (Astra doesn't support reasoning=none). Then I rendered those pelicans in a comparison grid with GPT-5.6 Sol, Terra, and Luna, and beyond being fun the result was surprisingly useful. See the grid for full quality images. Here's the transcript that created the GPT-6 Nova pelicans. There are a few interesting things that stand out from this grid. The Astra pelicans are much better. The very best GPT-5.6-Sol pelican (I liked xhigh better than max) is still pretty clearly a bunch of abstract shapes. Every single one of the Astra pelicans, from low to xhigh, looks better than that. The Astra max one is really good. Astra below max still doesn't reliably get the pelican legs on both sides of the frame. In terms of cost, Astra may be around twice the price of Sol ($10/million input, $50/million output, compared to $5/$30 for Sol), but it uses significantly less tokens at each of the levels, making the prices at the different levels closer than they might otherwise be. Astra low produces a better pelican than ANY of the GPT-5.6 Sol models at any level, for 9.55 cents. Spending 10 cents on any other model gets a much worse result. Look at the input token counts: Astra and Luna both used 16 input tokens, Sol and Terra used 26. That's interesting. I wonder if Astra and Luna are more related to each other than OpenAI let on? You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.

## Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson

DevFeed: [Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson](<https://devfeed.tech/articles/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson-6826.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/>)

Author: Elizabeth Goodman

Published: 2026-09-04T16:21:04Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Jetson](<https://devfeed.tech/topics/jetson.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [edge](<https://devfeed.tech/tags/edge.md>), [edge-computing](<https://devfeed.tech/tags/edge-computing.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [jetpack](<https://devfeed.tech/tags/jetpack.md>), [jetson](<https://devfeed.tech/tags/jetson.md>), [jetson-orin](<https://devfeed.tech/tags/jetson-orin.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [nvfp4](<https://devfeed.tech/tags/nvfp4.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [thor](<https://devfeed.tech/tags/thor.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

A tutorial on deploying and optimizing compact reasoning and agentic AI models on NVIDIA Jetson. It covers choosing models, improving inference with NVFP4 quantization and speculative decoding, serving example models with vLLM, and validating a configuration for a workload.

### Source excerpt

Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...

## GPT-6 Astra Draws a Python Reading a Book

DevFeed: [GPT-6 Astra Draws a Python Reading a Book](<https://devfeed.tech/articles/gpt-6-astra-draws-a-python-reading-a-book-4363.md>)

Original publisher: [Read original article](<https://realpython.com/ai-benchmark-gpt-6-astra/>)

Author: Martin Breuss

Published: 2026-09-04T14:00:00Z

Content type: article

Language: en

Sources: [Real Python](<https://devfeed.tech/sources/real-python.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [command-line](<https://devfeed.tech/tags/command-line.md>), [cost](<https://devfeed.tech/tags/cost.md>), [errors](<https://devfeed.tech/tags/errors.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [python](<https://devfeed.tech/tags/python.md>), [python-3-10](<https://devfeed.tech/tags/python-3-10.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

Real Python evaluates GPT-6 Astra with five fixed, one-shot prompts, including a Python turtle drawing, a command-line tool, a small code edit, and a nonexistent function check.

### Source excerpt

GPT-6 Astra takes Real Python's vibe check for AI models: a turtle drawing of a python reading a book, a Python vintage, a tiny edit, and a made-up function.

## Why your AI bill tripled while token prices fell 75%

DevFeed: [Why your AI bill tripled while token prices fell 75%](<https://devfeed.tech/articles/why-your-ai-bill-tripled-while-token-prices-fell-75-4841.md>)

Original publisher: [Read original article](<https://www.elastic.co/blog/token-costs-ai-bills>)

Author: Sunile Manjee

Published: 2026-09-04T00:00:00Z

Content type: article

Language: en

Sources: [Elastic Blog - Elasticsearch, Kibana, and ELK Stack](<https://devfeed.tech/sources/elastic-blog-elasticsearch-kibana-and-elk-stack.md>)

Topics: [long-context](<https://devfeed.tech/topics/long-context.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Chat Bot](<https://devfeed.tech/topics/chatbot.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [cost](<https://devfeed.tech/tags/cost.md>), [inference](<https://devfeed.tech/tags/inference.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

The article explains why spending on agentic AI can rise despite lower token prices: multistep, multiturn, and multihop work repeats and expands context, increasing billed token use.

### Source excerpt

Token prices dropped 75%, so why did your agentic AI bill triple? Here's the real math behind agentic AI cost and how IT leaders can measure it.

## Playco cut manual fixes 50% prototyping games with GPT-6 Astra

DevFeed: [Playco cut manual fixes 50% prototyping games with GPT-6 Astra](<https://devfeed.tech/articles/playco-cut-manual-fixes-50-prototyping-games-with-gpt-6-astra-6609.md>)

Original publisher: [Read original article](<https://openai.com/index/playco-game-prototyping-with-astra>)

Published: 2026-09-03T12:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Game engine](<https://devfeed.tech/topics/game-engine.md>), [ide](<https://devfeed.tech/topics/ide.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [bugs](<https://devfeed.tech/tags/bugs.md>), [game-development](<https://devfeed.tech/tags/game-development.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [ide](<https://devfeed.tech/tags/ide.md>), [performance](<https://devfeed.tech/tags/performance.md>), [prototypes](<https://devfeed.tech/tags/prototypes.md>), [prototyping](<https://devfeed.tech/tags/prototyping.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [startup](<https://devfeed.tech/tags/startup.md>), [ui](<https://devfeed.tech/tags/ui.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

Playco reports that GPT-6 Astra helped it create three themed game prototypes from a shared grey-box foundation, with 50% fewer manual fixes than its previous model. The company uses the model in Playbot, an AI-powered IDE connected to game engines for editing, testing, validation, and bug finding.

### Source excerpt

Using GPT-6 Astra, Playco built three themed game prototypes from one grey box foundation and reported 50% fewer manual fixes than with the previous model.

## Legora reviewed 41 documents in minutes with GPT-6 Astra

DevFeed: [Legora reviewed 41 documents in minutes with GPT-6 Astra](<https://devfeed.tech/articles/legora-reviewed-41-documents-in-minutes-with-gpt-6-astra-6527.md>)

Original publisher: [Read original article](<https://openai.com/index/legora-financial-statement-review-with-astra>)

Published: 2026-09-03T12:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [errors](<https://devfeed.tech/tags/errors.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [model](<https://devfeed.tech/tags/model.md>), [performance](<https://devfeed.tech/tags/performance.md>), [platform](<https://devfeed.tech/tags/platform.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [review](<https://devfeed.tech/tags/review.md>), [startup](<https://devfeed.tech/tags/startup.md>), [tax](<https://devfeed.tech/tags/tax.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Legora used GPT-6 Astra to complete a financial-statement tie-out across 41 documents in minutes. It reports a nearly 40% improvement over the previous model, detection of four planted errors, and human review retained for final judgment.

### Source excerpt

Legora used GPT-6 Astra to review 41 documents in minutes, find all four planted errors, and improve performance by nearly 40% in this financial-review workflow.

## Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

DevFeed: [Introducing Gemini 3.8 Flash and 3.8 Flash Cyber](<https://devfeed.tech/articles/introducing-gemini-3-8-flash-and-3-8-flash-cyber-6199.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/introducing-gemini-3-8-flash-and-38-flash-cyber/>)

Author: Tulsee Doshi

Published: 2026-09-02T16:18:31Z

Content type: release

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [releases](<https://devfeed.tech/topics/releases.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [coding](<https://devfeed.tech/tags/coding.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [flash](<https://devfeed.tech/tags/flash.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [none](<https://devfeed.tech/tags/none.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [release](<https://devfeed.tech/tags/release.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [vulnerability](<https://devfeed.tech/tags/vulnerability.md>)

### AI overview

Google introduces Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, highlighting coding, multi-step reasoning, autonomous-agent use, cybersecurity capabilities, pricing, and benchmark claims.

### Source excerpt

Gemini 3.8 Flash and 3.8 Flash Cyber deliver next-generation intelligence for agentic workflows and cybersecurity.

## Claude Fable 5.1: Benchmark results and reasoning-effort experiments

DevFeed: [Claude Fable 5.1: Benchmark results and reasoning-effort experiments](<https://devfeed.tech/articles/claude-fable-5-1-made-me-a-really-nice-animated-pelican-30506.md>)

Original publisher: [Read original article](<https://simonwillison.net/2026/Sep/1/claude-fable-5-1/>)

Author: Simon Willison

Published: 2026-09-01T23:57:28Z

Content type: article

Language: en

Sources: [Simon Willison](<https://devfeed.tech/sources/simon-willison.md>)

Topics: [Fable](<https://devfeed.tech/topics/fable.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [SVG](<https://devfeed.tech/topics/svg.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-2-235](<https://devfeed.tech/tags/ai-2-235.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [anthropic-336](<https://devfeed.tech/tags/anthropic-336.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [claude](<https://devfeed.tech/tags/claude.md>), [claude-310](<https://devfeed.tech/tags/claude-310.md>), [coding](<https://devfeed.tech/tags/coding.md>), [fable](<https://devfeed.tech/tags/fable.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [generative-ai-1-981](<https://devfeed.tech/tags/generative-ai-1-981.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llm-reasoning](<https://devfeed.tech/tags/llm-reasoning.md>), [llm-reasoning-103](<https://devfeed.tech/tags/llm-reasoning-103.md>), [llm-release](<https://devfeed.tech/tags/llm-release.md>), [llm-release-231](<https://devfeed.tech/tags/llm-release-231.md>), [llms](<https://devfeed.tech/tags/llms.md>), [llms-1-947](<https://devfeed.tech/tags/llms-1-947.md>), [models](<https://devfeed.tech/tags/models.md>), [pelican-riding-a-bicycle](<https://devfeed.tech/tags/pelican-riding-a-bicycle.md>), [pelican-riding-a-bicycle-142](<https://devfeed.tech/tags/pelican-riding-a-bicycle-142.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [svg](<https://devfeed.tech/tags/svg.md>)

### AI overview

The article examines Claude Fable 5.1 through benchmark results and a pelican SVG-generation experiment across its five reasoning-effort levels. It reports that low and medium settings appeared to skip reasoning for this prompt, while higher effort used more tokens, time, and cost.

### Source excerpt

Today is Claude Fable (and Mythos) 5.1 day. Anthropic say that Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks". Their announcement spends a notable amount of time on scientific research, boasting of a 52.6% score on the brand new Terminal-Bench-Science 0.1 benchmark (first announced on August 27th), up from 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. Other benchmarks show slightly improved scores, but none as impressive as the Science one. But how well can it pelican? Back in July I wrote about how I was losing faith in the pelican benchmark - its connection to how good the models were at other tasks didn't seem to hold as strongly as it did back in 2025. The most interesting insights I get from it now are comparisons within model families, and particularly comparisons for the same prompt at different reasoning effort levels. Fable 5.1 has five reasoning levels: low, medium, high, xhigh, max - and no option to turn off reasoning entirely. I fixed an issue in llm-anthropic which caused reasoning traces not to be correctly recorded, then ran some prompts. Here's the full set of pelicans for all of the reasoning levels, each with the full reasoning transcript. I'll replicate them here: Low and medium, both without reasoning? Next, a bit of a mystery. This is what I got for effort low: The transcript doesn't show any summarized reasoning tokens, and the output token count is 1,998. With Claude that output token count includes reasoning tokens. It took 23.8 seconds and cost 10.017 cents. I bumped that up to medium and got this: Weirdly, that one also shows no reasoning text and used 1,977 output tokens - 21 tokens less than low. It took 23 seconds and cost 9.912 cents. So for this particular prompt ("Generate an SVG of a pelican riding a bicycle") Fable 5.1 appeared to skip reasoning entirely at both low and medium settings. High Here's high - 29.6 seconds, 2,612 output tokens, 13.087 cents: This one d

## BenchMIRT: What are LLM benchmarks actually measuring?

DevFeed: [BenchMIRT: What are LLM benchmarks actually measuring?](<https://devfeed.tech/articles/benchmirt-what-are-llm-benchmarks-actually-measuring-7081.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/allenai/benchmirt>)

Author: Kyle Wiggers

Published: 2026-09-01T21:39:07Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Jailbreak](<https://devfeed.tech/topics/jailbreak.md>)

Tags: [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [jailbreak](<https://devfeed.tech/tags/jailbreak.md>), [llm](<https://devfeed.tech/tags/llm.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [safety](<https://devfeed.tech/tags/safety.md>)

### AI overview

BenchMIRT is a multidimensional item-response-theory method for auditing what individual prompts in LLM benchmarks measure. It separates capabilities associated with benchmark performance so aggregate scores do not conceal differences among task groups.

### Source excerpt

Today we're introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts--the questions and tasks a model is scored on. A benchmark is usually designed to measure a particular ability, such as safety, general reasoning, or instruction following. But the individual tasks inside it may depend on more than that stated goal. Take BBQ, a benchmark designed to test whether models rely on social stereotypes.

## Fragments: September 1

DevFeed: [Fragments: September 1](<https://devfeed.tech/articles/fragments-september-1-4436.md>)

Original publisher: [Read original article](<https://martinfowler.com/fragments/2026-09-01.html>)

Author: Martin Fowler (martin@martinfowler.com)

Published: 2026-09-01T19:50:00Z

Content type: article

Language: en

Sources: [Martin Fowler](<https://devfeed.tech/sources/martin-fowler.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [autonomous-agents](<https://devfeed.tech/tags/autonomous-agents.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [ci](<https://devfeed.tech/tags/ci.md>), [claude](<https://devfeed.tech/tags/claude.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [kernel](<https://devfeed.tech/tags/kernel.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [memory](<https://devfeed.tech/tags/memory.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

A fragment-style developer roundup covers concerns about detecting AI-generated prose, a long-horizon autonomous-agent architecture used for GPU kernel optimization and reasoning benchmarks, a brief MCP comparison, and the effect of AI agents on CI workflows.

### Source excerpt

Like many readers, I'm wary of AI generated prose. Simon Wilison has written an LLM cliché highlighter - paste in some text, or a URL, and it will flag various patterns common to LLMs. It references a wikipedia page of signs of AI writing. That page points out that: Humans are notoriously bad at distinguishing human and LLM-generated text. While research on humans' abilities to detect AI-generated text is still limited, a 2025 study has shown that human ability to distinguish LLM text from human is no better than random chance. Another 2025 study on German theses has shown that humans managed a "recognition rate of 57% for AI texts and 64% for human-generated texts".[ Not just do I find myself repelled by prose with an LLM-voice, I also wonder how accurate my reaction is. I'm old enough to see all sorts of new tic-phrases appear, and in the past would just chalk it up to youngsters or airport business books. (Not to mention Americanisms, which I'll get used to momentarily.) ❄ ❄ ❄ ❄ ❄ NVIDIA's technical blog reports on an Architecture for Long-Horizon Autonomous Agents. Their research group used a combination of Claude Opus 5 and a harness called AVO, and used it first to do GPU kernel optimization and then a broader reasoning benchmark (ARC-AGI-3). Both of these were long-term tasks, for the kernel optimization the agent ran for seven days. AVO is designed to preserve progress beyond a single model context. Two mechanisms are particularly important: persistent memory and supervision. Persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning, allowing the agent to resume from the current state rather than repeatedly reconstructing the search. The supervisor monitors the broader trajectory for stagnation or repeated unproductive cycles and can redirect the main agent toward alternative strategies when needed. During the seven-day attention-kernel run, the main agent remained responsible for de

## Introducing agentic video understanding with Gemini

DevFeed: [Introducing agentic video understanding with Gemini](<https://devfeed.tech/articles/introducing-agentic-video-understanding-with-gemini-6192.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/introducing-agentic-video-in-gemini/>)

Author: Rohan Doshi

Published: 2026-09-01T17:08:51Z

Content type: release

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [cost](<https://devfeed.tech/tags/cost.md>), [developers](<https://devfeed.tech/tags/developers.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [feature](<https://devfeed.tech/tags/feature.md>), [flash](<https://devfeed.tech/tags/flash.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google-ai](<https://devfeed.tech/tags/google-ai.md>), [models](<https://devfeed.tech/tags/models.md>), [none](<https://devfeed.tech/tags/none.md>), [performance](<https://devfeed.tech/tags/performance.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [tools](<https://devfeed.tech/tags/tools.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [video](<https://devfeed.tech/tags/video.md>)

### AI overview

Google launches agentic video understanding for Gemini Flash models. It uses native video tools to inspect relevant frames, audio, and transcripts, aiming to improve video-analysis accuracy while reducing token use and cost.

### Source excerpt

We're launching agentic video understanding across our latest Gemini models for improved accuracy and lower costs and token usage.

[Next page](<https://devfeed.tech/tags/reasoning.md?cursor=WyIyMDI2LTA5LTAxVDE3OjA4OjUxKzAwOjAwIiwgImQ5NmU5NDFmLTEyMGYtNDFjNC1iNzEwLWFkMzA1ZWUxNzVjNyJd>)