# Chain-of-thought

A prompting method that has language models produce intermediate reasoning steps before a final answer.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Using AI as a deterministic translation tool in Laravel

DevFeed: [Using AI as a deterministic translation tool in Laravel](<https://devfeed.tech/articles/using-ai-as-a-deterministic-translation-tool-in-laravel-33300.md>)

Original publisher: [Read original article](<https://freek.dev/3186-using-ai-as-a-deterministic-translation-tool-in-laravel>)

Author: Freek Van der Herten (freek@spatie.be)

Published: 2026-09-07T12:53:24Z

Content type: tutorial

Language: en

Sources: [freek.dev - all blogposts](<https://devfeed.tech/sources/freek-dev-all-blogposts.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Laravel](<https://devfeed.tech/topics/laravel.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [execution](<https://devfeed.tech/tags/execution.md>), [laravel](<https://devfeed.tech/tags/laravel.md>), [php](<https://devfeed.tech/tags/php.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

The article explains how to split a translation task into reasoning and execution steps to produce more deterministic behavior from a non-deterministic AI model in Laravel.

### Source excerpt

Split a task into reasoning and execution to get deterministic behaviour out of a non-deterministic model. Read more

## Researchers Found That Encrypted AI Reasoning Blocks Can Be Replayed to Reveal Hidden Reasoning

DevFeed: [Researchers Found That Encrypted AI Reasoning Blocks Can Be Replayed to Reveal Hidden Reasoning](<https://devfeed.tech/articles/how-to-steal-an-ai-model-s-private-thoughts-17995.md>)

Original publisher: [Read original article](<https://blog.bytebytego.com/p/how-to-steal-an-ai-models-private>)

Author: ByteByteGo

Published: 2026-08-25T15:31:09Z

Content type: article

Language: en

Sources: [ByteByteGo](<https://devfeed.tech/sources/bytebytego.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Security](<https://devfeed.tech/topics/security.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [Google](<https://devfeed.tech/topics/google.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [google](<https://devfeed.tech/tags/google.md>), [openai](<https://devfeed.tech/tags/openai.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

The article describes research testing whether encrypted reasoning blocks returned by Anthropic, OpenAI, and Google protect hidden model reasoning. It reports that the blocks can be replayed into a cheaper model from the same family, revealing the reasoning in plaintext, and outlines the extraction method, attack vectors, and proposed fixes.

### Source excerpt

In August 2026, a team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems wanted to test whether the encrypted reasoning blocks that Anthropic, OpenAI, and Google hand back to clients actually keep that reasoning private.

## How Reasoning Traces Work in Language Models

DevFeed: [How Reasoning Traces Work in Language Models](<https://devfeed.tech/articles/what-is-reasoning-30732.md>)

Original publisher: [Read original article](<https://lucumr.pocoo.org/2026/8/19/what-is-reasoning/>)

Author: Armin Ronacher

Published: 2026-08-19T00:00:00Z

Content type: article

Language: en

Sources: [Armin Ronacher](<https://devfeed.tech/sources/armin-ronacher.md>)

Topics: [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [gpt-oss](<https://devfeed.tech/topics/gpt-oss.md>), [Parser](<https://devfeed.tech/topics/parser.md>), [API](<https://devfeed.tech/topics/api.md>), [Cache](<https://devfeed.tech/topics/cache.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [token](<https://devfeed.tech/tags/token.md>)

### AI overview

The article explains reasoning traces as text emitted by a model into a scratchpad before its final answer. It discusses how GPT-OSS uses channel markers and a parser to route analysis into a separate stream, and argues that reasoning effort is shaped by system prompts and training rather than being solely a sampling-process property.

### Source excerpt

A few weeks ago a paper was shared that showed how to extract reasoning traces from closed-weight models. Together with online discussions about tricking models into leaking them, it made me investigate it more out of curiosity. Twitter seems full of half-truths and confusion about how this works, so perhaps this helps some to understand what is happening. Hiding Traces Reasoning traces are usually hidden from us. We have lamented this, but mostly have to accept it. Open-weight models thankfully reveal them, and from their behavior you can see that their traces can be long and confusing. This is probably a good reason to separate them from what is normally shown to users. At minimum, UIs need to detect them. The industry has done a good job at making reasoning traces sound special and exotic, but they really are just text: the model is trained to emit its thinking into a scratchpad as part of its response, before its final answer. GPT-OSS's Harmony response format makes this easy to see: <|channel|>analysis<|message|> I need to work this out ... <|end|><|start|>assistant<|channel|>final<|message|> The answer is ... <|return|> The markers are special tokens, but the reasoning between them uses "the same text" as the final answer (just that GPT chain-of-thought text sounds really funny). When the model samples the analysis channel token, a parser routes the following text into a separate stream exposed through the Responses API. For closed models, presumably a simple model redacts and summarizes it. Reasoning Effort How much budget goes to reasoning? Earlier APIs exposed reasoning token budgets, making it seem like a property of the sampling process. In reality, reasoning effort is baked into the system prompt. GPT-OSS puts this into the system prompt: Reasoning: low That's it. Training produces the resulting behavior, such as emitting the token sequence that switches to the analysis channel. This also explains why changing the effort invalidates the KV cache. I think

## Token-budget-aware LLM reasoning: cut costs in 2026

DevFeed: [Token-budget-aware LLM reasoning: cut costs in 2026](<https://devfeed.tech/articles/token-budget-aware-llm-reasoning-cut-costs-in-2026-4855.md>)

Original publisher: [Read original article](<https://redis.io/blog/token-budget-aware-llm-reasoning/>)

Author: Jeff Mills

Published: 2026-07-28T00:00:00Z

Content type: tutorial

Language: en

Sources: [Redis Blog](<https://devfeed.tech/sources/redis-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [caching](<https://devfeed.tech/tags/caching.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [claude](<https://devfeed.tech/tags/claude.md>), [cost](<https://devfeed.tech/tags/cost.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [routing](<https://devfeed.tech/tags/routing.md>), [tech-de](<https://devfeed.tech/tags/tech-de.md>), [techniques](<https://devfeed.tech/tags/techniques.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

This guide explains token-budget-aware LLM reasoning, a technique for matching a model's reasoning-token budget to problem complexity. It covers the cost of reasoning and output tokens, prompt-level methods such as chain-of-thought and Chain of Draft, and architectural approaches including caching, routing, and memory.

### Source excerpt

Reasoning models think before they answer, and those reasoning tokens are usually part of what you pay for. They're billed as output tokens, the expensive kind, and a single request can generate a few hundred of them depending on the problem. If your ...

## Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning

DevFeed: [Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning](<https://devfeed.tech/articles/lessons-from-the-leaderboard-what-5-000-kagglers-taught-us-about-improving-ai-reasoning-6875.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/lessons-from-the-leaderboard-what-5000-kagglers-taught-us-about-improving-ai-reasoning/>)

Author: Elizabeth Goodman

Published: 2026-07-14T18:20:32Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Google](<https://devfeed.tech/topics/google.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [cost](<https://devfeed.tech/tags/cost.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [featured](<https://devfeed.tech/tags/featured.md>), [google](<https://devfeed.tech/tags/google.md>), [kaggle](<https://devfeed.tech/tags/kaggle.md>), [lora](<https://devfeed.tech/tags/lora.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pre-trained-foundation-models](<https://devfeed.tech/tags/pre-trained-foundation-models.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [techniques](<https://devfeed.tech/tags/techniques.md>)

### AI overview

The article distills lessons from NVIDIA's Nemotron Model Reasoning Challenge, where more than 5,000 Kaggle participants tested ways to improve AI reasoning under shared model, infrastructure, and evaluation constraints. It highlights synthetic chain-of-thought data, trace quality, targeted solvers, validation beyond public leaderboards, and careful training and context-budget management.

### Source excerpt

The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when...

## How Layered Codebases and LLM-Assisted Reasoning Accumulate Complexity

DevFeed: [How Layered Codebases and LLM-Assisted Reasoning Accumulate Complexity](<https://devfeed.tech/articles/make-no-assumptions-35687.md>)

Original publisher: [Read original article](<https://lethain.com/make-no-assumptions/>)

Published: 2026-07-11T13:00:00Z

Content type: opinion

Language: en

Sources: [Will Larson - Irrational Exuberance](<https://devfeed.tech/sources/will-larson-irrational-exuberance.md>)

Topics: [coding](<https://devfeed.tech/topics/coding.md>), [Software](<https://devfeed.tech/topics/software.md>), [context](<https://devfeed.tech/topics/context.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [coding](<https://devfeed.tech/tags/coding.md>), [complexity](<https://devfeed.tech/tags/complexity.md>), [context](<https://devfeed.tech/tags/context.md>), [llms](<https://devfeed.tech/tags/llms.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [software](<https://devfeed.tech/tags/software.md>)

### AI overview

The article uses soil horizons as an analogy for software and organizational reasoning. It argues that codebases can accumulate distinct layers as architects and team leads change, especially in complex domains, and that relying on LLMs for conclusions rather than exploration or drafting can create similarly flawed reasoning layers.

### Source excerpt

I've recently been thinking a lot about the concept of "soil horizons", which is the idea that there are many distinct layers of soil, from topsoil all the way down to bedrock, which all combine into a soil horizon. Translating this idea into software, the ideal codebase would have a single uniform "code layer", but a surprisingly large percentage of production software has numerous, distinct code layers as the leading architect shifted over time. I've found this particularly true for software in problem-spaces with high essential complexity and low scale complexity, where the purifying challenges of scaling never create enough pressure to compact disjoint layers into a unified layer. Codebases with the most code layers tend to be created by small teams working on complex domains over a long period of time. In many companies this might be an identity, permissions or payments team: stuff that's permanently valuable, but usually not the central concern at any given time. On such teams, there is often only one architect who understands the nuances of the domain well enough to make tradeoffs. When that architect leaves, they are replaced by someone who aspires to operate in the same code layer, but simply cannot because they lack enough context to do so. As a result, that new replacement creates a new code layer, despite not intending to. If the team runs through a handful of folks as the new team leads struggle, it's easy to end up with a complex code horizon very quickly. The problem of messy code horizons is not a new one, and the general approach to addressing them is the same one I wrote about seven years ago in Reclaim unreasonable software, but with the proliferation of coding and non-coding harnesses, lately I'm running into the problem of messy code horizons more frequently. Even more concerning, I'm seeing this problem expand from impacting code horizons into impacting how organizations make decisions outside of software, e.g. the company's general reasoning h

## Thinking to recall: How reasoning unlocks parametric knowledge in LLMs

DevFeed: [Thinking to recall: How reasoning unlocks parametric knowledge in LLMs](<https://devfeed.tech/articles/thinking-to-recall-how-reasoning-unlocks-parametric-knowledge-in-llms-6896.md>)

Original publisher: [Read original article](<https://research.google/blog/thinking-to-recall-how-reasoning-unlocks-parametric-knowledge-in-llms/>)

Published: 2026-06-24T16:51:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Google](<https://devfeed.tech/topics/google.md>), [Large language models (LLMs)](<https://devfeed.tech/topics/large-language-models-llms.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>)

Tags: [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

Google Research examines why generating chain-of-thought reasoning can help large language models recall simple facts that are already encoded in their parametric memory. Controlled experiments identify two mechanisms: latent computation through generated reasoning tokens and factual priming through related facts.

### Source excerpt

Generative AI

## Claude Code Best Practices, Planning in 8 Tokens, and Why Reasoning Models Can't Control Their Own Thoughts - 📚 The Tokenizer Edition #20

DevFeed: [Claude Code Best Practices, Planning in 8 Tokens, and Why Reasoning Models Can't Control Their Own Thoughts - 📚 The Tokenizer Edition #20](<https://devfeed.tech/articles/claude-code-best-practices-planning-in-8-tokens-and-why-reasoning-models-can-t-control-their-own-thoughts-the-tokenizer-edition-20-18332.md>)

Original publisher: [Read original article](<https://newsletter.artofsaience.com/p/claude-code-best-practices-planning>)

Author: Sairam Sundaresan

Published: 2026-03-18T13:03:02Z

Content type: article

Language: en

Sources: [Gradient Ascent](<https://devfeed.tech/sources/gradient-ascent.md>)

Topics: [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [newsletter](<https://devfeed.tech/tags/newsletter.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

A newsletter edition curating AI and machine-learning resources, including Claude Code design workflows, reasoning-model research, retrieval systems, training efficiency, vision-language models, and developer tools.

### Source excerpt

This week's most valuable resources

## Improve AI output with continuous improvement

DevFeed: [Improve AI output with continuous improvement](<https://devfeed.tech/articles/improve-ai-output-with-continuous-improvement-16648.md>)

Original publisher: [Read original article](<https://firebase.blog/posts/2026/01/continuous-improvement>)

Author: Alexander Nohe; Dan Gruhl

Published: 2026-01-13T00:00:00Z

Content type: tutorial

Language: en

Sources: [Firebase Blog](<https://devfeed.tech/sources/firebase-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [adk](<https://devfeed.tech/tags/adk.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automated](<https://devfeed.tech/tags/automated.md>), [firebase](<https://devfeed.tech/tags/firebase.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [prompting](<https://devfeed.tech/tags/prompting.md>), [python](<https://devfeed.tech/tags/python.md>), [rewrite](<https://devfeed.tech/tags/rewrite.md>), [verify](<https://devfeed.tech/tags/verify.md>)

### AI overview

A tutorial explains continuous improvement, a prompting strategy in which one AI generates an output and another AI critiques it before the first AI revises the result. It demonstrates the approach with ADK and LoopAgent for designing gym workouts.

### Source excerpt

Learn how to improve your AI outputs to a more reliable output through continuous improvement.

## AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems

DevFeed: [AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems](<https://devfeed.tech/articles/aprielguard-a-guardrail-for-safety-and-adversarial-robustness-in-modern-llm-systems-7045.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ServiceNow-AI/aprielguard>)

Author: Jaykumar Kasundra

Published: 2025-12-23T14:07:35Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [prompt injection](<https://devfeed.tech/topics/prompt-injection.md>), [Adversarial attacks](<https://devfeed.tech/topics/adversarial-attacks.md>), [Security](<https://devfeed.tech/topics/security.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [attacks](<https://devfeed.tech/tags/attacks.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model](<https://devfeed.tech/tags/model.md>), [prompt-injection](<https://devfeed.tech/tags/prompt-injection.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [safety-security](<https://devfeed.tech/tags/safety-security.md>), [security](<https://devfeed.tech/tags/security.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

AprielGuard is an 8B-parameter safety and security safeguard model for modern LLM systems. It detects 16 categories of safety risks and a broad range of adversarial attacks, including prompt injection, jailbreaks, chain-of-thought corruption, context hijacking, memory poisoning, and multi-agent exploit sequences. It supports standalone prompts, multi-turn conversations, and agentic workflows containing tool calls, reasoning traces, memory, and system context. The model offers reasoning and non-reasoning modes for explainable or low-latency classification.

### Source excerpt

In this work, we introduce AprielGuard, an 8B parameter safety-security safeguard model designed to detect: - 16 categories of safety risks, spanning toxicity, hate, sexual content, misinformation, self-harm, illegal activities, and more. - Wide range of adversarial attacks, including prompt injection, jailbreaks, chain-of-thought corruption, context hijacking, memory poisoning, and multi-agent exploit sequences.

## Reflections on AI at the end of 2025

DevFeed: [Reflections on AI at the end of 2025](<https://devfeed.tech/articles/reflections-on-ai-at-the-end-of-2025-20648.md>)

Original publisher: [Read original article](<http://antirez.com/news/157>)

Published: 2025-12-20T08:58:29Z

Content type: opinion

Language: en

Sources: [Antirez](<https://devfeed.tech/sources/antirez.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [rlvr](<https://devfeed.tech/topics/rlvr.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Programming](<https://devfeed.tech/topics/programming.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [code](<https://devfeed.tech/tags/code.md>), [coding](<https://devfeed.tech/tags/coding.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [reflections](<https://devfeed.tech/tags/reflections.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>)

### AI overview

An end-of-2025 reflection on developments in AI, including changing views of LLM representations, chain-of-thought, reinforcement learning with verifiable rewards, and growing adoption of AI-assisted programming. The author presents these as observations and expectations, including the possibility that improved reinforcement learning could become a major direction in AI.

### Source excerpt

* For years, despite functional evidence and scientific hints accumulating, certain AI researchers continued to claim LLMs were stochastic parrots: probabilistic machines that would: 1. NOT have any representation about the meaning of the prompt. 2. NOT have any representation about what they were going to say. In 2025 finally almost everybody stopped saying so. * Chain of thought is now a fundamental way to improve LLM output. But, what is CoT? Why it improves output? I believe it is two things: 1. Sampling in the model representations (that is, a form of internal search). After information and concepts relevant to the prompt topic is in the context window, the model can better reply. 2. But if you mix this to reinforcement learning, the model also learns to put one token after the other (each token will change the model state) in order to converge to some useful reply. * The idea that scaling is limited to the number of tokens we have, is no longer true, because of reinforcement learning with verifiable rewards. We are still not at AlphaGo move 37 moment, but is this really impossible in the future? There are certain tasks, like improving a given program for speed, for instance, where in theory the model can continue to make progress with a very clear reward signal for a very long time. I believe improvements to RL applied to LLMs will be the next big thing in AI. * Programmers resistance to AI assisted programming has lowered considerably. Even if LLMs make mistakes, the ability of LLMs to deliver useful code and hints improved to the point most skeptics started to use LLMs anyway: now the return on the investment is acceptable for many more folks. The programming world is still split among who uses LLMs as colleagues (for instance, all my interaction is via the web interface of Gemini, Claude, ...), and who uses LLMs as independent coding agents. * A few well known AI scientists believe that what happened with Transformers can happen again, and better, following d

## Evaluating chain-of-thought monitorability

DevFeed: [Evaluating chain-of-thought monitorability](<https://devfeed.tech/articles/evaluating-chain-of-thought-monitorability-6396.md>)

Original publisher: [Read original article](<https://openai.com/index/evaluating-chain-of-thought-monitorability>)

Published: 2025-12-18T12:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

OpenAI introduces a framework and 13-evaluation suite spanning 24 environments to measure chain-of-thought monitorability. The study finds that monitoring internal reasoning is generally more effective than monitoring actions and final outputs alone, and that monitorability often improves when models reason for longer.

### Source excerpt

OpenAI introduces a new framework and evaluation suite for chain-of-thought monitorability, covering 13 evaluations across 24 environments. Our findings show that monitoring a model's internal reasoning is far more effective than monitoring outputs alone, offering a promising path toward scalable control as AI systems grow more capable.

## Deepening our partnership with the UK AI Security Institute

DevFeed: [Deepening our partnership with the UK AI Security Institute](<https://devfeed.tech/articles/deepening-our-partnership-with-the-uk-ai-security-institute-6146.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/deepening-our-partnership-with-the-uk-ai-security-institute/>)

Author: William Isaac; Owen Larter

Published: 2025-12-11T00:06:40Z

Content type: article

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [ai safety](<https://devfeed.tech/topics/ai-safety.md>), [Google](<https://devfeed.tech/topics/google.md>), [Responsibility & Safety](<https://devfeed.tech/topics/responsibility-safety.md>), [Security](<https://devfeed.tech/topics/security.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-safety](<https://devfeed.tech/tags/ai-safety.md>), [ai-security](<https://devfeed.tech/tags/ai-security.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [google](<https://devfeed.tech/tags/google.md>), [publications](<https://devfeed.tech/tags/publications.md>), [reports](<https://devfeed.tech/tags/reports.md>), [research](<https://devfeed.tech/tags/research.md>), [responsibility-safety](<https://devfeed.tech/tags/responsibility-safety.md>), [safety](<https://devfeed.tech/tags/safety.md>), [technical](<https://devfeed.tech/tags/technical.md>), [techniques](<https://devfeed.tech/tags/techniques.md>), [uk](<https://devfeed.tech/tags/uk.md>)

### AI overview

Google DeepMind announces an expanded partnership with the UK AI Security Institute focused on foundational AI security and safety research. The collaboration includes model testing, shared research resources, joint publications, and work on monitoring AI reasoning processes.

### Source excerpt

Google DeepMind and UK AI Security Institute (AISI) strengthen collaboration on critical AI safety and security research

## Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models

DevFeed: [Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models](<https://devfeed.tech/articles/apriel-h1-the-surprising-key-to-distilling-efficient-reasoning-models-7043.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ServiceNow-AI/apriel-h1>)

Author: Torsten Scholak; Oleksiy Ostapenko; Raymond Li; Luke Kumar; Joel Lamy-Poirier

Published: 2025-11-19T05:19:07Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Mamba](<https://devfeed.tech/topics/mamba.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [models](<https://devfeed.tech/tags/models.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

The article explains how the Apriel-H1 team distilled a strong 15B reasoning model into more efficient hybrids using Mamba layers. Distillation on pretraining data and generic SFT data reduced reasoning quality, while high-quality reasoning traces from the teacher's SFT dataset preserved performance. The flagship model achieved roughly 2.1x throughput with minimal quality loss across several benchmarks.

### Source excerpt

When MiniMax published their M2 post-mortem in October explaining why they abandoned efficient attention at 230B scale, the narrative briefly became "efficient attention is dead." Within days, Kimi Linear proved otherwise. The real lesson: it depends on your constraints. Our constraint was simple: we had a strong 15B reasoning model and needed to make it efficient without starting over. No infinite compute for 20T-token pretraining. No luxury of architectural co-design from day one.

## GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum

DevFeed: [GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum](<https://devfeed.tech/articles/gpt-5-1-instant-and-gpt-5-1-thinking-system-card-addendum-6436.md>)

Original publisher: [Read original article](<https://openai.com/index/gpt-5-system-card-addendum-gpt-5-1>)

Published: 2025-11-12T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [ChatGPT](<https://devfeed.tech/topics/chatgpt.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [health](<https://devfeed.tech/tags/health.md>), [mental-health](<https://devfeed.tech/tags/mental-health.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [publication](<https://devfeed.tech/tags/publication.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [safety](<https://devfeed.tech/tags/safety.md>)

### AI overview

This system card addendum presents updated safety metrics for GPT-5.1 Instant and GPT-5.1 Thinking. It expands pre-deployment evaluations to cover mental health concerns and emotional reliance on ChatGPT, while describing the models' adaptive reasoning and safety mitigations.

### Source excerpt

This GPT-5 system card addendum provides updated safety metrics for GPT-5.1 Instant and Thinking, including new evaluations for mental health and emotional reliance.

## gpt-oss-safeguard technical report

DevFeed: [gpt-oss-safeguard technical report](<https://devfeed.tech/articles/gpt-oss-safeguard-technical-report-6441.md>)

Original publisher: [Read original article](<https://openai.com/index/gpt-oss-safeguard-technical-report>)

Published: 2025-10-29T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [gpt-oss](<https://devfeed.tech/topics/gpt-oss.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [oss](<https://devfeed.tech/tags/oss.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [report](<https://devfeed.tech/tags/report.md>), [safety](<https://devfeed.tech/tags/safety.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

This technical report introduces gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, open-weight reasoning models fine-tuned from gpt-oss to classify content according to a supplied policy. It describes their capabilities, customization, chain-of-thought support, Responses API compatibility, and baseline safety evaluations.

### Source excerpt

gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard's capabilities and provide our baseline safety evaluations on the gpt-oss-safeguard models, using the underlying gpt-oss models as a baseline. For more information about the development and architecture of the underlying gpt-oss models, see the original gpt-oss model model card⁠.

## Introducing gpt-oss-safeguard

DevFeed: [Introducing gpt-oss-safeguard](<https://devfeed.tech/articles/introducing-gpt-oss-safeguard-6497.md>)

Original publisher: [Read original article](<https://openai.com/index/introducing-gpt-oss-safeguard>)

Published: 2025-10-29T00:00:00Z

Content type: release

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [gpt-oss](<https://devfeed.tech/topics/gpt-oss.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [developers](<https://devfeed.tech/tags/developers.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [openai](<https://devfeed.tech/tags/openai.md>), [oss](<https://devfeed.tech/tags/oss.md>), [product](<https://devfeed.tech/tags/product.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [safety](<https://devfeed.tech/tags/safety.md>)

### AI overview

OpenAI introduces gpt-oss-safeguard, a research preview of open-weight reasoning models for safety classification. The models let developers provide and revise their own policies at inference time, classify messages and conversations, and review the model's reasoning. They are available in 120b and 20b sizes under the Apache 2.0 license and can be downloaded from Hugging Face.

### Source excerpt

OpenAI introduces gpt-oss-safeguard--open-weight reasoning models for safety classification that let developers apply and iterate on custom policies.

## Introducing the Palmyra-mini family: Powerful, lightweight, and ready to reason!

DevFeed: [Introducing the Palmyra-mini family: Powerful, lightweight, and ready to reason!](<https://devfeed.tech/articles/introducing-the-palmyra-mini-family-powerful-lightweight-and-ready-to-reason-7060.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/Writer/announcing-palmyra-mini>)

Author: Rakshith; Tom Peres

Published: 2025-09-11T20:04:44Z

Content type: news

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [sglang](<https://devfeed.tech/topics/sglang.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [vllm](<https://devfeed.tech/topics/vllm.md>)

Tags: [announce](<https://devfeed.tech/tags/announce.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [inference](<https://devfeed.tech/tags/inference.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [release](<https://devfeed.tech/tags/release.md>), [rl](<https://devfeed.tech/tags/rl.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [tgi](<https://devfeed.tech/tags/tgi.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

WRITER announces three open Palmyra-mini models ranging from 1.5B to 1.7B parameters: a lightweight base model and two reasoning variants. The models target efficient inference across varied applications, with GGUF and MLX quantizations available. The article reports benchmark results, describes Chain of Thought training for the reasoning variants, and discusses inference-framework compatibility and reinforcement-learning trade-offs.

### Source excerpt

The team at WRITER is thrilled to announce the release of three new open models in the Palmyra-mini family. These models are designed to be powerful, lightweight, and highly performant for their size (1.5B to 1.7B), making them ideal for a wide range of applications with efficient inference. - palmyra-mini: A powerful, lightweight non-thinking base model. - palmyra-mini-thinking-a: A specialized variant optimized for complex reasoning and logic.

## Explainer: K2 & Math Olympiad Golds

DevFeed: [Explainer: K2 & Math Olympiad Golds](<https://devfeed.tech/articles/explainer-k2-math-olympiad-golds-33471.md>)

Original publisher: [Read original article](<https://timkellogg.me/blog/2025/07/19/olympiad>)

Published: 2025-07-19T00:00:00Z

Content type: article

Language: en

Sources: [Tim Kellogg](<https://devfeed.tech/sources/tim-kellogg.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [multi-agents](<https://devfeed.tech/topics/multi-agents.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [context-engineering](<https://devfeed.tech/tags/context-engineering.md>), [cost](<https://devfeed.tech/tags/cost.md>), [explainer](<https://devfeed.tech/tags/explainer.md>), [multi-agent](<https://devfeed.tech/tags/multi-agent.md>), [multi-agents](<https://devfeed.tech/tags/multi-agents.md>), [openai](<https://devfeed.tech/tags/openai.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

This explainer reviews major AI-agent developments from January through July 2025, focusing on K2 and the International Math Olympiad gold result. It argues that K2's strong agentic performance without a long chain-of-thought trace raises questions about whether extended thinking is necessary for effective agents, while noting that shorter reasoning can reduce token costs.

### Source excerpt

the best way to emphasize the importance of this week's developments is to go all the way back to January and see how we got here.

## Magistral

DevFeed: [Magistral](<https://devfeed.tech/articles/magistral-7030.md>)

Original publisher: [Read original article](<https://mistral.ai/news/magistral/>)

Published: 2025-06-10T12:00:00Z

Content type: news

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Mistral AI announces Magistral, its first reasoning model, in two variants: the open-source Magistral Small and enterprise-focused Magistral Medium. The article highlights domain-specific and multilingual reasoning, transparent chain-of-thought, enterprise use cases, benchmark results, and supporting research on training infrastructure and reinforcement learning.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## The 4 Things Qwen-3's Chat Template Teaches Us

DevFeed: [The 4 Things Qwen-3's Chat Template Teaches Us](<https://devfeed.tech/articles/the-4-things-qwen-3-s-chat-template-teaches-us-7451.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/qwen-3-chat-template-deep-dive>)

Author: Caleb Fahlgren

Published: 2025-04-30T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [qwen](<https://devfeed.tech/topics/qwen.md>), [Template](<https://devfeed.tech/topics/template.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [JSON](<https://devfeed.tech/topics/json.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [chat-template](<https://devfeed.tech/tags/chat-template.md>), [code](<https://devfeed.tech/tags/code.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [json](<https://devfeed.tech/tags/json.md>), [llms](<https://devfeed.tech/tags/llms.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [tool](<https://devfeed.tech/tags/tool.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

The article examines Qwen-3's more sophisticated Jinja chat template and compares it with earlier Qwen models. It explains optional reasoning control, rolling checkpoint preservation during tool calls, conditional JSON serialization, and the absence of a default system prompt in Qwen-3 and QwQ.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Prompt Engineering as a Developer Discipline

DevFeed: [Prompt Engineering as a Developer Discipline](<https://devfeed.tech/articles/prompt-engineering-as-a-developer-discipline-5750.md>)

Original publisher: [Read original article](<https://neon.com/blog/prompt-engineering-developer-discipline>)

Author: Andrew Tate

Published: 2025-04-21T19:19:58Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Prompt Engineering](<https://devfeed.tech/topics/prompt-engineering.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Programming](<https://devfeed.tech/topics/programming.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [coding](<https://devfeed.tech/tags/coding.md>), [developer](<https://devfeed.tech/tags/developer.md>), [llms](<https://devfeed.tech/tags/llms.md>), [product](<https://devfeed.tech/tags/product.md>), [prompt-engineering](<https://devfeed.tech/tags/prompt-engineering.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article presents prompt engineering as a core developer discipline for working effectively with AI and language models. It recommends treating prompts like modular, testable software components and using examples, structured instructions, and step-by-step reasoning to produce clearer, more reliable code.

### Source excerpt

AI is here. That might seem like a trite comment, but almost a quarter of developers still see AI as something they don't plan to use: But 'using AI' doesn't necessarily mean vibe coding your application into oblivion. Using AI as a developer means two things: The key to the seco...

## Open R1: Update #3

DevFeed: [Open R1: Update #3](<https://devfeed.tech/articles/open-r1-update-3-7423.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-r1/update-3>)

Author: Guilherme Penedo; Lewis Tunstall; Anton Lozhkov; Hynek Kydlicek; Edward Beeching; Loubna Ben Allal; Quentin Gallouédec; Leandro von Werra; Agustín Piqueres Lajarín; Nathan Habib

Published: 2025-03-11T20:40:47Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Programming](<https://devfeed.tech/topics/programming.md>), [Code](<https://devfeed.tech/topics/code.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [C++](<https://devfeed.tech/topics/c-plus-plus.md>), [Python](<https://devfeed.tech/topics/python.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [c-plus-plus](<https://devfeed.tech/tags/c-plus-plus.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [code](<https://devfeed.tech/tags/code.md>), [contests](<https://devfeed.tech/tags/contests.md>), [data](<https://devfeed.tech/tags/data.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [programming](<https://devfeed.tech/tags/programming.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [solutions](<https://devfeed.tech/tags/solutions.md>), [training](<https://devfeed.tech/tags/training.md>), [update](<https://devfeed.tech/tags/update.md>)

### AI overview

Open R1 Update #3 presents CodeForces-CoTs, a dataset of nearly 100,000 DeepSeek-R1 chain-of-thought samples for generating competitive-programming solutions in C++ and Python. It also introduces the IOI benchmark and OlympicCoder models fine-tuned on this data, reporting strong performance on challenging IOI problems.

### Source excerpt

Over the last few weeks, we have focused our efforts on reproducing the competitive programming (code reasoning) aspects of the DeepSeek-R1 recipe. In this post, we are excited to share: - The construction of CodeForces-CoTs: a dataset of nearly 100k high-quality samples distilled from R1 to produce solutions in C++ and Python. - The IOI benchmark: a new benchmark of challenging problems from the 2024 International Olympiad in Informatics (IOI).

## Open R1: Update #2

DevFeed: [Open R1: Update #2](<https://devfeed.tech/articles/open-r1-update-2-7422.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-r1/update-2>)

Author: Loubna Ben Allal; Lewis Tunstall; Anton Lozhkov; Elie Bakouch; Guilherme Penedo; Hynek Kydlicek; Gabriel Martín Blázquez

Published: 2025-02-10T16:10:47Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [deepseek](<https://devfeed.tech/topics/deepseek.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [math](<https://devfeed.tech/topics/math.md>), [llama](<https://devfeed.tech/topics/llama.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [sglang](<https://devfeed.tech/topics/sglang.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [math-verify](<https://devfeed.tech/topics/math-verify.md>), [Parser](<https://devfeed.tech/topics/parser.md>)

Tags: [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [llama](<https://devfeed.tech/tags/llama.md>), [math](<https://devfeed.tech/tags/math.md>), [math-verify](<https://devfeed.tech/tags/math-verify.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

Open R1 Update #2 presents OpenR1-Math-220k, a large-scale mathematical reasoning dataset created to help reconstruct parts of the DeepSeek R1 training pipeline and synthetic-data process. The article describes reasoning-trace generation, local inference with vLLM and SGLang, automated filtering with Math Verify, and the use of Llama3.3-70B-Instruct as a judge. It also discusses distillation and fine-tuning of Qwen and Llama models using reasoning traces.

### Source excerpt

We are now two weeks into the Open R1 project which aims to reconstruct the missing pieces of DeepSeek R1--specifically, the training pipeline and synthetic data. In this post, we are happy to share the construction of OpenR1-Math-220k: our first large-scale dataset for mathematical reasoning!

[Next page](<https://devfeed.tech/topics/chain-of-thought.md?cursor=WyIyMDI1LTAyLTEwVDE2OjEwOjQ3KzAwOjAwIiwgImQ2MWVkZmI0LTI4YzgtNGRiMy1iYzM3LTgwZTNiMThkZmVhMiJd>)