# rl

Published articles for rl.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## How to Fine-Tune LLMs in 2026

DevFeed: [How to Fine-Tune LLMs in 2026](<https://devfeed.tech/articles/how-to-fine-tune-llms-in-2026-31467.md>)

Original publisher: [Read original article](<https://blog.dailydoseofds.com/p/how-to-fine-tune-llms-in-2026-bf8>)

Author: Avi Chawla

Published: 2026-09-16T20:40:26Z

Content type: tutorial

Language: en

Sources: [Daily Dose of Data Science](<https://devfeed.tech/sources/daily-dose-of-data-science.md>)

Topics: [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [llms](<https://devfeed.tech/tags/llms.md>), [rl](<https://devfeed.tech/tags/rl.md>)

### AI overview

A developer newsletter explains how supervised fine-tuning differs from reinforcement fine-tuning for LLMs and describes GRPO and RULER as approaches for training agents through experience without manually written reward functions or labeled examples. It also briefly discusses Rowboat Spaces, an open-source shared workspace for personal AI assistants.

### Source excerpt

Reward-free RL is here!

## Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train

DevFeed: [Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train](<https://devfeed.tech/articles/bypassing-inference-bottlenecks-accelerating-complex-ai-search-with-retrieve-for-train-26972.md>)

Original publisher: [Read original article](<https://research.google/blog/bypassing-inference-bottlenecks-accelerating-complex-ai-search-with-retrieve-for-train/>)

Published: 2026-09-15T20:00:35Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [AI search](<https://devfeed.tech/topics/ai-search.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Algorithms & Theory](<https://devfeed.tech/topics/algorithms-theory.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [data-mining-modeling](<https://devfeed.tech/tags/data-mining-modeling.md>), [diffusion](<https://devfeed.tech/tags/diffusion.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [icml](<https://devfeed.tech/tags/icml.md>), [icml-2026](<https://devfeed.tech/tags/icml-2026.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [rl](<https://devfeed.tech/tags/rl.md>), [search](<https://devfeed.tech/tags/search.md>)

### AI overview

Google Research presents Retrieve-for-Train, a framework that uses offline reinforcement learning to compile reward-aligned query fan-outs into training data for a lightweight diffusion retriever. The approach is intended to produce diverse, complementary, and coherent search-result sets in a single inference pass, reducing reliance on expensive inference-time reasoning.

### Source excerpt

Algorithms & Theory

## Открываем претрейн Alice AI Search: как устроена модель быстрых ответов Алисы на Поиске

DevFeed: [Открываем претрейн Alice AI Search: как устроена модель быстрых ответов Алисы на Поиске](<https://devfeed.tech/articles/alice-ai-search-24897.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1080654/>)

Author: pet67 (Яндекс)

Published: 2026-09-11T06:05:13Z

Content type: article

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>), [Transformers](<https://devfeed.tech/topics/transformers.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [alice-ai](<https://devfeed.tech/tags/alice-ai.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [llm](<https://devfeed.tech/tags/llm.md>), [ml](<https://devfeed.tech/tags/ml.md>), [moe](<https://devfeed.tech/tags/moe.md>), [rl](<https://devfeed.tech/tags/rl.md>), [tag-178bc8f01f24](<https://devfeed.tech/tags/tag-178bc8f01f24.md>), [tag-4004cf5948d3](<https://devfeed.tech/tags/tag-4004cf5948d3.md>), [tag-61cd5a476b1d](<https://devfeed.tech/tags/tag-61cd5a476b1d.md>), [tag-d89cae10e887](<https://devfeed.tech/tags/tag-d89cae10e887.md>), [tag-e6d9cc1f0757](<https://devfeed.tech/tags/tag-e6d9cc1f0757.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

This developer article explains the Alice AI Search pipeline for generating fast answers, including its search and context-processing stages, shorter information contexts, a sparse Mixture-of-Experts architecture combined with an Encoder-Decoder, and online reinforcement learning from user behavior signals. It also announces the open release of the Alice AI-T5-35B-A0.6B Base model, with external inference available through Hugging Face Transformers while optimized production inference remains internal to Yandex.

### Source excerpt

Быстрый ответ Алисы AI -- это самый массовый генеративный продукт Яндекса и первое соприкосновение с Алисой для пользователей Поиска. Даже в час пиковой нагрузки пользователь должен получить лаконичный ответ за считаные секунды. Для этого мы, команда Alice AI Search, адаптируем весь пайплайн быстрых ответов -- от собственного претрейна с кастомной архитектурой до онлайн-rl-обучения на поведенческие сигналы пользователей. В статье разберём, как устроен генеративный ответ в Поиске, и расскажем про основные улучшения июньского релиза: как мы ускорили ответы за счёт коротких инфоконтекстов, зачем совместили Encoder-Decoder с разреженной MoE-архитектурой и как обучение на реальных пользовательских сигналах повлияло на качество и использование продукта. Кроме того, мы выложили в открытый доступ обученную с нуля модель Alice AI-T5-35B-A0.6B Base с тем ограничением, что внешним пользователям доступен инференс через Hugging Face Transformers, а оптимизированный production-инференс пока доступен только внутри Яндекса. Читать далее

## Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

DevFeed: [Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL](<https://devfeed.tech/articles/async-grpo-with-lora-across-hf-jobs-a-bucket-a-proxy-and-no-nccl-17376.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/asyncgrpo-lora-hfjobs>)

Author: Amine Dirhoussi; Quentin Gallouédec; Kashif Rasul; Sergio Paniego

Published: 2026-09-10T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [lora](<https://devfeed.tech/topics/lora.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [jobs](<https://devfeed.tech/topics/jobs.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>), [async](<https://devfeed.tech/topics/async.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>)

Tags: [async](<https://devfeed.tech/tags/async.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [hub](<https://devfeed.tech/tags/hub.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [llm](<https://devfeed.tech/tags/llm.md>), [lora](<https://devfeed.tech/tags/lora.md>), [nccl](<https://devfeed.tech/tags/nccl.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [rl](<https://devfeed.tech/tags/rl.md>), [storage](<https://devfeed.tech/tags/storage.md>), [trl](<https://devfeed.tech/tags/trl.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This article describes asynchronous GRPO training with a LoRA adapter across separate Hugging Face Jobs. The adapter is synchronized to vLLM replicas through a shared Storage Bucket, while a proxy handles authentication, rollout routing, and adapter-load broadcasts. Five runs reduced the time for 500 steps from 3 hours 27 minutes to 53 minutes.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Training a coding model to paint watercolours with TRL and OpenEnv

DevFeed: [Training a coding model to paint watercolours with TRL and OpenEnv](<https://devfeed.tech/articles/training-a-coding-model-to-paint-watercolours-with-trl-and-openenv-7531.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/train-to-paint-with-code>)

Author: Sergio Paniego

Published: 2026-09-03T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [openenv](<https://devfeed.tech/topics/openenv.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [jobs](<https://devfeed.tech/topics/jobs.md>), [dataset](<https://devfeed.tech/topics/dataset.md>)

Tags: [ai-art](<https://devfeed.tech/tags/ai-art.md>), [coding](<https://devfeed.tech/tags/coding.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [guide](<https://devfeed.tech/tags/guide.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [openenv](<https://devfeed.tech/tags/openenv.md>), [rl](<https://devfeed.tech/tags/rl.md>), [spaces](<https://devfeed.tech/tags/spaces.md>), [training](<https://devfeed.tech/tags/training.md>), [trl](<https://devfeed.tech/tags/trl.md>)

### AI overview

A tutorial describing an open reproduction of a reinforcement-learning pipeline that trains a coding model to create watercolor-like paintings by writing JavaScript with p5.brush. It uses TRL and OpenEnv, with datasets, environments, training scripts, models, and other artifacts published on Hugging Face.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

DevFeed: [Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps](<https://devfeed.tech/articles/fine-tuning-a-350m-model-for-better-structured-outputs-in-100-grpo-steps-7235.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/grpo-with-trl-ifstruct>)

Author: Leonie Monigatti; ben burtenshaw; Sergio Paniego

Published: 2026-09-03T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [guide](<https://devfeed.tech/tags/guide.md>), [json](<https://devfeed.tech/tags/json.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model](<https://devfeed.tech/tags/model.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [python](<https://devfeed.tech/tags/python.md>), [rl](<https://devfeed.tech/tags/rl.md>), [training](<https://devfeed.tech/tags/training.md>), [trl](<https://devfeed.tech/tags/trl.md>)

### AI overview

A tutorial on fine-tuning a 350M language model with GRPO to improve structured-output and JSON Schema compliance, then evaluating it on the IFStruct benchmark.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Training 100x Cheaper Retrieval models Neon and Castform

DevFeed: [Training 100x Cheaper Retrieval models Neon and Castform](<https://devfeed.tech/articles/training-100x-cheaper-retrieval-models-neon-and-castform-5343.md>)

Original publisher: [Read original article](<https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency>)

Author: Pranav Aurora

Published: 2026-08-05T12:00:00Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [cost](<https://devfeed.tech/tags/cost.md>), [embedding](<https://devfeed.tech/tags/embedding.md>), [infra](<https://devfeed.tech/tags/infra.md>), [llms](<https://devfeed.tech/tags/llms.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [product](<https://devfeed.tech/tags/product.md>), [rag](<https://devfeed.tech/tags/rag.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [rl](<https://devfeed.tech/tags/rl.md>), [scale](<https://devfeed.tech/tags/scale.md>), [search](<https://devfeed.tech/tags/search.md>), [systems](<https://devfeed.tech/tags/systems.md>), [tools](<https://devfeed.tech/tags/tools.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

The article explains how Castform uses reinforcement-learning post-training to improve open-weight models for agentic retrieval. It contrasts multi-step retrieval with one-shot embedding search, emphasizing the cost and latency of repeated frontier-model calls and the potential for smaller open models to perform specific search tasks more cheaply.

### Source excerpt

"Most teams' best training data is just sitting in their databases. The problem is that turning raw data into something usable is hard, and letting agents read, search, and mutate data cheaply at scale requires advanced infra. Pointing Castform at Neon skips both." -- Ying Hang Seah, cofounder, Castform

## Чем запомнилась ICRA 2026: Reinforcement Learning, генерация сложных сценариев поведения и будущее робототехники

DevFeed: [Чем запомнилась ICRA 2026: Reinforcement Learning, генерация сложных сценариев поведения и будущее робототехники](<https://devfeed.tech/articles/icra-2026-reinforcement-learning-24875.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1065938/>)

Author: egavolk (Яндекс)

Published: 2026-08-04T08:00:45Z

Content type: article

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [icra](<https://devfeed.tech/tags/icra.md>), [ml](<https://devfeed.tech/tags/ml.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [rl](<https://devfeed.tech/tags/rl.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [tag-511fbf58fd45](<https://devfeed.tech/tags/tag-511fbf58fd45.md>), [tag-6faff4be08e9](<https://devfeed.tech/tags/tag-6faff4be08e9.md>), [tag-d704a344cc75](<https://devfeed.tech/tags/tag-d704a344cc75.md>), [tag-dace475544fb](<https://devfeed.tech/tags/tag-dace475544fb.md>)

### AI overview

The article reviews notable trends, papers, and engineering trade-offs discussed at ICRA 2026, with emphasis on reinforcement learning, autonomous-vehicle perception and planning pipelines, simulation, rare edge-case generation, and robotic learning. It also discusses award-winning work on manipulation, humanoid robots, and camera-conditioned policy learning.

### Source excerpt

Привет, Хабр! В начале июня в Вене прошла главная международная конференция по робототехнике и автономным системам -- International Conference on Robotics and Automation (ICRA). В этом году среди участников была и наша команда автономного транспорта Яндекса. Топиков, которые обсуждаются на ICRA, много, потому что она не только об ML -- она скорее о робототехнике в целом. Например, есть секции о механизмах и дизайне, а также о медицинских роботах. Было немало и чисто инженерных работ. Ключевой топик докладов на конференции -- RL, он же Reinforcement Learning, обучение с подкреплением. Также нас интересовали статьи по классическому пайплайну автономного автомобиля: perception + prediction + planner + simulation. Новые подходы к Robotic Learning тоже интересны, так как их можно перенести на задачи автономного транспорта. Меня зовут Егор Волков, я занимаюсь претрейном модели планирования движения в автономном транспорте Яндекса. Вместе со мной на конференцию ездил Максим Спорышев -- руководитель службы поведения и предсказания движения. В этой статье мы собрали самые интересные тренды, доклады и инженерные развилки, которые заметили на ICRA 2026, -- от Reinforcement Learning и генерации редких edge-кейсов до того, куда вообще двигается ML в робототехнике. Читать далее

## Towards a quantum computer that learns from its errors

DevFeed: [Towards a quantum computer that learns from its errors](<https://devfeed.tech/articles/towards-a-quantum-computer-that-learns-from-its-errors-6906.md>)

Original publisher: [Read original article](<https://research.google/blog/towards-a-quantum-computer-that-learns-from-its-errors/>)

Published: 2026-07-22T18:40:21Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Neural Network](<https://devfeed.tech/topics/neural-network.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [errors](<https://devfeed.tech/tags/errors.md>), [learning](<https://devfeed.tech/tags/learning.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [quantum](<https://devfeed.tech/tags/quantum.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [rl](<https://devfeed.tech/tags/rl.md>)

### AI overview

Google Research describes a reinforcement-learning framework that uses quantum error detections to continuously adjust control parameters during computation, helping stabilize a quantum computer against drift. The article also discusses quantum error correction and the AlphaQubit neural-network decoder.

### Source excerpt

Machine Intelligence

## How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo

DevFeed: [How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo](<https://devfeed.tech/articles/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo-6853.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/>)

Author: Tanya Lenz

Published: 2026-07-14T16:00:00Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [debugging](<https://devfeed.tech/topics/debugging.md>)

Tags: [agent-skill](<https://devfeed.tech/tags/agent-skill.md>), [agent-skills](<https://devfeed.tech/tags/agent-skills.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [python](<https://devfeed.tech/tags/python.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [rl](<https://devfeed.tech/tags/rl.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [training](<https://devfeed.tech/tags/training.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

A tutorial on running a skill-based autoresearch workflow in which coding AI agents set up, debug, run, monitor, and iterate on reinforcement-learning experiments using NVIDIA NeMo RL and NeMo Gym.

### Source excerpt

Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve...

## 🗓 This Week In AI Research (1-8 July 26)

DevFeed: [🗓 This Week In AI Research (1-8 July 26)](<https://devfeed.tech/articles/this-week-in-ai-research-1-8-july-26-18283.md>)

Original publisher: [Read original article](<https://www.intoai.pub/p/this-week-in-ai-research-1-8-july>)

Author: Dr. Ashish Bamania

Published: 2026-07-12T11:25:32Z

Content type: article

Language: en

Sources: [Into AI](<https://devfeed.tech/sources/into-ai.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [Algorithms](<https://devfeed.tech/topics/algorithms.md>), [releases](<https://devfeed.tech/topics/releases.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [algorithm](<https://devfeed.tech/tags/algorithm.md>), [algorithms](<https://devfeed.tech/tags/algorithms.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [releases](<https://devfeed.tech/tags/releases.md>), [research](<https://devfeed.tech/tags/research.md>), [rl](<https://devfeed.tech/tags/rl.md>), [training](<https://devfeed.tech/tags/training.md>), [update](<https://devfeed.tech/tags/update.md>)

### AI overview

A weekly roundup of AI research papers and releases highlights findings that reinforcement-learning gains can be concentrated in a single transformer layer and presents LLM-as-a-Verifier, a framework for continuous scoring and ranking of agentic-task solutions.

### Source excerpt

The top 10 research papers and AI releases this week (SpaceXAI's Grok 4.5, OpenAI's GPT-Live voice models, Cognition's SWE-1.7, Meta's Muse Spark 1.1, and many more)

## NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads

DevFeed: [NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads](<https://devfeed.tech/articles/nvidia-vera-cpu-boosts-ai-factory-throughput-to-accelerate-agentic-workloads-6910.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/nvidia-vera-cpu-boosts-ai-factory-throughput-to-accelerate-agentic-workloads/>)

Author: Michelle Horton

Published: 2026-07-07T18:10:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Vera CPU](<https://devfeed.tech/topics/vera-cpu.md>), [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [Neural Network](<https://devfeed.tech/topics/neural-network.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [cache](<https://devfeed.tech/tags/cache.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [featured](<https://devfeed.tech/tags/featured.md>), [inference](<https://devfeed.tech/tags/inference.md>), [model](<https://devfeed.tech/tags/model.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [performance](<https://devfeed.tech/tags/performance.md>), [rl](<https://devfeed.tech/tags/rl.md>), [top-stories](<https://devfeed.tech/tags/top-stories.md>), [vera-cpu](<https://devfeed.tech/tags/vera-cpu.md>)

### AI overview

This NVIDIA developer article explains how the NVIDIA Vera CPU can improve AI factory throughput for agentic workloads. It emphasizes sustained per-core performance for CPU tasks between model steps, including tool calls, code execution, sandbox evaluations, data processing, orchestration, KV-cache coordination, and result handling. The article also describes how CPU performance affects reinforcement learning rollouts, user response time, and cached-context efficiency.

### Source excerpt

Agentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration, and...

## GenPage: Towards End-to-End Generative Homepage Construction at Netflix

DevFeed: [GenPage: Towards End-to-End Generative Homepage Construction at Netflix](<https://devfeed.tech/articles/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-136.md>)

Original publisher: [Read original article](<https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08?source=rss----2615bd06b42e---4>)

Author: Netflix Technology Blog

Published: 2026-06-29T13:01:02Z

Content type: article

Language: en

Sources: [Netflix](<https://devfeed.tech/sources/netflix.md>), [Netflix TechBlog - Medium](<https://devfeed.tech/sources/netflix-techblog-medium.md>)

Topics: [Netflix](<https://devfeed.tech/topics/netflix.md>), [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [app](<https://devfeed.tech/tags/app.md>), [diversity](<https://devfeed.tech/tags/diversity.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [feature-engineering](<https://devfeed.tech/tags/feature-engineering.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative](<https://devfeed.tech/tags/generative.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [model](<https://devfeed.tech/tags/model.md>), [netflix](<https://devfeed.tech/tags/netflix.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [recommendation-system](<https://devfeed.tech/tags/recommendation-system.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [rl](<https://devfeed.tech/tags/rl.md>)

### AI overview

This Netflix developer article introduces GenPage, a generative approach that uses a single autoregressive model to construct a personalized homepage by generating recommendation rows, entities, and layout together. It describes replacing a multi-stage recommendation pipeline with end-to-end modeling and using reinforcement learning to optimize whole-page rewards, including interactions such as diversity and the balance between rows.

### Source excerpt

Authors: Lequn Wang, Jiangwei Pan, and Linas Baltrunas Figure 1. Autoregressive homepage generation. GenPage builds a Netflix homepage one row or entity at a time, each one conditioned on what's already on the page and the user's context.Introduction The Netflix homepage is the first thing users see when they open the app and the primary way they discover content to enjoy. Almost every part of it is personalized, including which rows appear, which entities show up within those rows, and how everything is arranged on the page. Constructing that homepage is a genuinely hard problem. It is not simply producing one ranked list. The homepage is a structured, two-dimensional layout, made up of recommendation rows and the entities within them. Here, an entity can be a movie, show, game, live event, or other recommendable item. Each choice can affect the value of the others. Traditionally, it is built through a complex, multi-stage pipeline, with separate components for candidate generation and ranking at both the row and entity levels. We saw an opportunity to rethink this design. Large language models have shown that a single generative model can perform diverse tasks just by generating a response to a prompt. Inspired by this prompt-response paradigm, we trained a single generative model to build the homepage by directly answering one question: Given everything we know about this user and this request, what homepage should we generate to maximize user satisfaction? We call this approach GenPage. It treats the user history and request context as the prompt, and autoregressively generates the entire homepage as the response (Figure 1). Unlike most generative recommenders, such as TIGER, HSTU, and OneRec, which generate flat ranked lists, GenPage generates the rows, entities, and layout together. This shift is motivated by several goals: End-to-end modeling. A single transformer model that constructs the page from raw input signals can replace a complex multi-stage recommen

## MosaicLeaks: Can your research agent keep a secret?

DevFeed: [MosaicLeaks: Can your research agent keep a secret?](<https://devfeed.tech/articles/mosaicleaks-can-your-research-agent-keep-a-secret-7054.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ServiceNow/mosaicleaks>)

Author: Alexander Gurung; Rafael Pardinas

Published: 2026-06-18T18:13:13Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Security](<https://devfeed.tech/topics/security.md>), [Machine Learning, Security Attacks](<https://devfeed.tech/topics/machine-learning-security-attacks.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [external](<https://devfeed.tech/tags/external.md>), [healthcare](<https://devfeed.tech/tags/healthcare.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [local](<https://devfeed.tech/tags/local.md>), [migration](<https://devfeed.tech/tags/migration.md>), [models](<https://devfeed.tech/tags/models.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [research](<https://devfeed.tech/tags/research.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [rl](<https://devfeed.tech/tags/rl.md>), [security](<https://devfeed.tech/tags/security.md>), [tools](<https://devfeed.tech/tags/tools.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

MosaicLeaks examine how deep-research agents can expose private enterprise information through their external web queries. The proposed Privacy-Aware Deep Research training method improves strict multi-hop task success while substantially reducing answer and full-information leakage.

### Source excerpt

Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent's external queries may leak sensitive information. MosaicLeaks proposes a new deep-research task with multi-hop questions that interleave public and private information. Across the models we tested, agents frequently leaked private information, and training only for task performance made it worse.

## Introducing OpenRL: A self-hosted post-training API for fine-tuning LLMs

DevFeed: [Introducing OpenRL: A self-hosted post-training API for fine-tuning LLMs](<https://devfeed.tech/articles/introducing-openrl-a-self-hosted-post-training-api-for-fine-tuning-llms-34311.md>)

Original publisher: [Read original article](<http://opensource.googleblog.com/2026/06/introducing-openrl-a-self-hosted-post-training-api-for-fine-tuning-llms.html>)

Author: Google Open Source (noreply@blogger.com)

Published: 2026-06-11T18:30:00Z

Content type: release

Language: en

Sources: [Google Open Source Blog](<https://devfeed.tech/sources/google-open-source-blog.md>)

Topics: [API](<https://devfeed.tech/topics/api.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Self-hosted](<https://devfeed.tech/topics/self-hosted.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [reliability](<https://devfeed.tech/topics/reliability.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gke](<https://devfeed.tech/tags/gke.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [llms](<https://devfeed.tech/tags/llms.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [rl](<https://devfeed.tech/tags/rl.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>), [sft](<https://devfeed.tech/tags/sft.md>)

### AI overview

This article announces OpenRL, an open-source research preview from GKE Labs. OpenRL is a self-hosted training API for fine-tuning LLMs on a Kubernetes cluster, designed to separate post-training infrastructure from AI research workflows. The article describes potential benefits including concurrent reinforcement-learning jobs, improved GPU utilization, and simpler researcher workflows.

### Source excerpt

by Sunil Arora, Shuby Mishra & Chuang Wang, GKE We are pleased to share a research preview of OpenRL, a new open-source project coming out of GKE Labs. OpenRL is a self-hosted training API for fine-tuning LLMs on your own Kubernetes cluster. Why we built it If you look at agentic RL on LLMs, it is incredibly easy to get bogged down in system complexity. To run a single RL loop, you have to coordinate a dozen different things: selecting and cleaning datasets, choosing RL environments, debugging training loops, managing reward signals, handling inference mismatches, allocating hardware, and managing infrastructure. Picture looks something like this: Figure shows an AI researcher and an infrastructure engineer staring at the hurdles in post training along the way to the summit. Each of these is a hard problem. But what makes it more complex is how tightly AI research and infrastructure concerns are mixed together in today's tooling and frameworks. We believe decoupling the infrastructure from AI research can make these problems more tractable so that infrastructure engineers and AI researchers can independently tackle them. We have seen this pattern with Kubernetes where Kubernetes abstracted out the infrastructure and made application developers and SREs life easier. So, can you abstract out post training infrastructure? We believe so and drew huge inspiration/validation from Tinker (from Thinking Machines). The Tinker APIs for post training hit that Goldilocks zone where it hides all the post training infrastructure behind four key APIs: Figure shows high level components and their interaction in a OpenRL based RL workflow So the end result of this abstraction is that AI Researchers get full flexibility on their RL loop and infrastructure engineers can focus on scaling, orchestration, and reliability. OpenRL allows you to run the same training APIs but on your own infrastructure. And this decoupling has other interesting benefits. Sharing GPUs Traditional RL loops ar

## The Open Source Community is backing OpenEnv for Agentic RL

DevFeed: [The Open Source Community is backing OpenEnv for Agentic RL](<https://devfeed.tech/articles/the-open-source-community-is-backing-openenv-for-agentic-rl-7429.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/openenv-agentic-rl>)

Author: ben burtenshaw; Joseph Spisak; Lysandre; Davide Testuggine; will brown; Joy Liu; Peyton Walters; Chris Wing; Daniel (Unsloth); Andrew Zhou

Published: 2026-06-08T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [openenv](<https://devfeed.tech/topics/openenv.md>), [interoperability](<https://devfeed.tech/topics/interoperability.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [community](<https://devfeed.tech/tags/community.md>), [interoperability](<https://devfeed.tech/tags/interoperability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [releases](<https://devfeed.tech/tags/releases.md>), [rl](<https://devfeed.tech/tags/rl.md>)

### AI overview

OpenEnv is becoming a more open, community-governed interoperability layer for training agents with reinforcement learning. It standardizes how RL environments are published, deployed, and consumed while allowing different models, harnesses, inference engines, and training libraries to work together.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

DevFeed: [Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL](<https://devfeed.tech/articles/shipping-a-trillion-parameters-with-a-hub-bucket-delta-weight-sync-in-trl-7166.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/delta-weight-sync>)

Author: Amine Dirhoussi; Quentin Gallouédec; Kashif Rasul; Lewis Tunstall; Edward Beeching; Albert Villanova del Moral; Leandro von Werra; Sergio Paniego

Published: 2026-05-27T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [compute](<https://devfeed.tech/tags/compute.md>), [hub](<https://devfeed.tech/tags/hub.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nccl](<https://devfeed.tech/tags/nccl.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [payload](<https://devfeed.tech/tags/payload.md>), [policy](<https://devfeed.tech/tags/policy.md>), [rl](<https://devfeed.tech/tags/rl.md>), [space](<https://devfeed.tech/tags/space.md>), [storage](<https://devfeed.tech/tags/storage.md>), [sync](<https://devfeed.tech/tags/sync.md>), [train](<https://devfeed.tech/tags/train.md>), [training](<https://devfeed.tech/tags/training.md>), [trl](<https://devfeed.tech/tags/trl.md>), [update](<https://devfeed.tech/tags/update.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

The article describes delta weight synchronization for asynchronous reinforcement-learning training. A TRL change stores only modified model weights in sparse safetensors files and lets vLLM fetch them from a Hugging Face bucket, reducing transfer payloads and enabling disaggregated training without a shared cluster.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Harness, Scaffold, and the AI Agent Terms Worth Getting Right

DevFeed: [Harness, Scaffold, and the AI Agent Terms Worth Getting Right](<https://devfeed.tech/articles/harness-scaffold-and-the-ai-agent-terms-worth-getting-right-7067.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/agent-glossary>)

Author: Sergio Paniego; Aritra Roy Gosthipaty

Published: 2026-05-25T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [codex](<https://devfeed.tech/topics/codex.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [frameworks](<https://devfeed.tech/tags/frameworks.md>), [guide](<https://devfeed.tech/tags/guide.md>), [llms](<https://devfeed.tech/tags/llms.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [rl](<https://devfeed.tech/tags/rl.md>)

### AI overview

A practical glossary explaining commonly confused terms in AI agents, including models, harnesses, scaffolding, tools, memory, and context management. It presents a mental model for understanding how an LLM becomes an agent and notes that terminology varies across frameworks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## vLLM V0 to V1: Correctness Before Corrections in RL

DevFeed: [vLLM V0 to V1: Correctness Before Corrections in RL](<https://devfeed.tech/articles/vllm-v0-to-v1-correctness-before-corrections-in-rl-7047.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ServiceNow-AI/correctness-before-corrections>)

Author: Rafael Pardinas; Ehsan Kamalloo

Published: 2026-05-06T19:06:55Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vllm](<https://devfeed.tech/topics/vllm.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>)

Tags: [backend](<https://devfeed.tech/tags/backend.md>), [comparison](<https://devfeed.tech/tags/comparison.md>), [inference](<https://devfeed.tech/tags/inference.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [rl](<https://devfeed.tech/tags/rl.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This article explains how a vLLM V1 migration was brought to parity with a vLLM V0 reference for reinforcement-learning training. The authors fixed processed rollout logprobs, V1 runtime defaults, the inflight weight-update path, and final-projection precision before changing the RL objective.

### Source excerpt

TL;DR. vLLM V1 matched our vLLM V0 reference after we fixed four things: processed rollout logprobs, V1-specific runtime defaults, the inflight weight-update path, and the fp32 used for the final projection. We fixed the backend behavior before changing the RL objective. The reference run used vLLM ; the V1 runs used vLLM . Figure 1 shows the final result. The red run is the initial V1 attempt, and the green run is the final V1 run after the fixes described below.

## Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries

DevFeed: [Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries](<https://devfeed.tech/articles/keep-the-tokens-flowing-lessons-from-16-open-source-rl-libraries-7109.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/async-rl-training-landscape>)

Author: Amine Dirhoussi; Quentin Gallouédec; Kashif Rasul; Lewis Tunstall; Edward Beeching; Albert Villanova del Moral; Nouamane Tazi; Leandro von Werra; Sergio Paniego

Published: 2026-03-10T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [distributed-training](<https://devfeed.tech/topics/distributed-training.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [lora](<https://devfeed.tech/topics/lora.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [llm](<https://devfeed.tech/tags/llm.md>), [lora](<https://devfeed.tech/tags/lora.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [nccl](<https://devfeed.tech/tags/nccl.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [rl](<https://devfeed.tech/tags/rl.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

This article surveys 16 open-source libraries for asynchronous reinforcement-learning training. It explains how separating inference and training across GPU pools, using rollout buffers, and synchronizing weights asynchronously can reduce training-GPU idle time. The comparison covers orchestration, buffering, weight synchronization, staleness management, partial rollouts, LoRA, and distributed-training backends, highlighting Ray, NCCL broadcasts, limited LoRA support, and distributed MoE as an emerging differentiator.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## OpenEnv in Practice: Evaluating Tool-Using Agents in Real-World Environments

DevFeed: [OpenEnv in Practice: Evaluating Tool-Using Agents in Real-World Environments](<https://devfeed.tech/articles/openenv-in-practice-evaluating-tool-using-agents-in-real-world-environments-7430.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/openenv-turing>)

Author: Christian Washington; Ankit Jasuja; Santosh Sah; Lewis Tunstall; ben burtenshaw

Published: 2026-02-12T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [api](<https://devfeed.tech/tags/api.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [openenv](<https://devfeed.tech/tags/openenv.md>), [rl](<https://devfeed.tech/tags/rl.md>)

### AI overview

OpenEnv evaluates tool-using AI agents in real environments. The article presents a calendar-management environment for testing long-horizon reasoning, permissions, and multi-step workflows.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective

DevFeed: [Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective](<https://devfeed.tech/articles/unlocking-agentic-rl-training-for-gpt-oss-a-practical-retrospective-7015.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/LinkedIn/gpt-oss-agentic-rl>)

Author: Jason Zhu; Hejian Sang; Arup De; Rohit Jain; Yanning Chen

Published: 2026-01-27T01:53:15Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [gpt-oss](<https://devfeed.tech/topics/gpt-oss.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [agentic-coding](<https://devfeed.tech/topics/agentic-coding.md>), [Retool](<https://devfeed.tech/topics/retool.md>), [Frameworks](<https://devfeed.tech/topics/frameworks.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [coding](<https://devfeed.tech/topics/coding.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [applications](<https://devfeed.tech/tags/applications.md>), [blog](<https://devfeed.tech/tags/blog.md>), [code](<https://devfeed.tech/tags/code.md>), [coding](<https://devfeed.tech/tags/coding.md>), [company](<https://devfeed.tech/tags/company.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [linkedin](<https://devfeed.tech/tags/linkedin.md>), [models](<https://devfeed.tech/tags/models.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [rl](<https://devfeed.tech/tags/rl.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

LinkedIn's retrospective describes experiments to train GPT-OSS models for agentic reinforcement learning. It covers tool interaction, Harmony chat-template support, rollout and tool parsing, ReTool coding tasks, attention-sink fixes, and benchmark results using GPT-OSS-20B, GPT-OSS-120B, and Qwen-2.5-32B.

### Source excerpt

LinkedIn is an AI-first company that's built agents to help professionals be more successful. In this setting, models must reason over incomplete information, interact with structured services, and adapt to evolving user intent across multiple steps rather than produce a single static response.

## Building the Open Agent Ecosystem Together: Introducing OpenEnv

DevFeed: [Building the Open Agent Ecosystem Together: Introducing OpenEnv](<https://devfeed.tech/articles/building-the-open-agent-ecosystem-together-introducing-openenv-7428.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/openenv>)

Author: Joseph Spisak; Davide Testuggine; Zach Wentz; Pierre Andrews; Sanyam Bhutani; Hamid Shojanazeri; Pankit Thapar; Emre Guven; Lewis Tunstall; Vaibhav Srivastav

Published: 2025-10-23T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [openenv](<https://devfeed.tech/topics/openenv.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [trl](<https://devfeed.tech/topics/trl.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [building](<https://devfeed.tech/tags/building.md>), [community](<https://devfeed.tech/tags/community.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [openenv](<https://devfeed.tech/tags/openenv.md>), [rl](<https://devfeed.tech/tags/rl.md>)

### AI overview

OpenEnv is an open specification and ecosystem for secure, sandboxed environments that provide AI agents with the tools, APIs, credentials, and execution context needed to perform tasks. Meta-PyTorch and Hugging Face are introducing an Environment Hub where developers can create, share, inspect, and run OpenEnv-compatible environments for training and deployment.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Introducing the Palmyra-mini family: Powerful, lightweight, and ready to reason!

DevFeed: [Introducing the Palmyra-mini family: Powerful, lightweight, and ready to reason!](<https://devfeed.tech/articles/introducing-the-palmyra-mini-family-powerful-lightweight-and-ready-to-reason-7060.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/Writer/announcing-palmyra-mini>)

Author: Rakshith; Tom Peres

Published: 2025-09-11T20:04:44Z

Content type: news

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [sglang](<https://devfeed.tech/topics/sglang.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [vllm](<https://devfeed.tech/topics/vllm.md>)

Tags: [announce](<https://devfeed.tech/tags/announce.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [inference](<https://devfeed.tech/tags/inference.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [release](<https://devfeed.tech/tags/release.md>), [rl](<https://devfeed.tech/tags/rl.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [tgi](<https://devfeed.tech/tags/tgi.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

WRITER announces three open Palmyra-mini models ranging from 1.5B to 1.7B parameters: a lightweight base model and two reasoning variants. The models target efficient inference across varied applications, with GGUF and MLX quantizations available. The article reports benchmark results, describes Chain of Thought training for the reasoning variants, and discusses inference-framework compatibility and reinforcement-learning trade-offs.

### Source excerpt

The team at WRITER is thrilled to announce the release of three new open models in the Palmyra-mini family. These models are designed to be powerful, lightweight, and highly performant for their size (1.5B to 1.7B), making them ideal for a wide range of applications with efficient inference. - palmyra-mini: A powerful, lightweight non-thinking base model. - palmyra-mini-thinking-a: A specialized variant optimized for complex reasoning and logic.

[Next page](<https://devfeed.tech/tags/rl.md?cursor=WyIyMDI1LTA5LTExVDIwOjA0OjQ0KzAwOjAwIiwgImQ1NGFmNGUwLTQ3MmMtNDdiZi1iYmY3LTNlZGY0NDhhYzQzZSJd>)