# rlhf

Published articles for rlhf.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Paul Christiano joins OpenAI Foundation Board

DevFeed: [Paul Christiano joins OpenAI Foundation Board](<https://devfeed.tech/articles/paul-christiano-joins-openai-foundation-board-6604.md>)

Original publisher: [Read original article](<https://openai.com/index/paul-christiano-joins-openai-foundation-board>)

Published: 2026-09-09T17:00:00Z

Content type: news

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Responsibility & Safety](<https://devfeed.tech/topics/responsibility-safety.md>), [human feedback](<https://devfeed.tech/topics/human-feedback.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-alignment](<https://devfeed.tech/tags/ai-alignment.md>), [company](<https://devfeed.tech/tags/company.md>), [government](<https://devfeed.tech/tags/government.md>), [openai](<https://devfeed.tech/tags/openai.md>), [research](<https://devfeed.tech/tags/research.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [safety](<https://devfeed.tech/tags/safety.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

OpenAI announces Paul Christiano's appointment to the OpenAI Foundation Board and its Safety and Security Committee. The article highlights his experience in AI alignment, frontier AI model evaluation, safety and security risk mitigation, governance, and reinforcement learning from human feedback.

### Source excerpt

Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.

## Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

DevFeed: [Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original](<https://devfeed.tech/articles/quantization-aware-healing-a-compressed-4-bit-model-that-outperforms-its-full-precision-original-7023.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing>)

Author: Antonio Tiene; Iker García-Ferrero; Ali Hashemi; Bakbergen Ryskulov

Published: 2026-08-25T11:39:24Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [blog](<https://devfeed.tech/tags/blog.md>), [compression](<https://devfeed.tech/tags/compression.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [llms](<https://devfeed.tech/tags/llms.md>), [model](<https://devfeed.tech/tags/model.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>)

### AI overview

The article presents Quantization-Aware Healing (QAH), a method for recovering structurally compressed and 4-bit-quantized LLMs. It contrasts QAH with quantization-aware training and distillation, arguing that the latter can be limited when no independently trained full-precision version of the compressed architecture exists.

### Source excerpt

A Blog post by Multiverse Computing on Hugging Face

## What I Saw at ICML 2026

DevFeed: [What I Saw at ICML 2026](<https://devfeed.tech/articles/arxiv-icml-2026-24886.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1073774/>)

Author: zj-karina (Яндекс)

Published: 2026-08-25T07:01:28Z

Content type: article

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [icml](<https://devfeed.tech/tags/icml.md>), [icml-2026](<https://devfeed.tech/tags/icml-2026.md>), [llm-agents](<https://devfeed.tech/tags/llm-agents.md>), [ml](<https://devfeed.tech/tags/ml.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [tag-316edb31b5b3](<https://devfeed.tech/tags/tag-316edb31b5b3.md>), [tag-44d9110ce940](<https://devfeed.tech/tags/tag-44d9110ce940.md>)

### AI overview

A Yandex developer reports from ICML 2026 in Seoul, describing the conference format, its scale, Yandex research presented there, and discussions about AI agents.

### Source excerpt

Зачем тратить сутки на перелёты, мчаться на другой конец света и жить неделю в режиме нон-стоп на одной из главных ML-конференций планеты, когда пейпер уже на arXiv, код -- на GitHub, а краткие выжимки из выступлений -- мгновенно в соцсетях? Меня зовут Карина Романова, я разработчик в Яндексе и занимаюсь LLM-агентами в Алисе. В июле мы с командой прилетели в Сеул на ICML 2026, и я ответила себе на вопрос "зачем?". Для нас офлайн-конференции -- это единственный способ за несколько дней прочувствовать реальный фокус сообщества, встретиться с авторами работ и узнать детали, которых нет в опубликованных текстах. В этой статье расскажу, как устроена ICML изнутри, чем запомнилась программа этого года, какие наши исследования вызвали наибольший ажиотаж и почему заметная часть разговоров на конференции снова вращалась вокруг AI-агентов. Читать далее

## Mastering Agentic Techniques: AI Agent Reinforcement Learning

DevFeed: [Mastering Agentic Techniques: AI Agent Reinforcement Learning](<https://devfeed.tech/articles/mastering-agentic-techniques-ai-agent-reinforcement-learning-6879.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/mastering-agentic-techniques-ai-agent-reinforcement-learning/>)

Author: Elizabeth Goodman

Published: 2026-07-01T17:04:02Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Language models](<https://devfeed.tech/topics/language-models.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-platforms-deployment](<https://devfeed.tech/tags/ai-platforms-deployment.md>), [featured](<https://devfeed.tech/tags/featured.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [rag](<https://devfeed.tech/tags/rag.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [rlvr](<https://devfeed.tech/tags/rlvr.md>), [tool](<https://devfeed.tech/tags/tool.md>)

### AI overview

A guide to using reinforcement learning with verifiable rewards to post-train language models for specialized, long-running AI agents. It explains when prompting, RAG, tools, and agent harnesses are insufficient, and describes reward signals based on verifiers, execution, validation, models, and human feedback.

### Source excerpt

Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer...

## Using a Claude Code Stop Hook to Externally Verify Work Before Completion

DevFeed: [Using a Claude Code Stop Hook to Externally Verify Work Before Completion](<https://devfeed.tech/articles/the-stop-hook-that-won-t-let-claude-lie-to-you-28993.md>)

Original publisher: [Read original article](<https://codingwithroby.substack.com/p/the-stop-hook-that-wont-let-claude>)

Author: Eric Roby

Published: 2026-06-02T13:01:38Z

Content type: tutorial

Language: en

Sources: [Eric Roby](<https://devfeed.tech/sources/eric-roby.md>)

Topics: [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [rlhf](<https://devfeed.tech/topics/rlhf.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [claude-code](<https://devfeed.tech/tags/claude-code.md>), [pull-request](<https://devfeed.tech/tags/pull-request.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [shell](<https://devfeed.tech/tags/shell.md>), [tests](<https://devfeed.tech/tags/tests.md>)

### AI overview

This tutorial explains how a Claude Code Stop hook can externally verify work before Claude is allowed to declare a task complete. It frames the approach as a safeguard against unverified claims such as reporting that all tests pass when they have not been run.

### Source excerpt

How to make Claude prove the work is done before it claims to be done.

## No GPU left behind: Unlocking Efficiency with Co-located vLLM in TRL

DevFeed: [No GPU left behind: Unlocking Efficiency with Co-located vLLM in TRL](<https://devfeed.tech/articles/no-gpu-left-behind-unlocking-efficiency-with-co-located-vllm-in-trl-7558.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/vllm-colocate>)

Author: Mert Toslali; Yu Chin Fabian Lim; Quentin Gallouédec; Ed Snible; Raghu Ganti; Mudhakar Srivatsa

Published: 2025-06-03T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vllm](<https://devfeed.tech/topics/vllm.md>), [trl](<https://devfeed.tech/topics/trl.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [efficiency](<https://devfeed.tech/tags/efficiency.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [performance](<https://devfeed.tech/tags/performance.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [train](<https://devfeed.tech/tags/train.md>), [training](<https://devfeed.tech/tags/training.md>), [trl](<https://devfeed.tech/tags/trl.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This article explains how TRL integrates vLLM to accelerate GRPO training of LLMs. It describes the GPU inefficiencies caused by running training and generation on separate devices and introduces colocated vLLM, which allows both tasks to share GPUs within the same distributed process group.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## 🐯 Liger GRPO meets TRL

DevFeed: [🐯 Liger GRPO meets TRL](<https://devfeed.tech/articles/liger-grpo-meets-trl-7332.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/liger-grpo>)

Author: Shivam Sahni; Kashif Rasul; Salman Mohammadi; Shirin Yamani; Yanning Chen; Liberty

Published: 2025-05-25T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [grpo](<https://devfeed.tech/topics/grpo.md>), [trl](<https://devfeed.tech/topics/trl.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [rlhf](<https://devfeed.tech/topics/rlhf.md>), [coding](<https://devfeed.tech/topics/coding.md>), [math](<https://devfeed.tech/topics/math.md>)

Tags: [coding](<https://devfeed.tech/tags/coding.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [liger](<https://devfeed.tech/tags/liger.md>), [llm](<https://devfeed.tech/tags/llm.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [performance](<https://devfeed.tech/tags/performance.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [trl](<https://devfeed.tech/tags/trl.md>)

### AI overview

This article explains how Group Relative Policy Optimization (GRPO) can reduce the resource requirements of reinforcement learning fine-tuning for language models. It presents a TRL optimization based on chunked GRPO loss that reduces peak memory usage by 40% and discusses scaling GRPO across multiple GPUs and nodes while preserving performance and correctness.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Preference Optimization for Vision Language Models

DevFeed: [Preference Optimization for Vision Language Models](<https://devfeed.tech/articles/preference-optimization-for-vision-language-models-7175.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/dpo_vlm>)

Author: Quentin Gallouédec; Shengyi Costa Huang; merve; Kashif Rasul

Published: 2024-07-10T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [images](<https://devfeed.tech/tags/images.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [training](<https://devfeed.tech/tags/training.md>), [trl](<https://devfeed.tech/tags/trl.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

A tutorial on training vision-language models with TRL's direct preference optimization support, covering preference data formatting and memory considerations.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Welcome Gemma 2 - Google's new open LLM

DevFeed: [Welcome Gemma 2 - Google's new open LLM](<https://devfeed.tech/articles/welcome-gemma-2-google-s-new-open-llm-7211.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/gemma2>)

Author: Philipp Schmid; Omar Sanseviero; Pedro Cuenca; Lewis Tunstall; Tom Aarsen; Vaibhav Srivastav

Published: 2024-06-27T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [gemma](<https://devfeed.tech/topics/gemma.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Google](<https://devfeed.tech/topics/google.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [rlhf](<https://devfeed.tech/topics/rlhf.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>)

Tags: [community](<https://devfeed.tech/tags/community.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [google](<https://devfeed.tech/tags/google.md>), [llm](<https://devfeed.tech/tags/llm.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [tpu](<https://devfeed.tech/tags/tpu.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [warp](<https://devfeed.tech/tags/warp.md>)

### AI overview

The article introduces Gemma 2, Google's open large language model family available in 9-billion- and 27-billion-parameter sizes, with base and instruction-tuned variants. It describes the models' training data, permissive licensing, architectural improvements, TPU-based training, and instruction-tuning methods including supervised fine-tuning, distillation, RLHF, and model merging.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Putting RL back in RLHF

DevFeed: [Putting RL back in RLHF](<https://devfeed.tech/articles/putting-rl-back-in-rlhf-7446.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/putting_rl_back_in_rlhf_with_rloo>)

Author: Shengyi Costa Huang; Arash Ahmadian

Published: 2024-06-12T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [rlhf](<https://devfeed.tech/topics/rlhf.md>), [trl](<https://devfeed.tech/topics/trl.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [dpo](<https://devfeed.tech/topics/dpo.md>), [cohere](<https://devfeed.tech/topics/cohere.md>)

Tags: [cohere](<https://devfeed.tech/tags/cohere.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [rl](<https://devfeed.tech/tags/rl.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [trl](<https://devfeed.tech/tags/trl.md>)

### AI overview

This article introduces the RLOO Trainer in TRL, an online reinforcement learning algorithm for RLHF designed as a more accessible alternative to PPO. It explains that RLOO uses less GPU memory, converges faster, performs competitively with PPO, and outperforms offline methods such as DPO in the reported comparisons.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Predictive Human Preference: From Model Ranking to Model Routing

DevFeed: [Predictive Human Preference: From Model Ranking to Model Routing](<https://devfeed.tech/articles/predictive-human-preference-from-model-ranking-to-model-routing-31797.md>)

Original publisher: [Read original article](<https://huyenchip.com//2024/02/28/predictive-human-preference.html>)

Author: Chip Huyen

Published: 2024-02-28T00:00:00Z

Content type: article

Language: en

Sources: [Chip Huyen](<https://devfeed.tech/sources/chip-huyen.md>)

Topics: [Model Routing](<https://devfeed.tech/topics/model-routing.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Model Development](<https://devfeed.tech/topics/model-development.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [latency](<https://devfeed.tech/tags/latency.md>), [model](<https://devfeed.tech/tags/model.md>), [model-routing](<https://devfeed.tech/tags/model-routing.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>)

### AI overview

The article examines predictive human preference: predicting which AI model users will prefer for a specific prompt. It describes model routing as a use case, where prompts could be directed to a preferred model to potentially reduce cost and latency, and discusses using preference predictions to interpret model strengths and weaknesses. It also outlines evaluating predictions against Chatbot Arena and building a preference predictor.

### Source excerpt

A challenge of building AI applications is choosing which model to use. What if we don't have to? What if we can predict the best model for any prompt? Predictive human preference aims to predict which model users might prefer for a specific query. Human preference has emerged to be both the Northstar and a powerful tool for AI model development. Human preference guides post-training techniques including RLHF and DPO. Human preference is also used to rank AI models, as used by LMSYS's Chatbot Arena. Chatbot Arena aims to determine which model is generally preferred. I wanted to see if it's possible to predict which model is preferred for each query. One use case of predictive human preference is model routing. For example, if we know in advance that for a prompt, users will prefer Claude Instant's response over GPT-4, and Claude Instant is cheaper/faster than GPT-4, we can route this prompt to Claude Instant. Model routing has the potential to increase response quality while reducing costs and latency. Another use case of predictive human preference is interpretability. Mapping out a model's performance on different prompts can help us understand this model's strengths and weaknesses. See section Experiment results for examples. Here's what predictive human preference for different model pairs looks like for the prompt "What's the best way to cluster text embeddings?". The predictions were generated by my toy preference predictor. The bright yellow color for the (GPT-4, GPT-3.5-Turbo) cell means that my predictor thinks GPT-4's response is very likely to be preferred to that of GPT-3.5-Turbo's for this prompt. This post first discusses the correctness of Chatbot Arena, which will then be used as a baseline to evaluate the correctness of preference predictions. It then discusses how to build a preference predictor and the initial results. Ranking Models Using Human Preference Using preferential signals (comparisons) to rank models has grown in popularity in the last