# ICLR

Published articles for ICLR.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Towards demystifying the creativity of diffusion models

DevFeed: [Towards demystifying the creativity of diffusion models](<https://devfeed.tech/articles/towards-demystifying-the-creativity-of-diffusion-models-6909.md>)

Original publisher: [Read original article](<https://research.google/blog/towards-demystifying-the-creativity-of-diffusion-models/>)

Published: 2026-07-15T18:06:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Mathematics](<https://devfeed.tech/topics/mathematics.md>), [Neural Network](<https://devfeed.tech/topics/neural-network.md>), [Algorithms & Theory](<https://devfeed.tech/topics/algorithms-theory.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [diffusion](<https://devfeed.tech/tags/diffusion.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [iclr-2026](<https://devfeed.tech/tags/iclr-2026.md>), [images](<https://devfeed.tech/tags/images.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [research](<https://devfeed.tech/tags/research.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Google Research explains that diffusion models can generate novel data rather than merely memorize training examples. It attributes this creativity to neural networks learning a smoothed score function, which causes denoising to interpolate between training data points along a hidden data manifold.

### Source excerpt

Algorithms & Theory

## Diverse reasoning traces teach LLMs to make better decisions

DevFeed: [Diverse reasoning traces teach LLMs to make better decisions](<https://devfeed.tech/articles/diverse-reasoning-traces-teach-llms-to-make-better-decisions-7597.md>)

Original publisher: [Read original article](<https://www.amazon.science/blog/diverse-reasoning-traces-teach-llms-to-make-better-decisions>)

Author: Sheng Jia; Xiao Wang; Shiva Kasiviswanathan

Published: 2026-05-26T15:17:06Z

Content type: article

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [conversational-ai](<https://devfeed.tech/tags/conversational-ai.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [iclr-2026](<https://devfeed.tech/tags/iclr-2026.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [llms](<https://devfeed.tech/tags/llms.md>), [math-reasoning](<https://devfeed.tech/tags/math-reasoning.md>), [parallel-reasoning](<https://devfeed.tech/tags/parallel-reasoning.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [post-training-optimization](<https://devfeed.tech/tags/post-training-optimization.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

The article presents set-supervised fine tuning and global forking policy optimization to train LLMs on multiple distinct reasoning paths. It reports 5% to 7% single-shot accuracy gains on standard benchmarks.

### Source excerpt

How to train language models to generate diverse, accurate reasoning paths using tokens that control distinct reasoning strategies.

## ReasoningBank: Enabling agents to learn from experience

DevFeed: [ReasoningBank: Enabling agents to learn from experience](<https://devfeed.tech/articles/reasoningbank-enabling-agents-to-learn-from-experience-6854.md>)

Original publisher: [Read original article](<https://research.google/blog/reasoningbank-enabling-agents-to-learn-from-experience/>)

Published: 2026-04-21T16:42:22Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [llm](<https://devfeed.tech/tags/llm.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

ReasoningBank is an agent memory framework that distills generalizable reasoning strategies from successful and failed experiences. It uses a continuous retrieval, extraction, and consolidation loop to support test-time self-evolution, improving effectiveness and efficiency on web browsing and software engineering benchmarks.

### Source excerpt

Generative AI

## AI-generated synthetic neurons speed up brain mapping

DevFeed: [AI-generated synthetic neurons speed up brain mapping](<https://devfeed.tech/articles/ai-generated-synthetic-neurons-speed-up-brain-mapping-6748.md>)

Original publisher: [Read original article](<https://research.google/blog/ai-generated-synthetic-neurons-speed-up-brain-mapping/>)

Published: 2026-04-16T12:18:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Google](<https://devfeed.tech/topics/google.md>), [Point cloud](<https://devfeed.tech/topics/point-cloud.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [neuron](<https://devfeed.tech/topics/neuron.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [3d](<https://devfeed.tech/tags/3d.md>), [accelerate](<https://devfeed.tech/tags/accelerate.md>), [ai](<https://devfeed.tech/tags/ai.md>), [classification](<https://devfeed.tech/tags/classification.md>), [errors](<https://devfeed.tech/tags/errors.md>), [general-science](<https://devfeed.tech/tags/general-science.md>), [generation](<https://devfeed.tech/tags/generation.md>), [google](<https://devfeed.tech/tags/google.md>), [health-bioscience](<https://devfeed.tech/tags/health-bioscience.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [iclr-2026](<https://devfeed.tech/tags/iclr-2026.md>), [images](<https://devfeed.tech/tags/images.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [mapping](<https://devfeed.tech/tags/mapping.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [neuron](<https://devfeed.tech/tags/neuron.md>), [partners](<https://devfeed.tech/tags/partners.md>), [research](<https://devfeed.tech/tags/research.md>), [scale](<https://devfeed.tech/tags/scale.md>), [science](<https://devfeed.tech/tags/science.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [visualization](<https://devfeed.tech/tags/visualization.md>)

### AI overview

Google Research describes how MoGen generates synthetic neuronal shapes to improve AI models that reconstruct brain wiring maps. Adding synthetic training examples reduced reconstruction errors by 4.4%, potentially saving 157 person-years of manual proofreading for a complete mouse brain.

### Source excerpt

General Science

## TurboQuant: Redefining AI efficiency with extreme compression

DevFeed: [TurboQuant: Redefining AI efficiency with extreme compression](<https://devfeed.tech/articles/turboquant-redefining-ai-efficiency-with-extreme-compression-6917.md>)

Original publisher: [Read original article](<https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/>)

Published: 2026-03-24T19:54:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [large-language-models](<https://devfeed.tech/topics/large-language-models.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Algorithms, Complexity](<https://devfeed.tech/topics/algorithms-complexity.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [algorithms](<https://devfeed.tech/tags/algorithms.md>), [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [cache](<https://devfeed.tech/tags/cache.md>), [compression](<https://devfeed.tech/tags/compression.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [iclr-2026](<https://devfeed.tech/tags/iclr-2026.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [memory](<https://devfeed.tech/tags/memory.md>), [model](<https://devfeed.tech/tags/model.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [research](<https://devfeed.tech/tags/research.md>), [search](<https://devfeed.tech/tags/search.md>)

### AI overview

Google Research introduces TurboQuant, a theoretically grounded compression algorithm for large language models and vector search. It targets vector-quantization overhead and key-value cache bottlenecks, aiming to reduce model size and memory costs while preserving accuracy.

### Source excerpt

Algorithms & Theory

## Топовые работы на ICLR 2025

DevFeed: [Топовые работы на ICLR 2025](<https://devfeed.tech/articles/iclr-2025-24018.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/redmadrobot/articles/911228/>)

Author: redmadrobot (red\_mad\_robot)

Published: 2025-05-20T17:20:00Z

Content type: article

Language: ru

Sources: [Redmadrobot EN](<https://devfeed.tech/sources/redmadrobot-en.md>), [Redmadrobot RU](<https://devfeed.tech/sources/redmadrobot-ru.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Machine Intelligence](<https://devfeed.tech/topics/machine-intelligence.md>), [Learning](<https://devfeed.tech/topics/learning.md>), [Jailbreak](<https://devfeed.tech/topics/jailbreak.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [alignment](<https://devfeed.tech/tags/alignment.md>), [china](<https://devfeed.tech/tags/china.md>), [conference](<https://devfeed.tech/tags/conference.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [llm](<https://devfeed.tech/tags/llm.md>), [ml](<https://devfeed.tech/tags/ml.md>), [safety](<https://devfeed.tech/tags/safety.md>), [science](<https://devfeed.tech/tags/science.md>), [tag-fa773991fb10](<https://devfeed.tech/tags/tag-fa773991fb10.md>), [university](<https://devfeed.tech/tags/university.md>)

### AI overview

This Russian-language article reviews highly rated papers from ICLR 2025 on artificial intelligence and machine learning. It discusses deepening safety alignment to improve resistance to jailbreak attacks, findings about SFT and DPO fine-tuning dynamics, and AlphaEdit, a method for targeted knowledge editing in large language models.

### Source excerpt

Аналитический центр red_mad_robot продолжает обозревать топовые технологические конференции. В этот раз подготовили для вас инсайты с прошедшей в Сингапуре International Conference on Learning Representations (ICLR), посвящённой искусственному интеллекту и машинному обучению. На ICLR 2025 из более 3 тыс. работ наивысшие оценки получили 36 статей, из которых три были отмечены как "outstanding papers". Разберём выдающиеся, достойные упоминания и получившие высокие оценки работы этого года. Читать далее

## Introducing HELMET: Holistically Evaluating Long-context Language Models

DevFeed: [Introducing HELMET: Holistically Evaluating Long-context Language Models](<https://devfeed.tech/articles/introducing-helmet-holistically-evaluating-long-context-language-models-7237.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/helmet>)

Author: Howard Yen; Tianyu Gao; Minmin Hou; Ke Ding; Daniel Fleischer; Moshe Wasserblat; Danqi Chen

Published: 2025-04-16T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [long-context](<https://devfeed.tech/topics/long-context.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [intel](<https://devfeed.tech/tags/intel.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

HELMET is a comprehensive benchmark for evaluating long-context language models. The article presents its construction, findings from evaluating 59 models, and a quickstart guide for practitioners using it with HuggingFace.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Talk like a graph: Encoding graphs for large language models

DevFeed: [Talk like a graph: Encoding graphs for large language models](<https://devfeed.tech/articles/talk-like-a-graph-encoding-graphs-for-large-language-models-28565.md>)

Original publisher: [Read original article](<http://blog.research.google/2024/03/talk-like-graph-encoding-graphs-for.html>)

Author: Google AI (noreply@blogger.com)

Published: 2024-03-12T21:15:00Z

Content type: article

Language: en

Sources: [Google Research](<https://devfeed.tech/sources/google-research.md>)

Topics: [Graphs](<https://devfeed.tech/topics/graphs.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Google](<https://devfeed.tech/topics/google.md>), [Computer science](<https://devfeed.tech/topics/computer-science.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [graphs](<https://devfeed.tech/tags/graphs.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

Google Research scientists describe a study of how to encode graphs as text so large language models can reason about graph information. The work introduces the GraphQA benchmark and evaluates how encoding methods, task types, and graph structure affect performance.

### Source excerpt

Posted by Bahare Fatemi and Bryan Perozzi, Research Scientists, Google Research Imagine all the things around you -- your friends, tools in your kitchen, or even the parts of your bike. They are all connected in different ways. In computer science, the term graph is used to describe connections between objects. Graphs consist of nodes (the objects themselves) and edges (connections between two nodes, indicating a relationship between them). Graphs are everywhere now. The internet itself is a giant graph of websites linked together. Even the knowledge search engines use is organized in a graph-like way. Furthermore, consider the remarkable advancements in artificial intelligence -- such as chatbots that can write stories in seconds, and even software that can interpret medical reports. This exciting progress is largely thanks to large language models (LLMs). New LLM technology is constantly being developed for different uses. Since graphs are everywhere and LLM technology is on the rise, in "Talk like a Graph: Encoding Graphs for Large Language Models", presented at ICLR 2024, we present a way to teach powerful LLMs how to better reason with graph information. Graphs are a useful way to organize information, but LLMs are mostly trained on regular text. The objective is to test different techniques to see what works best and gain practical insights. Translating graphs into text that LLMs can understand is a remarkably complex task. The difficulty stems from the inherent complexity of graph structures with multiple nodes and the intricate web of edges that connect them. Our work studies how to take a graph and translate it into a format that an LLM can understand. We also design a benchmark called GraphQA to study different approaches on different graph reasoning problems and show how to phrase a graph-related problem in a way that enables the LLM to solve the graph problem. We show that LLM performance on graph reasoning tasks varies on three fundamental levels: 1) the

## Stanford AI Lab Papers and Talks at ICLR 2022

DevFeed: [Stanford AI Lab Papers and Talks at ICLR 2022](<https://devfeed.tech/articles/stanford-ai-lab-papers-and-talks-at-iclr-2022-7584.md>)

Original publisher: [Read original article](<https://ai.stanford.edu/blog/iclr-2022/>)

Author: Compiled by Drew A. Hudson

Published: 2022-04-25T07:00:00Z

Content type: article

Language: en

Sources: [The Stanford AI Lab Blog](<https://devfeed.tech/sources/the-stanford-ai-lab-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [Robotic manipulation](<https://devfeed.tech/topics/robotic-manipulation.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [iclr](<https://devfeed.tech/tags/iclr.md>), [knowledge-graph](<https://devfeed.tech/tags/knowledge-graph.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [neural-networks](<https://devfeed.tech/tags/neural-networks.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [robotic-manipulation](<https://devfeed.tech/tags/robotic-manipulation.md>), [robotics](<https://devfeed.tech/tags/robotics.md>)

### AI overview

Stanford AI Lab presents its work at ICLR 2022, including papers and talks on reinforcement learning, benchmark datasets, distribution shifts, in-context learning, language models, graph reasoning, model editing, robotics, and robotic manipulation.

### Source excerpt

The International Conference on Learning Representations (ICLR) 2022 is being hosted virtually from April 25th - April 29th. We're excited to share all the work from SAIL that's being presented, and you'll find links to papers, videos and blogs below. Feel free to reach out to the contact authors directly to learn more about the work that's happening at Stanford! List of Accepted Papers Autonomous Reinforcement Learning: Formalism and Benchmarking Authors: Archit Sharma*, Kelvin Xu*, Nikhil Sardana, Abhishek Gupta, Karol Hausman, Sergey Levine, Chelsea Finn Contact: architsh@stanford.edu Links: Paper | Website Keywords: reinforcement learning, continual learning, reset-free reinforcement learning MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training Conflicts Authors: Weixin Liang, James Zou Contact: wxliang@stanford.edu Links: Paper | Video | Website Keywords: benchmark dataset, distribution shift, out-of-domain generalization An Explanation of In-context Learning as Implicit Bayesian Inference Authors: Sang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu Ma Contact: xie@cs.stanford.edu Links: Paper | Video Keywords: gpt-3, in-context learning, pretraining, few-shot learning GreaseLM: Graph REASoning Enhanced Language Models for Question Answering Authors: Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D. Manning, Jure Leskovec Contact: xikunz2@cs.stanford.edu Award nominations: Spotlight Links: Paper | Website Keywords: knowledge graph, question answering, language model, commonsense reasoning, graph neural networks, biomedical qa Fast Model Editing at Scale Authors: Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, Christopher D. Manning Contact: eric.mitchell@cs.stanford.edu Links: Paper | Website Keywords: model editing; meta-learning; language models; continual learning; temporal generalization Vision-Based Manipulators Need to Also See from Their Hands Authors: Kyle H

## Selective Classification Can Magnify Disparities Across Groups

DevFeed: [Selective Classification Can Magnify Disparities Across Groups](<https://devfeed.tech/articles/selective-classification-can-magnify-disparities-across-groups-7589.md>)

Original publisher: [Read original article](<https://ai.stanford.edu/blog/sc-magnifies-disparities/>)

Author: A Href; Erik Jones

Published: 2021-10-13T07:00:00Z

Content type: article

Language: en

Sources: [The Stanford AI Lab Blog](<https://devfeed.tech/sources/the-stanford-ai-lab-blog.md>)

Topics: [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Natural language processing](<https://devfeed.tech/topics/nlp.md>)

Tags: [iclr](<https://devfeed.tech/tags/iclr.md>), [ml](<https://devfeed.tech/tags/ml.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [overview](<https://devfeed.tech/tags/overview.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

The article examines selective classification, in which machine-learning models can abstain when uncertain. It shows that although abstention can improve average accuracy, it may fail to help or may worsen accuracy for clinically important subgroups, such as patients with pleural effusion who have not yet received a chest drain. The article discusses the underlying failure mode, theoretical implications, and approaches for building more equitable selective classifiers.

### Source excerpt

Selective classification, where models are allowed to "abstain" when they are uncertain about a prediction, is a useful approach for deploying models in settings where errors are costly. For example, in medicine, model errors can have life-or-death ramifications, but abstentions can be easily handled by backing off to a doctor, who then makes a diagnosis. Across a range of applications from vision 123 and NLP 45, even simple selective classifiers, relying only on model logits, routinely and often dramatically improve accuracy by abstaining. This makes selective classification a compelling tool for ML practitioners 67. However, in our recent ICLR paper, we find that despite reliably improving average accuracy, selective classification can fail to improve and even hurt the accuracy over certain subpopulations of the data. As a motivating example, consider the task of diagnosing pleural effusion, or fluid in the lungs, from chest X-rays. Pleural effusion is often treated with a chest drain, so many pleural effusion cases also have chest drains, while most cases without pleural effusion do not have chest drains 8. While selective classification improves average accuracy for this task, we find that it does not appreciably improve accuracy on the most clinically relevant subgroup, or subpopulation, of the data: those that have pleural effusion but don't yet have a chest drain, i.e. those that have pleural effusion but have not yet been treated for it. Practitioners, thus, should be wary of these potential failure modes of using selective classification in the wild. Example of the spurious correlation setup. This patient has a pleural effusion (excess fluid in the lung), but does not yet have a chest drain. The model, relying on the presence of a chest drain to make a prediction, incorrectly predicts negative. To further outline this critical failure mode of selective classification, we'll first provide an overview of selective classification. We then demonstrate empirical