# icml

Published articles for icml.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train

DevFeed: [Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train](<https://devfeed.tech/articles/bypassing-inference-bottlenecks-accelerating-complex-ai-search-with-retrieve-for-train-26972.md>)

Original publisher: [Read original article](<https://research.google/blog/bypassing-inference-bottlenecks-accelerating-complex-ai-search-with-retrieve-for-train/>)

Published: 2026-09-15T20:00:35Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [AI search](<https://devfeed.tech/topics/ai-search.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Algorithms & Theory](<https://devfeed.tech/topics/algorithms-theory.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [data-mining-modeling](<https://devfeed.tech/tags/data-mining-modeling.md>), [diffusion](<https://devfeed.tech/tags/diffusion.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [icml](<https://devfeed.tech/tags/icml.md>), [icml-2026](<https://devfeed.tech/tags/icml-2026.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [rl](<https://devfeed.tech/tags/rl.md>), [search](<https://devfeed.tech/tags/search.md>)

### AI overview

Google Research presents Retrieve-for-Train, a framework that uses offline reinforcement learning to compile reward-aligned query fan-outs into training data for a lightweight diffusion retriever. The approach is intended to produce diverse, complementary, and coherent search-result sets in a single inference pass, reducing reliance on expensive inference-time reasoning.

### Source excerpt

Algorithms & Theory

## What I Saw at ICML 2026

DevFeed: [What I Saw at ICML 2026](<https://devfeed.tech/articles/arxiv-icml-2026-24886.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1073774/>)

Author: zj-karina (Яндекс)

Published: 2026-08-25T07:01:28Z

Content type: article

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [icml](<https://devfeed.tech/tags/icml.md>), [icml-2026](<https://devfeed.tech/tags/icml-2026.md>), [llm-agents](<https://devfeed.tech/tags/llm-agents.md>), [ml](<https://devfeed.tech/tags/ml.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [tag-316edb31b5b3](<https://devfeed.tech/tags/tag-316edb31b5b3.md>), [tag-44d9110ce940](<https://devfeed.tech/tags/tag-44d9110ce940.md>)

### AI overview

A Yandex developer reports from ICML 2026 in Seoul, describing the conference format, its scale, Yandex research presented there, and discussions about AI agents.

### Source excerpt

Зачем тратить сутки на перелёты, мчаться на другой конец света и жить неделю в режиме нон-стоп на одной из главных ML-конференций планеты, когда пейпер уже на arXiv, код -- на GitHub, а краткие выжимки из выступлений -- мгновенно в соцсетях? Меня зовут Карина Романова, я разработчик в Яндексе и занимаюсь LLM-агентами в Алисе. В июле мы с командой прилетели в Сеул на ICML 2026, и я ответила себе на вопрос "зачем?". Для нас офлайн-конференции -- это единственный способ за несколько дней прочувствовать реальный фокус сообщества, встретиться с авторами работ и узнать детали, которых нет в опубликованных текстах. В этой статье расскажу, как устроена ICML изнутри, чем запомнилась программа этого года, какие наши исследования вызвали наибольший ажиотаж и почему заметная часть разговоров на конференции снова вращалась вокруг AI-агентов. Читать далее

## Как проект из ШАДа попал в Spotlight статей на конференции ICML 2026

DevFeed: [Как проект из ШАДа попал в Spotlight статей на конференции ICML 2026](<https://devfeed.tech/articles/spotlight-icml-2026-24866.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1055232/>)

Author: mightyneighbor (Яндекс)

Published: 2026-07-06T07:04:35Z

Content type: article

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [ai](<https://devfeed.tech/tags/ai.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [graph](<https://devfeed.tech/tags/graph.md>), [icml](<https://devfeed.tech/tags/icml.md>), [icml-2026](<https://devfeed.tech/tags/icml-2026.md>), [ml](<https://devfeed.tech/tags/ml.md>), [research](<https://devfeed.tech/tags/research.md>), [spotlight](<https://devfeed.tech/tags/spotlight.md>), [tag-1e7f0701b819](<https://devfeed.tech/tags/tag-1e7f0701b819.md>), [tag-4004cf5948d3](<https://devfeed.tech/tags/tag-4004cf5948d3.md>)

### AI overview

The article explains why graph neural networks underutilize modern GPUs: their irregular, sparse memory access patterns leave the hardware waiting for data. It describes a project that investigated the problem and produced three families of specialized GPU kernels. The resulting paper, "On Efficient Scaling of GNNs via IO-Aware Layer Implementations," was accepted to ICML 2026 as a Spotlight paper.

### Source excerpt

Граф из миллионов вершин не загружает современную GPU на все 100%: видеокарта почти всё время не вычисляет, а ждёт загрузки данных из памяти. Графовые нейросети, или GNN, упираются в это давно: сами операции достаточно простые, но доступ к памяти нерегулярный и разреженный. И чем мощнее GPU, тем заметнее недостаточная её утилизация. Идея выросла из проектного курса в ШАДе. Толчком стало то, что один из самых популярных фреймворков для работы с графами, Deep Graph Library, на момент начала работы не обновлялся уже около года -- это знак того, что в области что-то застряло. Меня зовут Федя Великонивцев, я старший исследователь Yandex Research, руковожу группой, которая занимается эффективными вычислениями на GPU. На том курсе мы с коллегами -- Дарьей Фоминой из команды ML-инфраструктуры Яндекса, Вячеславом Ждановским из команды разработки инференса -- и студентами Даниилом Красильниковым, Алексеем Бойковым и Андреем Долговязовым взялись выяснить, почему графовые нейросети тормозят на современных GPU. Так появился проект, который мы оформили в отдельную статью -- On Efficient Scaling of GNNs via IO-Aware Layer Implementations. Её приняли на ICML-2026 со статусом Spotlight. Для контекста: из 23 918 поданных работ приняли 6 352 (26,6%), а Spotlight достался только 536 работам -- это 2,2% заявок с самыми высокими оценками программного комитета. Дальше расскажу, как мы прошли путь от этого вопроса до трёх семейств специализированных GPU-кернелов -- с парой неожиданных находок по дороге. Читать далее

## LLM reasoning and agentic safety at ICML 2026

DevFeed: [LLM reasoning and agentic safety at ICML 2026](<https://devfeed.tech/articles/llm-reasoning-and-agentic-safety-at-icml-2026-22575.md>)

Original publisher: [Read original article](<https://medium.com/capital-one-tech/llm-reasoning-and-agentic-safety-at-icml-2026-55f341e21caa?source=rss----3db3a67cb648---4>)

Author: Capital One Tech

Published: 2026-07-02T14:14:40Z

Content type: article

Language: en

Sources: [Capital One Tech](<https://devfeed.tech/sources/capital-one-tech.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Trustworthy AI](<https://devfeed.tech/topics/trustworthy-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [icml](<https://devfeed.tech/tags/icml.md>), [icml-2026](<https://devfeed.tech/tags/icml-2026.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llm-reasoning](<https://devfeed.tech/tags/llm-reasoning.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [safety](<https://devfeed.tech/tags/safety.md>), [science](<https://devfeed.tech/tags/science.md>), [trustworthy-ai](<https://devfeed.tech/tags/trustworthy-ai.md>)

### AI overview

Capital One presents research for ICML 2026 on critique-guided distillation for robust LLM reasoning and on safety risks in multi-turn tool-using agents. The article says its Critique-Guided Distillation framework trains models to refine flawed responses using teacher critiques, and reports higher mathematical-reasoning benchmark performance than critique fine-tuning and standard distillation, including a 7% average improvement and gains of up to 15.0% on AMC23 and 12.2% on MATH-500.

### Source excerpt

Explore our latest research in critique-guided distillation and multi-turn agent uncertainty in Seoul.Explore our latest research in critique-guided distillation and multi-turn agent uncertainty in Seoul. Capital One technologists are excited to participate in the 43rd International Conference on Machine Learning (ICML) taking place at the COEX Convention & Exhibition Center in Seoul, South Korea, July 6-11, 2026. As a premier global venue for machine learning research, ICML provides an essential forum for exploring foundational advancements, algorithmic innovations and cutting-edge deep learning systems. Capital One is excited to share advancements in large language model (LLM) scaling efficiencies, multi-turn tool-using agent safety and the development of robust, trustworthy AI frameworks. This work delivers the underlying engineering and algorithmic improvements crucial for deploying the next generation of safe financial technologies. Main conference research: Robust reasoning and agentic risk The following research, accepted to the ICML Main Conference, pushes the boundaries of how models self-correct, how trajectory-level risks can be proactively flagged, and how multi-turn agent interactions maintain reliable execution. This section features work led by Capital One researchers alongside deep collaborations with academic partners. Critique-Guided Distillation for Robust Reasoning via Refinement Capital One Authors: Berkcan Kapusuzoglu, Supriyo Chakraborty, Michael Lee, Sambit Sahu Supervised fine-tuning with expert demonstrations often produces models that imitate outputs without internalizing the reasoning processes needed for robust generalization. While critique-based approaches show promise, training models to generate critiques directly, such as Critique Fine-Tuning (CFT), can lead to output-format drift and degradation of general capabilities. We propose Critique-Guided Distillation (CGD), a training framework that decouples critique consumption from crit

## Intervening on early readouts for mitigating spurious features and simplicity bias

DevFeed: [Intervening on early readouts for mitigating spurious features and simplicity bias](<https://devfeed.tech/articles/intervening-on-early-readouts-for-mitigating-spurious-features-and-simplicity-bias-28549.md>)

Original publisher: [Read original article](<http://blog.research.google/2024/02/intervening-on-early-readouts-for.html>)

Author: Google AI (noreply@blogger.com)

Published: 2024-02-02T17:49:00Z

Content type: article

Language: en

Sources: [Google Research](<https://devfeed.tech/sources/google-research.md>)

Topics: [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Deep learning](<https://devfeed.tech/topics/deep-learning.md>), [responsible-ai](<https://devfeed.tech/topics/responsible-ai.md>), [generalization in machine learning](<https://devfeed.tech/topics/generalization-in-machine-learning.md>), [Model Development](<https://devfeed.tech/topics/model-development.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>)

Tags: [bias](<https://devfeed.tech/tags/bias.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [icml](<https://devfeed.tech/tags/icml.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [ml-fairness](<https://devfeed.tech/tags/ml-fairness.md>), [research](<https://devfeed.tech/tags/research.md>), [responsible-ai](<https://devfeed.tech/tags/responsible-ai.md>), [supervised-learning](<https://devfeed.tech/tags/supervised-learning.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Google Research describes methods for detecting and reducing spurious features and simplicity bias in deep learning models. Early readouts expose confidently wrong predictions associated with spurious features, while feature forgetting helps models identify more predictive features and generalize to unseen domains.

### Source excerpt

Posted by Rishabh Tiwari, Pre-doctoral Researcher, and Pradeep Shenoy, Research Scientist, Google Research Machine learning models in the real world are often trained on limited data that may contain unintended statistical biases. For example, in the CELEBA celebrity image dataset, a disproportionate number of female celebrities have blond hair, leading to classifiers incorrectly predicting "blond" as the hair color for most female faces -- here, gender is a spurious feature for predicting hair color. Such unfair biases could have significant consequences in critical applications such as medical diagnosis. Surprisingly, recent work has also discovered an inherent tendency of deep networks to amplify such statistical biases, through the so-called simplicity bias of deep learning. This bias is the tendency of deep networks to identify weakly predictive features early in the training, and continue to anchor on these features, failing to identify more complex and potentially more accurate features. With the above in mind, we propose simple and effective fixes to this dual challenge of spurious features and simplicity bias by applying early readouts and feature forgetting. First, in "Using Early Readouts to Mediate Featural Bias in Distillation", we show that making predictions from early layers of a deep network (referred to as "early readouts") can automatically signal issues with the quality of the learned representations. In particular, these predictions are more often wrong, and more confidently wrong, when the network is relying on spurious features. We use this erroneous confidence to improve outcomes in model distillation, a setting where a larger "teacher" model guides the training of a smaller "student" model. Then in "Overcoming Simplicity Bias in Deep Networks using a Feature Sieve", we intervene directly on these indicator signals by making the network "forget" the problematic features and consequently look for better, more predictive features. This substanti

## Exphormer: Scaling transformers for graph-structured data

DevFeed: [Exphormer: Scaling transformers for graph-structured data](<https://devfeed.tech/articles/exphormer-scaling-transformers-for-graph-structured-data-28542.md>)

Original publisher: [Read original article](<http://blog.research.google/2024/01/exphormer-scaling-transformers-for.html>)

Author: Google AI (noreply@blogger.com)

Published: 2024-01-23T22:27:00Z

Content type: article

Language: en

Sources: [Google Research](<https://devfeed.tech/sources/google-research.md>)

Topics: [Graphs](<https://devfeed.tech/topics/graphs.md>), [Transformers](<https://devfeed.tech/topics/transformers.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Natural language processing](<https://devfeed.tech/topics/nlp.md>)

Tags: [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [graph](<https://devfeed.tech/tags/graph.md>), [graphs](<https://devfeed.tech/tags/graphs.md>), [icml](<https://devfeed.tech/tags/icml.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [theory](<https://devfeed.tech/tags/theory.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

This article explains the scalability challenges of graph transformers, whose complete attention graphs create quadratic computational and memory costs. It introduces Exphormer, a sparse attention framework designed for graph data that uses expander graphs from spectral graph theory.

### Source excerpt

Posted by Ameya Velingker, Research Scientist, Google Research, and Balaji Venkatachalam, Software Engineer, Google Graphs, in which objects and their relations are represented as nodes (or vertices) and edges (or links) between pairs of nodes, are ubiquitous in computing and machine learning (ML). For example, social networks, road networks, and molecular structure and interactions are all domains in which underlying datasets have a natural graph structure. ML can be used to learn the properties of nodes, edges, or entire graphs. A common approach to learning on graphs are graph neural networks (GNNs), which operate on graph data by applying an optimizable transformation on node, edge, and global attributes. The most typical class of GNNs operates via a message-passing framework, whereby each layer aggregates the representation of a node with those of its immediate neighbors. Recently, graph transformer models have emerged as a popular alternative to message-passing GNNs. These models build on the success of Transformer architectures in natural language processing (NLP), adapting them to graph-structured data. The attention mechanism in graph transformers can be modeled by an interaction graph, in which edges represent pairs of nodes that attend to each other. Unlike message passing architectures, graph transformers have an interaction graph that is separate from the input graph. The typical interaction graph is a complete graph, which signifies a full attention mechanism that models direct interactions between all pairs of nodes. However, this creates quadratic computational and memory bottlenecks that limit the applicability of graph transformers to datasets on small graphs with at most a few thousand nodes. Making graph transformers scalable has been considered one of the most important research directions in the field (see the first open problem here). A natural remedy is to use a sparse interaction graph with fewer edges. Many sparse and efficient transformers