# AI research agents

AI systems that autonomously conduct iterative machine-learning research and optimization tasks.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Why don't machine learning research agents overfit?

DevFeed: [Why don't machine learning research agents overfit?](<https://devfeed.tech/articles/why-don-t-machine-learning-research-agents-overfit-7610.md>)

Original publisher: [Read original article](<https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit>)

Author: Martin Bertran Lopez; Aaron Roth

Published: 2026-09-10T15:03:39Z

Content type: article

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [machine learning overfitting](<https://devfeed.tech/topics/machine-learning-overfitting.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Occam's razor machine learning](<https://devfeed.tech/topics/occam-s-razor-machine-learning.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai-research-agents](<https://devfeed.tech/tags/ai-research-agents.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmark-overfitting-machine-learning](<https://devfeed.tech/tags/benchmark-overfitting-machine-learning.md>), [compressibility-and-memorization](<https://devfeed.tech/tags/compressibility-and-memorization.md>), [compression-and-generalization](<https://devfeed.tech/tags/compression-and-generalization.md>), [generalization-in-machine-learning](<https://devfeed.tech/tags/generalization-in-machine-learning.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [information-bottleneck-overfitting](<https://devfeed.tech/tags/information-bottleneck-overfitting.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [llm-compression-theory](<https://devfeed.tech/tags/llm-compression-theory.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [machine-learning-overfitting](<https://devfeed.tech/tags/machine-learning-overfitting.md>), [machine-learning-research](<https://devfeed.tech/tags/machine-learning-research.md>), [occam-s-razor-machine-learning](<https://devfeed.tech/tags/occam-s-razor-machine-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [validation](<https://devfeed.tech/tags/validation.md>), [why-don-t-ml-models-overfit-on-benchmarks](<https://devfeed.tech/tags/why-don-t-ml-models-overfit-on-benchmarks.md>)

### AI overview

The article explains why repeated evaluation on held-out benchmarks can cause overfitting, then frames the apparent contradiction in machine learning research, where benchmark-driven iteration is widespread. It also summarizes research suggesting that compressible models limit memorization.

### Source excerpt

New research indicates that AI agents learn compressible models of data, which don't have enough space to enable memorization.

## How GPT-5.6 Sol helps run quantum computing experiments

DevFeed: [How GPT-5.6 Sol helps run quantum computing experiments](<https://devfeed.tech/articles/how-gpt-5-6-sol-helps-run-quantum-computing-experiments-6352.md>)

Original publisher: [Read original article](<https://openai.com/index/codex-quantum-computing-experiments>)

Published: 2026-09-08T17:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [applied-ai](<https://devfeed.tech/tags/applied-ai.md>), [codex](<https://devfeed.tech/tags/codex.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [quantum](<https://devfeed.tech/tags/quantum.md>), [quantum-computing](<https://devfeed.tech/tags/quantum-computing.md>)

### AI overview

An MIT researcher connected GPT-5.6 Sol through Codex to laboratory software for superconducting-qubit experiments. The system ran routine measurements, analyzed results, and selected next steps, reducing the need for constant supervision during calibration workflows.

### Source excerpt

See how an MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits.

## Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science

DevFeed: [Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science](<https://devfeed.tech/articles/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science-6934.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/>)

Author: Michelle Horton

Published: 2026-08-31T16:30:00Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [OpenSSH](<https://devfeed.tech/topics/openssh.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [bionemo](<https://devfeed.tech/tags/bionemo.md>), [claude](<https://devfeed.tech/tags/claude.md>), [code](<https://devfeed.tech/tags/code.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [drug-discovery](<https://devfeed.tech/tags/drug-discovery.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [nim](<https://devfeed.tech/tags/nim.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [science](<https://devfeed.tech/tags/science.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

A tutorial for running NVIDIA BioNeMo NIM microservices with Claude Science to perform protein-structure prediction using multiple-sequence alignment and multiple folding models.

### Source excerpt

Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next....

## Raising machine-checked security benchmarks to advance hash-based SNARKs through agentic collaboration

DevFeed: [Raising machine-checked security benchmarks to advance hash-based SNARKs through agentic collaboration](<https://devfeed.tech/articles/raising-machine-checked-security-benchmarks-to-advance-hash-based-snarks-through-agentic-collaboration-17233.md>)

Original publisher: [Read original article](<https://blog.ethereum.org/en/2026/08/20/better-codes-challenge>)

Author: Ethereum Foundation Formal Verification team

Published: 2026-08-20T00:00:00Z

Content type: article

Language: en

Sources: [Ethereum Foundation Blog](<https://devfeed.tech/sources/ethereum-foundation-blog.md>)

Topics: [Formal verification](<https://devfeed.tech/topics/formal-verification.md>), [Lean](<https://devfeed.tech/topics/lean.md>), [Ethereum](<https://devfeed.tech/topics/ethereum.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Library](<https://devfeed.tech/topics/library.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [ethereum](<https://devfeed.tech/tags/ethereum.md>), [formal-verification](<https://devfeed.tech/tags/formal-verification.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [paper](<https://devfeed.tech/tags/paper.md>), [research](<https://devfeed.tech/tags/research.md>), [research-development](<https://devfeed.tech/tags/research-development.md>), [security](<https://devfeed.tech/tags/security.md>), [verification](<https://devfeed.tech/tags/verification.md>)

### AI overview

The Ethereum Foundation's Formal Verification team launched better.codes, an open autoresearch challenge focused on raising the machine-checked soundness bound of the Lean-formalized koalaIRS12 Reed-Solomon proximity problem toward a fixed 128-bit target. Submissions are checked by the Lean kernel and promoted proofs are shared publicly.

### Source excerpt

better.codes, an open autoresearch challenge built by the Ethereum Foundation Formal Verification team in collaboration with Yukon and zkSecurity, is now live. better.codes takes a self-contained problem from the Proximity Prize research, formalized in Lean, and puts its soundness bound on a public leaderboard that anyone can push forward....

## What We Learned by Reproducing 2,200 papers from ICML

DevFeed: [What We Learned by Reproducing 2,200 papers from ICML](<https://devfeed.tech/articles/what-we-learned-by-reproducing-2-200-papers-from-icml-7271.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/icml-2026-open-reproductions>)

Author: Abubakar Abid

Published: 2026-08-13T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [codex](<https://devfeed.tech/topics/codex.md>), [cursor](<https://devfeed.tech/topics/cursor.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [community](<https://devfeed.tech/tags/community.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [science](<https://devfeed.tech/tags/science.md>)

### AI overview

The article reports lessons from the ICML 2026 Open Reproductions challenge, in which the community used coding agents to reproduce research papers at scale. It discusses how agents can read papers, write code, run experiments, and report findings, while examining the continuing role of human oversight in AI research reproducibility.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Can AI invent new attack techniques? New research from James Kettle and PortSwigger Research

DevFeed: [Can AI invent new attack techniques? New research from James Kettle and PortSwigger Research](<https://devfeed.tech/articles/can-ai-invent-new-attack-techniques-new-research-from-james-kettle-and-portswigger-research-7702.md>)

Original publisher: [Read original article](<https://portswigger.net/blog/can-ai-invent-new-attack-techniques-new-research-from-james-kettle-and-portswigger-research>)

Author: Kieron Hughes

Published: 2026-08-12T09:04:45Z

Content type: article

Language: en

Sources: [PortSwigger Blog](<https://devfeed.tech/sources/portswigger-blog.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Vulnerabilities](<https://devfeed.tech/topics/vulnerabilities.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [attacks](<https://devfeed.tech/tags/attacks.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [bug-bounty](<https://devfeed.tech/tags/bug-bounty.md>), [http](<https://devfeed.tech/tags/http.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [vulnerabilities](<https://devfeed.tech/tags/vulnerabilities.md>)

### AI overview

PortSwigger Research describes the HTTP Terminator, an autonomous system used to generate and test new HTTP desync attack techniques against authorized bug-bounty targets. The research emphasizes that human expertise remained important in steering follow-on discoveries.

### Source excerpt

We already know AI can find vulnerabilities. James Kettle, PortSwigger's Director of Research, wanted to answer a harder question: can an autonomous system invent genuinely new attack techniques? To f

## Can AI do novel security research? Meet the HTTP Terminator

DevFeed: [Can AI do novel security research? Meet the HTTP Terminator](<https://devfeed.tech/articles/can-ai-do-novel-security-research-meet-the-http-terminator-7670.md>)

Original publisher: [Read original article](<https://portswigger.net/research/can-ai-do-novel-security-research>)

Author: James Kettle

Published: 2026-08-05T19:30:00Z

Content type: article

Language: en

Sources: [PortSwigger Research](<https://devfeed.tech/sources/portswigger-research.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [web applications](<https://devfeed.tech/topics/web-applications.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [exploits](<https://devfeed.tech/tags/exploits.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [http](<https://devfeed.tech/tags/http.md>), [research](<https://devfeed.tech/tags/research.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

The article examines whether autonomous AI systems can discover novel web-security attack techniques. It presents the HTTP Terminator, describes HTTP desynchronization research and exploits, and explores the limits of fully autonomous versus human/AI research loops.

### Source excerpt

Abstract We all know AI can find bugs. After a decade of research, I asked a harder question: can an autonomous system invent new attack techniques, and use them to hack live websites at scale? Buildi

## SceneSmith uses collaborative AI agents to create 3D environments for robot training

DevFeed: [SceneSmith uses collaborative AI agents to create 3D environments for robot training](<https://devfeed.tech/articles/ai-agents-create-virtual-playgrounds-to-help-robots-get-crucial-training-data-37940.md>)

Original publisher: [Read original article](<https://news.mit.edu/2026/ai-agents-create-virtual-playgrounds-to-help-robots-get-crucial-training-data-0713>)

Author: Alex Shipps | MIT CSAIL

Published: 2026-07-13T18:50:00Z

Content type: news

Language: en

Sources: [MIT AI News](<https://devfeed.tech/sources/mit-ai-news.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [robot grasping simulation](<https://devfeed.tech/topics/robot-grasping-simulation.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Computer Science and Artificial Intelligence Laboratory (CSAIL)](<https://devfeed.tech/topics/computer-science-and-artificial-intelligence-laboratory-csail.md>)

Tags: [3-d](<https://devfeed.tech/tags/3-d.md>), [3d](<https://devfeed.tech/tags/3d.md>), [adversarial-machine-learning](<https://devfeed.tech/tags/adversarial-machine-learning.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [computer-science-and-artificial-intelligence-laboratory-csail](<https://devfeed.tech/tags/computer-science-and-artificial-intelligence-laboratory-csail.md>), [computer-science-and-technology](<https://devfeed.tech/tags/computer-science-and-technology.md>), [electrical-engineering-and-computer-science-eecs](<https://devfeed.tech/tags/electrical-engineering-and-computer-science-eecs.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [general-purpose-robotics](<https://devfeed.tech/tags/general-purpose-robotics.md>), [gpt-5-2](<https://devfeed.tech/tags/gpt-5-2.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [mit-csail](<https://devfeed.tech/tags/mit-csail.md>), [mit-eecs](<https://devfeed.tech/tags/mit-eecs.md>), [mit-schwarzman-college-of-computing](<https://devfeed.tech/tags/mit-schwarzman-college-of-computing.md>), [national-science-foundation-nsf](<https://devfeed.tech/tags/national-science-foundation-nsf.md>), [nicholas-pfaff](<https://devfeed.tech/tags/nicholas-pfaff.md>), [research](<https://devfeed.tech/tags/research.md>), [robot-simulations](<https://devfeed.tech/tags/robot-simulations.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [robots](<https://devfeed.tech/tags/robots.md>), [russ-tedrake](<https://devfeed.tech/tags/russ-tedrake.md>), [scene-generation](<https://devfeed.tech/tags/scene-generation.md>), [scenesmith](<https://devfeed.tech/tags/scenesmith.md>), [school-of-engineering](<https://devfeed.tech/tags/school-of-engineering.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [simulation-ready-indoor-scenes](<https://devfeed.tech/tags/simulation-ready-indoor-scenes.md>), [virtual-playgrounds](<https://devfeed.tech/tags/virtual-playgrounds.md>), [vision-language-models-vlms](<https://devfeed.tech/tags/vision-language-models-vlms.md>), [zero-shot-policy](<https://devfeed.tech/tags/zero-shot-policy.md>)

### AI overview

MIT CSAIL and Toyota Research Institute researchers developed SceneSmith, a system that uses three collaborative AI agents to create realistic 3D environments for robot training. The scenes can be loaded into physics simulation software, allowing robots to practice tasks before real-world testing.

### Source excerpt

"SceneSmith" system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores.

## The triage is the product: running AI agents against Ethereum's protocol code

DevFeed: [The triage is the product: running AI agents against Ethereum's protocol code](<https://devfeed.tech/articles/the-triage-is-the-product-running-ai-agents-against-ethereum-s-protocol-code-17227.md>)

Original publisher: [Read original article](<https://blog.ethereum.org/en/2026/07/09/triage-is-the-product>)

Author: Nikos Baxevanis

Published: 2026-07-09T00:00:00Z

Content type: article

Language: en

Sources: [Ethereum Foundation Blog](<https://devfeed.tech/sources/ethereum-foundation-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Ethereum](<https://devfeed.tech/topics/ethereum.md>), [Security](<https://devfeed.tech/topics/security.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Code](<https://devfeed.tech/topics/code.md>), [P2P](<https://devfeed.tech/topics/p2p.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [cve](<https://devfeed.tech/tags/cve.md>), [ethereum](<https://devfeed.tech/tags/ethereum.md>), [protocol](<https://devfeed.tech/tags/protocol.md>), [research](<https://devfeed.tech/tags/research.md>), [research-development](<https://devfeed.tech/tags/research-development.md>), [security](<https://devfeed.tech/tags/security.md>), [security-research](<https://devfeed.tech/tags/security-research.md>)

### AI overview

The Ethereum Foundation's Protocol Security team describes how it uses coordinated AI agents to audit Ethereum protocol code. The article focuses on organizing agent work, validating candidate bugs, and reducing false positives. It reports a remotely triggerable panic in libp2p's gossipsub, disclosed as CVE-2026-34219.

### Source excerpt

Notes from the Ethereum Foundation's Protocol Security team on running coordinated AI agents against real protocol code, including how we organize the work, what holds up under scrutiny, and what client teams and security researchers can take from it. This post stands on its own; later posts will go deeper...

## Building an LLM Wiki as Persistent Memory for AI Agents

DevFeed: [Building an LLM Wiki as Persistent Memory for AI Agents](<https://devfeed.tech/articles/your-second-brain-is-a-graveyard-make-it-agent-memory-18300.md>)

Original publisher: [Read original article](<https://www.decodingai.com/p/llm-wiki-agent-memory>)

Author: Paul Iusztin

Published: 2026-07-07T05:01:21Z

Content type: tutorial

Language: en

Sources: [Decoding ML](<https://devfeed.tech/sources/decoding-ml.md>)

Topics: [Wiki](<https://devfeed.tech/topics/wiki.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [article](<https://devfeed.tech/tags/article.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [context](<https://devfeed.tech/tags/context.md>), [google](<https://devfeed.tech/tags/google.md>), [llm](<https://devfeed.tech/tags/llm.md>), [memory](<https://devfeed.tech/tags/memory.md>), [notes](<https://devfeed.tech/tags/notes.md>)

### AI overview

This tutorial describes building an AI Research OS that turns notes and web research into a queryable, maintainable LLM wiki. The wiki is designed to provide persistent memory for AI agents used in research, coding, and content creation.

### Source excerpt

Turn dead notes into a living LLM wiki your AI agents can query, maintain, and grow

## Introducing GeneBench-Pro

DevFeed: [Introducing GeneBench-Pro](<https://devfeed.tech/articles/introducing-genebench-pro-6487.md>)

Original publisher: [Read original article](<https://openai.com/index/introducing-genebench-pro>)

Published: 2026-06-30T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [complexity](<https://devfeed.tech/tags/complexity.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [diagnostics](<https://devfeed.tech/tags/diagnostics.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [explore](<https://devfeed.tech/tags/explore.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [research](<https://devfeed.tech/tags/research.md>), [skills](<https://devfeed.tech/tags/skills.md>), [support](<https://devfeed.tech/tags/support.md>), [testing](<https://devfeed.tech/tags/testing.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

GeneBench-Pro is a research-level benchmark for evaluating whether AI agents can make judgment-heavy decisions during computational-biology analysis. It uses messy datasets and downstream decision targets to assess ambiguity handling, analytical-path selection, iterative experimentation, and revision of initial plans.

### Source excerpt

Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.

## Gemini for Science: AI experiments and tools for a new era of discovery

DevFeed: [Gemini for Science: AI experiments and tools for a new era of discovery](<https://devfeed.tech/articles/gemini-for-science-ai-experiments-and-tools-for-a-new-era-of-discovery-6167.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/gemini-for-science-ai-experiments-and-tools-for-a-new-era-of-discovery/>)

Author: Pushmeet Kohli

Published: 2026-05-17T13:50:34Z

Content type: article

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [code](<https://devfeed.tech/tags/code.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [none](<https://devfeed.tech/tags/none.md>), [research](<https://devfeed.tech/tags/research.md>), [science](<https://devfeed.tech/tags/science.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

Google introduces Gemini for Science, a collection of experimental tools for scientific exploration. Its prototypes support hypothesis generation, computational discovery through parallel code variations, and literature insights.

### Source excerpt

A collection of science tools and experiments to expand the scale and precision of scientific exploration.

## Co-Scientist: A multi-agent AI partner to accelerate research

DevFeed: [Co-Scientist: A multi-agent AI partner to accelerate research](<https://devfeed.tech/articles/co-scientist-a-multi-agent-ai-partner-to-accelerate-research-6143.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/>)

Author: Co-Scientist team

Published: 2026-05-12T14:40:07Z

Content type: article

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>)

Tags: [accelerate](<https://devfeed.tech/tags/accelerate.md>), [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [research](<https://devfeed.tech/tags/research.md>), [science](<https://devfeed.tech/tags/science.md>)

### AI overview

Introduces Co-Scientist, a collaborative multi-agent AI partner built with Gemini to help researchers accelerate scientific breakthroughs.

### Source excerpt

Introducing Co-Scientist, a collaborative AI partner built with Gemini to help researchers accelerate scientific breakthroughs.

## How NVIDIA engineers and researchers build with Codex

DevFeed: [How NVIDIA engineers and researchers build with Codex](<https://devfeed.tech/articles/how-nvidia-engineers-and-researchers-build-with-codex-6556.md>)

Original publisher: [Read original article](<https://openai.com/index/nvidia>)

Published: 2026-05-12T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [production](<https://devfeed.tech/tags/production.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

NVIDIA engineers use Codex with GPT-5.5 for complex engineering work and end-to-end machine learning experiments. The article describes autonomous coding, testing, and research workflows, including building production systems and running experiments remotely.

### Source excerpt

Teams use Codex with GPT-5.5 to ship production systems and turn research ideas into runnable experiments.

## Improving the academic workflow: Introducing two AI agents for better figures and peer review

DevFeed: [Improving the academic workflow: Introducing two AI agents for better figures and peer review](<https://devfeed.tech/articles/improving-the-academic-workflow-introducing-two-ai-agents-for-better-figures-and-peer-review-6820.md>)

Original publisher: [Read original article](<https://research.google/blog/improving-the-academic-workflow-introducing-two-ai-agents-for-better-figures-and-peer-review/>)

Published: 2026-04-08T20:01:33Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

Google introduces PaperVizAgent and ScholarPeer, two AI agents designed to support academic research. PaperVizAgent generates publication-ready figures from academic text, while ScholarPeer evaluates papers and inlined diagrams through literature-grounded peer review.

### Source excerpt

Generative AI

## Cockroach Labs announces CockroachDB capabilities for AI agents

DevFeed: [Cockroach Labs announces CockroachDB capabilities for AI agents](<https://devfeed.tech/articles/cockroachdb-is-built-for-ai-agents-23759.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/cockroachdb-ai-agents-agent-ready-database>)

Author: Lakshmi Kannan

Published: 2026-03-25T00:00:00Z

Content type: release

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [CockroachDB](<https://devfeed.tech/topics/cockroachdb.md>), [Cockroach Labs](<https://devfeed.tech/topics/cockroach-labs.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [Agent Skills](<https://devfeed.tech/topics/agent-skills.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [agent-skills](<https://devfeed.tech/tags/agent-skills.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [auditability](<https://devfeed.tech/tags/auditability.md>), [backend](<https://devfeed.tech/tags/backend.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [cockroach-labs](<https://devfeed.tech/tags/cockroach-labs.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [database](<https://devfeed.tech/tags/database.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [operations](<https://devfeed.tech/tags/operations.md>), [provisioning](<https://devfeed.tech/tags/provisioning.md>), [scale](<https://devfeed.tech/tags/scale.md>), [upgrades](<https://devfeed.tech/tags/upgrades.md>)

### AI overview

Cockroach Labs announces capabilities intended to make CockroachDB agent-ready, including a fully managed MCP Server, a redesigned ccloud CLI, and a public library of CockroachDB Agent Skills. The article describes requirements such as structured interfaces, scoped permissions, and auditability for agents operating across the database lifecycle.

### Source excerpt

Today Cockroach Labs is announcing new capabilities that make CockroachDB agent-ready, giving AI agents a secure, structured way to work with your database.

## Karpathy's Autonomous ML Lab, Sleeper Cells in LLMs, and Andrew Ng's Context Hub - 📚 The Tokenizer Edition #19

DevFeed: [Karpathy's Autonomous ML Lab, Sleeper Cells in LLMs, and Andrew Ng's Context Hub - 📚 The Tokenizer Edition #19](<https://devfeed.tech/articles/karpathy-s-autonomous-ml-lab-sleeper-cells-in-llms-and-andrew-ng-s-context-hub-the-tokenizer-edition-19-18340.md>)

Original publisher: [Read original article](<https://newsletter.artofsaience.com/p/karpathys-autonomous-ml-lab-sleeper>)

Author: Sairam Sundaresan

Published: 2026-03-11T12:03:21Z

Content type: article

Language: en

Sources: [Gradient Ascent](<https://devfeed.tech/sources/gradient-ascent.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Security](<https://devfeed.tech/topics/security.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>), [long-context](<https://devfeed.tech/topics/long-context.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

The Tokenizer newsletter curates AI and machine learning resources covering autonomous ML experiments, reasoning graphs, temporal backdoors in tool-using LLMs, spatial reasoning benchmarks, long-context inference costs, open-weight models, and developer tools.

### Source excerpt

This week's most valuable AI resources

## Building a Deep Research Agent with Neon and Durable Endpoints

DevFeed: [Building a Deep Research Agent with Neon and Durable Endpoints](<https://devfeed.tech/articles/building-a-deep-research-agent-with-neon-and-durable-endpoints-5084.md>)

Original publisher: [Read original article](<https://neon.com/blog/building-a-deep-research-agent-with-neon-and-durable-endpoints>)

Author: Charly Poly

Published: 2026-02-24T17:14:34Z

Content type: tutorial

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [AI-generated research reports](<https://devfeed.tech/topics/ai-generated-research-reports.md>), [web applications](<https://devfeed.tech/topics/web-applications.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [building](<https://devfeed.tech/tags/building.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [community](<https://devfeed.tech/tags/community.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [next-js](<https://devfeed.tech/tags/next-js.md>), [postgres](<https://devfeed.tech/tags/postgres.md>), [product](<https://devfeed.tech/tags/product.md>), [research](<https://devfeed.tech/tags/research.md>), [serverless](<https://devfeed.tech/tags/serverless.md>)

### AI overview

A tutorial on building a recursive AI research agent with Neon, Inngest durable endpoints, semantic memory, and a Next.js progress UI.

### Source excerpt

Every AI lab is shipping research agents. OpenAI's Deep Research, Perplexity, and Gemini's research mode. These products are not simple RAG pipelines. Recent papers like DeepResearcher and Step-DeepResearch formalize what makes them work: a recursive loop of planning, searching,...

## Atoms is Out, a Multi-Agent AI Team that Builds Full-Stack Apps for You

DevFeed: [Atoms is Out, a Multi-Agent AI Team that Builds Full-Stack Apps for You](<https://devfeed.tech/articles/atoms-is-out-a-multi-agent-ai-team-that-builds-full-stack-apps-for-you-4989.md>)

Original publisher: [Read original article](<https://neon.com/blog/atoms-is-out-a-multi-agent-ai-team-that-builds-full-stack-apps-for-you>)

Author: Carlota Soto

Published: 2026-01-13T17:48:26Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [App](<https://devfeed.tech/topics/app.md>), [Multitenancy](<https://devfeed.tech/topics/multitenancy.md>), [Per-user Database](<https://devfeed.tech/topics/per-user-database.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [Front end](<https://devfeed.tech/topics/frontend.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>), [Iris](<https://devfeed.tech/topics/iris.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [apps](<https://devfeed.tech/tags/apps.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [automation](<https://devfeed.tech/tags/automation.md>), [backend](<https://devfeed.tech/tags/backend.md>), [case-studies](<https://devfeed.tech/tags/case-studies.md>), [code](<https://devfeed.tech/tags/code.md>), [data](<https://devfeed.tech/tags/data.md>), [database](<https://devfeed.tech/tags/database.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [product](<https://devfeed.tech/tags/product.md>), [product-development](<https://devfeed.tech/tags/product-development.md>), [saas](<https://devfeed.tech/tags/saas.md>), [scale](<https://devfeed.tech/tags/scale.md>), [software](<https://devfeed.tech/tags/software.md>)

### AI overview

Atoms is a multi-agent AI platform that builds full-stack software products from a user prompt. Specialized agents handle research, product requirements, architecture, engineering, data analysis, testing, and deployment. Its multi-tenant architecture gives each user app a dedicated Postgres database provisioned dynamically through Neon.

### Source excerpt

"We chose Neon as our backend because of scale, cost-efficiency, and superior developer and user experience, all critical for Atoms' multi-tenant architecture. Each user app needs its own dedicated database, and Neon lets us spin up Postgres instances dynamically while only payin...