# NVIDIA Research

Published articles for NVIDIA Research.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents

DevFeed: [NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents](<https://devfeed.tech/articles/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents-6887.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/>)

Author: Tanya Lenz

Published: 2026-08-21T13:00:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [autonomous-agents](<https://devfeed.tech/tags/autonomous-agents.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [coding](<https://devfeed.tech/tags/coding.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [kernel](<https://devfeed.tech/tags/kernel.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-research](<https://devfeed.tech/tags/nvidia-research.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [top-stories](<https://devfeed.tech/tags/top-stories.md>), [trustworthy-ai-cybersecurity](<https://devfeed.tech/tags/trustworthy-ai-cybersecurity.md>)

### AI overview

NVIDIA introduces AVO, a general-purpose coding-agent architecture intended for sustained autonomous work on long, multistep tasks. The article describes its use in GPU-kernel optimization and its adaptation to the ARC-AGI-3 benchmark through different task-specific tools and evaluation.

### Source excerpt

A frontier language model is only one component of an AI agent. The surrounding agent system--often called a harness--determines how the model receives...

## Where Security Fits in an AI Agent Stack

DevFeed: [Where Security Fits in an AI Agent Stack](<https://devfeed.tech/articles/where-security-fits-in-an-ai-agent-stack-6946.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/where-security-fits-in-an-ai-agent-stack/>)

Author: Michelle Horton

Published: 2026-08-21T13:00:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Application Security](<https://devfeed.tech/topics/application-security.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-safety](<https://devfeed.tech/tags/ai-safety.md>), [ai-security](<https://devfeed.tech/tags/ai-security.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia-research](<https://devfeed.tech/tags/nvidia-research.md>), [openshell](<https://devfeed.tech/tags/openshell.md>), [security](<https://devfeed.tech/tags/security.md>), [trustworthy-ai-cybersecurity](<https://devfeed.tech/tags/trustworthy-ai-cybersecurity.md>)

### AI overview

The article explains where security controls fit in an emerging AI agent stack. It emphasizes runtime boundaries, scoped access, authorization, isolation, auditability, and defense in depth rather than relying solely on prompts, model safeguards, or harness logic.

### Source excerpt

As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important....

## Beyond VLAs: How World Action Models Reshape Robot Manipulation

DevFeed: [Beyond VLAs: How World Action Models Reshape Robot Manipulation](<https://devfeed.tech/articles/beyond-vlas-how-world-action-models-reshape-robot-manipulation-6764.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/beyond-vlas-how-world-action-models-reshape-robot-manipulation/>)

Author: Michelle Horton

Published: 2026-08-04T16:00:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Robotics](<https://devfeed.tech/topics/robotics.md>), [World models](<https://devfeed.tech/topics/world-models.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [Cosmos](<https://devfeed.tech/topics/cosmos.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [NVIDIA Research](<https://devfeed.tech/topics/nvidia-research.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [cosmos](<https://devfeed.tech/tags/cosmos.md>), [featured](<https://devfeed.tech/tags/featured.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-research](<https://devfeed.tech/tags/nvidia-research.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [robot-manipulation](<https://devfeed.tech/tags/robot-manipulation.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [thor](<https://devfeed.tech/tags/thor.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [world-model](<https://devfeed.tech/tags/world-model.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

The article explains how World Action Models (WAMs) use video world models as backbones for robot policies, addressing the physical-generalization limitations of vision-language-action models. It discusses post-training WAMs into specialized policies and presents NVIDIA Cosmos 3 as a foundation for building them.

### Source excerpt

A central challenge in robotics is building policies that generalize beyond the demonstrations they're trained on. A policy that succeeds in a training scene...

## How to Evaluate General-Purpose Robot Policies for Real-World Deployment

DevFeed: [How to Evaluate General-Purpose Robot Policies for Real-World Deployment](<https://devfeed.tech/articles/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment-6849.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment/>)

Author: Brad Nemire

Published: 2026-07-12T01:08:17Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Robotics](<https://devfeed.tech/topics/robotics.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [Simulation and Design](<https://devfeed.tech/topics/simulation-and-design.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [featured](<https://devfeed.tech/tags/featured.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [nvidia-research](<https://devfeed.tech/tags/nvidia-research.md>), [physics](<https://devfeed.tech/tags/physics.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This developer article examines the challenge of rigorously evaluating general-purpose robot policies for real-world deployment. It discusses simulation as a scalable proxy for expensive real-world testing and identifies limitations in current benchmarks, including shared visual sources between training and evaluation, costly Real2sim reconstruction, static task sets, performance saturation, limited failure diagnostics, and uncertainty in success-rate estimates.

### Source excerpt

Robotics foundation models have made remarkable progress. Today's best systems can follow natural language instructions to pick, place, sort, and manipulate a...