# Synthetic Data Generation

Published articles for Synthetic Data Generation.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

DevFeed: [Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video](<https://devfeed.tech/articles/skild-ai-taps-nvidia-physical-ai-to-teach-robots-new-tasks-from-a-single-video-6961.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/>)

Author: Sasa Docca

Published: 2026-09-10T16:30:35Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [Cosmos](<https://devfeed.tech/topics/cosmos.md>), [Simulation and Design](<https://devfeed.tech/topics/simulation-and-design.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cosmos](<https://devfeed.tech/tags/cosmos.md>), [customer-stories](<https://devfeed.tech/tags/customer-stories.md>), [industrial-and-manufacturing](<https://devfeed.tech/tags/industrial-and-manufacturing.md>), [isaac](<https://devfeed.tech/tags/isaac.md>), [model](<https://devfeed.tech/tags/model.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [omniverse](<https://devfeed.tech/tags/omniverse.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [robots](<https://devfeed.tech/tags/robots.md>), [simulation-and-design](<https://devfeed.tech/tags/simulation-and-design.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [tensorrt](<https://devfeed.tech/tags/tensorrt.md>), [video](<https://devfeed.tech/tags/video.md>)

### AI overview

Skild AI's S1 robot foundation model learns new long-horizon physical tasks from a single video demonstration through in-context learning, without task-specific retraining. The article describes its development on NVIDIA AI infrastructure and use of NVIDIA Isaac Lab and Cosmos technologies.

### Source excerpt

Manufacturing floors, warehouses and production lines rarely stay fixed -- tasks change, layouts shift and new products arrive, and most robots can't keep up without significant reprogramming. Skild AI's new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. The model, launched last week, uses [...]

## 34 Amazon Research Awards Build on Trainium recipients announced

DevFeed: [34 Amazon Research Awards Build on Trainium recipients announced](<https://devfeed.tech/articles/34-amazon-research-awards-build-on-trainium-recipients-announced-7614.md>)

Original publisher: [Read original article](<https://www.amazon.science/research-awards/latest-news/34-amazon-research-awards-build-on-trainium-recipients-announced>)

Author: Amazon Research Awards team

Published: 2026-08-05T15:00:00Z

Content type: news

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [responsible-ai](<https://devfeed.tech/topics/responsible-ai.md>), [AWS AI chips](<https://devfeed.tech/topics/aws-ai-chips.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [academic-ai-funding](<https://devfeed.tech/tags/academic-ai-funding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [ai-research-grants](<https://devfeed.tech/tags/ai-research-grants.md>), [ai-safety-and-alignment](<https://devfeed.tech/tags/ai-safety-and-alignment.md>), [amazon-research-awards](<https://devfeed.tech/tags/amazon-research-awards.md>), [ara](<https://devfeed.tech/tags/ara.md>), [aws-ai-chips](<https://devfeed.tech/tags/aws-ai-chips.md>), [aws-trainium](<https://devfeed.tech/tags/aws-trainium.md>), [build-on-trainium](<https://devfeed.tech/tags/build-on-trainium.md>), [data](<https://devfeed.tech/tags/data.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [inference](<https://devfeed.tech/tags/inference.md>), [internal-ara-program-updates](<https://devfeed.tech/tags/internal-ara-program-updates.md>), [llm](<https://devfeed.tech/tags/llm.md>), [machine-learning-research](<https://devfeed.tech/tags/machine-learning-research.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [research](<https://devfeed.tech/tags/research.md>), [responsible-ai](<https://devfeed.tech/tags/responsible-ai.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Amazon announces 34 recipients of its Build on Trainium program, a $110 million credit initiative supporting AI research and university education. The awards fund work in areas including Responsible AI, language models, synthetic data, distributed systems, model architectures, libraries, and optimization on AWS Trainium.

### Source excerpt

Amazon announces 34 recipients of the Build on Trainium program, a $110 million credit initiative supporting AI research at 30 universities including Stanford, UC Berkeley, UIUC, UCLA, CMU, and MIT, with a focus on Responsible AI.

## NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics

DevFeed: [NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics](<https://devfeed.tech/articles/nvidia-cosmos-h-dreams-bringing-real-time-generative-simulation-to-surgical-robotics-7378.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/cosmos-h-dreams>)

Author: Lukas Zbinden; Javier Gamazo; Mostafa Toloui; Sean Huver

Published: 2026-07-27T09:32:20Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Cosmos](<https://devfeed.tech/topics/cosmos.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [cosmos](<https://devfeed.tech/tags/cosmos.md>), [data](<https://devfeed.tech/tags/data.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative](<https://devfeed.tech/tags/generative.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

NVIDIA introduces Cosmos-H-Dreams, a real-time, action-conditioned generative simulator for surgical robotics. The model generates future surgical video from an initial RGB frame and live robot kinematics, enabling interactive closed-loop control, offline policy evaluation, and synthetic data generation.

### Source excerpt

World foundation models offer a different path. Instead of manually authoring every object and physical interaction, they learn visual dynamics directly from synchronized video and robot kinematics. NVIDIA's Cosmos-H-Surgical-Simulator demonstrated this approach by generating future surgical video from an initial scene and a sequence of robot actions. It enabled faster-than-physical evaluation and synthetic data generation across the Open-H-Embodiment ecosystem.

## How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo

DevFeed: [How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo](<https://devfeed.tech/articles/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo-6853.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/>)

Author: Tanya Lenz

Published: 2026-07-14T16:00:00Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [debugging](<https://devfeed.tech/topics/debugging.md>)

Tags: [agent-skill](<https://devfeed.tech/tags/agent-skill.md>), [agent-skills](<https://devfeed.tech/tags/agent-skills.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [python](<https://devfeed.tech/tags/python.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [rl](<https://devfeed.tech/tags/rl.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [training](<https://devfeed.tech/tags/training.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

A tutorial on running a skill-based autoresearch workflow in which coding AI agents set up, debug, run, monitor, and iterate on reinforcement-learning experiments using NVIDIA NeMo RL and NeMo Gym.

### Source excerpt

Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve...

## Synthetic Data Generation for Financial AI Research with NVIDIA NeMo

DevFeed: [Synthetic Data Generation for Financial AI Research with NVIDIA NeMo](<https://devfeed.tech/articles/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo-6943.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo/>)

Author: Elizabeth Goodman

Published: 2026-07-09T19:40:37Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Natural language processing](<https://devfeed.tech/topics/nlp.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [datasets](<https://devfeed.tech/topics/datasets.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-ready-data](<https://devfeed.tech/tags/ai-ready-data.md>), [cloud-services](<https://devfeed.tech/tags/cloud-services.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data](<https://devfeed.tech/tags/data.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [featured](<https://devfeed.tech/tags/featured.md>), [financial-services](<https://devfeed.tech/tags/financial-services.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [generation](<https://devfeed.tech/tags/generation.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llms](<https://devfeed.tech/tags/llms.md>), [models](<https://devfeed.tech/tags/models.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [research](<https://devfeed.tech/tags/research.md>), [structured-generation](<https://devfeed.tech/tags/structured-generation.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [vllm](<https://devfeed.tech/tags/vllm.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

This developer article presents an iterative pipeline for generating a diverse synthetic dataset of more than 500,000 financial news headlines. It combines NeMo Data Designer for structured generation, NeMo Curator for semantic deduplication, Nemotron models for synthesis, and a farthest-from-centroid few-shot strategy to reduce repetition and correct category imbalance.

### Source excerpt

Fine-tuning LLMs for financial natural language processing (NLP) is constrained by limited, imbalanced data. Real-world financial news overrepresents earnings...

## Optimizing a Neural Reconstruction Pipeline Using NVIDIA Nsight Developer Tools

DevFeed: [Optimizing a Neural Reconstruction Pipeline Using NVIDIA Nsight Developer Tools](<https://devfeed.tech/articles/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools-6918.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools/>)

Author: Tanya Lenz

Published: 2026-06-30T16:00:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Omniverse](<https://devfeed.tech/topics/omniverse.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Tooling](<https://devfeed.tech/topics/tooling.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [CUDA](<https://devfeed.tech/topics/cuda.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>), [Physical AI](<https://devfeed.tech/topics/physical-ai.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>)

Tags: [3d](<https://devfeed.tech/tags/3d.md>), [ai](<https://devfeed.tech/tags/ai.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [autonomous-vehicles](<https://devfeed.tech/tags/autonomous-vehicles.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [developer-tools](<https://devfeed.tech/tags/developer-tools.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [driving](<https://devfeed.tech/tags/driving.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [lidar](<https://devfeed.tech/tags/lidar.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [omniverse](<https://devfeed.tech/tags/omniverse.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

This article explains how NVIDIA Nsight Developer Tools can optimize the NVIDIA Omniverse NuRec neural reconstruction pipeline. It focuses on reducing GPU-intensive reconstruction and rendering costs to improve engineering iteration and move toward real-time performance.

### Source excerpt

NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high-fidelity 3D representations of real-world environments from multisensor data such...

## Datadog acquires Adaptive ML

DevFeed: [Datadog acquires Adaptive ML](<https://devfeed.tech/articles/datadog-acquires-adaptive-ml-2255.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/datadog-acquires-adaptive-ml/>)

Author: Alexis Lê-Quôc

Published: 2026-06-30T00:00:00Z

Content type: news

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [agent observability](<https://devfeed.tech/topics/agent-observability.md>), [Frontier AI](<https://devfeed.tech/topics/frontier-ai.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [A/B Testing](<https://devfeed.tech/topics/a-b-testing.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Security](<https://devfeed.tech/topics/security.md>), [real-time](<https://devfeed.tech/topics/real-time.md>)

Tags: [a-b-testing](<https://devfeed.tech/tags/a-b-testing.md>), [acquisition](<https://devfeed.tech/tags/acquisition.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [data](<https://devfeed.tech/tags/data.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [frontier-ai](<https://devfeed.tech/tags/frontier-ai.md>), [models](<https://devfeed.tech/tags/models.md>), [observability](<https://devfeed.tech/tags/observability.md>), [platform](<https://devfeed.tech/tags/platform.md>), [production](<https://devfeed.tech/tags/production.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [security](<https://devfeed.tech/tags/security.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Datadog announces the acquisition of Adaptive ML, whose Adaptive Engine helps enterprises build, own, and deploy specialized AI agents and models. The platform supports fine-tuning open models with reinforcement learning and synthetic data, evaluating them with AI judges and A/B testing, and using production signals to improve subsequent training.

### Source excerpt

Datadog has acquired Adaptive ML, a platform for building, owning, and deploying specialized AI agents and models.

## Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles

DevFeed: [Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles](<https://devfeed.tech/articles/designing-synthetic-datasets-for-the-real-world-mechanism-design-and-reasoning-from-first-principles-6758.md>)

Original publisher: [Read original article](<https://research.google/blog/designing-synthetic-datasets-for-the-real-world-mechanism-design-and-reasoning-from-first-principles/>)

Published: 2026-04-16T14:41:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [datasets](<https://devfeed.tech/topics/datasets.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [machine learning research](<https://devfeed.tech/topics/machine-learning-research.md>), [Test coverage](<https://devfeed.tech/topics/coverage.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [machine-learning-research](<https://devfeed.tech/tags/machine-learning-research.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Google Research introduces Simula, a framework that treats synthetic data generation as dataset-level mechanism design. It uses reasoning from first principles to control coverage, diversity, complexity, and quality for scalable generation in data-scarce or privacy-sensitive domains.

### Source excerpt

Generative AI

## Build a Domain-Specific Embedding Model in Under a Day

DevFeed: [Build a Domain-Specific Embedding Model in Under a Day](<https://devfeed.tech/articles/build-a-domain-specific-embedding-model-in-under-a-day-7379.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/domain-specific-embedding-finetune>)

Author: Steve Han; Rucha Apte; Sean Sodha; Oliver Holworthy

Published: 2026-03-20T19:38:16Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [Model Development](<https://devfeed.tech/topics/model-development.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [NeMo](<https://devfeed.tech/topics/nemo.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [NVIDIA NIM](<https://devfeed.tech/topics/nvidia-nim.md>), [TensorRT](<https://devfeed.tech/topics/tensorrt.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nim](<https://devfeed.tech/tags/nim.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-nim](<https://devfeed.tech/tags/nvidia-nim.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

A tutorial showing how to fine-tune a general-purpose embedding model for a specific domain in less than a day using synthetic question-answer pairs generated from domain documents. It covers data generation, contrastive training, retrieval evaluation, and deployment, using NVIDIA NeMo components and a Llama-Nemotron embedding model.

### Source excerpt

With a single GPU and less than a day of training time, you can transform a general-purpose embedding model into one that truly understands your domain, no manual labeling required. To help you hit the ground running, we are also releasing a ready-to-use synthetic training dataset generated from NVIDIA's public documentation using this exact pipeline.

## Mistral AI partners with NVIDIA to accelerate open frontier models

DevFeed: [Mistral AI partners with NVIDIA to accelerate open frontier models](<https://devfeed.tech/articles/mistral-ai-partners-with-nvidia-to-accelerate-open-frontier-models-7044.md>)

Original publisher: [Read original article](<https://mistral.ai/news/mistral-ai-and-nvidia-partner-to-accelerate-open-frontier-models/>)

Published: 2026-03-16T20:00:00Z

Content type: news

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Model Development](<https://devfeed.tech/topics/model-development.md>), [NVIDIA DGX](<https://devfeed.tech/topics/nvidia-dgx.md>), [DGX Cloud](<https://devfeed.tech/topics/dgx-cloud.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [NeMo](<https://devfeed.tech/topics/nemo.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [dgx-cloud](<https://devfeed.tech/tags/dgx-cloud.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open](<https://devfeed.tech/tags/open.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Mistral AI announces a partnership with NVIDIA and its founding membership in the NVIDIA Nemotron Coalition. The collaboration will develop open frontier AI models using Mistral AI's model expertise and NVIDIA's compute, development tools, and synthetic-data pipelines. The coalition's first initiative will support the NVIDIA Nemotron 4 family, while Mistral AI also releases Mistral Small 4 for developers, researchers, and organizations.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## Teaching AI to read a map

DevFeed: [Teaching AI to read a map](<https://devfeed.tech/articles/teaching-ai-to-read-a-map-6884.md>)

Original publisher: [Read original article](<https://research.google/blog/teaching-ai-to-read-a-map/>)

Published: 2026-02-17T21:37:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Machine Perception](<https://devfeed.tech/topics/machine-perception.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [large-language-models](<https://devfeed.tech/topics/large-language-models.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [generation](<https://devfeed.tech/tags/generation.md>), [google](<https://devfeed.tech/tags/google.md>), [images](<https://devfeed.tech/tags/images.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [machine-perception](<https://devfeed.tech/tags/machine-perception.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-models-datasets](<https://devfeed.tech/tags/open-source-models-datasets.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Google Research presents MapTrace, a task, dataset, and synthetic data generation pipeline for teaching multimodal large language models to trace valid routes on maps. The work addresses models' difficulty with spatial, geometric, and topological reasoning and releases 2 million generated question-answer pairs using Gemini 2.5 Pro and Imagen-4 Models.

### Source excerpt

Machine Perception

## A picture's worth a thousand (private) words: Hierarchical generation of coherent synthetic photo albums

DevFeed: [A picture's worth a thousand (private) words: Hierarchical generation of coherent synthetic photo albums](<https://devfeed.tech/articles/a-picture-s-worth-a-thousand-private-words-hierarchical-generation-of-coherent-synthetic-photo-albums-6742.md>)

Original publisher: [Read original article](<https://research.google/blog/a-pictures-worth-a-thousand-private-words-hierarchical-generation-of-coherent-synthetic-photo-albums/>)

Published: 2025-10-20T21:54:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Image](<https://devfeed.tech/topics/image.md>), [Google](<https://devfeed.tech/topics/google.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative](<https://devfeed.tech/tags/generative.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [image](<https://devfeed.tech/tags/image.md>), [images](<https://devfeed.tech/tags/images.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [security-privacy-and-abuse-prevention](<https://devfeed.tech/tags/security-privacy-and-abuse-prevention.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Google Research introduces a method for generating differentially private synthetic photo albums. The approach translates image data into an intermediate text representation and generates albums hierarchically to preserve thematic coherence and character consistency across multiple photos. It uses differentially private fine-tuning, such as DP-SGD, to produce representative synthetic data without unique details from individual users.

### Source excerpt

Generative AI

## Nemotron-Personas-India: Synthesized Data for Sovereign AI

DevFeed: [Nemotron-Personas-India: Synthesized Data for Sovereign AI](<https://devfeed.tech/articles/nemotron-personas-india-synthesized-data-for-sovereign-ai-7397.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/nemotron-personas-india>)

Author: Kiran Praveen; Utkarsh Vaidya; Evan A; Lipika Ramaswamy; Dhruv Nathawani; Dane Corneil; Yev Meyer

Published: 2025-10-13T23:00:42Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Sovereign AI](<https://devfeed.tech/topics/sovereign-ai.md>), [NeMo](<https://devfeed.tech/topics/nemo.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [LLMs](<https://devfeed.tech/topics/llms.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-adoption](<https://devfeed.tech/tags/ai-adoption.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [generation](<https://devfeed.tech/tags/generation.md>), [india](<https://devfeed.tech/tags/india.md>), [language](<https://devfeed.tech/tags/language.md>), [llms](<https://devfeed.tech/tags/llms.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

NVIDIA releases Nemotron-Personas-India, an open synthetic dataset of Indic personas designed to address the lack of multilingual and culturally representative data for Indian AI systems. Built with NeMo Data Designer and licensed under CC BY 4.0, it contains 21 million personas across English and Hindi in Devanagari and Latin scripts, with demographic, geographic, occupational, and cultural attributes.

### Source excerpt

India represents one of the world's largest AI opportunities -- with over 700 million internet users, a multitude of languages, and a rapidly growing developer ecosystem. Yet, most open datasets reflect Western norms and English-only contexts, creating a data gap that limits AI adoption in India's multilingual, multi-script environment.

## NVIDIA's GTC 2025 Announcement for Physical AI Developers: New Open Models and Datasets

DevFeed: [NVIDIA's GTC 2025 Announcement for Physical AI Developers: New Open Models and Datasets](<https://devfeed.tech/articles/nvidia-s-gtc-2025-announcement-for-physical-ai-developers-new-open-models-and-datasets-7370.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia-physical-ai>)

Author: Ming-Yu Liu; Hanzi Mao; Jinwei Gu; Pranjali Joshi; Asawaree

Published: 2025-03-18T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Cosmos](<https://devfeed.tech/topics/cosmos.md>), [Physical AI](<https://devfeed.tech/topics/physical-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [autonomous vehicles](<https://devfeed.tech/topics/autonomous-vehicles.md>), [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [Isaac](<https://devfeed.tech/topics/isaac.md>), [Omniverse](<https://devfeed.tech/topics/omniverse.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [autonomous-vehicles](<https://devfeed.tech/tags/autonomous-vehicles.md>), [community](<https://devfeed.tech/tags/community.md>), [cosmos](<https://devfeed.tech/tags/cosmos.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [isaac](<https://devfeed.tech/tags/isaac.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [omniverse](<https://devfeed.tech/tags/omniverse.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [research](<https://devfeed.tech/tags/research.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

NVIDIA announces Cosmos Transfer, an open world foundation model that generates controllable, photorealistic video scenes from structured inputs such as depth maps, trajectories, LiDAR scans, and 3D bounding boxes. Together with NVIDIA Omniverse, it supports synthetic data generation for robotics and autonomous vehicle development. NVIDIA also introduces an open Physical AI Dataset on Hugging Face and the Isaac GR00T N1 foundation model for humanoid robot reasoning.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Vibe Coding With AI to Generate Synthetic Data: Part 1

DevFeed: [Vibe Coding With AI to Generate Synthetic Data: Part 1](<https://devfeed.tech/articles/vibe-coding-with-ai-to-generate-synthetic-data-part-1-5846.md>)

Original publisher: [Read original article](<https://neon.com/blog/vibe-coding-with-ai-to-generate-synthetic-data-part-1>)

Author: Paul Scanlon

Published: 2025-03-13T16:11:50Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [GitHub Actions](<https://devfeed.tech/topics/github-actions.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [github-actions](<https://devfeed.tech/tags/github-actions.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [product](<https://devfeed.tech/tags/product.md>), [sql](<https://devfeed.tech/tags/sql.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [vibe-coding](<https://devfeed.tech/tags/vibe-coding.md>)

### AI overview

The article describes an experiment using vibe coding and an Anthropic API prompt to generate synthetic data from PostgreSQL schemas. It proposes GitHub Actions to dump a production or staging schema, pass it to the model, and automate data generation for development and testing workflows.

### Source excerpt

In this blog post, I'll share my experience of vibe coding as I explore AI-driven synthetic data generation, the tests I ran, and the challenges I faced along the way. What is vibe coding? Vibe coding is all about skipping the boilerplate and getting straight to the good stuff. I...

## Open Preference Dataset for Text-to-Image Generation by the 🤗 Community

DevFeed: [Open Preference Dataset for Text-to-Image Generation by the 🤗 Community](<https://devfeed.tech/articles/open-preference-dataset-for-text-to-image-generation-by-the-community-7274.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/image-preferences>)

Author: David Berenstein; ben burtenshaw; Daniel Vila; Daniel van Strien; Sayak Paul; Ame Vi; Linoy Tsaban

Published: 2024-12-09T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [flux](<https://devfeed.tech/topics/flux.md>), [stable-diffusion](<https://devfeed.tech/topics/stable-diffusion.md>), [argilla](<https://devfeed.tech/topics/argilla.md>), [distilabel](<https://devfeed.tech/topics/distilabel.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>)

Tags: [argilla](<https://devfeed.tech/tags/argilla.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [data](<https://devfeed.tech/tags/data.md>), [data-is-better-together](<https://devfeed.tech/tags/data-is-better-together.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [distilabel](<https://devfeed.tech/tags/distilabel.md>), [flux](<https://devfeed.tech/tags/flux.md>), [generation](<https://devfeed.tech/tags/generation.md>), [github](<https://devfeed.tech/tags/github.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [image](<https://devfeed.tech/tags/image.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [stable-diffusion](<https://devfeed.tech/tags/stable-diffusion.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [text-to-image](<https://devfeed.tech/tags/text-to-image.md>)

### AI overview

The article describes an open community effort to create an image-preference dataset for text-to-image generation. It covers prompt preparation with distilabel, synthetic data generation, image generation with Flux and Stable Diffusion, and filtering with text- and image-based classifiers plus manual review. The resulting dataset and related code are available through the Hugging Face Hub and GitHub.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.