# synthetic-data

Artificial data created from seed data that retain some statistical characteristics of that data.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Scale AV Perception Across Vehicle Platforms with NVIDIA Omniverse NuRec

DevFeed: [Scale AV Perception Across Vehicle Platforms with NVIDIA Omniverse NuRec](<https://devfeed.tech/articles/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec-6936.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/>)

Author: Michelle Horton

Published: 2026-08-31T16:00:00Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Simulation and Design](<https://devfeed.tech/topics/simulation-and-design.md>), [data](<https://devfeed.tech/topics/data.md>), [configuration](<https://devfeed.tech/topics/configuration.md>)

Tags: [3d](<https://devfeed.tech/tags/3d.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [autonomous-vehicles](<https://devfeed.tech/tags/autonomous-vehicles.md>), [data](<https://devfeed.tech/tags/data.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [developers](<https://devfeed.tech/tags/developers.md>), [development](<https://devfeed.tech/tags/development.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [omniverse](<https://devfeed.tech/tags/omniverse.md>), [performance](<https://devfeed.tech/tags/performance.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [post](<https://devfeed.tech/tags/post.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [scale](<https://devfeed.tech/tags/scale.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [training](<https://devfeed.tech/tags/training.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>)

### AI overview

This tutorial explains how to adapt an autonomous-vehicle perception stack across carline and sensor-rig variants using existing real-world drives. It presents a four-step workflow with NVIDIA Omniverse NuRec: pair a reconstructed drive with a target rig, render target camera views, refine the frames with NVIDIA Harmonizer, and train a perception model on the output.

### Source excerpt

A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline--for example, from an SUV to a sedan or another vehicle...

## How Bits Database Optimization proves a query rewrite is faster

DevFeed: [How Bits Database Optimization proves a query rewrite is faster](<https://devfeed.tech/articles/how-bits-database-optimization-proves-a-query-rewrite-is-faster-2278.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/how-bits-database-optimization-proves-a-query-rewrite-is-faster/>)

Author: Alex Weisberger; Nenad Noveljić; Bowen Chen

Published: 2026-08-31T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Optimization](<https://devfeed.tech/topics/optimization.md>), [database monitoring](<https://devfeed.tech/topics/database-monitoring.md>), [Database](<https://devfeed.tech/topics/database.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [IO](<https://devfeed.tech/topics/io.md>), [Security](<https://devfeed.tech/topics/security.md>), [Statistics](<https://devfeed.tech/topics/statistics.md>)

Tags: [cpu](<https://devfeed.tech/tags/cpu.md>), [data](<https://devfeed.tech/tags/data.md>), [database](<https://devfeed.tech/tags/database.md>), [database-monitoring](<https://devfeed.tech/tags/database-monitoring.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [security](<https://devfeed.tech/tags/security.md>), [software](<https://devfeed.tech/tags/software.md>), [sql](<https://devfeed.tech/tags/sql.md>), [statistics](<https://devfeed.tech/tags/statistics.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>)

### AI overview

The article explains how Bits Database Optimization validates that a proposed query rewrite is faster. It describes controlled benchmarking with simulated production-like datasets, accounting for cache state, CPU and I/O contention, execution time, and database work.

### Source excerpt

Learn how Bits generates synthetic data, measures simulation fidelity, and uses execution time and database work to determine whether an optimization is truly faster.

## 34 Amazon Research Awards Build on Trainium recipients announced

DevFeed: [34 Amazon Research Awards Build on Trainium recipients announced](<https://devfeed.tech/articles/34-amazon-research-awards-build-on-trainium-recipients-announced-7614.md>)

Original publisher: [Read original article](<https://www.amazon.science/research-awards/latest-news/34-amazon-research-awards-build-on-trainium-recipients-announced>)

Author: Amazon Research Awards team

Published: 2026-08-05T15:00:00Z

Content type: news

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [responsible-ai](<https://devfeed.tech/topics/responsible-ai.md>), [AWS AI chips](<https://devfeed.tech/topics/aws-ai-chips.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [academic-ai-funding](<https://devfeed.tech/tags/academic-ai-funding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [ai-research-grants](<https://devfeed.tech/tags/ai-research-grants.md>), [ai-safety-and-alignment](<https://devfeed.tech/tags/ai-safety-and-alignment.md>), [amazon-research-awards](<https://devfeed.tech/tags/amazon-research-awards.md>), [ara](<https://devfeed.tech/tags/ara.md>), [aws-ai-chips](<https://devfeed.tech/tags/aws-ai-chips.md>), [aws-trainium](<https://devfeed.tech/tags/aws-trainium.md>), [build-on-trainium](<https://devfeed.tech/tags/build-on-trainium.md>), [data](<https://devfeed.tech/tags/data.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [inference](<https://devfeed.tech/tags/inference.md>), [internal-ara-program-updates](<https://devfeed.tech/tags/internal-ara-program-updates.md>), [llm](<https://devfeed.tech/tags/llm.md>), [machine-learning-research](<https://devfeed.tech/tags/machine-learning-research.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [research](<https://devfeed.tech/tags/research.md>), [responsible-ai](<https://devfeed.tech/tags/responsible-ai.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Amazon announces 34 recipients of its Build on Trainium program, a $110 million credit initiative supporting AI research and university education. The awards fund work in areas including Responsible AI, language models, synthetic data, distributed systems, model architectures, libraries, and optimization on AWS Trainium.

### Source excerpt

Amazon announces 34 recipients of the Build on Trainium program, a $110 million credit initiative supporting AI research at 30 universities including Stanford, UC Berkeley, UIUC, UCLA, CMU, and MIT, with a focus on Responsible AI.

## NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics

DevFeed: [NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics](<https://devfeed.tech/articles/nvidia-cosmos-h-dreams-bringing-real-time-generative-simulation-to-surgical-robotics-7378.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/cosmos-h-dreams>)

Author: Lukas Zbinden; Javier Gamazo; Mostafa Toloui; Sean Huver

Published: 2026-07-27T09:32:20Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Cosmos](<https://devfeed.tech/topics/cosmos.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [cosmos](<https://devfeed.tech/tags/cosmos.md>), [data](<https://devfeed.tech/tags/data.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative](<https://devfeed.tech/tags/generative.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

NVIDIA introduces Cosmos-H-Dreams, a real-time, action-conditioned generative simulator for surgical robotics. The model generates future surgical video from an initial RGB frame and live robot kinematics, enabling interactive closed-loop control, offline policy evaluation, and synthetic data generation.

### Source excerpt

World foundation models offer a different path. Instead of manually authoring every object and physical interaction, they learn visual dynamics directly from synchronized video and robot kinematics. NVIDIA's Cosmos-H-Surgical-Simulator demonstrated this approach by generating future surgical video from an initial scene and a sequence of robot actions. It enabled faster-than-physical evaluation and synthetic data generation across the Open-H-Embodiment ecosystem.

## Beyond Redaction: Anatomy of a Privacy-Safe Data Platform

DevFeed: [Beyond Redaction: Anatomy of a Privacy-Safe Data Platform](<https://devfeed.tech/articles/beyond-redaction-anatomy-of-a-privacy-safe-data-platform-18252.md>)

Original publisher: [Read original article](<https://www.dataengineeringweekly.com/p/beyond-redaction-anatomy-of-a-privacy>)

Author: Ananth Packkildurai

Published: 2026-07-10T10:20:21Z

Content type: tutorial

Language: en

Sources: [Data Engineering Weekly](<https://devfeed.tech/sources/data-engineering-weekly.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [pii](<https://devfeed.tech/topics/pii.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Query (disambiguation)](<https://devfeed.tech/topics/query.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [encryption](<https://devfeed.tech/tags/encryption.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [pii](<https://devfeed.tech/tags/pii.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [tokenization](<https://devfeed.tech/tags/tokenization.md>)

### AI overview

This article explains why privacy-safe data platforms must preserve only the utility required for an approved purpose while controlling linkability and generating evidence that safeguards operated. It compares redaction, encryption, tokenization, governed views, aggregation, synthetic data, and clean-room access, emphasizing that deterministic tokens are generally pseudonymous rather than automatically anonymous.

### Source excerpt

Why privacy engineering is about governing data in use--not simply hiding it.

## Synthetic Data Generation for Financial AI Research with NVIDIA NeMo

DevFeed: [Synthetic Data Generation for Financial AI Research with NVIDIA NeMo](<https://devfeed.tech/articles/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo-6943.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/synthetic-data-generation-for-financial-ai-research-with-nvidia-nemo/>)

Author: Elizabeth Goodman

Published: 2026-07-09T19:40:37Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Natural language processing](<https://devfeed.tech/topics/nlp.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [datasets](<https://devfeed.tech/topics/datasets.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-ready-data](<https://devfeed.tech/tags/ai-ready-data.md>), [cloud-services](<https://devfeed.tech/tags/cloud-services.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data](<https://devfeed.tech/tags/data.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [featured](<https://devfeed.tech/tags/featured.md>), [financial-services](<https://devfeed.tech/tags/financial-services.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [generation](<https://devfeed.tech/tags/generation.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llms](<https://devfeed.tech/tags/llms.md>), [models](<https://devfeed.tech/tags/models.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [research](<https://devfeed.tech/tags/research.md>), [structured-generation](<https://devfeed.tech/tags/structured-generation.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [vllm](<https://devfeed.tech/tags/vllm.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

This developer article presents an iterative pipeline for generating a diverse synthetic dataset of more than 500,000 financial news headlines. It combines NeMo Data Designer for structured generation, NeMo Curator for semantic deduplication, Nemotron models for synthesis, and a farthest-from-centroid few-shot strategy to reduce repetition and correct category imbalance.

### Source excerpt

Fine-tuning LLMs for financial natural language processing (NLP) is constrained by limited, imbalanced data. Real-world financial news overrepresents earnings...

## Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles

DevFeed: [Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles](<https://devfeed.tech/articles/designing-synthetic-datasets-for-the-real-world-mechanism-design-and-reasoning-from-first-principles-6758.md>)

Original publisher: [Read original article](<https://research.google/blog/designing-synthetic-datasets-for-the-real-world-mechanism-design-and-reasoning-from-first-principles/>)

Published: 2026-04-16T14:41:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [datasets](<https://devfeed.tech/topics/datasets.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [machine learning research](<https://devfeed.tech/topics/machine-learning-research.md>), [Test coverage](<https://devfeed.tech/topics/coverage.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [machine-learning-research](<https://devfeed.tech/tags/machine-learning-research.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Google Research introduces Simula, a framework that treats synthetic data generation as dataset-level mechanism design. It uses reasoning from first principles to control coverage, diversity, complexity, and quality for scalable generation in data-scarce or privacy-sensitive domains.

### Source excerpt

Generative AI

## Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents

DevFeed: [Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents](<https://devfeed.tech/articles/granite-4-0-3b-vision-compact-multimodal-intelligence-for-enterprise-documents-7259.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ibm-granite/granite-4-vision>)

Author: Madison Lee; Rogerio Feris; Eli Schwartz; Dhiraj Joshi; Pengyuan Li; Isaac Sanchez

Published: 2026-03-31T15:10:41Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [multimodal](<https://devfeed.tech/topics/multimodal.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [lora](<https://devfeed.tech/topics/lora.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [enterprise deployment](<https://devfeed.tech/topics/enterprise-deployment.md>)

Tags: [code](<https://devfeed.tech/tags/code.md>), [data](<https://devfeed.tech/tags/data.md>), [data-augmentation](<https://devfeed.tech/tags/data-augmentation.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [enterprise-deployment](<https://devfeed.tech/tags/enterprise-deployment.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [lora](<https://devfeed.tech/tags/lora.md>), [model](<https://devfeed.tech/tags/model.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [training](<https://devfeed.tech/tags/training.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

Granite 4.0 3B Vision is a compact multimodal model for enterprise document understanding. It extracts tables, interprets charts, and identifies semantic key-value pairs, using a LoRA adapter, visual-language capabilities, and the ChartNet dataset to support structured chart reasoning and document-processing pipelines.

### Source excerpt

- Table Extraction: Accurately parsing complex table structures (e.g., multi-row, multi-column, etc.) from document images - Chart Understanding: Converting charts and figures into structured machine-readable formats, summaries, or executable code - Semantic Key-Value Pair (KVP) Extraction: Identifying and grounding semantically meaningful key-value field pairs across diverse document layouts The model ships as a LoRA adapter on top of Granite 4.0 Micro, our dense language model, keeping...

## Build a Domain-Specific Embedding Model in Under a Day

DevFeed: [Build a Domain-Specific Embedding Model in Under a Day](<https://devfeed.tech/articles/build-a-domain-specific-embedding-model-in-under-a-day-7379.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/domain-specific-embedding-finetune>)

Author: Steve Han; Rucha Apte; Sean Sodha; Oliver Holworthy

Published: 2026-03-20T19:38:16Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [Model Development](<https://devfeed.tech/topics/model-development.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [NeMo](<https://devfeed.tech/topics/nemo.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [NVIDIA NIM](<https://devfeed.tech/topics/nvidia-nim.md>), [TensorRT](<https://devfeed.tech/topics/tensorrt.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nim](<https://devfeed.tech/tags/nim.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-nim](<https://devfeed.tech/tags/nvidia-nim.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

A tutorial showing how to fine-tune a general-purpose embedding model for a specific domain in less than a day using synthetic question-answer pairs generated from domain documents. It covers data generation, contrastive training, retrieval evaluation, and deployment, using NVIDIA NeMo components and a Llama-Nemotron embedding model.

### Source excerpt

With a single GPU and less than a day of training time, you can transform a general-purpose embedding model into one that truly understands your domain, no manual labeling required. To help you hit the ground running, we are also releasing a ready-to-use synthetic training dataset generated from NVIDIA's public documentation using this exact pipeline.

## Workload Capture & Replay

DevFeed: [Workload Capture & Replay](<https://devfeed.tech/articles/workload-capture-replay-30845.md>)

Original publisher: [Read original article](<https://hookrace.net/blog/workload-capture-replay/>)

Published: 2026-02-09T23:00:00Z

Content type: article

Language: en

Sources: [Dennis Felsing](<https://devfeed.tech/sources/dennis-felsing.md>)

Topics: [Tooling](<https://devfeed.tech/topics/tooling.md>), [Docker Compose](<https://devfeed.tech/topics/docker-compose.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [blog-post](<https://devfeed.tech/tags/blog-post.md>), [docker-compose](<https://devfeed.tech/tags/docker-compose.md>), [regression](<https://devfeed.tech/tags/regression.md>), [state](<https://devfeed.tech/tags/state.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>)

### AI overview

This article introduces workload capture and replay tooling for reproducing production issues. It records a Materialize instance's state, recent queries, and ingestion rates, then replays them in a Docker Compose environment using synthetic data.

### Source excerpt

When customers hit issues in production, it can be an effort to locally reproduce them, especially when external sources are involved. Reproducing issues is useful not just to figure out the root cause, but also to verify the fix and add a regression test. The newly introduced workload capture & replay tooling records a Materialize instance's state as well as recent queries and ingestion rates, then replays them in a Docker Compose environment with synthetic data. In this blog post I'll show how it works and talk about some of the challenges and future work. Read the rest of the blog post over on the Materialize blog.

## Building a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac

DevFeed: [Building a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac](<https://devfeed.tech/articles/building-a-healthcare-robot-from-simulation-to-deployment-with-nvidia-isaac-7330.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/lerobotxnvidia-healthcare>)

Author: Steven Palma; Andres Diaz-Pinto

Published: 2025-10-29T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Isaac for Healthcare](<https://devfeed.tech/topics/isaac-for-healthcare.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [lerobot](<https://devfeed.tech/topics/lerobot.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [healthcare](<https://devfeed.tech/tags/healthcare.md>), [inference](<https://devfeed.tech/tags/inference.md>), [isaac](<https://devfeed.tech/tags/isaac.md>), [lerobot](<https://devfeed.tech/tags/lerobot.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>)

### AI overview

This hands-on guide presents an end-to-end workflow for building an autonomous surgical assistant robot with NVIDIA Isaac for Healthcare. It covers collecting real and synthetic data with LeRobot, fine-tuning GR00T N1.5, evaluating in IsaacLab, and deploying real-time inference on physical hardware.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## How to Build a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac for Healthcare

DevFeed: [How to Build a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac for Healthcare](<https://devfeed.tech/articles/how-to-build-a-healthcare-robot-from-simulation-to-deployment-with-nvidia-isaac-for-healthcare-7402.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/nvidia-isaac-for-healthcare>)

Author: Asawaree

Published: 2025-10-28T20:42:35Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Isaac for Healthcare](<https://devfeed.tech/topics/isaac-for-healthcare.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>), [Simulation and Design](<https://devfeed.tech/topics/simulation-and-design.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [lerobot](<https://devfeed.tech/topics/lerobot.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Medical imaging](<https://devfeed.tech/topics/medical-imaging.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [healthcare](<https://devfeed.tech/tags/healthcare.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [inference](<https://devfeed.tech/tags/inference.md>), [isaac-for-healthcare](<https://devfeed.tech/tags/isaac-for-healthcare.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [training](<https://devfeed.tech/tags/training.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>)

### AI overview

This hands-on guide presents NVIDIA Isaac for Healthcare v0.4 and its SO-ARM starter workflow for building autonomous surgical-assistance robots. It covers collecting real and synthetic data with SO-ARM and LeRobot, post-training GR00T N1.5, evaluation in Isaac Lab, and deployment with real-time inference on physical hardware.

### Source excerpt

A hands-on guide to collecting data, training policies, and deploying autonomous medical robotics workflows on real hardware Simulation has been a cornerstone in medical imaging to address the data gap. However, in healthcare robotics until now, it's often been too slow, siloed, or difficult to translate into real-world systems. That's now changing.

## A picture's worth a thousand (private) words: Hierarchical generation of coherent synthetic photo albums

DevFeed: [A picture's worth a thousand (private) words: Hierarchical generation of coherent synthetic photo albums](<https://devfeed.tech/articles/a-picture-s-worth-a-thousand-private-words-hierarchical-generation-of-coherent-synthetic-photo-albums-6742.md>)

Original publisher: [Read original article](<https://research.google/blog/a-pictures-worth-a-thousand-private-words-hierarchical-generation-of-coherent-synthetic-photo-albums/>)

Published: 2025-10-20T21:54:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Image](<https://devfeed.tech/topics/image.md>), [Google](<https://devfeed.tech/topics/google.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [generation](<https://devfeed.tech/tags/generation.md>), [generative](<https://devfeed.tech/tags/generative.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [image](<https://devfeed.tech/tags/image.md>), [images](<https://devfeed.tech/tags/images.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [security-privacy-and-abuse-prevention](<https://devfeed.tech/tags/security-privacy-and-abuse-prevention.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Google Research introduces a method for generating differentially private synthetic photo albums. The approach translates image data into an intermediate text representation and generates albums hierarchically to preserve thematic coherence and character consistency across multiple photos. It uses differentially private fine-tuning, such as DP-SGD, to produce representative synthetic data without unique details from individual users.

### Source excerpt

Generative AI

## Nemotron-Personas-India: Synthesized Data for Sovereign AI

DevFeed: [Nemotron-Personas-India: Synthesized Data for Sovereign AI](<https://devfeed.tech/articles/nemotron-personas-india-synthesized-data-for-sovereign-ai-7397.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/nemotron-personas-india>)

Author: Kiran Praveen; Utkarsh Vaidya; Evan A; Lipika Ramaswamy; Dhruv Nathawani; Dane Corneil; Yev Meyer

Published: 2025-10-13T23:00:42Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Sovereign AI](<https://devfeed.tech/topics/sovereign-ai.md>), [NeMo](<https://devfeed.tech/topics/nemo.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [LLMs](<https://devfeed.tech/topics/llms.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-adoption](<https://devfeed.tech/tags/ai-adoption.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [generation](<https://devfeed.tech/tags/generation.md>), [india](<https://devfeed.tech/tags/india.md>), [language](<https://devfeed.tech/tags/language.md>), [llms](<https://devfeed.tech/tags/llms.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

NVIDIA releases Nemotron-Personas-India, an open synthetic dataset of Indic personas designed to address the lack of multilingual and culturally representative data for Indian AI systems. Built with NeMo Data Designer and licensed under CC BY 4.0, it contains 21 million personas across English and Hindi in Devanagari and Latin scripts, with demographic, geographic, occupational, and cultural attributes.

### Source excerpt

India represents one of the world's largest AI opportunities -- with over 700 million internet users, a multitude of languages, and a rapidly growing developer ecosystem. Yet, most open datasets reflect Western norms and English-only contexts, creating a data gap that limits AI adoption in India's multilingual, multi-script environment.

## Fighting Unwanted Notifications with Machine Learning in Chrome

DevFeed: [Fighting Unwanted Notifications with Machine Learning in Chrome](<https://devfeed.tech/articles/fighting-unwanted-notifications-with-machine-learning-in-chrome-4193.md>)

Original publisher: [Read original article](<https://blog.chromium.org/2025/05/fighting-unwanted-notifications-with.html>)

Author: Chromium Blog (noreply@blogger.com)

Published: 2025-05-08T16:59:00Z

Content type: article

Language: en

Sources: [Chromium Blog](<https://devfeed.tech/sources/chromium-blog.md>)

Topics: [Chrome](<https://devfeed.tech/topics/chrome.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Android](<https://devfeed.tech/topics/android.md>), [online privacy](<https://devfeed.tech/topics/online-privacy.md>), [Security](<https://devfeed.tech/topics/security.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [android](<https://devfeed.tech/tags/android.md>), [chrome](<https://devfeed.tech/tags/chrome.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [none](<https://devfeed.tech/tags/none.md>), [notifications](<https://devfeed.tech/tags/notifications.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [security](<https://devfeed.tech/tags/security.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>)

### AI overview

Chrome is introducing warnings for potentially deceptive or spammy notifications on Android. A local on-device machine learning model analyzes notification text, flags likely unwanted messages, and lets users unsubscribe, view the content, or allow future notifications. Notification contents remain on the device to protect privacy, and the model was trained with synthetic data generated by the Gemini large language model.

### Source excerpt

Notifications in Chrome are a useful feature to keep up with updates from your favorite sites. However, we know that some notifications may be spammy or even deceptive. We've received reports of notifications diverting you to download suspicious software, tricking you into sharing personal information or asking you to make purchases on potentially fraudulent online store fronts. To defend against these threats, Chrome is launching warnings of unwanted notifications on Android. This new feature uses on-device machine learning to detect and warn you about potentially deceptive or spammy notifications, giving you an extra level of control over the information displayed on your device. When a notification is flagged by Chrome, you'll see the name of the site sending the notification, a message warning that the contents of the notification are potentially deceptive or spammy, and the option to either unsubscribe from the site or see the flagged content. An example of a notification flagged as possibly spam. If you choose to see the notification you will still see the option to unsubscribe or you can choose to always allow notifications from that site and not see warnings in the future. What you see when viewing a flagged notification. How It Works Chrome uses a local, on-device machine learning model to analyze notification content. This model identifies notifications that are likely to be unwanted. The model is trained on the textual contents of the notification, like the title, body, and action button texts. Notifications are end to end encrypted. The analysis of each message is done on-device and notification contents are not sent to Google, to protect user privacy. Due to the sensitive nature of notifications content, the model was trained using synthetic data generated by the Gemini large language model (LLM). The training data was evaluated against real notifications Chrome security team collected by subscribing to a variety of websites that were then classified by

## A Field Guide to Improving AI Products Through Measurement and Iteration

DevFeed: [A Field Guide to Improving AI Products Through Measurement and Iteration](<https://devfeed.tech/articles/a-field-guide-to-rapidly-improving-ai-products-18792.md>)

Original publisher: [Read original article](<https://hamel.dev/blog/posts/field-guide/>)

Author: Hamel Husain

Published: 2025-03-24T07:00:00Z

Content type: tutorial

Language: en

Sources: [Hamel Husain](<https://devfeed.tech/sources/hamel-husain.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-development](<https://devfeed.tech/tags/ai-development.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [guide](<https://devfeed.tech/tags/guide.md>), [llms](<https://devfeed.tech/tags/llms.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>)

### AI overview

This field guide argues that AI teams should prioritize measurement and iteration over tools and frameworks. It presents error analysis as a high-value practice and discusses data viewers, domain experts, synthetic data, evaluation trust, and experiment-focused roadmaps.

### Source excerpt

Most AI teams focus on the wrong things. Here's a common scene from my consulting work: AI TEAM Here's our agent architecture - we've got RAG here, a router there, and we're using this new framework for... ME [Holding up my hand to pause the enthusiastic tech lead.] "Can you show me how you're measuring if any of this actually works?" ... Room goes quiet This scene has played out dozens of times over the last two years. Teams invest weeks building complex AI systems, but can't tell me if their changes are helping or hurting. This isn't surprising. With new tools and frameworks emerging weekly, it's natural to focus on tangible things we can control - which vector database to use, which LLM provider to choose, which agent framework to adopt. But after helping 30+ companies build AI products, I've discovered the teams who succeed barely talk about tools at all. Instead, they obsess over measurement and iteration. In this post, I'll show you exactly how these successful teams operate. You'll learn: How error analysis consistently reveals the highest-ROI improvements Why a simple data viewer is your most important AI investment How to empower domain experts (not just engineers) to improve your AI Why synthetic data is more effective than you think How to maintain trust in your evaluation system Why your AI roadmap should count experiments, not features I'll explain each of these topics with real examples. While every situation is unique, you'll see patterns that apply regardless of your domain or team size. Let's start by examining the most common mistake I see teams make - one that derails AI projects before they even begin. 1. The Most Common Mistake: Skipping Error Analysis The "tools first" mindset is the most common mistake in AI development. Teams get caught up in architecture diagrams, frameworks, and dashboards while neglecting the process of actually understanding what's working and what isn't. One client proudly showed me this evaluation dashboard: The kind of das

## Open R1: Update #2

DevFeed: [Open R1: Update #2](<https://devfeed.tech/articles/open-r1-update-2-7422.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-r1/update-2>)

Author: Loubna Ben Allal; Lewis Tunstall; Anton Lozhkov; Elie Bakouch; Guilherme Penedo; Hynek Kydlicek; Gabriel Martín Blázquez

Published: 2025-02-10T16:10:47Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [deepseek](<https://devfeed.tech/topics/deepseek.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [math](<https://devfeed.tech/topics/math.md>), [llama](<https://devfeed.tech/topics/llama.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [sglang](<https://devfeed.tech/topics/sglang.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [math-verify](<https://devfeed.tech/topics/math-verify.md>), [Parser](<https://devfeed.tech/topics/parser.md>)

Tags: [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [llama](<https://devfeed.tech/tags/llama.md>), [math](<https://devfeed.tech/tags/math.md>), [math-verify](<https://devfeed.tech/tags/math-verify.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

Open R1 Update #2 presents OpenR1-Math-220k, a large-scale mathematical reasoning dataset created to help reconstruct parts of the DeepSeek R1 training pipeline and synthetic-data process. The article describes reasoning-trace generation, local inference with vLLM and SGLang, automated filtering with Math Verify, and the use of Llama3.3-70B-Instruct as a judge. It also discusses distillation and fine-tuning of Qwen and Llama models using reasoning traces.

### Source excerpt

We are now two weeks into the Open R1 project which aims to reconstruct the missing pieces of DeepSeek R1--specifically, the training pipeline and synthetic data. In this post, we are happy to share the construction of OpenR1-Math-220k: our first large-scale dataset for mathematical reasoning!

## Neon's Instant Branches: Schema-Only or With Data, the Choice Is Yours

DevFeed: [Neon's Instant Branches: Schema-Only or With Data, the Choice Is Yours](<https://devfeed.tech/articles/neon-s-instant-branches-schema-only-or-with-data-the-choice-is-yours-5448.md>)

Original publisher: [Read original article](<https://neon.com/blog/instant-branches-schema-only-or-with-data-the-choice-is-yours>)

Author: Bryan Clark

Published: 2025-02-05T14:51:44Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Database](<https://devfeed.tech/topics/database.md>), [pii](<https://devfeed.tech/topics/pii.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [database](<https://devfeed.tech/tags/database.md>), [develop](<https://devfeed.tech/tags/develop.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [pii](<https://devfeed.tech/tags/pii.md>), [product](<https://devfeed.tech/tags/product.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

Neon introduces instant schema-only database branches for its Early Access Program, alongside existing branches that copy production data and schema. Schema-only branches help teams work within PII restrictions by allowing them to seed isolated environments with synthetic data, while data-inclusive branches remain useful for testing production-like behavior. Both branch types can be created through the Neon console or API.

### Source excerpt

If you've been keeping up with the Neon story, we introduced database branching in December 2022. Our instant branches--complete copies of production, including data and schema--have enabled thousands of developers to work and test efficiently within the safety of their own isolate...

## Open-R1: Update #1

DevFeed: [Open-R1: Update #1](<https://devfeed.tech/articles/open-r1-update-1-7420.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-r1/update-1>)

Author: Leandro von Werra; Lewis Tunstall; Quentin Gallouédec; Guilherme Penedo; Edward Beeching; Anton Lozhkov; Brigitte Tousignant; Daniel van Strien

Published: 2025-02-02T00:04:28Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [deepseek](<https://devfeed.tech/topics/deepseek.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [trl](<https://devfeed.tech/topics/trl.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [follow](<https://devfeed.tech/tags/follow.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [trl](<https://devfeed.tech/tags/trl.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This update reports early progress on the open-r1 project to reproduce the DeepSeek-R1 training pipeline and dataset. It covers matching MATH-500 evaluation results, the unusually long model responses and their GPU-memory implications, a public evaluation leaderboard, and the integration of GRPO into TRL with DeepSpeed and vLLM support.

### Source excerpt

It's been two weeks since the release of DeepSeek R1 and just a week since we started the open-r1 project to replicate the missing pieces, namely the training pipeline and the synthetic data. This post summarizes: - the progress of Open-R1 to replicate the DeepSeek-R1 pipeline and dataset - what we learned about DeepSeek-R1 and discussions around it - cool projects the community has built since the release of DeepSeek-R1 It should serve both as an update on the project and as a collection of...

## Introducing the Synthetic Data Generator - Build Datasets with Natural Language

DevFeed: [Introducing the Synthetic Data Generator - Build Datasets with Natural Language](<https://devfeed.tech/articles/introducing-the-synthetic-data-generator-build-datasets-with-natural-language-7497.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/synthetic-data-generator>)

Author: David Berenstein; Sara Han Díaz; Leire Aguirre; Daniel Vila; Ame Vi; ben burtenshaw

Published: 2024-12-16T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [argilla](<https://devfeed.tech/topics/argilla.md>), [distilabel](<https://devfeed.tech/topics/distilabel.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [API](<https://devfeed.tech/topics/api.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [argilla](<https://devfeed.tech/tags/argilla.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [distilabel](<https://devfeed.tech/tags/distilabel.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [llm](<https://devfeed.tech/tags/llm.md>), [rag](<https://devfeed.tech/tags/rag.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>)

### AI overview

Hugging Face introduces a Synthetic Data Generator that creates text classification and chat datasets from natural-language descriptions. The tool uses distilabel and the Hugging Face text-generation API, supports configurable models and providers, and can upload generated datasets after authentication.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## "Show Me What's Wrong!": Enhancing Fraud Detection Analysis by Combining Charts and Text

DevFeed: ["Show Me What's Wrong!": Enhancing Fraud Detection Analysis by Combining Charts and Text](<https://devfeed.tech/articles/show-me-what-s-wrong-enhancing-fraud-detection-analysis-by-combining-charts-and-text-26299.md>)

Original publisher: [Read original article](<https://medium.com/feedzaitech/show-me-whats-wrong-enhancing-fraud-detection-analysis-by-combining-charts-and-text-22ecfb342fb0?source=rss----e11168e7fe6b---4>)

Author: Beatriz Feliciano

Published: 2024-11-22T18:35:14Z

Content type: article

Language: en

Sources: [Feedzai](<https://devfeed.tech/sources/feedzai.md>)

Topics: [Transactions](<https://devfeed.tech/topics/transactions.md>), [data](<https://devfeed.tech/topics/data.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Tool](<https://devfeed.tech/topics/tool.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-analysis](<https://devfeed.tech/tags/data-analysis.md>), [data-visualization](<https://devfeed.tech/tags/data-visualization.md>), [financial-fraud](<https://devfeed.tech/tags/financial-fraud.md>), [fraud](<https://devfeed.tech/tags/fraud.md>), [fraud-detection](<https://devfeed.tech/tags/fraud-detection.md>), [fraud-investigation](<https://devfeed.tech/tags/fraud-investigation.md>), [image](<https://devfeed.tech/tags/image.md>), [interface](<https://devfeed.tech/tags/interface.md>), [research](<https://devfeed.tech/tags/research.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [transactions](<https://devfeed.tech/tags/transactions.md>)

### AI overview

The article describes a fraud-analysis tool that combines charts and text to help analysts review suspicious financial transactions. It explains that tabular review can make it difficult to identify patterns and anomalies within a 1-to-5-minute review window, and presents an interface using synthetic data to help analysts prioritize investigation areas.

### Source excerpt

Every year, millions of people fall victim to financial fraud. In 2023, the losses tied to this type of crime were estimated at US$159 billion just in the US, with some people losing all of their retirement savings to scammers. However, the impacts of this issue stretch beyond someone's finances. It can also impact a victim's life in many dimensions. Detecting and quickly acting upon suspicious transactions is essential to tackle this problem. Finding Fraud Through Data Tables To review the data of alerted transactions, analysts look at information in tabular format (similar to what is presented in Figure 1), scrolling through it to assess past activity patterns of the alerted person and comparing those with the alerted event. "How much money was spent on average on past transactions?" or "Is that significantly different from the amount on the current alert?" are some questions they might try to answer during their review. Figure 1: Image of a table that analysts typically use to review the data of alerted transactions. The issue with this approach is that finding groups of patterns and anomalies in tabular data can be overwhelming since it requires an increased cognitive load from analysts to interpret the data effectively. This becomes even more complex since these professionals must review and classify the alerted transaction in a short time -- between 1 and 5 minutes. Revamping the analysis To solve this problem, we present a tool that combines charts and text to guide the analysis of financial transactions. As presented in Figure 2, the tool (populated with synthetic data) is divided into three regions that provide different levels of information detail -- from the most high-level to the most detailed. The goal is that the analyst can scan the charts and prioritize their review towards specific areas of the alert. Figure 2: Proposed interface composed of multiple regions: the Knowledge Area Console (A) to detect suspicious areas of the analysis; the Knowledge Are

## A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality

DevFeed: [A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality](<https://devfeed.tech/articles/a-deepdive-into-aya-expanse-advancing-the-frontier-of-multilinguality-7114.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/aya-expanse>)

Author: John Dang; Shivalika Singh; Daniel D'souza; Arash Ahmadian

Published: 2024-10-24T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [aya](<https://devfeed.tech/topics/aya.md>), [cohere](<https://devfeed.tech/topics/cohere.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [aya](<https://devfeed.tech/tags/aya.md>), [cohere](<https://devfeed.tech/tags/cohere.md>), [comparison](<https://devfeed.tech/tags/comparison.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llms](<https://devfeed.tech/tags/llms.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [technical](<https://devfeed.tech/tags/technical.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

The Cohere For AI team presents Aya Expanse, an 8B and 32B multilingual model family designed to improve performance across languages. The article describes data arbitrage, multilingual preference training, safety tuning, model merging, evaluations across 23 languages, and the release of both models as open weights.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## The GANfather: Using Malicious GenAI Agents to Combat Money Laundering

DevFeed: [The GANfather: Using Malicious GenAI Agents to Combat Money Laundering](<https://devfeed.tech/articles/the-ganfather-using-malicious-genai-agents-to-combat-money-laundering-26300.md>)

Original publisher: [Read original article](<https://medium.com/feedzaitech/the-ganfather-using-malicious-genai-agents-to-combat-money-laundering-1666908113fc?source=rss----e11168e7fe6b---4>)

Author: Ricardo Ribeiro Pereira

Published: 2024-10-04T13:49:42Z

Content type: article

Language: en

Sources: [Feedzai](<https://devfeed.tech/sources/feedzai.md>)

Topics: [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [banking](<https://devfeed.tech/tags/banking.md>), [data](<https://devfeed.tech/tags/data.md>), [feedzai](<https://devfeed.tech/tags/feedzai.md>), [financial-sector](<https://devfeed.tech/tags/financial-sector.md>), [gans](<https://devfeed.tech/tags/gans.md>), [genai](<https://devfeed.tech/tags/genai.md>), [generative](<https://devfeed.tech/tags/generative.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [legacy](<https://devfeed.tech/tags/legacy.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [money-laundering](<https://devfeed.tech/tags/money-laundering.md>), [research](<https://devfeed.tech/tags/research.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [transactions](<https://devfeed.tech/tags/transactions.md>)

### AI overview

Feedzai describes a method that uses Generative AI to create synthetic data simulating realistic money-laundering activity. The generated examples are intended to support machine-learning approaches to detection and help identify vulnerabilities in banks' defenses, addressing the limited availability of labeled data.

### Source excerpt

Digital systems have become deeply integrated into many aspects of modern life, particularly within the financial sector. While digital banking simplifies day-to-day operations for clients, it also creates new opportunities for malicious actors to exploit these systems. As a result, money laundering has grown particularly prevalent due to this digital expansion. Banks are required to monitor for money laundering activities and issue alerts when suspicious transactions are detected. Typically, monitoring is performed by rules-based legacy systems. A better approach would be to use Machine Learning models, but these usually require labeled data to train, which are mostly unavailable in this use case. To tackle this problem, we employ advanced Generative AI (GenAI) techniques to generate synthetic data that simulates realistic money laundering activities. These synthetic examples help us identify vulnerabilities and strengthen the defense mechanisms used by banks and other financial institutions In this blog post, we will explore the method developed by Feedzai, which leverages GenAI to tackle the challenges of detecting and preventing money laundering in today's digital landscape. This blog post is the first of a series dedicated to the work done on GenAI by Feedzai Research in the last few years. Problem Statement First, let's briefly introduce the concepts behind money laundering and the difficulties that banks face when trying to prevent it. Money laundering is the process of concealing the origins of illegally obtained funds. Criminals cannot directly spend "dirty" money without risking exposure of their illegal activities. Therefore, they want to disguise the origins of funds before using them. Money laundering typically involves three stages: Placement: the money is introduced into the financial system, often in small amounts spread across various banks. Layering: the money launderer moves the funds through a series of transactions, typically across multiple fin

## Postgres 17 is Now Available on Neon

DevFeed: [Postgres 17 is Now Available on Neon](<https://devfeed.tech/articles/postgres-17-is-now-available-on-neon-5725.md>)

Original publisher: [Read original article](<https://neon.com/blog/postgres-17>)

Author: Heikki Linnakangas

Published: 2024-09-26T13:32:32Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [JSON](<https://devfeed.tech/topics/json.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [database](<https://devfeed.tech/tags/database.md>), [features](<https://devfeed.tech/tags/features.md>), [free](<https://devfeed.tech/tags/free.md>), [generate](<https://devfeed.tech/tags/generate.md>), [json](<https://devfeed.tech/tags/json.md>), [postgres](<https://devfeed.tech/tags/postgres.md>), [random](<https://devfeed.tech/tags/random.md>), [series](<https://devfeed.tech/tags/series.md>), [sql](<https://devfeed.tech/tags/sql.md>), [updates](<https://devfeed.tech/tags/updates.md>)

### AI overview

Neon announces that Postgres 17 is available to all Neon users, including Free-plan users. The article highlights MERGE with RETURNING support, ranged random integer generation, bulk test-data generation with generate_series(), and JSON_TABLE for working with JSON data.

### Source excerpt

Postgres 17 is out today, and we're here for it. In keeping with our version policy, the latest major version is now available for all Neon users including those on the Free plan. Get your Postgres 17 database and start exploring the latest features immediately. How to speed run...

[Next page](<https://devfeed.tech/topics/synthetic-data.md?cursor=WyIyMDI0LTA5LTI2VDEzOjMyOjMyKzAwOjAwIiwgIjA0MDhhODY2LWMzYmEtNDdiYy05YzFjLTIyOTA1NzhhYTFhZCJd>)