# NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset

DevFeed: [NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset](<https://devfeed.tech/articles/nvidia-releases-6-million-multi-lingual-reasoning-dataset-7390.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/multilingual-reasoning-v1>)

Author: Jane Polak Scowcroft; Dhruv Nathawani; Shuoyang Ding; Oleksii Kuchaiev; Vitaly Lavrukhin

Published: 2025-08-20T22:13:18Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [datasets](<https://devfeed.tech/topics/datasets.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [datasets](<https://devfeed.tech/tags/datasets.md>), [japanese](<https://devfeed.tech/tags/japanese.md>), [llama](<https://devfeed.tech/tags/llama.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [model-development](<https://devfeed.tech/tags/model-development.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open](<https://devfeed.tech/tags/open.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

## AI overview

NVIDIA announces a 6-million-example multilingual reasoning dataset translated into French, Spanish, German, Italian, and Japanese. The article also presents Nemotron Nano 2 9B, an edge-oriented model using a hybrid Transformer-Mamba architecture, configurable thinking budgets, and open model weights and training resources.

## Source excerpt

NVIDIA continues releasing permissive datasets in support of the open ecosystem with 6 Million Multilingual Reasoning Dataset. Continuing the success of the recent Nemotron Post-Training Dataset v1 release used in Llama Nemotron Super model, and our Llama Nemotron Post-Training Dataset release earlier this year, we're excited to release the reasoning dataset translated into five target languages: French, Spanish, German, Italian, and Japanese.