# Mamba

A state-space model architecture for linear-time sequence modeling with selective state spaces.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer

DevFeed: [Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer](<https://devfeed.tech/articles/developing-nemotron-3-5-lightning-nvfp4-with-qad-using-nvidia-model-optimizer-6811.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/developing-nemotron-3-5-lightning-nvfp4-with-qad-using-nvidia-model-optimizer/>)

Author: Tanya Lenz

Published: 2026-08-17T18:12:48Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [NVFP4](<https://devfeed.tech/topics/nvfp4.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [Post-training optimization](<https://devfeed.tech/topics/post-training-optimization.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [compute](<https://devfeed.tech/tags/compute.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [developers](<https://devfeed.tech/tags/developers.md>), [edge-computing](<https://devfeed.tech/tags/edge-computing.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [featured](<https://devfeed.tech/tags/featured.md>), [latency](<https://devfeed.tech/tags/latency.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [megatron](<https://devfeed.tech/tags/megatron.md>), [memory](<https://devfeed.tech/tags/memory.md>), [model](<https://devfeed.tech/tags/model.md>), [model-optimizer](<https://devfeed.tech/tags/model-optimizer.md>), [models](<https://devfeed.tech/tags/models.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvfp4](<https://devfeed.tech/tags/nvfp4.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open](<https://devfeed.tech/tags/open.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [speed](<https://devfeed.tech/tags/speed.md>), [training](<https://devfeed.tech/tags/training.md>), [training-ai-models](<https://devfeed.tech/tags/training-ai-models.md>)

### AI overview

This tutorial explains how quantization-aware distillation (QAD) creates the Nemotron 3.5 Lightning NVFP4 checkpoint using NVIDIA Model Optimizer. It covers post-training quantization, teacher-student distillation, and evaluation, showing how QAD can recover accuracy while reducing memory usage and increasing throughput.

### Source excerpt

Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models, developers can find...

## Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents

DevFeed: [Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents](<https://devfeed.tech/articles/introducing-nvidia-nemotron-3-nano-omni-long-context-multimodal-intelligence-for-documents-audio-and-video-agents-7395.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-intelligence>)

Author: Tuomas Rintamaki; Amala Sanjay Deshmukh; Nabin Mulepati; Collin McCarthy; Pritam Biswas; Arushi Goel; Alexandre Milesi; Danial Mohseni Taheri; Kateryna Chumachenko; Isabel Hulseman; Zhehuai Chen; Kara

Published: 2026-04-28T15:58:57Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [asr](<https://devfeed.tech/topics/asr.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [NVFP4](<https://devfeed.tech/topics/nvfp4.md>)

Tags: [alternatives](<https://devfeed.tech/tags/alternatives.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [computer-use](<https://devfeed.tech/tags/computer-use.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [speed](<https://devfeed.tech/tags/speed.md>)

### AI overview

NVIDIA introduces Nemotron 3 Nano Omni, an omni-modal model for document analysis, image reasoning, speech recognition, long audio-video understanding, computer use, and general reasoning. It combines a hybrid Mamba-Transformer Mixture-of-Experts backbone with vision and audio encoders, supports long multimodal contexts, and reports strong benchmark accuracy, throughput, reasoning speed, and system efficiency.

### Source excerpt

Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents - NVIDIA Nemotron 3 Nano Omni is a new omni-modal understanding model built for real-world document analysis, multiple image reasoning, automatic speech recognition, long audio-video understanding, agentic computer use, and general reasoning. - It extends the Nemotron multimodal line from a strong vision-language system to a broader text + image + video + audio model.

## Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture

DevFeed: [Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture](<https://devfeed.tech/articles/introducing-falcon-h1-arabic-pushing-the-boundaries-of-arabic-language-ai-with-hybrid-architecture-7509.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tiiuae/falcon-h1-arabic>)

Author: Basma Boussaha; Mohammed Alyafeai; Ahmed Alzubaidi; Leen AlQadi; Shaikha Alsuwaidi; Omar saif alkaabi; Hamza Alobeidli; Hakim Hacid

Published: 2026-01-05T09:16:51Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Natural language processing](<https://devfeed.tech/topics/nlp.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [arabic](<https://devfeed.tech/tags/arabic.md>), [blog](<https://devfeed.tech/tags/blog.md>), [building](<https://devfeed.tech/tags/building.md>), [community](<https://devfeed.tech/tags/community.md>), [design](<https://devfeed.tech/tags/design.md>), [developers](<https://devfeed.tech/tags/developers.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [innovation](<https://devfeed.tech/tags/innovation.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [learning](<https://devfeed.tech/tags/learning.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

The article introduces Falcon-H1-Arabic, a family of 3B, 7B, and 34B Arabic language models. It describes a hybrid architecture that combines Mamba State Space Models with Transformer attention in parallel, aiming to improve long-context coherence, reasoning, efficiency, and deployment across edge devices and enterprise applications.

### Source excerpt

A Blog post by Technology Innovation Institute on Hugging Face

## NVIDIA Nemotron Nano 2 VL 12B: Architecture and Improvements for Multimodal Reasoning

DevFeed: [NVIDIA Nemotron Nano 2 VL 12B: Architecture and Improvements for Multimodal Reasoning](<https://devfeed.tech/articles/small-vlms-will-soon-compete-with-frontier-ai-models-10x-their-size-35020.md>)

Original publisher: [Read original article](<https://read.theaimerge.com/p/small-vlms-will-soon-compete-with>)

Author: Alex Razvant

Published: 2025-11-22T14:58:39Z

Content type: article

Language: en

Sources: [Neural Bits](<https://devfeed.tech/sources/neural-bits.md>)

Topics: [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [context window](<https://devfeed.tech/topics/context-window.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>)

Tags: [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [ml-engineering](<https://devfeed.tech/tags/ml-engineering.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>)

### AI overview

This article examines NVIDIA's Nemotron Nano 2 VL 12B, an open small vision-language model designed for document understanding, long-video comprehension, and multimodal reasoning. It discusses the model's architecture, encoders, training and inference improvements, reasoning modes, and reported benchmark performance, including a 128k-token context window.

### Source excerpt

What makes NVIDIA Nemotron Nano 2 VL a breakthrough in small, fast, long-context visual reasoning.

## Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models

DevFeed: [Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models](<https://devfeed.tech/articles/apriel-h1-the-surprising-key-to-distilling-efficient-reasoning-models-7043.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ServiceNow-AI/apriel-h1>)

Author: Torsten Scholak; Oleksiy Ostapenko; Raymond Li; Luke Kumar; Joel Lamy-Poirier

Published: 2025-11-19T05:19:07Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Mamba](<https://devfeed.tech/topics/mamba.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [models](<https://devfeed.tech/tags/models.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

The article explains how the Apriel-H1 team distilled a strong 15B reasoning model into more efficient hybrids using Mamba layers. Distillation on pretraining data and generic SFT data reduced reasoning quality, while high-quality reasoning traces from the teacher's SFT dataset preserved performance. The flagship model achieved roughly 2.1x throughput with minimal quality loss across several benchmarks.

### Source excerpt

When MiniMax published their M2 post-mortem in October explaining why they abandoned efficient attention at 230B scale, the narrative briefly became "efficient attention is dead." Within days, Kimi Linear proved otherwise. The real lesson: it depends on your constraints. Our constraint was simple: we had a strong 15B reasoning model and needed to make it efficient without starting over. No infinite compute for 20T-token pretraining. No luxury of architectural co-design from day one.

## NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset

DevFeed: [NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset](<https://devfeed.tech/articles/nvidia-releases-6-million-multi-lingual-reasoning-dataset-7390.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/multilingual-reasoning-v1>)

Author: Jane Polak Scowcroft; Dhruv Nathawani; Shuoyang Ding; Oleksii Kuchaiev; Vitaly Lavrukhin

Published: 2025-08-20T22:13:18Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [datasets](<https://devfeed.tech/topics/datasets.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [datasets](<https://devfeed.tech/tags/datasets.md>), [japanese](<https://devfeed.tech/tags/japanese.md>), [llama](<https://devfeed.tech/tags/llama.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [model-development](<https://devfeed.tech/tags/model-development.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open](<https://devfeed.tech/tags/open.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

NVIDIA announces a 6-million-example multilingual reasoning dataset translated into French, Spanish, German, Italian, and Japanese. The article also presents Nemotron Nano 2 9B, an edge-oriented model using a hybrid Transformer-Mamba architecture, configurable thinking budgets, and open model weights and training resources.

### Source excerpt

NVIDIA continues releasing permissive datasets in support of the open ecosystem with 6 Million Multilingual Reasoning Dataset. Continuing the success of the recent Nemotron Post-Training Dataset v1 release used in Llama Nemotron Super model, and our Llama Nemotron Post-Training Dataset release earlier this year, we're excited to release the reasoning dataset translated into five target languages: French, Spanish, German, Italian, and Japanese.

## Bamba: Inference-Efficient Hybrid Mamba2 Model

DevFeed: [Bamba: Inference-Efficient Hybrid Mamba2 Model](<https://devfeed.tech/articles/bamba-inference-efficient-hybrid-mamba2-model-7119.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/bamba>)

Author: LINSONG CHU; Divya Kumari; Tri Dao; Albert Gu; Raghu Ganti; Dakshi Agrawal; Mudhakar Srivatsa; Davis Wertheimer; Yu Chin Fabian Lim; Antoni Viros

Published: 2024-12-18T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Mamba](<https://devfeed.tech/topics/mamba.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [ibm](<https://devfeed.tech/topics/ibm.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [community](<https://devfeed.tech/tags/community.md>), [data](<https://devfeed.tech/tags/data.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [model](<https://devfeed.tech/tags/model.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reproducibility](<https://devfeed.tech/tags/reproducibility.md>), [research](<https://devfeed.tech/tags/research.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

The article introduces Bamba-9B, an inference-efficient hybrid Mamba2 model trained by IBM, Princeton, CMU, and UIUC on open data. It reports higher throughput and lower latency than standard transformers in vLLM, and releases training resources, checkpoints, and reproducibility materials for community experimentation.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Welcome to the Falcon 3 Family of Open Models!

DevFeed: [Welcome to the Falcon 3 Family of Open Models!](<https://devfeed.tech/articles/welcome-to-the-falcon-3-family-of-open-models-7190.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/falcon3>)

Author: Falcon LLM TII UAE

Published: 2024-12-17T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Open Source](<https://devfeed.tech/topics/open-source.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Post-training optimization](<https://devfeed.tech/topics/post-training-optimization.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [code](<https://devfeed.tech/tags/code.md>), [community](<https://devfeed.tech/tags/community.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [models](<https://devfeed.tech/tags/models.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [releases](<https://devfeed.tech/tags/releases.md>), [research](<https://devfeed.tech/tags/research.md>), [science](<https://devfeed.tech/tags/science.md>), [techniques](<https://devfeed.tech/tags/techniques.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Falcon3 is a family of open models designed to improve science, mathematics, and code capabilities across small and medium language-model scales. The article describes transformer-based and pure state-space variants, training and scaling methods, compact models produced with pruning and knowledge distillation, multiple quantized formats, and benchmark results against comparable models.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Codestral Mamba

DevFeed: [Codestral Mamba](<https://devfeed.tech/articles/codestral-mamba-6988.md>)

Original publisher: [Read original article](<https://mistral.ai/news/codestral-mamba/>)

Published: 2024-07-16T08:00:00Z

Content type: news

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [Mamba](<https://devfeed.tech/topics/mamba.md>), [code productivity](<https://devfeed.tech/topics/code-productivity.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Local AI](<https://devfeed.tech/topics/local-ai.md>), [SDK](<https://devfeed.tech/topics/sdk.md>), [TensorRT-LLM](<https://devfeed.tech/topics/tensorrt-llm.md>), [llama.cpp](<https://devfeed.tech/topics/llama-cpp.md>)

Tags: [efficiency](<https://devfeed.tech/tags/efficiency.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [mixtral](<https://devfeed.tech/tags/mixtral.md>), [open](<https://devfeed.tech/tags/open.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [tensorrt-llm](<https://devfeed.tech/tags/tensorrt-llm.md>)

### AI overview

Mistral introduces Codestral Mamba, a freely usable and distributable code-focused model based on the Mamba architecture. It offers linear-time inference, supports very long sequences, and is intended for efficient local code-assistant use. The article describes deployment through the mistral-inference SDK and TensorRT-LLM, with anticipated llama.cpp support, and notes its Apache 2.0 license.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.