# MiniMax M3 Running Fastest on SambaCloud

DevFeed: [MiniMax M3 Running Fastest on SambaCloud](<https://devfeed.tech/articles/minimax-m3-running-fastest-on-sambacloud-81508.md>)

Original publisher: [Read original article](<https://sambanova.ai/blog/minimax-m3-running-fastest-on-sambacloud>)

Author: justin.woo@sambanova.ai (Justin Woo)

Published: 2026-08-24T22:41:42Z

Content type: release

Language: en

Sources: [Sambanova](<https://devfeed.tech/sources/sambanova.md>)

Topics: [Frontier AI](<https://devfeed.tech/topics/frontier-ai.md>), [Loop Engineering](<https://devfeed.tech/topics/loop-engineering.md>), [llm-reasoning](<https://devfeed.tech/topics/llm-reasoning.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [blog](<https://devfeed.tech/tags/blog.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [for-developers](<https://devfeed.tech/tags/for-developers.md>), [minimax](<https://devfeed.tech/tags/minimax.md>)

## AI overview

MiniMax M3 is an open-weight model presented for long-context agent workflows and software engineering. The article reports a 1M-token context window using MiniMax Sparse Attention, benchmark results for coding and agent tasks, and internal tests in which the model worked autonomously on a CUDA kernel optimization task and reproduced research experiments. It also describes the model's multimodal capabilities and availability on SambaCloud.

## Source excerpt

Frontier Coding, 1M Context, and Native Multimodality. Build Long-Horizon Agents on SambaCloud TL;DR MiniMax M3 runs fastest on SambaCloud, where it is available today for developers building long-horizon, long-context agents. On coding and agent benchmarks, MiniMax M3 scores 59.0% on SWE-Bench Pro, 66.0% on Terminal-Bench 2.1, and 74.2% on MCP Atlas, making it one of the strongest open-weight models available. MiniMax M3 introduces MiniMax Sparse Attention (MSA), delivering a 1M-token context window with more than 9x faster prefill and more than 15x faster decoding than the previous generation. In MiniMax's internal testing, MiniMax M3 ran autonomously for roughly 24 hours to optimize an FP8 GEMM CUDA kernel, lifting hardware peak utilization from 7.6% to 71.3%, a 9.4x speedup, with zero human intervention. MiniMax M3 is natively multimodal, trained on text, image, and video from step 0, so it can read charts, formulas, and screenshots and operate a desktop computer. MiniMax M3 is the successor to MiniMax M2.7 and their most capable model to date. It is built for long-horizon agent workflows, complex software engineering, and multimodal tasks that were previously the exclusive territory of closed-source frontier models. It hits 59.0% on SWE-Bench Pro, 66.0% on Terminal-Bench 2.1, and 74.2% on MCP Atlas, making it one of the strongest open-weight options available for coding and agent harnesses.