# Ollama's highest performance on Apple Silicon yet with MLX

DevFeed: [Ollama's highest performance on Apple Silicon yet with MLX](<https://devfeed.tech/articles/ollama-s-highest-performance-on-apple-silicon-yet-with-mlx-83389.md>)

Original publisher: [Read original article](<https://ollama.com/blog/mlx-performance>)

Published: 2026-06-11T00:00:00Z

Content type: release

Language: en

Sources: [Ollama](<https://devfeed.tech/sources/ollama.md>)

Topics: [MLX](<https://devfeed.tech/topics/mlx.md>), [Ollama](<https://devfeed.tech/topics/ollama.md>), [scaling laws](<https://devfeed.tech/topics/scaling-laws.md>), [GPU optimization](<https://devfeed.tech/topics/gpu-optimization.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [agent-workflows](<https://devfeed.tech/tags/agent-workflows.md>), [apple-silicon](<https://devfeed.tech/tags/apple-silicon.md>), [coding-agent](<https://devfeed.tech/tags/coding-agent.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [local](<https://devfeed.tech/tags/local.md>), [mlx](<https://devfeed.tech/tags/mlx.md>), [nvfp4](<https://devfeed.tech/tags/nvfp4.md>), [ollama](<https://devfeed.tech/tags/ollama.md>), [performance](<https://devfeed.tech/tags/performance.md>)

## AI overview

Ollama's updated MLX engine for Apple Silicon supports NVIDIA's NVFP4 format, with the release describing higher quality responses, faster output, and lower memory use. It also adds selective model-state snapshots to reduce repeated prompt processing in agent workflows, including handoffs, branching, and retries.

## Source excerpt

Ollama's MLX engine has been updated to deliver its highest performance on Apple Silicon yet. Models output higher quality responses, respond faster, and use less memory.