# mixture-of-experts

Published articles for mixture-of-experts.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Cohere's new translation model is open weights -- but not for commercial use

DevFeed: [Cohere's new translation model is open weights -- but not for commercial use](<https://devfeed.tech/articles/cohere-s-new-translation-model-is-open-weights-but-not-for-commercial-use-8474.md>)

Original publisher: [Read original article](<https://thenewstack.io/cohere-translation-commercial-licensing/>)

Author: Meredith Shubel

Published: 2026-09-11T17:50:11Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [moe](<https://devfeed.tech/topics/moe.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-strategy](<https://devfeed.tech/tags/ai-strategy.md>), [api](<https://devfeed.tech/tags/api.md>), [cohere](<https://devfeed.tech/tags/cohere.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [production](<https://devfeed.tech/tags/production.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

Cohere released North Small Translate 1.0 as open weights under CC BY-NC 4.0, allowing download, evaluation, and study but requiring a commercial agreement for production use. Commercial deployment requires a license and use of Cohere's managed Model Vault platform.

### Source excerpt

This week, Cohere released North Small Translate 1.0 under a CC BY-NC 4.0 license: the weights are there to download, The post Cohere's new translation model is open weights -- but not for commercial use appeared first on The New Stack.

## How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra

DevFeed: [How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra](<https://devfeed.tech/articles/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra-6840.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/>)

Author: Elizabeth Goodman

Published: 2026-09-10T16:55:32Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [build-ai-agents](<https://devfeed.tech/tags/build-ai-agents.md>), [cache](<https://devfeed.tech/tags/cache.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [latency](<https://devfeed.tech/tags/latency.md>), [mamba](<https://devfeed.tech/tags/mamba.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nim](<https://devfeed.tech/tags/nim.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [optimization](<https://devfeed.tech/tags/optimization.md>)

### AI overview

The article explains how NVIDIA NIM bundles serving optimizations to improve throughput for Nemotron 3 Ultra while meeting latency targets on GPU infrastructure.

### Source excerpt

Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as...

## How XPUs Meet a World-Class AI Factory

DevFeed: [How XPUs Meet a World-Class AI Factory](<https://devfeed.tech/articles/how-xpus-meet-a-world-class-ai-factory-6958.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/nvlink-fusion-xpu-ai-factory/>)

Author: Jesse Clayton

Published: 2026-08-24T15:00:54Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Network](<https://devfeed.tech/topics/network.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ecosystem](<https://devfeed.tech/tags/ecosystem.md>), [ethernet](<https://devfeed.tech/tags/ethernet.md>), [gb300-nvl72](<https://devfeed.tech/tags/gb300-nvl72.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia-dsx](<https://devfeed.tech/tags/nvidia-dsx.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [software](<https://devfeed.tech/tags/software.md>), [time](<https://devfeed.tech/tags/time.md>), [xpu](<https://devfeed.tech/tags/xpu.md>)

### AI overview

The article explains how NVLink Fusion combines custom XPUs with NVIDIA's established AI infrastructure to help build semi-custom AI factories. It focuses on scale-up networking, performance, resiliency, telemetry, platform maturity, and the economics of large-scale AI workloads.

### Source excerpt

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. Hyperscalers and AI-native companies building custom XPUs must consider [...]

## Thinking Machines' Inkling: Architecture and Customization Choices

DevFeed: [Thinking Machines' Inkling: Architecture and Customization Choices](<https://devfeed.tech/articles/the-new-american-ai-model-designed-to-be-customized-17999.md>)

Original publisher: [Read original article](<https://blog.bytebytego.com/p/the-new-american-ai-model-designed>)

Author: ByteByteGo

Published: 2026-08-18T15:30:36Z

Content type: article

Language: en

Sources: [ByteByteGo](<https://devfeed.tech/sources/bytebytego.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [context window](<https://devfeed.tech/topics/context-window.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [interfaces](<https://devfeed.tech/topics/interfaces.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [model](<https://devfeed.tech/tags/model.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

The article examines the architecture and design choices behind Thinking Machines' Inkling model, including its mixture-of-experts structure, local and global attention, position encoding, multimodal inputs, and adjustable thinking effort. It also notes that Inkling is the company's first model trained from scratch and that its weights are available on Hugging Face under an Apache 2.0 license.

### Source excerpt

In this article, we will work through the various choices Thinking Machines made while building Inkling.

## 🗓 This Week In AI Research (1-7 August 26)

DevFeed: [🗓 This Week In AI Research (1-7 August 26)](<https://devfeed.tech/articles/this-week-in-ai-research-1-7-august-26-18282.md>)

Original publisher: [Read original article](<https://www.intoai.pub/p/this-week-in-ai-research-1-7-august>)

Author: Dr. Ashish Bamania

Published: 2026-08-13T19:29:25Z

Content type: article

Language: en

Sources: [Into AI](<https://devfeed.tech/sources/into-ai.md>)

Topics: [AI Research](<https://devfeed.tech/topics/ai-research.md>), [releases](<https://devfeed.tech/topics/releases.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [qwen](<https://devfeed.tech/topics/qwen.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [algorithm](<https://devfeed.tech/tags/algorithm.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [releases](<https://devfeed.tech/tags/releases.md>)

### AI overview

A weekly roundup of AI research and model releases. It highlights Pathway, Bielik AI, and NYU's BDH-CQ reasoning model, which uses in-context learning with recurrent memory and latent-state reasoning, reports ARC-AGI-1 cost-efficiency results, and describes Alibaba's Qwen3.8-Max release and the U-OPSD self-distillation algorithm.

### Source excerpt

The top 10 AI research papers and releases that you must know about this week.

## How Keras 3 Helped Modernise Expedia Group's Lodging Ranking Stack

DevFeed: [How Keras 3 Helped Modernise Expedia Group's Lodging Ranking Stack](<https://devfeed.tech/articles/how-keras-3-helped-modernise-expedia-group-s-lodging-ranking-stack-19734.md>)

Original publisher: [Read original article](<https://medium.com/expedia-group-tech/how-keras-3-helped-modernise-expedia-groups-lodging-ranking-stack-7fec96f052fd?source=rss----38998a53046f---4>)

Author: Conor Worthington

Published: 2026-08-11T11:01:02Z

Content type: article

Language: en

Sources: [Expedia](<https://devfeed.tech/sources/expedia.md>)

Topics: [Keras](<https://devfeed.tech/topics/keras.md>), [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [data-science](<https://devfeed.tech/tags/data-science.md>), [features](<https://devfeed.tech/tags/features.md>), [framework](<https://devfeed.tech/tags/framework.md>), [improvements](<https://devfeed.tech/tags/improvements.md>), [inference](<https://devfeed.tech/tags/inference.md>), [keras](<https://devfeed.tech/tags/keras.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llms](<https://devfeed.tech/tags/llms.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [model-training](<https://devfeed.tech/tags/model-training.md>), [neural-networks](<https://devfeed.tech/tags/neural-networks.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [recommendation-system](<https://devfeed.tech/tags/recommendation-system.md>)

### AI overview

Expedia Group describes modernising its lodging-ranking stack around Keras 3. The changes included rewriting parts of its pipelines, making model training 30% faster and reducing P99 inference latency by two-thirds. The article discusses limitations of the previous Keras 2, eager-training and combined-serving setup, along with the use of newer APIs and accelerator-oriented optimisations.

### Source excerpt

Expedia Group Technology -- DataWhat happened when we treated a framework migration as an architecture modernisation -- and cut P99 inference latency by two-thirdsSt Paul's and millennium bridge, London Expedia Group™ has always been a market leader in providing personalised search experiences for travellers. As our ranking models evolved, we saw an opportunity not just to migrate to Keras 3, but to modernise the broader stack around it so we can better serve travellers. This led us to rewrite key parts of our pipelines that made model training 30% faster and cut P99 inference latency by two-thirds. Our main focus in this blog is to discuss our improvements to lodging ranking -- our service which handles users' search requests and returns a personalised property ranking on top of a lightweight candidate generator. This service is frequently retrained to improve customer experience as we get new signals over time. Example of personalised lodging ranking on Expedia search result page for a London search As such, our velocity to make changes needs to be fast, but more importantly we need to be able to build models which are state-of-the-art, enabling customers to easily find and book the most relevant property for their trip. The problem with Keras 2, eager training and combined serving The challenge was not that the old stack was broken. It was that it had gradually become a limiting factor. On the training side, we were carrying dependencies on older components and missing out on modern Keras APIs, new optimisers, new layers and cleaner support for accelerator-oriented training. On the serving side, we had a path that was simple and operationally familiar, but not especially well suited to aggressively optimised inference for heavier ranking architectures. This blockage on utilising new layers and optimisers has become more pronounced as LLMs see blistering amounts of innovation. Naturally, these new features are only available in Keras 3 or as separate backends from ou

## 34 Amazon Research Awards Build on Trainium recipients announced

DevFeed: [34 Amazon Research Awards Build on Trainium recipients announced](<https://devfeed.tech/articles/34-amazon-research-awards-build-on-trainium-recipients-announced-7614.md>)

Original publisher: [Read original article](<https://www.amazon.science/research-awards/latest-news/34-amazon-research-awards-build-on-trainium-recipients-announced>)

Author: Amazon Research Awards team

Published: 2026-08-05T15:00:00Z

Content type: news

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [responsible-ai](<https://devfeed.tech/topics/responsible-ai.md>), [AWS AI chips](<https://devfeed.tech/topics/aws-ai-chips.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [academic-ai-funding](<https://devfeed.tech/tags/academic-ai-funding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [ai-research-grants](<https://devfeed.tech/tags/ai-research-grants.md>), [ai-safety-and-alignment](<https://devfeed.tech/tags/ai-safety-and-alignment.md>), [amazon-research-awards](<https://devfeed.tech/tags/amazon-research-awards.md>), [ara](<https://devfeed.tech/tags/ara.md>), [aws-ai-chips](<https://devfeed.tech/tags/aws-ai-chips.md>), [aws-trainium](<https://devfeed.tech/tags/aws-trainium.md>), [build-on-trainium](<https://devfeed.tech/tags/build-on-trainium.md>), [data](<https://devfeed.tech/tags/data.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [inference](<https://devfeed.tech/tags/inference.md>), [internal-ara-program-updates](<https://devfeed.tech/tags/internal-ara-program-updates.md>), [llm](<https://devfeed.tech/tags/llm.md>), [machine-learning-research](<https://devfeed.tech/tags/machine-learning-research.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [research](<https://devfeed.tech/tags/research.md>), [responsible-ai](<https://devfeed.tech/tags/responsible-ai.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Amazon announces 34 recipients of its Build on Trainium program, a $110 million credit initiative supporting AI research and university education. The awards fund work in areas including Responsible AI, language models, synthetic data, distributed systems, model architectures, libraries, and optimization on AWS Trainium.

### Source excerpt

Amazon announces 34 recipients of the Build on Trainium program, a $110 million credit initiative supporting AI research at 30 universities including Stanford, UC Berkeley, UIUC, UCLA, CMU, and MIT, with a focus on Responsible AI.

## Laguna S 2.1 is now available on AI Gateway

DevFeed: [Laguna S 2.1 is now available on AI Gateway](<https://devfeed.tech/articles/laguna-s-2-1-is-now-available-on-ai-gateway-995.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/laguna-s-2-1-is-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-21T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [AI Models](<https://devfeed.tech/topics/ai-models.md>), [agentic-coding](<https://devfeed.tech/topics/agentic-coding.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [context window](<https://devfeed.tech/topics/context-window.md>), [API](<https://devfeed.tech/topics/api.md>), [MLOps](<https://devfeed.tech/topics/mlops.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>)

Tags: [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [api-keys](<https://devfeed.tech/tags/api-keys.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [cost](<https://devfeed.tech/tags/cost.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [mlops](<https://devfeed.tech/tags/mlops.md>), [routing](<https://devfeed.tech/tags/routing.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [support](<https://devfeed.tech/tags/support.md>)

### AI overview

Poolside's Laguna S 2.1 is now available through Vercel AI Gateway in free and paid versions, with context windows of 256K and 1M tokens. The open-weight Mixture-of-Experts model supports thinking and no-thinking modes and is designed for agentic coding, long-running tasks, browser tooling, MLOps pipelines, and AI research.

### Source excerpt

Laguna S 2.1 from Poolside is now available on AI Gateway. There are 2 versions of the model available: Free version (256K context window): poolside/laguna-s-2.1-free Paid version (1M context window): poolside/laguna-s-2.1 Laguna S 2.1 is an open-weight Mixture-of-Experts model that supports a context window of up to 1M tokens and runs in thinking and no-thinking modes. The model specializes in agentic coding and long-running tasks, including writing and debugging code, running tests, building browser-based tooling, and working on MLOps pipelines and AI research. In thinking mode, Laguna S 2.1 reports 70.2% on Terminal-Bench 2.1, 78.5% on SWE-bench Multilingual, and 59.4% on SWE-Bench Pro. To use Laguna S 2.1, set model to poolside/laguna-s-2.1-free or poolside/laguna-s-2.1 in the AI SDK: AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Laguna S 2.1 in the model playground. Read more

## Welcome Inkling by Thinking Machines

DevFeed: [Welcome Inkling by Thinking Machines](<https://devfeed.tech/articles/welcome-inkling-by-thinking-machines-7502.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/thinkingmachines-inkling>)

Author: ben burtenshaw; merve; Pedro Cuenca; Aritra Roy Gosthipaty; Andres Marafioti

Published: 2026-07-15T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [sglang](<https://devfeed.tech/topics/sglang.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [audio](<https://devfeed.tech/tags/audio.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [generation](<https://devfeed.tech/tags/generation.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [moe](<https://devfeed.tech/tags/moe.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

Thinking Machines Lab's Inkling is presented as a large open multimodal language model that accepts image, text, and audio inputs. The article covers its mixture-of-experts architecture, million-token context window, reasoning across modalities, fine-tuning use cases, model variants, and deployment through Hugging Face Inference Endpoints and inference frameworks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Native-speed vLLM transformers modeling backend

DevFeed: [Native-speed vLLM transformers modeling backend](<https://devfeed.tech/articles/native-speed-vllm-transformers-modeling-backend-7364.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/native-speed-vllm-transformers-backend>)

Author: Harry Mellor; Lysandre

Published: 2026-07-08T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vllm](<https://devfeed.tech/topics/vllm.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [batching](<https://devfeed.tech/tags/batching.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [mlx](<https://devfeed.tech/tags/mlx.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This article describes the transformers modeling backend for vLLM and its role in serving Hugging Face models without requiring model authors to port their implementations. It explains how transformers supplies modeling code while vLLM contributes optimized inference techniques, and outlines comparisons with vLLM's native implementations across several Qwen3 model configurations. It also discusses inference-performance factors including GPU parallelization, compilation, fused kernels, and custom attention kernels.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Groq on Endless Compute, Inside Claude's Mind, and GLM-5.2 Open Weights - The Tokenizer Edition #32

DevFeed: [Groq on Endless Compute, Inside Claude's Mind, and GLM-5.2 Open Weights - The Tokenizer Edition #32](<https://devfeed.tech/articles/groq-on-endless-compute-inside-claude-s-mind-and-glm-5-2-open-weights-the-tokenizer-edition-32-18337.md>)

Original publisher: [Read original article](<https://newsletter.artofsaience.com/p/groq-on-endless-compute-inside-claudes>)

Author: Sairam Sundaresan

Published: 2026-06-21T16:45:54Z

Content type: article

Language: en

Sources: [Gradient Ascent](<https://devfeed.tech/sources/gradient-ascent.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [groq](<https://devfeed.tech/topics/groq.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [github](<https://devfeed.tech/tags/github.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>)

### AI overview

The Tokenizer Edition #32 curates AI and machine learning resources covering model interpretability, compute demand, open-weight models, multimodal video processing, speculative decoding, reinforcement learning, agent testing, and tools for cheaper or safer inference.

### Source excerpt

This week's most valuable AI resources

## Nemotron 3 Ultra now available on AI Gateway

DevFeed: [Nemotron 3 Ultra now available on AI Gateway](<https://devfeed.tech/articles/nemotron-3-ultra-now-available-on-ai-gateway-1025.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/nemotron-3-ultra-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-06-04T07:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [API](<https://devfeed.tech/topics/api.md>), [context window](<https://devfeed.tech/topics/context-window.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [performance](<https://devfeed.tech/tags/performance.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [tool](<https://devfeed.tech/tags/tool.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Nemotron 3 Ultra from Nvidia is now available through Vercel AI Gateway. The open Mixture-of-Experts reasoning model supports long-running, multi-turn agent workflows with a 1M token context window, tool use, delegation, and error recovery. The gateway provides model access, usage and cost tracking, retries, failover, provider sorting, and performance optimizations.

### Source excerpt

Nemotron 3 Ultra from Nvidia is now available on Vercel AI Gateway. Nemotron 3 Ultra is an open Mixture-of-Experts reasoning model built for orchestrating long-running agent workflows, with a 1M token context window. The model targets multi-turn agent workflows: planning, tool use, sub-agent delegation, and error recovery. Throughput reaches up to 350 tokens per second, with up to 30% lower cost on agentic tasks. To use Nemotron 3 Ultra, set model to nvidia/nemotron-3-ultra-550b-a55b in the AI SDK. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, dynamic provider sorting by latency and cost, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Learn more about AI Gateway, view the AI Gateway model leaderboard or try it in our model playground. Read more

## Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains

DevFeed: [Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains](<https://devfeed.tech/articles/introducing-mellum2-a-12b-mixture-of-experts-model-by-jetbrains-7013.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/JetBrains/mellum2-launch>)

Author: Nikita Pavlichenko

Published: 2026-06-01T15:45:17Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [code-completion](<https://devfeed.tech/topics/code-completion.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [coding](<https://devfeed.tech/tags/coding.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency-inference](<https://devfeed.tech/tags/low-latency-inference.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [model](<https://devfeed.tech/tags/model.md>), [rag](<https://devfeed.tech/tags/rag.md>), [routing](<https://devfeed.tech/tags/routing.md>), [software](<https://devfeed.tech/tags/software.md>)

### AI overview

JetBrains introduces Mellum2, a 12B-parameter Mixture-of-Experts model trained from scratch on natural language and code. It activates 2.5B parameters per token to provide efficient, low-latency inference for routing, retrieval-augmented generation, summarization, sub-agent tasks, and high-throughput coding features. The model is released under the Apache 2.0 license and is reported to deliver competitive benchmark performance with more than twice the inference speed of similarly sized models.

### Source excerpt

- Mellum2 is a 12B-parameter Mixture-of-Experts model trained from scratch on natural language and code. - The model activates only 2.5B parameters per token, making it efficient for high-throughput, low-latency inference. Mellum2 is can be used for routing, RAG, summarization, sub-agents, high-throughput coding features, and private deployments. - It is released under the Apache 2.0 license.

## Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries

DevFeed: [Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries](<https://devfeed.tech/articles/keep-the-tokens-flowing-lessons-from-16-open-source-rl-libraries-7109.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/async-rl-training-landscape>)

Author: Amine Dirhoussi; Quentin Gallouédec; Kashif Rasul; Lewis Tunstall; Edward Beeching; Albert Villanova del Moral; Nouamane Tazi; Leandro von Werra; Sergio Paniego

Published: 2026-03-10T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [distributed-training](<https://devfeed.tech/topics/distributed-training.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [NCCL](<https://devfeed.tech/topics/nccl.md>), [post-training](<https://devfeed.tech/topics/post-training.md>), [lora](<https://devfeed.tech/topics/lora.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [grpo](<https://devfeed.tech/tags/grpo.md>), [llm](<https://devfeed.tech/tags/llm.md>), [lora](<https://devfeed.tech/tags/lora.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [nccl](<https://devfeed.tech/tags/nccl.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [rl](<https://devfeed.tech/tags/rl.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

This article surveys 16 open-source libraries for asynchronous reinforcement-learning training. It explains how separating inference and training across GPU pools, using rollout buffers, and synchronizing weights asynchronously can reduce training-GPU idle time. The comparison covers orchestration, buffering, weight synchronization, staleness management, partial rollouts, LoRA, and distributed-training backends, highlighting Ray, NCCL broadcasts, limited LoRA support, and distributed MoE as an emerging differentiator.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Mixture of Experts (MoEs) in Transformers

DevFeed: [Mixture of Experts (MoEs) in Transformers](<https://devfeed.tech/articles/mixture-of-experts-moes-in-transformers-7358.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/moe-transformers>)

Author: Aritra Roy Gosthipaty; Pedro Cuenca; merve; Ilyas Moutawwakil; Arthur Zucker; Sergio Paniego; Pablo Montalvo

Published: 2026-02-26T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [quantization](<https://devfeed.tech/topics/quantization.md>)

Tags: [compute](<https://devfeed.tech/tags/compute.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [router](<https://devfeed.tech/tags/router.md>), [speed](<https://devfeed.tech/tags/speed.md>), [systems](<https://devfeed.tech/tags/systems.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [transformers-internal](<https://devfeed.tech/tags/transformers-internal.md>)

### AI overview

This article explains how Mixture of Experts models extend Transformer-based language models by replacing selected dense feed-forward layers with multiple learnable expert subnetworks. A router activates only a subset of experts for each token, allowing high total model capacity while reducing the number of active parameters used during inference. It also discusses compute efficiency, inference speed, memory requirements, parallelization, and quantization.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Technical Deep Dive: How DigitalOcean and AMD Delivered a 2x Production Inference Performance Increase for Character.ai

DevFeed: [Technical Deep Dive: How DigitalOcean and AMD Delivered a 2x Production Inference Performance Increase for Character.ai](<https://devfeed.tech/articles/technical-deep-dive-how-digitalocean-and-amd-delivered-a-2x-production-inference-performance-increase-for-character-ai-19948.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/technical-deep-dive-character-ai-amd>)

Author: Karnik Modi

Published: 2026-01-13T12:30:00Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [GPU optimization](<https://devfeed.tech/topics/gpu-optimization.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>), [qwen](<https://devfeed.tech/topics/qwen.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [news](<https://devfeed.tech/tags/news.md>), [performance](<https://devfeed.tech/tags/performance.md>), [platforms](<https://devfeed.tech/tags/platforms.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

This technical deep dive describes how Character.ai, AMD, and DigitalOcean optimized inference for the Qwen3-235B Instruct FP8 model on AMD Instinct MI300X and MI325X GPUs. The reported optimizations, including parallelization, FP8 execution paths, optimized kernels, topology-aware allocation, and Kubernetes orchestration, produced up to a 2x improvement in production request throughput under stated latency and concurrency constraints.

### Source excerpt

Background: How Character.ai worked with DigitalOcean and AMD to optimize performance Character.ai, a leading AI entertainment platform with about 20 million worldwide users, wanted to optimize GPU performance and achieve lower inference costs for its application, which requires low-latency performance at large scale. They approached DigitalOcean and AMD in order to achieve this goal. Working closely together, the Character.ai, AMD, and DigitalOcean teams optimized AMD Instinct™ MI300X and MI325X GPU platforms, resulting in a 2x production inference throughput. In optimized configurations, DigitalOcean delivered high request density per node while maintaining exceptional p90 responsiveness for initial token and sustained token generation throughput, outperforming prior deployments on generic, non-optimized GPU infrastructure. These gains were achieved through platform-level optimizations, including clever parallelization strategies for large Mixture-of-Experts models, efficient FP8 execution paths, optimized kernels with AITER, topology-aware GPU allocation, and production-ready Kubernetes orchestration through DigitalOcean Kubernetes (DOKS). Together, these capabilities allowed Character.ai to scale inference predictably without increasing operational burden. In this post, we will explore the specific orchestration and tuning strategies that made these gains possible. Technical deep dive overview Character.ai leverages multiple models like Qwen, Mistral and more to power their applications. This document is focused on how we optimized the Qwen3-235B Instruct FP8 model on a cluster of DigitalOcean featuring AMD Instinct GPUs. This workload was migrated from a generic, non-optimized setup on other providers to AMD Instinct™ MI325X platform on DigitalOcean, and following the outlined optimizations we were able to achieve up to a 2x improvement in request throughput (QPS) under strict latency and concurrency constraints. The Character.ai team has a demanding workload,

## Introducing Mistral 3

DevFeed: [Introducing Mistral 3](<https://devfeed.tech/articles/introducing-mistral-3-7040.md>)

Original publisher: [Read original article](<https://mistral.ai/news/mistral-3/>)

Published: 2025-12-02T16:00:00Z

Content type: release

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [NVFP4](<https://devfeed.tech/topics/nvfp4.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [Blackwell](<https://devfeed.tech/topics/blackwell.md>)

Tags: [apache](<https://devfeed.tech/tags/apache.md>), [blackwell](<https://devfeed.tech/tags/blackwell.md>), [mistral](<https://devfeed.tech/tags/mistral.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [models](<https://devfeed.tech/tags/models.md>), [nvfp4](<https://devfeed.tech/tags/nvfp4.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

Mistral announces Mistral 3, a family of open models including compact dense models and Mistral Large 3, a sparse mixture-of-experts model with 41B active and 675B total parameters. The models are released under Apache 2.0, with compressed formats and optimized checkpoints intended to improve accessibility, customization, and deployment across developer and enterprise environments.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## Cheaper, Better, Faster, Stronger

DevFeed: [Cheaper, Better, Faster, Stronger](<https://devfeed.tech/articles/cheaper-better-faster-stronger-7091.md>)

Original publisher: [Read original article](<https://mistral.ai/news/mixtral-8x22b/>)

Published: 2024-04-17T14:00:00Z

Content type: release

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [mixtral](<https://devfeed.tech/topics/mixtral.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Application Development](<https://devfeed.tech/topics/application-development.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [coding](<https://devfeed.tech/tags/coding.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mixtral](<https://devfeed.tech/tags/mixtral.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>)

### AI overview

Mistral introduces Mixtral 8x22B, an open sparse Mixture-of-Experts model with 141B total parameters and 39B active parameters. It emphasizes cost-efficient performance, multilingual capabilities, mathematics and coding, function calling, a 64K-token context window, fine-tuning, and Apache 2.0 licensing.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## Mixtral of experts

DevFeed: [Mixtral of experts](<https://devfeed.tech/articles/mixtral-of-experts-7094.md>)

Original publisher: [Read original article](<https://mistral.ai/news/mixtral-of-experts/>)

Published: 2023-12-11T07:00:00Z

Content type: release

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [mixtral](<https://devfeed.tech/topics/mixtral.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Code generation](<https://devfeed.tech/topics/code-generation.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [llama](<https://devfeed.tech/topics/llama.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [apache](<https://devfeed.tech/tags/apache.md>), [architectures](<https://devfeed.tech/tags/architectures.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [code](<https://devfeed.tech/tags/code.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llama](<https://devfeed.tech/tags/llama.md>), [mixtral](<https://devfeed.tech/tags/mixtral.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

Mistral AI releases Mixtral 8x7B, an open-weight sparse mixture-of-experts model licensed under Apache 2.0. The article reports strong benchmark results, faster inference, efficient cost and latency characteristics, multilingual support, and strong code-generation performance.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.