# 📚 3LM: A Benchmark for Arabic LLMs in STEM and Code

DevFeed: [📚 3LM: A Benchmark for Arabic LLMs in STEM and Code](<https://devfeed.tech/articles/3lm-a-benchmark-for-arabic-llms-in-stem-and-code-7503.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tiiuae/3lm-benchmark>)

Author: Basma Boussaha; Leen AlQadi; Mughaira; Shaikha Alsuwaidi; Giulia Campesan; Ahmed Alzubaidi; Mohammed Alyafeai; Hakim Hacid

Published: 2025-08-01T14:25:21Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Large language models (LLMs)](<https://devfeed.tech/topics/large-language-models-llms.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Code generation](<https://devfeed.tech/topics/code-generation.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [llms](<https://devfeed.tech/tags/llms.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [quality-assurance](<https://devfeed.tech/tags/quality-assurance.md>), [stem](<https://devfeed.tech/tags/stem.md>)

## AI overview

The article introduces 3LM, a benchmark for evaluating Arabic Large Language Models on STEM subjects, structured reasoning, formal logic, and code generation. It combines native educational multiple-choice questions, synthetic high-difficulty STEM questions, and translated code-generation tasks.

## Source excerpt

A Blog post by Technology Innovation Institute on Hugging Face