# Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard

DevFeed: [Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard](<https://devfeed.tech/articles/rethinking-llm-evaluation-with-3c3h-aragen-benchmark-and-leaderboard-7308.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/leaderboard-3c3h-aragen>)

Author: Ali El Filali; Neha Sengupta; Abouelseoud; Preslav Nakov; Clémentine Fourrier

Published: 2024-12-04T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [open-llm-leaderboard](<https://devfeed.tech/topics/open-llm-leaderboard.md>)

Tags: [arabic](<https://devfeed.tech/tags/arabic.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [llm](<https://devfeed.tech/tags/llm.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-llm-leaderboard](<https://devfeed.tech/tags/open-llm-leaderboard.md>), [performance](<https://devfeed.tech/tags/performance.md>), [research](<https://devfeed.tech/tags/research.md>)

## AI overview

The article introduces the AraGen Benchmark and Leaderboard for evaluating Arabic large language models. It presents the 3C3H Measure, which uses LLM-as-judge assessment across correctness, completeness, conciseness, helpfulness, honesty, and harmlessness. AraGen also uses private three-month blind testing cycles and a multi-turn and single-turn Arabic evaluation dataset to reduce data contamination and assess both factuality and practical usability.

## Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.