# Improving Evaluation Practices in Natural Language Generation

DevFeed: [Improving Evaluation Practices in Natural Language Generation](<https://devfeed.tech/articles/improving-evaluation-practices-in-natural-language-generation-28020.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2022-03-31-improving-evaluation-practices-in-natural-language-generation/>)

Author: Saad Mahamood NLG Expert; Lead Data Scientist

Published: 2022-03-31T00:00:00Z

Content type: article

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [Natural language processing](<https://devfeed.tech/topics/nlp.md>), [recommendations](<https://devfeed.tech/topics/recommendations.md>)

Tags: [data-science](<https://devfeed.tech/tags/data-science.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [language](<https://devfeed.tech/tags/language.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [recommendations](<https://devfeed.tech/tags/recommendations.md>), [report](<https://devfeed.tech/tags/report.md>), [reproducibility](<https://devfeed.tech/tags/reproducibility.md>), [research](<https://devfeed.tech/tags/research.md>)

## AI overview

This article reviews research into evaluation practices for Natural Language Generation. It discusses weaknesses in automated metrics, variation in human evaluation methods, and reproducibility concerns, including findings from trivago's HumEval 2021 work and its proposed Commonsense Evaluation Card.

## Source excerpt

Throughout last year I had the opportunity to participate and collaborate on multiple research initiatives in the field of Nat...