Imagine asking a conversational AI, like ChatGPT, a legal question in Greek about local traffic laws, only to receive an answer based on UK law, despite its fluency in Greek. This highlights a significant challenge faced by large language models (LLMs) — their proficiency in various languages often does not extend to understanding regional, cultural, and legal contexts.
The Emergence of INCLUDE
To address this gap, the Natural Language Processing Lab at EPFL, in collaboration with Cohere Labs and international partners, has developed INCLUDE — a groundbreaking multilingual benchmark. INCLUDE is designed to evaluate whether an AI not only understands a language but also effectively integrates the corresponding cultural and sociocultural nuances. This initiative is part of the broader Swiss AI Initiative aimed at creating AI models reflective of local languages and values.
According to Angelika Romanou, a key researcher at EPFL, “LLMs must absorb cultural and regional nuances to remain relevant and relatable.” Without this critical awareness, AI models fall short of genuinely meeting user needs, demonstrating a considerable oversight in AI’s multilingual capabilities.
Addressing a Blind Spot in AI
LLMs, such as GPT-4 and LLaMA-3, have achieved remarkable success in processing multiple languages. However, they frequently underperform in culturally rich or underrepresented languages like Urdu or Armenian, primarily due to insufficient high-quality training data. Traditionally used benchmarks, often derived from English, omit the distinct linguistic and regional attributes vital for genuine understanding, thereby perpetuating cultural biases.
INCLUDE diverges from translation-heavy benchmarks by collecting over 197,000 questions, directly sourced from native academic and professional exams, in 44 languages and 15 scripts. This comprehensive dataset covers a spectrum of areas—from literature and law to regional social norms—highlighting discrepancies in AI’s grasp of local contexts.
Toward More Inclusive AI
Initial evaluations of top LLMs revealed substantial gaps. While models performed adequately in French and Spanish, they faced significant challenges with languages less represented in datasets. This disparity underscores the critical need for AI systems to adapt to globally diverse worldviews, particularly as these technologies find applications in vital sectors such as education, healthcare, and governance.
Antoine Bosselut of EPFL underscores the importance of local understanding, “The democratization of AI necessitates that models align with the diverse realities of communities worldwide.”
Key Takeaways
The INCLUDE benchmark represents a pioneering effort to push AI beyond mere language proficiency toward a deeper cultural understanding. By focusing on region-specific contexts, INCLUDE offers a pathway to developing more inclusive, fair, and capable AI systems that truly reflect the diverse fabric of human societies. As the benchmark evolves to cover more languages and regional characteristics, it sets a new standard for AI evaluation, underpinning efforts to create AI that is both technically and culturally competent.