While the AI Act mandates high levels of accuracy and robustness, current oversight risks inheriting a blind spot by prioritizing English-first performance. Research from MuBench and P3B3 confirms that fluency often masks declining reliability, showing that models frequently falter when prompted in lower-resource languages or specific regional varieties. If regulators continue to rely on aggregate scores that conceal where performance deteriorates, they will fail to protect users across the union's 24 official languages.
To bridge this gap, the EU must move toward a multilingual evaluation commons. This requires shifting away from translated datasets toward native, domain-specific test modules that account for local legal and administrative contexts. Standardized performance cards should force providers to disclose how models behave across different languages, ensuring that safety, instruction-following, and factual accuracy are verified independently of linguistic fluency. Without these granular disclosures, incident reporting will remain fragmented, treating systemic inequalities as mere isolated errors. For the AI Act to succeed, the evidence used to judge these systems must be as diverse as the citizens they serve.

Comments (0)
No comments yet. Be the first!