NLPLanguage IdentificationLow-ResourceBenchmarkACL 2026

CommonLID: language identification on noisy web data

ACL 2026 · Contributor

End-to-end pipeline diagram: CommonCrawl text (code-switched · romanized) → Production LID stack (fastText · GlotLID · OpenLID) → Stratified eval (noise regimes isolated) → Failure map (guidance for LLM corpora)
Pipeline overview

Language identification supports translation routing, content moderation, dataset building, and search indexing. These pipelines often treat it as a solved problem. FastText LangID, GlotLID, and OpenLID regularly score above 95% on clean benchmarks, but those scores create false confidence.

Southeast Asian web text mixes languages within a sentence, uses Roman spellings alongside standard scripts, and includes phonetic spelling and social-media shorthand. Systems trained on curated text perform much worse on this data. Existing benchmarks miss the decline because they do not represent the web conditions where the systems operate.

We re-tested fastText LangID, GlotLID, and OpenLID on CommonCrawl text with code-switching, Romanised scripts, and spelling variation. We mapped each system's failures across language families and proposed a stricter protocol for claims about low-resource Southeast Asian coverage.

The results challenge published accuracy claims and show why pipelines using language identification must be tested again. This includes the CommonCrawl-based data used to train major language models.

These failures affect content moderation, translation routing, and language model datasets across Southeast Asia. The evaluation protocol also gives future low-resource benchmarks a clear method for testing real web data.

fastText LangIDGlotLIDOpenLIDCommonCrawlLow-resource SEA languages (11+)