Vision-Language ModelsBenchmarkDatasetCrowdsourcingCultural AIACL 2025

SEA-VL: multicultural vision-language benchmark for Southeast Asia

ACL 2025 (Main Conference, Long Paper) · Core Contributor, Data Pipeline Lead

End-to-end pipeline diagram: Community annotators (across SEA countries) → Dedup + HITL review (pHash · CLIP · SigLIP) → SEA-VL dataset (culturally grounded VQA) → Frontier audit (GPT-4V · Gemini · Claude)
Pipeline overview

Leading vision-language models rely mainly on English and Western image-text data. Southeast Asia spans 11 countries, more than 700 million people, hundreds of languages, and distinct cultural practices. Yet the region is underrepresented in training data and benchmarks. Before SEA-VL, no rigorous community benchmark measured culturally grounded visual reasoning across it.

A valid cultural benchmark needs more than scraped images. Its questions must reflect local food, festivals, buildings, scripts, and social practices. Standard crowdsourcing platforms lack enough regional knowledge. Using GPT-4 or Gemini to generate the test would also make the evaluation circular. Quality control becomes harder across 11 languages, including low-resource scripts.

I co-led more than 100 community annotators across all 11 Southeast Asian countries. They wrote image, question, and answer sets from lived cultural knowledge without using web search as a substitute.

I designed and built the quality pipeline. It covered annotation rules, automated filters, human review, language checks, and cultural validation. For image deduplication, I compared several similarity methods against a human-checked reference set and selected the most reliable one. The final contribution included more than 10,000 image-question pairs in 11 languages.

The benchmark measures consistent cultural reasoning gaps in GPT-4V, Gemini, and Claude. AI labs can use the results to set localisation priorities for Southeast Asian markets.

SEA-VL is the first rigorous, community-built vision-language benchmark for Southeast Asia. AI labs use its results to identify localisation gaps across a region of more than 700 million people. ACL 2025 accepted the work as a main-conference paper.

GPT-4VGemini 1.5Claude 3LLaVAInternVLpHashCLIP-ViTSigLIPNomic Embed Vision100+ annotators, 11 countries10,000+ image-question pairs