I'm an applied scientist working on calibration under safety alignment. Training a model to be safe has a cost, and some of it lands on calibration, meaning how well a model's confidence tracks whether it is right. A model that loses that defers badly, so you keep it on small problems. I want safe models to be more useful, and my starting position is that the current tradeoff is worse than it needs to be.
At Monash University, I'm working with Associate Professor Risqi Saputra and Professor Taufiq Asyhari on conformal prediction for vision-language navigation. That work left me with one habit. When I see a guarantee, I ask what it depends on, and which of those things breaks first. It is also the tool I want to point at this problem. A model whose calibration has slipped still needs some way to defer, and that deferral should come with a bound.
The other input I bring is access. Southeast Asia is one of the most linguistically and visually diverse regions in the world and it is almost absent from the data and benchmarks that define what modern AI can do. I've spent years building datasets and adaptation methods there with SEACrowd and through my own research. Safety training data is overwhelmingly English text. If its cost falls unevenly, this is where you would see it first. Measuring that takes data and native fluency. I have both.
This Month · updated July 2026
- Writing up my M.Sc. thesis on conformal prediction for vision-language navigation, sharpening the abstention and set-efficiency results
- Replicating the published calibration cost of safety alignment in a setup I can run end to end, before building anything on top of it
- Designing the experiment that follows: measure that cost per language rather than averaged, starting with one I speak natively so I can audit the eval data myself
- Presented the multilingual VLM abstention study at AI Safety India's Hackathon Winners event
Tech Stack
Research Agenda
Making a model safer costs something, and the literature calls it the alignment tax. It gets reported as one number averaged over a test set, but I think it is a distribution over inputs whose shape nobody has measured. If the cost is largest where the safety data was thinnest, then a model we call safety aligned is aligned unevenly, and we would not see it, because we measure where the data already was.
Calibration is where the safety tax gets paid
Safety training shifts what a model outputs, and its confidence is a property of that output. The capability cost is well studied, from Lin et al. (EMNLP 2024) on the alignment tax to Huang et al. on reasoning. The calibration cost has had less attention. Leng et al. (ICLR 2025) show RLHF drives models to verbalise overconfidence, and Hu et al. (ACL 2026 Findings) call the loss severe. That is the cost I care about, because calibration decides how much you can delegate.
The cost is probably not spread evenly
Safety training data is overwhelmingly English text. There is no particular reason its cost should fall evenly across inputs that were unevenly represented in it. So I measure the calibration change per language and per modality rather than averaged, because an average over a distribution you never sampled is not really a measurement. This is the experiment I am running now, and it is useful either way. If the cost is uniform, that is worth knowing and it makes the problem simpler. If it is not, then some users are getting a worse-calibrated model than the benchmark suggests.
Getting the deference back, with something you can bound
Measuring a problem is half a contribution. The other half is recovery. Hu et al. (ACL 2026 Findings) restore some calibration by merging model weights from before and after alignment. I want to know whether you can do it at the other end, by putting an abstention layer with distribution-free coverage (ICML 2024) on top of a model whose calibration has already been degraded, and what that costs in usefulness. This is where uncertainty quantification earns its place, as a tool aimed at a specific failure rather than a subject of its own.
Where it gets tested
These ideas have to survive contact with systems that were not built for the test. I start by replicating a published result in a setup I can run end to end, because building on numbers I have not reproduced myself is how you waste a year. The inputs come from work I already know. Agent trajectories from my thesis on vision-language navigation. Multimodal models for earth observation, published in IEEE and Remote Sensing of Environment. Open multilingual models for Southeast Asia, built with SEACrowd and SEA-VL. These are the inputs that English-first, benchmark-first evaluation never sees.
Track Record
WORK EXPERIENCE
-
OCT 2024 – PRESENT
SEACrowd - Researcher, Multimodal & Vision-Language
Open-science research collective · seacrowd.github.io
-
FEB 2025 – NOV 2025
Artefact - Senior Data Scientist
French AI consulting · Founding technical member, Jakarta office
-
DEC 2022 – JAN 2025
Monash University - Research Associate
Global research consortium: Monash, UQ, UCL, Nottingham
-
JUN 2021 – JUN 2023
GDP Labs (GLAIR.ai) - Senior Data Scientist / ML Engineer
AI consulting, backed by a major Indonesian conglomerate
-
JAN 2021 – JUN 2021
Jakarta Smart City - Data Scientist
Indonesia's smart city government initiative
PATENT
-
ISSUED JUNE 2025
Fish & Shrimp Pond Detection via Satellite Imagery
IDS000010594
TALKS
-
JUL 2026
Do Multilingual Vision-Language Models Abstain under Cross-Modal Conflict in Low-Resource Languages?
Hackathon Winners Present · AI Safety India Community Events
-
MAY 2026
PyPalu, Python for Localized Context
Python community meetup · Sulawesi Tengah, Indonesia
-
2025
MUSE, How Data Science Differs in Each Sector
Panel speaker · Monash University Indonesia
-
OCT 2024
Bank of Indonesia, Data Synthesis, Privacy & Responsible Data Management
Invited talk · Bank of Indonesia
-
Q3 2026 OPEN
Available for conference talks, podcasts, and panel invitations
Topics: trustworthy AI, conformal prediction, AI for Southeast Asia
EDUCATION
-
EXPECTED SEPT 2026
Monash University
Master of Data Science · GPA: 4.0/4.0
-
JUN 2026
BlueDot Impact, Technical AI Safety
Cohort intensive · alignment, interpretability, red-teaming, AI control
-
DECEMBER 2019
Monash University
Bachelor of Computer Science
-
MAR – DEC 2019
Monash CURIE Compass
Mentee · Centre for Undergraduate Research Initiatives and Excellence
-
JUN – DEC 2018
Nanyang Technological University
Computer Science Exchange Programme · Singapore
-
2019
Udacity, Deep Learning Nanodegree
Neural networks, CNNs, RNNs, GANs · Facebook AI scholarship recipient
TEACHING
-
MAY 2026
Monash University PGIE, Industry Judge
Faculty of IT Postgraduate Industry Experience
-
MAY 2026 – JUN 2026
IBM SkillsBuild
Capstone Project Advisor
-
MAY 2026 – JUN 2026
Coding Camp powered by DBS Foundation
Capstone Project Advisor
-
JAN 2024 – JAN 2025
Bangkit Academy (Google, Gojek, Traveloka)
ML Instructor & Capstone Advisor
-
NOV 2024
Bina Nusantara University
Guest Lecturer, Computer Vision
Recognition
- CHAMPION Microsoft Azure Virtual Hackathon APAC
- REGIONAL WINNER Asia Pacific Regional Winner, Global South AI Safety Hackathon (Apart Research)
- TOP 3 CamvsCovid, University of Cambridge
- BEST COMMUNITY TRACK hack:now, Cal Hacks UC Berkeley
- TOP 10 Reboot the Earth, United Nations Technology Innovation Labs
- MOST VISIONARY RESEARCH Mining Spatial Data Intelligence Research Hub, IISF 2024
- SCHOLARSHIP Monash Indonesia Inaugural Welcome Scholarship
- SCHOLARSHIP Deep Learning Nanodegree, Facebook & Udacity
- SCHOLARSHIP Jeffrey Cheah Entrance Scholarship
- TOP 3 Monash Hackathon
- SCHOLARSHIP Monash Travel Grant
- HONORABLE MENTION Monash University Coding Competition
PROFESSIONAL SERVICE
- Industry Judge: Monash FIT PGIE 2026, evaluated 6 cross-disciplinary teams (36 students) from Master of AI, Data Science, and IT presenting real-world industry projects.
- Peer Reviewer: IEEE IGARSS 2026, Premier global remote sensing and geoscience symposium.
- Technical Judge: Cal Hacks 8.0, CruzHacks 2022, iNTUition v8.0, evaluated 50+ projects across three international hackathons.
- Open Source Contributor: SEACrowd collective, building AI infrastructure and data democratization for Southeast Asia.