ABOUT VICKY FELIREN

I look for what breaks before a system is trusted.

I’m Vicky Feliren, an applied scientist working where AI safety, uncertainty, and underrepresented data meet. My work asks a practical question: when a model becomes safer, does it remain useful for the people and inputs its training represented least?

Explore the research View curriculum vitae
  • 7peer-reviewed papers
  • 5+ yrsproduction ML
  • IEEE · ACL · RSEpublication venues
Vicky Feliren
Applied Scientist · AI safety and reliable multimodal systems
01

The question

Safety is not only about changing an answer.

It is also about knowing when that answer should be trusted.

Safety training can change how well a model’s confidence matches its accuracy. Once that calibration slips, the model may continue when it should defer—or abstain so often that it is no longer useful. Average benchmark scores can hide where that trade-off is being paid.

I study the distribution beneath the average: which languages, input types, and communities absorb the largest cost. I am especially interested in Southeast Asia, where the world’s linguistic and visual variety is still poorly represented in mainstream datasets and evaluations.

02

The path

Research shaped by systems that had consequences.

From production constraints to research guarantees.

Public systems

Jakarta Smart City

Forecasting municipal waste taught me that model quality matters only when it changes a real allocation decision.

Production ML

Finance and identity

Biometrics, credit, and fraud systems made calibration, auditability, and failure costs operational—not theoretical.

Multimodal research

Earth observation

Satellite systems across sensors and regions made distribution shift visible in every map.

Open science

Southeast Asian AI

SEACrowd connected the technical problem to the missing languages, cultures, and visual worlds behind it.

AI safety

Calibration under alignment

I bring those threads together: measure the hidden cost, then recover useful deference with guarantees.

The complete chronology lives in the CV
03

The method

How I decide what deserves attention.

Start with the failure boundary, then build back toward use.

01

Find the hidden average

Disaggregate the result until the users and inputs carrying the cost become visible.

02

Make uncertainty legible

Turn confidence into a measurable decision variable—not a decorative score.

03

Test outside the comfortable case

Use multilingual, multicultural, and multimodal inputs that expose brittle assumptions.

04

Recover with a bound

Prefer interventions whose limits can be stated clearly enough for someone else to trust.

04

The direction

The next question is already in motion.

Can aligned models keep calibrated judgment beyond English and beyond text?

Current research direction

I am measuring how alignment changes calibration across languages and modalities, then testing whether distribution-free abstention can recover reliable deference without erasing usefulness.

Measure
Calibration tax by input group
Intervene
Bounded abstention after alignment
Evaluate
Multilingual + multimodal systems

The work is public. The complete record is separate. Choose the depth you need.