I research current and future AI that are safe and empowering for more people.
My current interests include:
- making AI go well for all
- developing AI for low-resource languages and cultures
- designing cooperative AI for humans and other AI agents
- using AI for good in learning sciences, climate science, and knowledge work
My technical work focuses on data and evaluation to make grounded, predictive, and specific claims about AI capability and safety.
I'm doing research for Lida Safety, and I contribute to community research with Cohere Labs Community and BenchFlow. Recent work includes SEATauBench and chain-of-thought monitorability.
I'm seeking a research Master's in CS, AI, or NLP for Fall 2027 and am open to research collaborations.
For more information, visit my website at mychiffonn.com or see my CV.
- SEATauBench: Extended Tau2-Bench for low-resource Southeast Asian languages and localized agentic evaluation.
- GPS-Bench: Evidence-grounded benchmark and simulation framework for forecasting AI-governance policy outcomes.
- Multicultural Riddles Benchmarking: Analysis of LLM factuality, hallucination, topic patterns, and cross-cultural reasoning failures.
- BenchFlow: Open-source runtime environment for multi-turn AI agent research.
- MyScholar: Open-source Astro theme for academic research portfolios, publications, projects, and blogs.




