I build large language models end to end—data curation, pretraining, fine-tuning, and evaluation—with a focus on code generation beyond English, beyond Python, and beyond the benchmarks the field already optimizes for.
I am a Provost's Postdoctoral Fellow at the University of Notre Dame, working with Dr. Joanna C. S. Santos in the S2E Lab on safety-by-construction guardrails and multilingual program synthesis for Code LLMs. I completed my Ph.D. in Computer Science at George Mason University in 2026, advised by Dr. Marcos Zampieri and Dr. Antonios Anastasopoulos. My open releases include mHumanEval, TigerLLM, TigerCoder, and MojoBench. I am the lead organizer of LangCode 2027, the first workshop on language and code, proposed for the 2027 ACL-community workshop cycle.
Code LLMs & program synthesis
Adapting code models to low-resource programming languages and underexplored domains.
Read more →Multilingual & low-resource NLP
Evaluation and modeling across natural languages, with a focus on Bangla and code-mixed text.
Read more →LLM safety & AI in CS education
Guardrails for code assistants and the role of LLMs in introductory computing.
Read more →Benchmarks & datasets
Large, openly released resources that make under-tested settings measurable.
Read more →