Research

I approach NLP with a clear objective: to make language technology effective for low-resource languages, specifically Burmese. I enjoy the challenge of overcoming the data scarcity and technical limitations that hold these languages back from performing at a high level. A few directions I focus on:

Language Modeling

The primary challenge is adapting foundation models to the linguistic nuances of Burmese without relying on massive, language-specific corpora. This matters because it moves us past the dependency on English-centric data, making advanced language technology accessible even in resource-constrained environments.

Evaluation & Benchmarking

Defining what "success" actually looks like in low-resource contexts is an ongoing effort. Standard metrics often mask poor practical performance; moving toward benchmarks that prioritize real-world reasoning and reliability is essential to ensuring that our progress is genuine rather than just an artifact of overfitting.

Data-Centric NLP

The focus here is on the relationship between intentional, human-curated data and model behavior. Since sheer volume is not always an option, understanding how strategic selection — rather than just scale — shapes performance is the most efficient path toward building reliable models.

Information Extraction

This direction explores how we can accurately parse complex, multi-layered information in Burmese. Developing specialized resources, such as nested NER corpora, offers a way to address the structural complexities of the language to improve the precision of information retrieval.

Foundational Resource Creation

Building a sustainable ecosystem for Burmese NLP requires creating core linguistic assets that are often missing. Without systematic, high-quality infrastructure, research remains fragmented; creating these assets provides a foundation for more consistent, large-scale experimentation.

See publications for specific papers, or get in touch to collaborate.