MMDT NER
Overview & Abstract
MMDT NER is a foundational language technology project that transforms unstructured Burmese text into structured, usable data by identifying people, organizations, locations, dates, times, and numerical expressions. Built on a 2.14-million-token annotated corpus, it combines transformer-based modeling and responsible evaluation to advance context-aware AI for Burmese—an underrepresented, low-resource language—and support research, journalism, digital archives, humanitarian initiatives, and other public-interest applications.