The major goal is to ease out the manual annotation process by having large language models (LLMs) annotate the morphosyntactic analysis pipeline. We are aiming to do Lemmatization, part of speech tagging, morphosyntactic feature specification and English glossing of individual lemmas. The core idea is with zero shot learning, LLMs can bypass the need for massive datasets, thereby helping in linguistic documentation. Overall, we are looking to see that if we are able to automate the whole process, the endangered and low resource language development would speed up and will reach a major milestone.

This team is recruiting until June 1, 2026.

Team Leads: Tianming Liu (Computer Science), Keith Langston (Linguistics)