We present a Universal Dependencies (UD) annotated dataset of the Book of Ezra from the Old Georgian Oshki Bible, based on the TITUS edition. Our semi-automated workflow leverages transfer learning from Modern Georgian (UD_Georgian-GNC) and few-shot prompting of LLMs, followed by manual corrections by human linguists. We discuss the challenges of creating UD-conformed guidelines tailored for Old Georgian’s unique linguistic structure and the limitations and errors of LLM-based annotation. This work aims to establish a gold standard for Old Georgian UD, laying the groundwork for future morphosyntactic analysis of Old Georgian.
Building a Treebank for the Book of Ezra in the Old Georgian Oshki Bible
Luinetti, Diego
2026-01-01
Abstract
We present a Universal Dependencies (UD) annotated dataset of the Book of Ezra from the Old Georgian Oshki Bible, based on the TITUS edition. Our semi-automated workflow leverages transfer learning from Modern Georgian (UD_Georgian-GNC) and few-shot prompting of LLMs, followed by manual corrections by human linguists. We discuss the challenges of creating UD-conformed guidelines tailored for Old Georgian’s unique linguistic structure and the limitations and errors of LLM-based annotation. This work aims to establish a gold standard for Old Georgian UD, laying the groundwork for future morphosyntactic analysis of Old Georgian.File in questo prodotto:
Non ci sono file associati a questo prodotto.
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

