Description
Sino-Nom documents preserve Vietnamese cultural, literary, administrative, and religious heritage, but remain difficult to access because of complex scripts, heterogeneous sources, and scarce computational resources. This paper presents a framework for building a digitalized Sino-Nom lexical database for scholarly use and artificial intelligence applications. The database is organized as three complementary JSON datasets: a Chinese-focused set with 49,237 entries, a SinoNom-focused set with 56,274 entries, and a consolidated set with 57,059 character-level entries. Each record structures orthographic metadata, Unicode, stroke count, radicals, variants, glyph-image paths, multilingual pronunciations, definitions, source references, compounds, type codes, glosses, and classical-text examples. By combining automated processing with manual cross-checking against source dictionaries, the framework supports lexicographic search, linguistic research, education, machine translation between Sino-Nom and modern Vietnamese, and OCR for historical scripts.
Từ khóa
Sino-Nom; Cultural Heritage Digitization; Machine Translation; OCR; Natural Language Processing (NLP); Linguistic Database.
Thông tin các tác giả
1/Khoa D. Nguyen: Undergraduate, Studying at VNUHCM-University of Science, No. 227, Nguyen Van Cu St., Cho Quan Ward, Ho Chi Minh City, email: caogiap2408@gmail.com