水果派AV解说

AI UNLOCKS H脕N N脭M HERITAGE: OPENING THE DOOR TO A MILLENNIUM OF KNOWLEDGE FOR WIDER PUBLIC ACCESS

AI UNLOCKS H脕N N脭M HERITAGE: OPENING THE DOOR TO A MILLENNIUM OF KNOWLEDGE FOR WIDER PUBLIC ACCESS

After more than a thousand years preserving the history, culture and knowledge of the Vietnamese nation, much of the H谩n N么m documentary heritage remains inaccessible to the wider public. By successfully developing an artificial intelligence (AI)-powered H谩n N么m automatic translation system, researchers at HCMUS are steadily opening the door to this invaluable repository of knowledge, while laying the groundwork for the digital preservation and broader utilisation of 水果派AV解说鈥檚 cultural heritage.

AI opens new pathways to H谩n N么m heritage

According to the Association for Preservation of H谩n N么m Heritage, more than 90% of H谩n N么m documents have yet to be translated into Qu峄慶 ng峄 (the modern Vietnamese writing system based on the Latin alphabet). Millions of pages of royal decrees, imperial records, genealogies, cadastral registers, stone inscriptions, ancient manuscripts, horizontal lacquered boards, parallel sentences and traditional medical texts remain available primarily as scanned images or digitised copies, while the number of people able to read H谩n N么m continues to decline.

To address this challenge, a research team led by Associate Professor 膼inh 膼i峄乶, Director of the Centre for Computational Linguistics at HCMUS, has carried out the project Research and Development of an Automatic H谩n N么m-to-Qu峄慶 ng峄 Translation System 鈥 Phase 2. The project has recently been evaluated as Outstanding by the Acceptance Council of the Ho Chi Minh City Department of Science and Technology.

Associate Professor 膼inh 膼i峄乶 explained: 鈥淭he N么m script was created by Vietnamese scholars around the 10th century based on Chinese characters and remained in use for more than a millennium. An enormous body of literature, historical records, geographical works, traditional medicine and many other fields was recorded using this writing system. Today, however, very few people are able to read H谩n N么m, while widely used AI systems such as ChatGPT, Gemini and DeepSeek can process modern Chinese characters but cannot interpret the N么m script.鈥

Homepage of the N么m transliteration website: https://tools.clc.hcmus.edu.vn/

Without timely technological solutions for recognising and converting historical documents, many valuable materials risk continuing deterioration due to ageing, storage conditions and the diminishing number of H谩n N么m specialists. The research team aims to develop a system that enables anyone to access, search and understand H谩n N么m documents conveniently and free of charge.

A major breakthrough in Phase 2 has been the successful development of optical character recognition (OCR) technology for H谩n N么m documents. While the first phase could process only text-based digital documents, the current system can recognise documents from images, which represent the vast majority of surviving H谩n N么m materials.

Another significant achievement has been the creation of the largest H谩n N么m dataset ever assembled in 水果派AV解说. The dataset includes 200,000 scanned pages of H谩n N么m documents, 200,000 manually annotated images for OCR training, 750,000 bilingual H谩n N么m鈥換u峄慶 ng峄 sentence pairs, and more than one million H谩n N么m monolingual sentences, comprising approximately 16 million characters.

The dataset has been developed through collaboration with numerous domestic and international partners, including the Nom Foundation, the Tran Nhan Tong Institute, several H谩n N么m research projects, together with scholars, lecturers and students specialising in H谩n N么m studies. Most source materials were originally available only in raw form, requiring the research team to establish comprehensive workflows for data standardisation, annotation and cross-validation by both AI models and domain experts to ensure high-quality training data.

Towards a shared digital platform for H谩n N么m heritage

Building upon the dataset and AI models, the research team has developed the Kim H谩n N么m ecosystem, available through web, Android and iOS platforms. The system can recognise H谩n N么m text from images, perform two-way transliteration between H谩n N么m and Qu峄慶 ng峄, support translation of Classical Chinese texts, and provide APIs for integration into other applications. For outdoor inscriptions such as horizontal lacquered boards and parallel sentences displayed in temples, pagodas and historical sites, the system can overlay the corresponding Qu峄慶 ng峄 text directly onto the original image, allowing users to compare the original script with the transliteration.

According to Associate Professor 膼inh 膼i峄乶, system performance depends largely on the quality of the input data. For text-based documents covering widely represented fields such as literature, history and geography, transliteration accuracy exceeds 99%. For high-quality scanned images, particularly those printed in the Kh岷 script, OCR accuracy exceeds 95%. More challenging materials, including heavily weathered stone inscriptions, handwritten texts and unusual calligraphic styles, still require expert review and correction by H谩n N么m specialists.

The system has been designed not only for researchers but also for the wider community. Members of the public can use a mobile phone to photograph horizontal lacquered boards, parallel sentences, royal decrees, genealogies, contracts or other historical family documents, allowing the system to recognise the original text automatically and convert the content into Qu峄慶 ng峄. As a result, historical materials that were previously difficult to access can become significantly easier to understand and explore.

Associate Professor 膼inh 膼i峄乶, Director of the Centre for Computational Linguistics at HCMUS, leads the research team undertaking the project 鈥淩esearch and Development of an Automatic H谩n N么m-to-Qu峄慶 ng峄 Translation System 鈥 Phase 2鈥.

Associate Professor 膼inh 膼i峄乶 believes the system will become a valuable resource for universities, libraries, museums, archival institutions and heritage conservation organisations, supporting document recognition, classification, summarisation and retrieval. Looking further ahead, the technology is expected to enable new research opportunities across history, geography, traditional medicine and studies related to 水果派AV解说鈥檚 maritime sovereignty.

Following the completion of Phase 2, the research team is preparing to launch Phase 3, focusing on semantic translation of H谩n texts. This stage is regarded as the most demanding because AI must move beyond language processing to incorporate knowledge of history, culture, religion and numerous specialised disciplines. At the same time, the team will continue expanding the dataset across subject areas including history, religion, traditional medicine, stone inscriptions and royal decrees, providing the foundation for increasingly specialised AI models.

Leave a Reply

Your email address will not be published.