Linguistic and orthographical classic Portuguese variants. Challenges for NLP
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
CEUR-WP org.
Abstract
In recent times, it was made a great investment in transfer from physical ancient Portuguese texts to digital support. This support transfer allows not
only the access to the texts, bringing them to the public in general, but also the possibility of texts to be readable and processed by machines. NLP tools are
addressed, mainly, to contemporary Portuguese and the application of NLP to
classic texts has several difficulties. The elaboration of big lexical corpora of
forms previous to modern Portuguese is an opportunity for multidisciplinary
field of studies allowing the enlargement of linguistic studies and also the possibility of obtaining, by NLP, validated corpora, collections and ontologies, that can be input in NLP tools for ancient Portuguese texts. In this work we will present, briefly, the problem of lexical variation of forms in processing classic Portuguese texts, the challenges that emerge from them and future perspectives of work.
Description
Keywords
Citation
Cameron, Helena Freire; Gonçalves, Maria Filomena; Quaresma, Paulo (2020): "Linguistic and orthographical classic Portuguese variants. Challenges for NLP". In: Maria José Finatto, Renata Vieira, Senja Pollak and Saturnino Luz (ed.), Proceedings of the Workshop on Digital Humanities and Natural Language Processing, co-located with International Conference on the Computational Processing of Portuguese (PROPOR 2020), vol. 2607. Évora (Portugal): CEUR-WP org, 43-48.