http://www.cnr.it/ontology/cnr/individuo/prodotto/ID220785
Automatic Creation of Quality Multi-Word Lexica from Noisy Text Data (Contributo in atti di convegno)
- Type
- Label
- Automatic Creation of Quality Multi-Word Lexica from Noisy Text Data (Contributo in atti di convegno) (literal)
- Anno
- 2012-01-01T00:00:00+01:00 (literal)
- Alternative label
Francesca Frontini, Valeria Quochi, Francesco Rubino (2012)
Automatic Creation of Quality Multi-Word Lexica from Noisy Text Data
in AND 2012, Mumbai, India, December 9, 2012
(literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#autori
- Francesca Frontini, Valeria Quochi, Francesco Rubino (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#altreInformazioni
- ID_PUMA: /cnr.ilc/2012-A3-008 (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#url
- http://www.kde.cs.tut.ac.jp/~aono/pdf/COLING2012/AND/pdf/AND04.pdf (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#titoloVolume
- Proceedings of the Sixth Workshop on Analytics for Noisy Unstructured Text Data (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#affiliazioni
- Titolo
- Automatic Creation of Quality Multi-Word Lexica from Noisy Text Data (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#isbn
- 978-1-4503-1919-5 (literal)
- Abstract
- This paper describes the design of a tool for the automatic creation of multi-word lexica that is deployed as a web service and runs on automatically web-crawled data within the framework of the PANACEA platform. The main purpose of our task is to provide a (computationally \"light\") tool that creates a full high quality lexical resource of multi-word items. Within the platform, this tool is typically inserted in a work flow whose first step is automatic web-crawling. Therefore, the input data of our lexical extractor is intrinsically noisy. The paper evaluates the capacity of the tool to deal with noisy data, and in particular with texts containing a significant amount of duplicated paragraphs. The accuracy of the extraction of multi-word expressions from the original crawled corpus is compared to the accuracy of the extraction from a later \"de-duplicated\" version of the corpus. The paper shows how our method can extract with sufficiently good precision also from the original, noisy crawled data. The output of our tool is a multi-word lexicon formatted and encoded in XML according to the Lexical Mark-up Framework. (literal)
- Editore
- Prodotto di
- Autore CNR
- Insieme di parole chiave
Incoming links:
- Prodotto
- Autore CNR di
- Editore di
- Insieme di parole chiave di