Using linguistic information to classify Portuguese text documents
| dc.contributor.author | Teresa, Gonçalves | |
| dc.contributor.author | Paulo, Quaresma | |
| dc.date.accessioned | 2009-04-06T15:49:04Z | |
| dc.date.available | 2009-04-06T15:49:04Z | |
| dc.date.issued | 2008-10 | |
| dc.description.abstract | This paper examines the role of various linguistic structures on text classification applying the study to the Portuguese language. Besides using a bag-of-words representation where we evaluate different measures and use linguistic knowledge for term selection, we do several experiments using syntactic information representing documents as strings of words and strings of syntactic parse trees. To build the classifier we use the Support Vector Machine (SVM) algorithm which is known to produce good results on text classification tasks and apply the study to a dataset of articles from the Público newspaper. The results show that sentences' syntactic structure is not useful for text classification (as initially expected), but part-of-speech information can be used as a term selection technique to construct the bag-of-words representation of documents. | en |
| dc.format.extent | 251581 bytes | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.accesstype | restrito_ue | en |
| dc.identifier.authoremail | tcg@di.uevora.pt | |
| dc.identifier.authoremail | pq@di.uevora.pt | |
| dc.identifier.editorperson | Gelbukh, Alexander | |
| dc.identifier.editorperson | Morales, Eduardo | |
| dc.identifier.isbn | 978-0-7695-3441-1 | en |
| dc.identifier.pagina | 94-100 | en |
| dc.identifier.principalpublicationtitle | 7th Mexican International Conference on Artificial Intelligence | en |
| dc.identifier.scientificarea | 283 | en |
| dc.identifier.uri | http://hdl.handle.net/10174/1410 | |
| dc.identifier.volume | 1 | en |
| dc.language.iso | eng | |
| dc.peerreviewed | yes | en |
| dc.publisher | IEEE Computer Society | en |
| dc.rights | openAccess | en |
| dc.subject | Text classification | en |
| dc.subject | Support vector machines | en |
| dc.subject | Linguistic Information | en |
| dc.title | Using linguistic information to classify Portuguese text documents | en |
| dc.type | article | en |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- goncalves-classifyPortuguesedocs.pdf
- Size:
- 245.68 KB
- Format:
- Adobe Portable Document Format
- Description:
- documento principal
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 3.61 KB
- Format:
- Item-specific license agreed upon to submission
- Description: