File Information

File: 05-lr/acl_arc_1_sum/cleansed_text/xml_by_section/abstr/06/p06-1070_abstr.xml

Size: 1,256 bytes

Last Modified: 2025-10-06 13:45:00

<?xml version="1.0" standalone="yes"?>
<Paper uid="P06-1070">
  <Title>Exploiting Comparable Corpora and Bilingual Dictionaries for Cross-Language Text Categorization</Title>
  <Section position="1" start_page="0" end_page="0" type="abstr">
    <SectionTitle>
Abstract
</SectionTitle>
    <Paragraph position="0"> Cross-language Text Categorization is the task of assigning semantic classes to documents written in a target language (e.g. English) while the system is trained using labeled documents in a source language (e.g.</Paragraph>
    <Paragraph position="1"> Italian).</Paragraph>
    <Paragraph position="2"> In this work we present many solutions according to the availability of bilingual resources, and we show that it is possible to deal with the problem even when no such resources are accessible. The core technique relies on the automatic acquisition of Multilingual Domain Models from comparable corpora.</Paragraph>
    <Paragraph position="3"> Experiments show the effectiveness of our approach, providing a low cost solution for the Cross Language Text Categorization task. In particular, when bilingual dictionaries are available the performance of the categorization gets close to that of mono-lingual text categorization.</Paragraph>
  </Section>
class="xml-element"></Paper>
Download Original XML