REPOSITORY > RESULTS

Doctoral dissertation

Cross-lingual text annotation

Author(s): Tadej Štajner (Author), Dunja Mladenić (Supervisor)

Thesis defense date: 17.05.2019

Organization: MPŠ - Mednarodna podiplomska šola Jožefa Stefana

PID: 20.500.12556/ReVIS-14446

Views: 18 | Downloads: 10

Abstract

This dissertation discusses the challenges of integrating a new language in an intelligent
system. A major issue in natural language processing is mixing multi-lingual content within
the same system. We introduce the notion of cross-linguality, a strategy of handling multilingual
content without machine translation pipelines. We demonstrate the possibilities
and weaknesses of individual proposed solutions and what kind of abstractions they let us
use when dealing with multilingual content.
The main motivation for the dissertation is learning how can we use knowledge bases
in improving access to multilingual information, and how can we better understand usergenerated
content.
First, we focus on the different roles of knowledge bases in multilingual information access.
We tackle the challenge of using multilingual knowledge bases as comparable corpora
to solve cross-lingual tasks, such as bilingual dictionary induction, similarity estimation
and vector space translation, tasks that are central to multilingual information access.
We propose a method using word embeddings and kernel approximation to train scalable
non-linear transformations. We demonstrate that this novel method works better on a majority
of evaluated language pairs. We also obtain results in posing vector space translation
as a sparse signal reconstruction problem, treating the cross-linguality as a hypothetical
noisy channel, where our task is to reconstruct the signal in the target language from
measurements in our source language.
As a second use case of knowledge bases, we present a novel method for named entity
disambiguation based on a notion of relatedness among entities. We present complementary
approaches of representing relatedness and show their respective contributions to the
quality of the result. We also discuss the cross-lingual version of this problem, along with
the challenges when developing a named entity recognition system for Slovene.
We also discuss the notion of language in social media as a genre of user-generated
language that does not conform to assumptions in typical natural language processing
systems. We present a novel method for summarizing streams of interesting responses to
news in social media in the form of a sampling algorithm. We show its performance on
real-life scenarios, as well as use it to demonstrate interesting responses, shedding light on
the subjectivity around the selection process and the perception of interestingness.
As another use case of user generated content, we present a novel method for sentiment
analysis that uses meta-modeling and integration of background knowledge. We show its
performance across domains and languages, showing that it can provide good performance
even on small training sets.

Attachments

Cite this work