Title:
|
A multi-layered annotated corpus of scientific papers
|
Author:
|
Fisas Elizalde, Beatriz; Ronzano, Francesco; Saggion, Horacio
|
Abstract:
|
Comunicació presentada a la Tenth International Conference on Language Resources and Evaluation (LREC 2016), celebrada els dies 23 a 28 de maig de 2016 a Portorož, Eslovènia. |
Abstract:
|
Scientific literature records the research process with a standardized structure and provides the clues to track the progress in a scientific
field. Understanding its internal structure and content is of paramount importance for natural language processing (NLP) technologies.
To meet this requirement, we have developed a multi-layered annotated corpus of scientific papers in the domain of Computer Graphics.
Sentences are annotated with respect to their role in the argumentative structure of the discourse. The purpose of each citation is specified.
Special features of the scientific discourse such as advantages and disadvantages are identified. In addition, a grade is allocated to each
sentence according to its relevance for being included in a summary.To the best of our knowledge, this complex, multi-layered collection
of annotations and metadata characterizing a set of research papers had never been grouped together before in one corpus and therefore
constitutes a newer, richer resource with respect to those currently available in the field. |
Abstract:
|
The research leading to these results has received funding from the European Project Dr. Inventor (FP7-ICT-2013.8.1 - grant agreement no 611383) and is partly supported by the Spanish Ministry of Economy and Competitiveness under the Maria de Maeztu Units of Excellence Programme (MDM-2015-0502). |
Subject(s):
|
-Multi-layered annotated corpus -Scientific discourse -Citations -Summarization gold standard |
Rights:
|
© any ELRA - European Language Resources Association. All rights reserved. The LREC 2016 Proceedings are licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
http://creativecommons.org/licenses/by-nc/4.0/ |
Document type:
|
Conference Object Article - Published version |
Published by:
|
ELRA (European Language Resources Association)
|
Share:
|
|