Integrating lexical and prosodic features for automatic paragraph segmentation

Lai, Catherine; Farrús, Mireia; Moore, Johanna D.; Lai, Catherine; Farrús, Mireia; Moore, Johanna D.

Integrating lexical and prosodic features for automatic paragraph segmentation

Autor/a

Lai, Catherine

Farrús, Mireia

Moore, Johanna D.

Data de publicació

2022-02-07T17:01:06Z

2022-08-31T05:10:25Z

2020-08

2022-02-07T17:01:06Z

Resum

Spoken documents, such as podcasts or lectures, are a growing presence in everyday life. Being able to automatically identify their discourse structure is an important step to understanding what a spoken document is about. Moreover, finer-grained units, such as paragraphs, are highly desirable for presenting and analyzing spoken content. However, little work has been done on discourse based speech segmentation below the level of broad topics. In order to examine how discourse transitions are cued in speech, we investigate automatic paragraph segmentation of TED talks using lexical and prosodic features. Experiments using Support Vector Machines, AdaBoost, and Neural Networks show that models using supra-sentential prosodic features and induced cue words perform better than those based on the type of lexical cohesion measures often used in broad topic segmentation. Moreover, combining a wide range of individually weak lexical and prosodic predictors improves performance, and modelling contextual information using recurrent neural networks outperforms other approaches by a large margin. Our best results come from using late fusion methods that integrate representations generated by separate lexical and prosodic models while allowing interactions between these features streams rather than treating them as independent information sources. Application to ASR outputs shows that adding prosodic features, particularly using late fusion, can significantly ameliorate decreases in performance due to transcription errors.

Tipus de document

Article

Versió acceptada

Llengua

Anglès

Matèries i paraules clau

Anàlisi prosòdica (Lingüística); Marcadors del discurs; Dicció; Prosodic analysis (Linguistics); Discourse markers; Diction

Publicat per

Elsevier B.V.

Documents relacionats

Versió postprint del document publicat a: https://doi.org/10.1016/j.specom.2020.04.007

Speech Communication, 2020, vol. 121, p. 44-57

https://doi.org/10.1016/j.specom.2020.04.007

info:eu-repo/grantAgreement/EC/H2020/645012/EU//KRISTINA

Citació recomanada

Aquesta citació s'ha generat automàticament.

Exportar

DIDL MARC MARC_CCUC METS OAI_DC ORE QDC RDF

Drets

cc-by-nc-nd (c) Elsevier B.V., 2020

https://creativecommons.org/licenses/by-nc-nd/4.0/

Aquest element apareix en la col·lecció o col·leccions següent(s)

Filologia Catalana i Lingüística General [951]

ISGlobal - Institut de Salut Global de Barcelona [61335]

Integrating lexical and prosodic features for automatic paragraph segmentation

Autor/a

Data de publicació

Compartir

Resum

Tipus de document

Llengua

Matèries i paraules clau

Publicat per

Documents relacionats

Citació recomanada

Exportar

Drets

Aquest element apareix en la col·lecció o col·leccions següent(s)