On the distribution of deep clausal embeddings: a large Cross-linguistic Study

Blasi, Damián E.; Cotterell, Ryan; Wolf-Sonkin, Lawrence; Stoll, Sabine; Bickel, Balthasar; Baroni, Marco

Lokale TagsFreigabegeschichteDetailsÜbersicht

On the distribution of deep clausal embeddings: a large Cross-linguistic Study

Blasi, D. E., Cotterell, R., Wolf-Sonkin, L., Stoll, S., Bickel, B., & Baroni, M. (2019). On the distribution of deep clausal embeddings: a large Cross-linguistic Study. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3938-3943). Florence, Italy: Association for Computational Linguistics. Retrieved from https://www.aclweb.org/anthology/P19-1384.

Item is Freigegeben

einblenden: alle ausblenden: alle

Basisdaten

einblenden: ausblenden:

Datensatz-Permalink: https://hdl.handle.net/21.11116/0000-0004-628D-F Versions-Permalink: https://hdl.handle.net/21.11116/0000-0004-628E-E

Genre: Konferenzbeitrag

Dateien

einblenden: Dateien

ausblenden: Dateien

:

shh2341.pdf (Verlagsversion), 816KB

Datei-Permalink:
-

Name:
shh2341.pdf

Beschreibung:
-

OA-Status:

Sichtbarkeit:
Privat

MIME-Typ / Prüfsumme:
application/pdf

Technische Metadaten:

Copyright Datum:
-

Copyright Info:
-

Lizenz:
-

Externe Referenzen

einblenden:

Urheber

einblenden:

ausblenden:

Urheber:
Blasi, Damián E.¹, Autor
Cotterell, Ryan, Autor
Wolf-Sonkin, Lawrence, Autor
Stoll, Sabine, Autor
Bickel, Balthasar, Autor
Baroni, Marco, Autor

Affiliations:
1Linguistic and Cultural Evolution, Max Planck Institute for the Science of Human History, Max Planck Society, ou_2074311

Inhalt

einblenden:

ausblenden:

Schlagwörter: -

Zusammenfassung: Embedding a clause inside another (``}the girl [who likes cars [that run fast]] has arrived{'') is a fundamental resource that has been argued to be a key driver of linguistic expressiveness. As such, it plays a central role in fundamental debates on what makes human language unique, and how they might have evolved. Empirical evidence on the prevalence and the limits of embeddings has however been based on either laboratory setups or corpus data of relatively limited size. We introduce here a collection of large, dependency-parsed written corpora in 17 languages, that allow us, for the first time, to capture clausal embedding through dependency graphs and assess their distribution. Our results indicate that there is no evidence for hard constraints on embedding depth: the tail of depth distributions is heavy. Moreover, although deeply embedded clauses tend to be shorter, suggesting processing load issues, complex sentences with many embeddings do not display a bias towards less deep embeddings. Taken together, the results suggest that deep embeddings are not disfavoured in written language. More generally, our study illustrates how resources and methods from latest-generation big-data NLP can provide new perspectives on fundamental questions in theoretical linguistics.

Details

einblenden:

ausblenden:

Sprache(n): eng - English

Datum: Online veröffentlicht: 2019-07Erschienen: 2019-07

Publikationsstatus: Erschienen

Seiten: 6

Ort, Verlag, Ausgabe: -

Inhaltsverzeichnis: -

Art der Begutachtung: Expertenbegutachtung

Identifikatoren: Anderer: shh2341
URI: https://www.aclweb.org/anthology/P19-1384

Art des Abschluß: -

ausblenden:

Titel: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics

Genre der Quelle: Konferenzband

Urheber:

Affiliations:

Ort, Verlag, Ausgabe: Florence, Italy : Association for Computational Linguistics

Seiten: - Band / Heft: - Artikelnummer: - Start- / Endseite: 3938 - 3943 Identifikator: -

Datensatz

Basisdaten

Dateien

Externe Referenzen

Urheber

Inhalt

Details

Veranstaltung

Entscheidung

Projektinformation

Quelle 1