The risks of mixing dependency lengths from sequences of different length

Mixing dependency lengths from sequences of different length is a common practice in language research. However, the empirical distribution of dependency lengths of sentences of the same length differs from that of sentences of varying length. The distribution of dependency lengths depends on senten...

Full description

Bibliographic Details
Authors: Ferrer Cancho, Ramon|||0000-0002-7820-923X, Liu, Haitao
Format: article
Publication Date:2014
Country:España
Institution:Universitat Politècnica de Catalunya (UPC)
Repository:UPCommons. Portal del coneixement obert de la UPC
Language:English
OAI Identifier:oai:upcommons.upc.edu:2117/28279
Online Access:https://hdl.handle.net/2117/28279
https://dx.doi.org/10.1515/glot-2014-0014
Access Level:Open access
Keyword:Computational linguistics
Syntactic dependency
Syntax
Dependency length
Lingüística computacional
Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Llenguatge natural
Description
Summary:Mixing dependency lengths from sequences of different length is a common practice in language research. However, the empirical distribution of dependency lengths of sentences of the same length differs from that of sentences of varying length. The distribution of dependency lengths depends on sentence length for real sentences and also under the null hypothesis that dependencies connect vertices located in random positions of the sequence. This suggests that certain results, such as the distribution of syntactic dependency lengths mixing dependencies from sentences of varying length, could be a mere consequence of that mixing. Furthermore, differences in the global averages of dependency length (mixing lengths from sentences of varying length) for two different languages do not simply imply a priori that one language optimizes dependency lengths better than the other because those differences could be due to differences in the distribution of sentence lengths and other factors.