Universitat Politècnica de Catalunya. Departament de Ciències de la Computació
Universitat Politècnica de Catalunya. LQMC - Lingüística Quantitativa, Matemàtica i Computacional
2025-09-05
Here we present a new class of optimality for coding systems. Members of that class are displaced linearly from optimal coding and thus exhibit Zipf's law, namely a power-law distribution of frequency ranks. Within that class, Zipf's law, the size-rank law and the size-probability law form a group-like structure. We identify human languages that are members of the class. All languages showing sufficient agreement with Zipf's law are potential members of the class. In contrast, there are communication systems in other species that cannot be members of that class for exhibiting an exponential distribution instead but dolphins and humpback whales might. We provide a new insight into plots of frequency vs. rank in double logarithmic scale. For any system, a straight line in that scale indicates that the lengths of optimal codes under non-singular coding and under uniquely decodable encoding are displaced by a linear function whose slope is the exponent of Zipf's law. For systems under compression and constrained to be uniquely decodable, such a straight line may indicate that the system is coding close to optimality. We provide support for the hypothesis that Zipf's law originates from compression and define testable conditions for the emergence of Zipf's law in compressing systems.
This research is supported by a recognition 2021SGR-Cat (01266 LQMC) from AGAUR (Generalitat de Catalunya) and the grant AGRUPS-2025 from Universitat Politècnica de Catalunya.
Peer Reviewed
Postprint (author's final draft)
Article
Anglès
Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Llenguatge natural; Exponential distribution; Probability; Optimal coding; Zipf's law; Human languages
Institute of Physics (IOP)
https://iopscience.iop.org/article/10.1209/0295-5075/adfa3e
http://creativecommons.org/licenses/by-nc-nd/4.0/
Restricted access - publisher's policy
Attribution-NonCommercial-NoDerivatives 4.0 International
E-prints [72986]