Zipf - Measuring Japanese by the Token
Tokens & stuff.
Measurements made on Japanese corpora for a personal survey of Zipf's law and corpus linguistics. Published on GitHub (kenrinzero/zipf-measurements). These are unedited notes excerpted from a private research dossier, computed and written by Claude Fable 5 and 5.1; references to sections and source documents point to that dossier. Human contribution was minimal. Two known caveats: the agreement in M2 follows algebraically from the Heaps fit rather than being an independent test, and vocabulary growth curves at large N, so the β in M5 is a global average. On general prose the local exponent drops below 0.5 at the top of the range, where Guiraud's R declines from its peak.
These measurements were computed for a private dossier, not taken from the literature. Two corpora were used: a fixed 198.7 MB stream of Japanese texts from Aozora Bunko, tokenised into short-unit words (SUW) with MeCab (Kudo et al., 2004) and the UniDic dictionary (Den et al., 2008), giving 38,618,336 tokens, and a technical extract of Japanese Wikipedia built for the purpose from the engineering and chemistry categories (工学, 化学; 5,685,976 SUW tokens). The Aozora files were taken in archive order with no filter on copyright status. The Wikipedia extract holds 9,657 distinct articles; 48 of them, reachable from both categories, were collected twice, so 0.37% of its text appears twice. Diversity measures were computed with the lexical-diversity reference implementation by Kristopher Kyle (version 0.1.1) as the oracle, with faster exact re-implementations checked against it to within $10^{-6}$ on 1,000-, 10,000- and 100,000-token prefixes, then run over prefixes of the running text at log-spaced lengths from $10^{3}$ tokens up to the full corpus (eleven points for the general corpus, nine for the technical one). Measurements are labelled M1–M5.
M1. Technical text carries far more hapaxes, and lemmatisation removes few of them. At equal size (5,685,976 tokens each), the technical corpus has 134,602 surface types against 88,116 for general prose, 53% more. Its hapaxes are 47.0% of types against 30.6% (dis legomena 13.7% against 14.2%). Lemmatisation merges away 39% of the general corpus's types (88,116 to 54,113; hapax share falls to 26.8%) but only 9.6% of the technical corpus's (134,602 to 121,742; hapax share 46.7%, almost unchanged), so at lemma grain the technical corpus has 2.25 times as many types. Hapax types fall from 26,930 to 14,517 in the general corpus but only from 63,206 to 56,800 in the technical one. The measurement does not show why. UniDic's lemma merges both inflected forms and spelling variants. Words the dictionary does not know keep their surface form as their lemma, and they make up 11.2% of technical tokens against 0.66% of general ones. How much of the technical corpus's resistance to lemmatisation comes from its vocabulary, and how much from dictionary coverage, is not separated. The surface-level figures do not depend on lemmas.
M2. TTR collapses at exactly the Heaps rate. Over the general grid, raw TTR falls from 0.377 at 1,000 tokens to 0.0053 at the full corpus, a relative range of 315%. Its log-log slope against $N$ is $-0.4095$. The corpus's Heaps (1978) exponent, fitted independently, is $\beta = 0.592$, so $\beta - 1 = -0.408$: the two agree to 0.002, as Section 4.1 says they must. On the technical corpus the slope is $-0.326$, implying $\beta \approx 0.67$.
M3. Guiraud's index fails upward. Because $\beta$ is 0.59 and 0.67 rather than 0.5, Guiraud's (1954) $R = V/\sqrt{N}$ does not hold steady: it rises by 17.0% of its mean per tenfold increase in corpus size on general prose (from 11.9 to a peak of 37.2) and by 34.1% per decade on technical text. The failure is confined to measures that hard-code a growth exponent.
M4. MATTR, HD-D, and MTLD are all effectively length-independent at corpus scale. Drift per tenfold increase in corpus size, as a percentage of each measure's mean (general prose / technical text): MATTR (Covington & McFall, 2010) with $W = 500$, $-0.52$ / $+0.37$; MATTR with $W = 1000$, $-0.71$ / $+0.30$; HD-D (McCarthy & Jarvis, 2007), $+1.00$ / $+1.87$; MTLD (McCarthy & Jarvis, 2010), $+1.79$ / $-0.13$. Full-corpus values on general prose: MATTR-1000 0.380, MTLD 60.2, HD-D 0.819. All three drift by less than 2% per tenfold increase on both corpora, against over 60% for raw TTR. No one measure is flattest on both: MATTR with $W = 500$ drifts least on general prose, MTLD on technical text. McCarthy and Jarvis (2010) had found MTLD alone length-invariant on short English texts; at corpus scale in Japanese, MATTR and HD-D pass as well. HD-D's drift on general prose comes from its rise up to $10^{5}$ tokens (0.784 to 0.826); from there to the full corpus it stays between 0.811 and 0.826. At equal length HD-D separates the two domains clearly (0.811 general against 0.874 technical at $3 \times 10^{6}$ tokens), though it also has the largest technical drift of the three.
M5. Heaps' exponent for Japanese is not 0.5. From M2: $\beta = 0.59$ for general prose (fitted directly) and about 0.67 for technical text (from the TTR slope). The source document had listed "does $\beta \approx 0.5$ hold for morphologically rich languages?" as an open question; for these two corpora, segmented into short-unit words, the answer is that it fails in the upward direction. Guiraud's $R$, which assumes $\beta = 0.5$, fails with it; raw TTR would be length-independent only if $\beta$ were 1. The measures of M4 assume no growth rate and pass.
References
- Covington, M. A., & McFall, J. D. (2010). Cutting the Gordian knot: The moving-average type–token ratio (MATTR). Journal of Quantitative Linguistics, 17(2), 94–100.
- Den, Y., Nakamura, J., Ogiso, T., & Ogura, H. (2008). A proper approach to Japanese morphological analysis: Dictionary, model, and evaluation. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08) (pp. 1019–1024). Marrakech: European Language Resources Association. https://aclanthology.org/L08-1535/
- Guiraud, P. (1954). Les caractères statistiques du vocabulaire: Essai de méthodologie. Paris: Presses Universitaires de France. [B: WorldCat OCLC 8416329]
- Heaps, H. S. (1978). Information Retrieval: Computational and Theoretical Aspects. New York: Academic Press. [L, F]
- Kudo, T., Yamamoto, K., & Matsumoto, Y. (2004). Applying conditional random fields to Japanese morphological analysis. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing (pp. 230–237). Barcelona: Association for Computational Linguistics. https://aclanthology.org/W04-3230/
- McCarthy, P. M., & Jarvis, S. (2007). vocd: A theoretical and empirical evaluation. Language Testing, 24, 459–488.
- McCarthy, P. M., & Jarvis, S. (2010). MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, 42(2), 381–392. [L, F]
Data and software
- Aozora Bunko (青空文庫). https://www.aozora.gr.jp/
- Japanese Wikipedia (ウィキペディア日本語版), categories 工学 and 化学. https://ja.wikipedia.org/
- Kyle, K. lexical-diversity (version 0.1.1) [Python package]. https://github.com/kristopherkyle/lexical_diversity
- UniDic. National Institute for Japanese Language and Linguistics. https://clrd.ninjal.ac.jp/unidic/
- kenrinzero. zipf-measurements [code and data]. https://github.com/kenrinzero/zipf-measurements