Zipf - Unnecessarily Obsessive Citation Hunt

August 30, 2026  ·  Kenrin  ·  #zipf#statistics#axtell#firms

Talking about Axtell, Zips, firms and numbers.

Citation Forensics

My inquiry into this subject commenced with an unforeseen discordance in the written record. Seemingly inconsequential, evidently still enough to spike my cortisol. I came across a while ago a sentence from Axtell's 2001 paper "Zipf Distribution of U.S. Firm Sizes,"[1] which prompted me to look further into the topic, yet it appears nowhere in the version published in Science, whose publication history runs "30 April 2001; accepted 9 August 2001." However, I wasn't hallucinating, because a distinct manuscript, "U.S. Firm Sizes are Zipf Distributed,"[2] submitted to Nature and dated 22 February 2001, does contain the sentence in question: "The stability of this distribution over time makes it, along with the distribution of city sizes, perhaps the most robust statistical regularity in all the social sciences."

As my awakened inner Kindaichi Hajime later discovered, this version, too, proved derivative: the line is presented as a citation drawn from Ijiri and Simon's 1977 monograph Skew Distributions and the Sizes of Business Firms.[3] Drawn, but not copied verbatim, because that does not contain it either. Instead, it includes two different comments of relevance.

The first (Ijiri and Simon, p. 2, the one actually cited by Axtell) presents Pareto or rank-size behavior as "a regularity in social phenomena that is both striking and observable in a number of quite diverse situations." The second (Ijiri and Simon, p. 13, relevance inferred by me) notes that "the degree of concentration in large-scale American industry has changed very little over a considerable period of years." Axtell's formulation can, with near precision, be reconstructed as a compression of these two claims.

None of this proved germane to the actual investigation, but I spent an inordinate number of hours in pursuit of a sentence of apocryphal provenance; so now, by rights, you must also bear this knowledge henceforth.

Contrasting the two versions, I believe it's most likely that the sentence was excised during revision. In the Nature draft the assertion stands at the end of the second paragraph, with footnote 1 pointing to Ijiri and Simon (p. 2); in the Science version the corresponding passage is simply a list that runs from percolation and immune response to word usage, city sizes and even Internet traffic (note to self: would be fun to look into this later, particularly into the differences between 2001 and today). The superlative has vanished without a trace. Ijiri and Simon lingers as reference 1 but now supports mergers-and-acquisitions immunity and the Yule distribution. What this presumably suggests is that an editor flagged the claim as unsupported. But, was it really unsupported? We shall see.

A further point of interest concerns the treatment of Herbert Simon. The Nature draft contains a formal acknowledgments section that credits Simon with having first stimulated Axtell's engagement with the topic. The Science article, accepted on 9 August and published on 7 September 2001, instead carries a 31st reference that conveys essentially the same, but now refers to "the late H. A. Simon." Given that Simon died on 9 February 2001, the Nature manuscript was plausibly composed prior to that date and submitted on 22 February, and at some point during the revision for Science, the living acknowledgment was converted into a posthumous reference note. Although, since the earlier version is more of a dedication than a conventional acknowledgment, the matter may simply be one of improved phrasing.

Simon's name will resurface in subsequent articles (probably). His 1955 stochastic model ("On a Class of Skew Distribution Functions")[4] explains how Zipf-like distributions can arise: new groups occasionally enter a system, while existing groups grow at rates proportional to their current size; a larger firm, city, or word category is more likely to receive the next increment simply because it's already large. Repeated over time, this produces many small groups, fewer medium-sized ones, and a small number of large ones, getting close to power-law distribution.

Zipf

Let me actually explain what we're trying to talk about. Zipfian distribution describes a rank-frequency relationship of the form f(r) \propto r^{-\alpha}, where f(r) denotes the frequency (or probability) of the item occupying rank r in a descending frequency-ordered list, and \alpha is a positive exponent typically close to unity. Imagine you open a book, count how often each word appears, and list them from most common to least. Zipf's law says the second-most-common word appears about half as often as the most common one, and the hundredth-most-common word about one-hundredth as often.

The law is named after George Kingsley Zipf (1902–1950), a linguist at Harvard University who systematically documented this regularity in his 1932 publication Selected Studies of the Principle of Relative Frequency in Language,[5] expanded on it in his 1935 monograph The Psycho-Biology of Language: An Introduction to Dynamic Philology,[6] and elaborated its theoretical foundations in Human Behavior and the Principle of Least Effort: An Introduction to Human Ecology (1949).[7] The empirical observation, however, predates Zipf. The French stenographer Jean-Baptiste Estoup noted the inverse rank-frequency pattern as early as 1916 in Gammes sténographiques.[8] Nobody is ever going to read that paragraph, so it is safe to mention that I'm listening to Kurusu Natsume ASMR while writing this. In 1913 the German physicist Felix Auerbach recorded a similar inverse proportionality between city population sizes.[9] Godfrey Dewey's Relativ Frequency of English Speech Sounds (1923)[10] and E. U. Condon's "Statistics of Vocabulary" (1928)[11] also documented the regularity, independently. Zipf himself, as far as I know, never claimed priority.

Axtell

Returning to Axtell, the main contribution of "Zipf Distribution of U.S. Firm Sizes" was to replace the prevailing lognormal account of firm sizes with a Zipf distribution. That earlier consensus was based on COMPUSTAT, but that covers only about 10,800 publicly traded U.S. firms, and thus it produces an approximately lognormal size distribution. By contrast, the 1997 U.S. Census recorded some 5.5 million employer firms (so, roughly five hundred times COMPUSTAT's coverage) and a distribution in which the number of firms rises as size falls. That's not a lognormal shape. That's logcrazy. Mean employment is 4,605 in COMPUSTAT versus 19.0 in the Census population.

When using data on the entire population of tax-paying U.S. firms, their sizes appear to become Zipf distributed: Pr[larger than s] ∝ 1/s, for multiple years and for various definitions of firm size.

So, was Axtell right in the original manuscript? Does the robustness claim hold, or was it better omitted? To assess this, we recomputed the headline 1997 results on current Census Statistics of U.S. Businesses (SUSB) public size tables using the binned-data maximum-likelihood estimator of Virkar and Clauset (2014),[12] and by "we," I mean AI, because I wouldn't be high IQ enough. The bulk of the work was done by GLM-5.3-Flash and GLM-5.3 with GPT-5.6 Sol and Claude Fable 5 providing assistance. We were trying to answer whether Zipf persists from 1997 into the 2020s, and whether the gap between ordinary least squares (OLS) and maximum-likelihood estimates (MLE) matters at Census bin widths. Cheap, deterministic, one data fetch, I assumed, blissfully unaware what this would actually entail (as said, low IQ). Our verdict, as GLM-5.3 put more aptly than I ever could, is as follows: "modern public SUSB data support a near-Zipf upper tail (ζ̂ ≈ 0.96–0.98, ζ = 1 within total-procedure uncertainty every year) and contradict the 'all the way down' reading."

The analysis is easily reproducible using the data from the project repository. Source code is under the MIT License; the accompanying analysis document (REPORT.md)[13] and interactive explorer[14] are under CC-BY-4.0. The raw Census SUSB files are U.S. government works in the public domain and can be freely redistributed, but they amount to around 1.3 GBs, so they're not vendored. Estimation uses only U.S.-total rows.

The Test, Briefly

One-employee upward the exponent is 0.55, not 1. Starting at five employees instead raises it to 0.87-0.92. The line does hold later on, but later is around 500 to 2,000 employees. From there up, a power law passes goodness-of-fit in 25 of 26 years, with exponent at 0.96-0.98. 1 is inside the uncertainty interval every year from 1997 to 2022. The lognormal, the old COMPUSTAT consensus, is favored against the power law in no year on that tail. The data supports Zipf from a few hundred employees up, but below that, it's visibly flat, like the Earth.

The Science paper never ran a goodness-of-fit test and never compared Zipf against any alternative on its own data. I would not have either without AI, to be fair. The quantitative basis is one binned histogram with a least-squares line drawn through all of it, and the line as declared to fit "all the way down," but nobody actually checked the bottom.

Looking back at where the excised sentence came from, Ijiri & Simon was talking about the stability of concentration among large firms. That's precisely what a quarter century of SUSB data show: the largest single-year move in the series is the 2020-2021 step, recovered by 2022. The sentence Axtell for whatever reason decided to cut claimed stability over time, but the replication test finds that while the tail is stable 1997-2022, the full distribution isn't Zipf at all. As such, our conclusion is the following: the excised superlative was half right, and the half that was wrong is the half that Axtell ultimately kept.

Footnotes

[1] Robert L. Axtell, "Zipf Distribution of U.S. Firm Sizes," Science 293, no. 5536 (September 7, 2001): 1818–20, https://doi.org/10.1126/science.1062081.

[2] Robert L. Axtell, "U.S. Firm Sizes are Zipf Distributed" (manuscript submitted to Nature, February 22, 2001), https://faculty.sites.iastate.edu/tesfatsi/archive/tesfatsi/USFirmSizesAreZipfDistributed.RAxtell2001.pdf.

[3] Yuji Ijiri and Herbert A. Simon, Skew Distributions and the Sizes of Business Firms (Amsterdam: North-Holland, 1977).

[4] Herbert A. Simon, "On a Class of Skew Distribution Functions," Biometrika 42, no. 3/4 (December 1955): 425–40, https://doi.org/10.2307/2333389.

[5] George Kingsley Zipf, Selected Studies of the Principle of Relative Frequency in Language (Cambridge, MA: Harvard University Press, 1932).

[6] George Kingsley Zipf, The Psycho-Biology of Language: An Introduction to Dynamic Philology (Boston: Houghton Mifflin, 1935).

[7] George Kingsley Zipf, Human Behavior and the Principle of Least Effort: An Introduction to Human Ecology (Cambridge, MA: Addison-Wesley, 1949).

[8] J.-B. Estoup, Gammes sténographiques, 4th ed. (Paris: Institut Sténographique, 1916). The rank-frequency observations Zipf later cited appear in this edition's theoretical fascicule; earlier editions of the drill book (from 1907/1912) are often cited in error.

[9] Felix Auerbach, "Das Gesetz der Bevölkerungskonzentration," Petermanns Geographische Mitteilungen 59 (1913): 74–76. English translation: Antonio Ciccone, trans., "The Law of Population Concentration," Environment and Planning B: Urban Analytics and City Science 50, no. 2 (2023), https://doi.org/10.1177/23998083221147139.

[10] Godfrey Dewey, Relativ Frequency of English Speech Sounds (Cambridge, MA: Harvard University Press, 1923).

[11] E. U. Condon, "Statistics of Vocabulary," Science 67, no. 1733 (March 16, 1928): 300, https://doi.org/10.1126/science.67.1733.300.

[12] Yogesh Virkar and Aaron Clauset, "Power-Law Distributions in Binned Empirical Data," Annals of Applied Statistics 8, no. 1 (2014): 89–119, https://doi.org/10.1214/13-AOAS710.

[13] "REPORT.md," kenrinzero/axtell-zipf-susb, GitHub, accessed August 30, 2026, https://github.com/kenrinzero/axtell-zipf-susb/blob/main/REPORT.md.

[14] "axtell-zipf-susb," GitHub Pages, accessed August 30, 2026, https://kenrinzero.github.io/axtell-zipf-susb/.