Words - Indefinably Obvious

September 21, 2026  ·  Kenrin  ·  #linguistics#morphology#word#typology

Talking about words.

Bloomfield's definition of the "word," that it is a minimum free form, the smallest unit that can occur on its own as an utterance,[1] is more or less applicable to English, but falls apart when we consider other languages. How, then, can we define "word" cross-linguistically? Dixon and Aikhenvald argue that most confusion about "word" comes from not separating three things: the lexeme from its inflected forms (look, looks, looked, looking are forms of one lexeme); the orthographic word, "something written between two spaces," from linguistic units; and, within linguistic units, the phonological word from the grammatical word.[2]

The phonological word is defined by no single criterion in every language. Dixon and Aikhenvald explain it as "a phonological unit larger than the syllable (in some languages it may minimally be just one syllable) which has at least one (and generally more than one) phonological defining property chosen from the following areas": (a) segmental features, such as internal syllable structure, word-boundary phenomena, and pause phenomena; (b) prosodic features, such as stress or tone assignment, nasalisation, retroflexion, and vowel harmony; and (c) phonological rules, some of which apply only within a phonological word and others (external sandhi) only across a boundary. Which criteria matter, and how much, differs from language to language.

The grammatical word is defined as follows: "A grammatical word consists of a number of grammatical elements which: (a) always occur together, rather than scattered through the clause (the criterion of cohesiveness); (b) occur in a fixed order; (c) have a conventionalised coherence and meaning."

Dixon and Aikhenvald distinguish four (well, three + "complex") kinds of relationship between these two:

Phonological and grammatical criteria may converge on the same unit. They interpret Newman's description of Yokuts this way: Newman separately gives phonological and morphological criteria for delimiting what he treats as a single "word unit," although he does not himself formulate this as a distinction between phonological and grammatical words.[3]

A phonological word may contain more than one grammatical word. This happens with clitics: grammatical words which don't constitute phonological words on their own, but attach phonologically to a word associated with another grammatical word.

One grammatical word may contain more than one phonological word. This is common in compounds and reduplicated forms in some languages: grammatical criteria treat the whole construction as one word, while phonological rules reveal an internal phonological-word boundary. Dixon and Aikhenvald cite Yidiɲ as one such case, where a phonological rule counting syllables shows a single grammatical word split into two phonological words.

More complex relationships are also possible. In Fijian, the derivational prefix i- is grammatically part of the noun i-sele 'knife', but phonologically it combines with the preceding article a to form the phonological word ai. The grammatical boundary therefore falls before i-, while the phonological boundary falls after it, so neither kind of word consists of a whole number of instances of the other.

Writing conventions place spaces where phonological and grammatical unity are felt to coincide; Pike's ideal, which Dixon and Aikhenvald quote, is that spaces should fall between grammatical words.[4] But what if the two kinds of unit diverge? Well, in that case, orthography has to choose.

The Word Comes First

"Why is it that the element of language which the naive speaker feels that he knows best is the one about which linguists say the least?" asks Bolinger.[5] "To the untutored person, speaking is putting words together, writing is a matter of correct word-spelling and word-spacing, translating is getting words to match words, meaning is a question of word definitions, and linguistic change is merely the addition or loss or corruption of words."

On one side of the coin, words can't be described in a "linguistic vacuum": phonemes and phrase structures are few enough to inventory exhaustively, but the stock of words is open-ended, so attempts at total accountability have been "discouraging": the phonologists' endeavor to set up a "phonological word" has been "rough indeed," and the generativists' rewrite rules often "left with no recourse except to list a few examples and add "etc.". On the other side, words "seek a cultural relevance:" the average person knows his culture and its designations, so crude reference ("word A associated with event A'") is the intuitive basis of the word concept and a good starting point.

According to Bolinger, during early language acquisition, an infant does not learn phonemes; rather, communication begins with approximations of words. "A word is recognized as a mother's face is recognized, not in terms of features that are repeated or not repeated from face to face, but as a whole." (Modern work on infant phonetic categorization has made the picture substantially more complicated, but for the sake of the argument, we'll conveniently just ignore that.) Bolinger's conclusion is that practical phonemicization is a gradual crystallization of repeated muscular habits out of word substance. Nobody rewards a child for better phonemes; the child's phonemic improvement is incidental to word improvement. The word comes first and "never yields quite all its distinctiveness to its later phonemic mapping," and traces of this fundamentally amorphous, word-level priority persist into adult language.

Word, Word-form and Lexeme

Take the opening of Yeats's Sailing to Byzantium:[6] "That is no country for old men." You can count seven items built of syllables and phonemes; that is the word in the phonological sense (not the same as Dixon and Aikhenvald's phonological word, though (word-form as phonological string versus phonological word as prosodic domain); "we are describing a 'word' in terms of phonological units: syllables and ultimately letters or phonemes"). Matthews poses that morphology goes wrong when we conflate three things the word "word" involves.[7]

First, the word-form (sense 1): the concrete phonological or orthographic unit, analysable into syllables, phonemes, or letters: dies and died are two different word-forms. Second, the lexeme (sense 2): the abstract dictionary unit that subsumes its varying forms: DIE, MAN, Latin AMO. Written in small capitals, a lexeme has no phonological composition at all; its properties are syntactic class and meaning, and the convention by which we cite it (Nominative Singular for Latin nouns, Infinitive for French verbs, root for Sanskrit) is mere symbolism, irrelevant to the lexeme's identity. Third, the grammatical word (sense 3): a lexeme specified for particular morphosyntactic categories, identified by a verbal formula like "the Past Participle of BEGET." The strict word/word-form/lexeme triplet is what matters; Matthews says it would be pedantic to police senses 1 vs. 3 everywhere, since context usually disambiguates.

These distinctions let him define homonymy precisely: one word-form corresponding to more than one grammatical word. Within a lexeme this is syncretism: English Past Tense vs. Past Participle (tried/tried, or the odder pattern in come), Latin Neuter Nom./Acc. identity, Singular/Plural in sheep. Across lexemes, when entire paradigms coincide, you get lexical homonymy: MATCH₁ "contest" vs. MATCH₂ "firestick," with forms homonymous at every point.

The distinctions matter in: (1) lexicography, where the OED's structure implicitly rests on them; (2) frequency counting, where "counting words" can mean counting lexemes, word-forms, or grammatical words, with very different results; his parable of counting verbs in Henry James by computer shows a naive type/token program conflating the verb means with the noun means and long the adjective with LONG the verb; (3) concordances, where headings should ideally be lexemes, but grammatical information (Tense, Transitivity, Object) may also be what users want; (4) collocation, where sometimes lexeme-pairs are right (SAUTE with POTATO, regardless of text distance) but sometimes tense matters, so collocation isn't always lexeme-to-lexeme; and (5) semantic theory, where collocational restrictions are evidence for meaning.

Matthews cares only that we stop saying "word" when we mean lexeme or word-form, he's indifferent to Bolinger's psychological gestalt. Together they separate the psychological question (is the word real?) from the analytic one (which sense of the "word" is doing the work?).

White Space

One of the main sources of confusion in establishing what a "word" is is the assumption that what we see on paper corresponds to words. The orthographic word is the unit delimited by spaces or their equivalent in writing, but spaces are not a direct reflection of linguistic structure, spaces are a convention.

For example, according to the opening's Bloomfield definition, "quick" is a word because it's free and indivisible. "Quickly" is also a word because, while it can be split into "quick" + "-ly," the suffix "-ly" is bound and cannot be uttered alone with meaning. "Blackbird," on the other hand, contains "black" and "bird," both free forms, however, its initial stress distinguishes it from the phrase "black bird." That prosodic pattern means "blackbird" cannot be fully decomposed into free parts while preserving the same form and meaning. However, Bloomfield's criterion that explicitly states that "A speaker may pause between words but not within a word," should be treated with caution as it only partially accounts for the suprasegmental features; while pausing between words is more likely than within words, a pause between morphemes is not unimaginable (e.g., "un- suitable," as per Dixon and Aikhenvald), and these limitations are even clearer cross-linguistically.

In agglutinative languages like Turkish or Finnish, a phonological and grammatical word can convey what English expresses as a whole phrase. The Turkish form ev-ler-im-den ("from my houses") consists of a single root and three bound suffixes. None of the suffixes can stand alone, but together they constitute one word. (As a native Hungarian speaker, I could also add Hungarian, since a compound like könyvtárban, where könyv alone would take -ben but the suffix harmonizes with tár in the compound, is a single grammatical word containing two vowel-harmony domains; however Hungarian is just kinda boring to me.)

In the case of polysynthetic languages like Inuktitut (ᐃᓄᒃᑎᑐᑦ), that can combine a lexical base with numerous bound morphemes to express what English would distribute across an entire clause.

German illustrates how orthography buries complexity, since in German, the rule of writing compounds conceals their internal morphological makeup. A reader sees Feierabend and intuitively treats it as one word, which it is phonologically and grammatically, but the morphological reality of two free stems, is not visible in the spelling, because orthography suppresses the boundary, even if the morphological identity of Feier and Abend stays recognizable.

Inuktitut syllabics (qaniujaaqpait), on the other hand, is an abugida in which most syllabic characters represent CV (or V) units, while standalone coda consonants are represented by smaller 'final' characters. A polysynthetic word is written as a long, internally unspaced string of syllable glyphs. For example, tusaatsiarunnanngittualuujunga (ᑐᓵᑦᓯᐊᕈᓐᓇᖖᒋᑦᑐᐊᓘᔪᖓ, "I cannot hear very well," example copied from Wikipedia, because I don't actually know Inuktitut) is one grammatical and phonological word whose many characters can look, to an outsider, like a sequence of visual units, however, that's not to be confused with wordhood.

Mandarin has no word-delimiting spaces at all, and that makes the point even more obvious. The characters of its logographic script prototypically correspond to syllable-morphemes, but not necessarily a word, and it provides no explicit morphological boundaries. 北京大学 (Běijīng Dàxué, "Peking University") is generally segmented as 北京 + 大学, two words. Other divisions of the same characters are not equally viable: 北大 is a real form, but 北大 + 学 is not a plausible parse of the string. In something like 美国会, which can be read as 美国 + 会 ("the United States will") or as 美 + 国会 ("the U.S. Congress"), the ambiguity becomes very easy to see. Cues (phonological and syntactic) help; third-tone sandhi is a plausible diagnostic, as in nǐ hǎo becoming ní hǎo, but the process depends on the domain. It applies inside lexical words but also, depending on speech rate and phrasing, inside larger phonological phrases. Japanese doesn't use spaces either, but in the case of Japanese, orthography encodes morphology without ever delimiting words with alternation between kanji and okurigana.

English is somewhere in the middle. Spacing is conventional, as in ice cream versus ice-cream, and even space-delimited units can be ambiguous, as in pick up as one word or two, since it can constitute a single lexical unit (potentially analysed as a multiword lexeme) while still consisting of two orthographic and usually phonological words. It is the default assumption that writing reflects speech, and that assumption fails wherever the two systems are misaligned. Therefore, we can (for now) conclude that the word is psychologically and culturally salient, but linguistically it decomposes into several only partially overlapping objects, and writing tricks us into treating those objects as identical. There may be no single universal linguistic object called the word, but languages repeatedly produce clusters of phonological, grammatical, morphological, semantic, psychological and orthographic cohesion that humans experience as word-like. "Word" may therefore be less a natural kind than a recurrent convergence point. So, is this just Haspelmath?[8] Well, kinda close, but not quite. Haspelmath's concern is the validity of word as a comparative morphosyntactic concept, but we have another thread to follow, or a slightly different approach, a more Bolingerian one, I suppose: why the concept remains phenomenologically obvious even when analytical attempts to define it break apart. Either way, words are pretty cool, right?

Footnotes

[1] Leonard Bloomfield, Language (New York: Holt, Rinehart and Winston, 1933). The minimum free form definition is the usual textbook citation of Bloomfield's account of the word.

[2] R. M. W. Dixon and Alexandra Y. Aikhenvald, "Word: A Typological Framework," in Word: A Cross-linguistic Typology, ed. R. M. W. Dixon and Alexandra Y. Aikhenvald (Cambridge: Cambridge University Press, 2002), 1–41.

[3] Stanley Newman, "Yokuts," Lingua 17 (1967): 182–99.

[4] Kenneth L. Pike, Phonemics: A Technique for Reducing Languages to Writing (Ann Arbor: University of Michigan Press, 1947).

[5] Dwight L. Bolinger, "The Uniqueness of the Word," Lingua 12, no. 2 (1963): 113–36, https://doi.org/10.1016/0024-3841(63)90022-6.

[6] W. B. Yeats, "Sailing to Byzantium" (1927); collected in The Tower (1928).

[7] P. H. Matthews, Morphology, 2nd ed. (Cambridge: Cambridge University Press, 1991), chap. 2; first edition 1974. The word-form / lexeme / grammatical-word distinctions developed here are Matthews's.

[8] Martin Haspelmath, "The Indeterminacy of Word Segmentation and the Nature of Morphology and Syntax," Folia Linguistica 45, no. 1 (2011): 31–80, https://doi.org/10.1515/flin.2011.002.