Japanese Typewriters

Thursday, April 22nd, 2010

With several thousand characters to contend with, how were the Japanese able to use typewriters before the advent of digital technology?

The answer is the kanji typewriter, which was invented by Kyota Sugimoto in 1915. This invention was deemed so important that it was selected as one of the ten greatest Japanese inventions by the Japanese Patent Office during their 100th anniversary celebrations in 1985.

The kanji typewriter used separate metal pieces as strikers, somewhat like movable type. They were arranged in a grid in the tray beneath the typewriter:

One of the things that made the typewriter difficult to use was getting the strike lever force correct. If struck with even just regular force, characters such as decimal points or punctuation would pierce the ribbon and paper, becoming stuck in the rubber platen. On the other hand, very complex characters required striking with additional force to compensate for the large surface area of the typeface. This combined with the huge number of characters (which makes hunt and peck typing on a QWERTY keyboard seem trivial) meant that only experienced operators could use these typewriters.

How do you type in Japanese?

Tuesday, April 20th, 2010

The Japanese language is infamous for using thousands of Chinese pictographs, or kanji, combined with two separate phonetic systems, or kana — along with plenty of Latin characters, or romaji. So, how do you type in Japanese?

Because of the large number of characters in languages such as Japanese and Chinese (Japanese education standards define around 2000 characters for high school students and JIS standards define more than 6000 characters), characters cannot be entered directly, but instead need to be entered via some kind of conversion process. Modern computers and electronic equipment therefore employ a system whereby words are entered phonetically, and then converted into kanji through an interactive conversion process. I’m going to describe this process by using the Japanese IME included in Microsoft Windows XP.

Romaji-kana conversion is the process of converting Roman alphabetic key presses into Japanese phonetic characters. This is a relatively straightforward process in which the IME looks up Romaji character sequences in a simple table, and replaces sequences with the corresponding kana as soon as they are matched.

Now that the target word is spelled out phonetically in the composition string, the next step is to enter the kana-kanji stage of the process. This is achieved by simply hitting the space bar. What actually goes on inside the IME at this point is actually quite sophisticated and largely beyond the scope of this article. However, the basics are that the IME analyzes the grammar of our text, attempts to identify the separate words in the text (a process known as segmentation that is necessary because there are no spaces in Japanese), and then perform a context-sensitive look up of each of those words in its built-in dictionaries. The IME then picks the best matches for each segment, and displays them like so:

In this case, we’ve used such a common phrase that the IME has no trouble identifying the segmentation and the best candidates for each word. At this point, the IME has still not accepted our input, but is waiting for our approval of the suggested candidate characters. At this point we can either hit enter to accept, escape to return to the kana composition string, look through the other candidates that the IME has dug out of its dictionary and choose alternatives if necessary, or even adjust the location of the break between the two words.

Rosetta Stone

Friday, April 16th, 2010

The Rosetta Stone is a Ptolemaic-era stele discovered by Napoleon’s troops during their occupation of Egypt and famous for its role in deciphering ancient Egyptian script, because it includes the same message in hieroglyphs, Demotic script, and classical Greek.

But what kind of message deserves to be inscribed in granite-hard rock in three different writing systems?

In essence, the Rosetta Stone is a tax amnesty given to the temple priests of the day, restoring the tax privileges they had traditionally enjoyed from more ancient times. Some scholars speculate that several copies of the Rosetta Stone must exist, as yet undiscovered, since this proclamation must have been made at many temples. The complete Greek portion, translated into English, is about 1600–1700 words in length, and is about 20 paragraphs long (average of 80 words per paragraph):

In the reign of the new king who was Lord of the diadems, great in glory, the stabilizer of Egypt, but also pious in matters relating to the gods, superior to his adversaries, rectifier of the life of men, Lord of the thirty-year periods like Hephaestus the Great, King like the Sun, the Great King of the Upper and Lower Lands, offspring of the Parent-loving gods, whom Hephaestus has approved, to whom the Sun has given victory, living image of Zeus, Son of the Sun, Ptolemy the ever-living, beloved by Ptah;

In the ninth year, when Aëtus, son of Aëtus, was priest of Alexander and of the Savior gods and the Brother gods and the Benefactor gods and the Parent-loving gods and the god Manifest and Gracious; Pyrrha, the daughter of Philinius, being athlophorus for Bernice Euergetis; Areia, the daughter of Diogenes, being canephorus for Arsinoë Philadelphus; Irene, the daughter of Ptolemy, being priestess of Arsinoë Philopator: on the fourth of the month Xanicus, or according to the Egyptians the eighteenth of Mecheir.

THE DECREE: The high priests and prophets, and those who enter the inner shrine in order to robe the gods, and those who wear the hawk’s wing, and the sacred scribes, and all the other priests who have assembled at Memphis before the king, from the various temples throughout the country, for the feast of his receiving the kingdom, even that of Ptolemy the ever-living, beloved by Ptah, the god Manifest and Gracious, which he received from his Father, being assembled in the temple in Memphis this day, declared: Since King Ptolemy, the ever-living, beloved by Ptah, the god Manifest and Gracious, the son of King Ptolemy and Queen Arsinoë, the Parent-loving gods, has done many benefactions to the temples and to those who dwell in them, and also to all those subject to his rule, being from the beginning a god born of a god and a goddess — like Horus, the son of Isis and Osirus, who came to the help of his Father Osirus; being benevolently disposed toward the gods, has concentrated to the temples revenues both of silver and of grain, and has generously undergone many expenses in order to lead Egypt to prosperity and to establish the temples… the gods have rewarded him with health, victory, power, and all other good things, his sovereignty to continue to him and his children forever.

If I had tax amnesty, I too would want it written in stone, in triplicate.

Thou, Thee, Ye, You

Thursday, April 15th, 2010

I was recently reading The Civilising Mission, and I couldn’t help but notice that its header includes a line from Kipling’s “The White Man’s Burden”:

The blame of those ye better, The hate of those ye guard

I suppose most modern English speakers vaguely recognize ye as an archaic form of you:

You is the second-person personal pronoun in Modern English. Ye was the original nominative form; the oblique/objective form is you (functioning originally as both accusative and dative), and the possessive is your or yours.

I know you. Ye know me.

Ye is not the only archaic form of you we’ve got, as thou already knowest:

In standard English, you is both singular and plural; it always takes a verb form that originally marked the word as plural, such as you are. This was not always so. Early Modern English distinguished between the plural you and the singular thou.

This distinction was lost in modern English due to the importation from France of a Romance linguistic feature which is commonly called the T-V distinction. This distinction made the plural forms more respectful and deferential; they were used to address strangers and social superiors. This distinction ultimately led to familiar thou becoming obsolete in standard English, although this did not happen in other languages such as French.

Because thou is now seen primarily in literary sources such as the King James Bible (often directed to God, who is traditionally addressed in the familiar) or Shakespeare (often in dramatic dialogs, e.g. “Wherefore art thou Romeo?”), many modern anglophones erroneously perceive it as more formal, rather than familiar (case in point: in Star Wars: The Empire Strikes Back, Darth Vader addresses the Emperor saying, “What is thy bidding, my master?”).

Naturally, once y’all use you for both singular and plural, y’all need a new plural:

Because you is both singular and plural, various English dialects have attempted to revive the distinction between a singular and plural you to avoid confusion between the two uses. This is typically done by adding a new plural form; examples of new plurals sometimes seen and heard are y’all, or you-all (primarily in the southern United States and African American Vernacular English), you guys (in the U.S., particularly in Midwest, Northeast, and West Coast, in Canada, and in Australia; regardless of the genders of those referred to), you lot (in the UK), youse (in Scotland), youse guys (in the U.S., particularly in New York City region, Philadelphia, Michigan’s Upper Peninsula and rural Canada; also spelt without the E), and you-uns/yinz (Western Pennsylvania, The Appalachians).

I Pledge Allegiance To Linguistic Obfuscation

Saturday, April 3rd, 2010

I pledge allegiance to linguistic obfuscation, says Geoff Nunberg:

Obscurity has been built into the pledge since Francis Bellamy created it in 1892. It was ostensibly designed to rouse the patriotic attachments of schoolchildren, particularly the recent immigrants who might need extra encouragement. But Bellamy obviously wasn’t thinking of all those little Solomons, Svens and Sergios when he chose to start with the words “I pledge allegiance.” That was an arcane scrap of feudal English that had made its last appearance in the loyalty oath that Confederate soldiers had to sign to recover their rights after the Civil War. But the reference was obscure to most people even in Bellamy’s time, and the words have always been utterly opaque to schoolchildren.

In fact, “pledge allegiance” is what linguists call a hapax legomenon, or hapax for short — an expression that only occurs in a single place in the language, like wardrobe malfunction, Corinthian leather or satisfactual. Or let’s not leave out my favorite, ginchiest. People don’t pledge allegiance to Hadassah or the U.S. Marines or Kappa Kappa Gamma, much less to other inanimate objects. We only use the words when we’re either quoting the flag pledge or riffing on it. So there’s no independent reference point, no way to know what you’ve just signed on for that you weren’t down for already.

Of course kids are mystified by most everything in the pledge. But “one nation under God” has the distinction of being a phrase that not even grown-ups are clear on. Congress inserted the words at the height of the Cold War in 1954 to underscore the difference between American values and those of the atheistic Communists. But its actual meaning is up for grabs. Does it affirm our faith in God or assert that we have his special protection? Is it a ceremonial deist formula with no especial religious character? Or is it merely a historical nod to the beliefs of the founders, as the 9th Circuit majority said? You can take this wherever you like, because “under God” is another hapax legomenon that doesn’t occur anywhere else in modern English. People don’t say things like “Western Europe isn’t under God anymore,” or “She only goes out with men who are under God.”

That ambiguity has certain advantages. But it actually came about because of a linguistic misunderstanding. The words were taken from the Gettysburg Address, where Lincoln asked his listeners to resolve that “this nation, under God, shall have a new birth of freedom.” Except that in the Gettysburg Address, “under God” didn’t modify “this nation” but the following phrase, “have a new birth of freedom.” In Lincoln’s time, “under God” was a common idiom that meant “with God’s help” or “the Lord willing.” People used it to qualify a bald prediction or promise, mindful of the admonition against vainglory in the book of James.

Actually, my guess is that Lincoln would have inserted the words “under God” if he had written the Pledge of Allegiance, too, although he probably would have put them at the end. He would have been uncomfortable about describing the country as indivisible, just and free without adding a “God willing” somewhere.

I doubt if the people who pushed for inserting “under God” in the pledge realized they were changing the meaning of Lincoln’s words. Most of them would have had to learn the Gettysburg Address by heart back then, but nobody ever stopped to parse it, no more than children parse the Pledge of Allegiance now. Anyway, it doesn’t matter — what’s important is that Lincoln sanctified the words, however we’ve repurposed them. Whatever you take the phrase to mean, it gets the “G” word in there, which is enough to satisfy some people and offend others.

And it doesn’t matter much what schoolchildren make of the phrase, either. As Eric Hobsbawm once said, patriotic rituals exist to instill a sense of membership in a club, not to enumerate its bylaws.

Tongue twisters

Thursday, January 14th, 2010

The Economist looks at difficult languages — or, rather, languages that are difficult for English-speakers to learn — and raises an issue I pondered ages ago while studying first-, second-, and third-person conjugations in singular and plural:

A truly boggling language is one that requires English speakers to think about things they otherwise ignore entirely. Take “we”. In Kwaio, spoken in the Solomon Islands, “we” has two forms: “me and you” and “me and someone else (but not you)”. And Kwaio has not just singular and plural, but dual and paucal too. While English gets by with just “we”, Kwaio has “we two”, “we few” and “we many”. Each of these has two forms, one inclusive (“we including you”) and one exclusive. It is not hard to imagine social situations that would be more awkward if you were forced to make this distinction explicit.

Monkey Talk

Thursday, December 10th, 2009

I would not say that Campbell’s monkeys have a language with syntax simply because they can string three sounds — boom, krak, and hok — together:

“Krak” is a call that warns of leopards in the vicinity. The monkeys gave it in response to real leopards and to model leopards or leopard growls broadcast by the researchers. The monkeys can vary the call by adding the suffix “-oo”: “krak-oo” seems to be a general word for predator, but one given in a special context — when monkeys hear but do not see a predator, or when they hear the alarm calls of another species known as the Diana monkey.

The “boom-boom” call invites other monkeys to come toward the male making the sound. Two booms can be combined with a series of “krak-oos,” with a meaning entirely different to that of either of its components. “Boom boom krak-oo krak-oo krak-oo” is the monkey’s version of “Timber!” — it warns of falling trees.

There is yet another variation on this theme, Dr. Zuberbühler’s team reports. Into the “Timber!” call, the Campbell’s monkeys insert a series of up to seven “hok-oo” calls. The combined call indicates the presence of other monkey groups and is heard most often when the monkeys are on the edge of their home range.

The meaning of monkey calls was first worked out with vervet monkeys, which have distinct alarm calls for each of their three main predators: the martial eagle, leopards and snakes. But the vervets did not combine their alarm calls to generate new meanings, unlike human words that can be combined in an infinite number of different sentences.

No R Sounds At All

Thursday, October 29th, 2009

The first thing John Derbyshire noticed about his Wikipedia entry was that they got his name wrong:

Not the spelling — they at least managed to get that right — but the pronunciation. Their rendering in the International Phonetic Alphabet is /?d?rb???r / That includes two fricative-lingual r sounds. In fact there are no r sounds at all in the pronunciation of my name, fricative-lingual or otherwise. It is pronounced with pure vowels: /?d??b???/ (DAH-bi-shuh). I refer interested readers to §773 of Daniel Jones’ classic Outline of English Phonetics: “[I]n London English the r is never sounded when final or followed by a consonant.” The following §774, “Words for practising the omission of r,” is also helpful. Prof. Jones does not give a phonetic transcription of “Derbyshire” in standard English but he does, in §287, show /?d??b?/ for “Derby.”

What’s wrong with "employees"?

Sunday, October 11th, 2009

After hearing ordinary workers referred to as team members, cast members, associates, etc., you may ask, What’s wrong with employees?

The introduction to Charles Nordoff’s The Communistic Societies of the United States (1875) reminds us that even employees began as a euphemism:

Though it is probable that for a long time to come the mass of mankind in civilized countries will find it both necessary and advantageous to labor for wages, and to accept the condition of hired laborers (or, as it has absurdly become the fashion to say, employés), every thoughtful and kind-hearted person must regard with interest any device or plan which promises to enable at least the more intelligent, enterprising, and determined part of those who are not capitalists to become such, and to cease to labor for hire.

Roots of Language Run Deeper Than Speech

Thursday, July 9th, 2009

Researchers found that speakers of subject-verb-object languages — “Bill eats cake” — reverted to a subject-object-verb form when asked to communicate with their hands, which implies that the roots of language run deeper than speech:

Goldin-Meadow’s team asked forty people — ten speakers apiece of English, Mandarin Chinese and Spanish, each of which follows the SVO order, and ten speakers of Turkish, which follows an SOV order — to describe a series of simple actions, such as a girl turning a knob, with gestures.

Regardless of their native language, the subjects almost universally preceded object with verb: girl knob turns.

“We expected that the language they spoke would influence the language of their gestures, but it didn’t,” said Goldin-Meadow.

To test whether the subjects used subject-object-verb as a communicative strategy, they were given a series of illustrated transparencies, each depicting one element of a scene, and told that order was irrelevant: the final layered montage would look the same regardless of its assembly. The subjects still put object before verb.

“It almost speaks to the independence of language from thought,” said Goldin-Meadow.

Does language affect thought?

Friday, July 3rd, 2009

In discussing the language of clear thinking recently, I brought up the notion that language may affect thought. I found this example amusing:

“We gave people sets of pictures that showed some kind of temporal progression (e.g., pictures of a man aging, or a crocodile growing, or a banana being eaten). Their job was to arrange the shuffled photos on the ground to show the correct temporal order. We tested each person in two separate sittings, each time facing in a different cardinal direction. If you ask English speakers to do this, they’ll arrange the cards so that time proceeds from left to right. Hebrew speakers will tend to lay out the cards from right to left, showing that writing direction in a language plays a role. So what about folks like the Kuuk Thaayorre, who don’t use words like “left” and “right”? What will they do?

The Kuuk Thaayorre did not arrange the cards more often from left to right than from right to left, nor more toward or away from the body. But their arrangements were not random: there was a pattern, just a different one from that of English speakers. Instead of arranging time from left to right, they arranged it from east to west. That is, when they were seated facing south, the cards went left to right. When they faced north, the cards went from right to left. When they faced east, the cards came toward the body and so on. This was true even though we never told any of our subjects which direction they faced. The Kuuk Thaayorre not only knew that already (usually much better than I did), but they also spontaneously used this spatial orientation to construct their representations of time.”

The Language of Clear Thinking

Saturday, June 27th, 2009

Alfred Korzybski famously said that the map is not the territory. This is the key point of his general semantics: we should be conscious of the abstractions we use.

If we try to reason from the “essence” of something, in true Aristotelian style, we might abstract away meaningful complexity. If we apply binary logic, we may label things true or false when they are largely true or largely false, or likely true or likely false. Korzybski thus recommended what he called null-A, or non-Aristotelian logic.

The language we use also introduces many questionable abstractions, and Korzybski believed that ambiguous language lent itself to unclear thinking. Most infamously, he railed against unclear use of the verb to be, which led a former student of his to suggest a modified form of English, E-Prime, which eliminated to be entirely:

To exist or not to exist,
I ask this question.
— modified from Shakespeare’s Hamlet

Proponents of E-Prime believed that it would do more than clarify communication; they believed it would clarify thought. This is an example of the Sapir-Whorf Hypothesis, which suggests that language influences thought, and that some languages might lead to clearer thinking. That was the rationale behind Loglan, the logical language, with its grammar based on predicate logic.

These ideas soon found their way into Golden Age science fiction. A.E. von Vogt had his protagonists overcome their totalitarian foes through clever use of intuitive, inductive logic in The World of Null-A. (Peter Chung’s animated Æon Flux shares many motifs with The World of Null-A.)

Robert Heinlein also embraced many elements of general semantics, especially the notion of languages designed to improve thinking. In The Moon Is a Harsh Mistress, the self-aware computer receives its precise instructions in Loglan — which makes Loglan sound like a variant of Prolog. Heinlein took the Sapir-Whorf Hypothesis much further in his short story, Gulf, which posits a new language used by a race of supermen. The language, Speedtalk, uses every phoneme (sound) used in any human language, not the small subset that belongs to any one language, and maps every word in Basic English to its own phoneme.

Basic English also shows up in H.G. Wells’ The Shape of Things to Come as the lingua franca of the future. Similarly, it inspired the Newspeak of Orwell’s Nineteen Eighty-Four. (It doesn’t take much to turn Wells’ utopian ideas upside-down.)

Nyrath has much more to say about future languages, but I thought I’d end with this amusing bit of geekery:

Raphaël Poss (AKA “Kena”) took the obvious step and adapted the Tengwar alphabet to the Lojban set of phonemes. As Mr. Poss puts it: “…it is far more natural to write Lojban with a logical writing system…. the tengwar system inherently contains some main Lojban morphology rules, making Lojban easier to learn when it is written with tengwar.”

Consonants

Wednesday, June 24th, 2009

I remember thinking, when I was first learning the alphabet, that p and b represented similar sounds, and p and b were similar symbols, but t and d also represented similar sounds — similar in the same way as p and b, in fact — and they were not similar symbols. Why wasn’t the t sound represented by a flipped d, like a q?

Similarly, why weren’t k and g flipped versions of one another, more like ? and g? And why weren’t f and v flipped versions of one another, like f and ?, or ? and v?

At least the letters s and z were clearly related, even if they weren’t flipped versions of one another.

I was looking for the logic behind a not-so-logical, evolved system of writing, and it wasn’t there.

The quality I recognized, by the way, was voiceless versus voiced articulation. The sounds p, t, k, and f are voiceless — the larynx does not vibrate — while the sounds b, d, g, and v are voiced — the larynx does vibrate.

It’s actually pretty straightforward to sort the consonants by whether they’re made with both lips (bilabials like p and b), teeth and lip (labiodentals like f and v), teeth (dentals like t and d), or back of the roof of the mouth (velars like k and g).

Once you put the consonants into an organized table like that, you can’t help but think that the symbols should reflect the nature of each sound and how it’s made — which is what Alexander Graham Bell’s father effectively did, with his visible speech alphabet for the deaf:

When I first mentioned the visible speech alphabet, I noted that you could easily imagine it as some obscure elven written language from Tolkien’s Middle Earth. What I didn’t realize at the time was that Tolkien’s tengwar alphabet is also a logically laid out system:

In tengwar, similar sounds have similar symbols. It all would have made so much more sense to my four-year-old self than our own Latin alphabet.

(For geeky fun, you might enjoy this tengwar transcriber, by the way.)

History of Mathematical Notation

Tuesday, October 21st, 2008

Stephen Wolfram, creator of Mathmatica, discusses the History of Mathematical Notation and shares some interesting factoids:

  • The first representations for numbers that we recognize are notches in bones made 25,000 years ago, which worked in unary: to represent 7, you made 7 notches, and so on.
  • More than 5000 years ago, the Babylonians — and probably the Sumerians before them — had the idea of positional notation for numbers, but they used base 60 — not base 10 — which is presumably where our hours, minutes, seconds scheme comes from. But they had the idea of using the same digits to represent multiples of different powers of 60.
  • Neither the Babylonians nor the Egyptians had the idea to use characters for digits though: not to make up a 7 digit with 7 of something, and so on.
  • The Greeks — perhaps following the Phoenicians — did have this idea, though, but their version of the idea was to label the sequence of numbers by the sequence of letters in their alphabet. So alpha was 1, beta was 2, and so on. But this creates a serious versioning problem: even if you decide to drop letters from your alphabet, you have to leave them in your numbers, or else all your previously-written numbers get messed up. So that means that there are various obsolete Greek letters left in their number system: like koppa for 90 and sampi for 900.
  • In Roman numerals, the length of the representation of a number increases fractally with the size of the number.
  • There was a serious conceptual problem with letters as numbers: it made it difficult to invent the concept of symbolic variables, because any letter one might use for a symbolic variable could be confused with a piece of the number.
  • There are a few hints of Hindu-Arabic notation in the mid-first-millennium AD, but it didn’t get really set up until about 1000 AD, and it didn’t really come to the West until Fibonacci wrote his book about calculating around 1200 AD.
  • The idea of breaking digits up into groups of three to make big numbers more readable is already in Fibonacci’s book from 1202, though he suggested using overparens on top of the numbers, not commas in the middle.
  • Algebraic variables didn’t get started until Vieta at the very end of the 1500s, and they weren’t common until way into the 1600s. So that means people like Copernicus didn’t have them. Nor for the most part did Kepler.
  • Even though math notation hadn’t gotten going very well by their time, the kind of symbolic notation used in alchemy, astrology, and music pretty much had been developed. So, for example, Kepler ended up using what looks like modern musical notation to explain his “music of the spheres” for ratios of planetary orbits in the early 1600s.
  • Starting with Vieta and friends, letters routinely got used for algebraic variables. Usually, by the way, he used vowels for unknowns, and consonants for knowns.
  • Vieta wrote out polynomials in a symbolic algebra scheme he called zetetics, using words for the operations, partly so the operations wouldn’t be confused with the variables.
  • The Babylonians didn’t usually use operation symbols: for addition they juxtaposed things, and they tended to put things into tables so they didn’t have to write out operations.
  • The Egyptians did have some notation for operations — they used a pair of legs walking forwards for plus, and walking backwards for minus.
  • The modern + sign — which was probably a shorthand for the Latin et for and — doesn’t seem to have arisen until the end of the 1400s.
  • In the early to mid-1600s there was kind of revolution in math notation, and things very quickly started looking quite modern. Square root signs got invented: previously Rx — the symbol we use now for medical prescriptions — was what was usually used.
  • A fellow called William Oughtred, who taught Christopher Wren, invented the cross for multiplication.
  • Newton invented the idea that you can write negative powers of things instead of one over things and so on.
  • Leibniz had been using omn., presumably standing for omnium, for integrals, but in 1675 he created the modern integral sign, the elongated S or ?. Then on Thursday November 11 of the same year, he wrote down the d for derivative. Actually, he said he didn’t think it was a terribly good notation, and he hoped he could think of a better one soon. But as we all know, that didn’t happen.
  • Euler, in the 1700s, was the first serious user of Greek as well as Roman letters for variables.
  • Euler popularized the letter ? (pi) for the famous constant — a notation that had originally been suggested by a character called William Jones, who thought of it as a shorthand for perimeter.

The 25 Most Commonly Used English Words

Wednesday, October 15th, 2008

The 25 Most Commonly Used English Words make up about one-third of the words in all printed material in English:

  1. the
  2. of
  3. and
  4. a
  5. to
  6. in
  7. is
  8. you
  9. that
  10. it
  11. he
  12. was
  13. for
  14. on
  15. are
  16. as
  17. with
  18. his
  19. they
  20. I
  21. at
  22. be
  23. this
  24. have
  25. from

Word frequencies follow a power law (like Pareto’s famous 80-20 rule) — Zipf’s Law, in this case, named after the linguist George Kingsley Zipf who first proposed it:

Zipf’s law states that given some corpus of natural language utterances, the frequency of any word is inversely proportional to its rank in the frequency table. Thus the most frequent word will occur approximately twice as often as the second most frequent word, which occurs twice as often as the fourth most frequent word, etc. For example, in the Brown Corpus “the” is the most frequently occurring word, and by itself accounts for nearly 7% of all word occurrences (69971 out of slightly over 1 million). True to Zipf’s Law, the second-place word “of” accounts for slightly over 3.5% of words (36411 occurrences), followed by “and” (28852). Only 135 vocabulary items are needed to account for half the Brown Corpus.