Communication in human language can be signed, spoken, or written. Let’s take a look at the spoken language.
Phonology
Spoken language is made up of sounds, which are organized into patterns that convey meaning.
Phonemes are the smallest units of sound that can distinguish meaning. Phonology is the study of how phonemes are organized and used in language.
The human anatomy is such that we can distinguish several hundred phonemes. No other animal has anywhere near this capability.
Exercise: But what about parrots?
Exercise: Research how the evolution of the human vocal tract and brain enabled the production and perception of such as massive number of phonemes. (A good place to start is Mithen’s The Language Puzzle, Chapter 5.) What were the selective pressures that led to this capability? How does this compare to other animals, such as birds, that can produce a wide variety of sounds?
Some spoken languages have only a few dozen phonemes. Some have over a hundred. They are impossible to count exactly, given that some counters may not distinguish short and long vowels, tones, morae, broadness, or other distinctions. Additionally, some dialects introduce new phonemes. Therefore most reports of phoneme inventory have a range rather than a precise number:
Language
Approx. Phoneme Count
Notes
Rotokas
~11–12
Hawaiian
~13–18
Depends on whether ā, ē, ī, ō, ū are counted as separate phonemes from a, e, i, o, u
Spanish
~25
20 consonants, 5 vowels
Italian
~30
23 consonants, 7 vowels
Japanese
~20–26
Depending on mora/long vowel count and whether palatalized consonants are counted separately
German
~36–48
English
~36–46
Depending on dialect and vowel analysis
Russian
~40–45
Large consonant count comes from palatalized/non-palatalized pairs, e.g., tʲ vs. t
Arabic (Standard)
~34
Irish
~49–72
High count due to broad/slender consonant distinctions
Here’s some vocabulary for talking about sounds, roughly ordered from the smallest units up to larger structures:
Phone: any distinct speech sound, regardless of whether it distinguishes meaning in a given language.
Phoneme: the smallest unit of sound that can distinguish meaning in a language.
Allophone: one of the different phones that can realize the same phoneme without changing meaning, e.g., the aspirated pʰ in pin and the unaspirated p in spin are both allophones of the English phoneme p.
Minimal pair: a pair of words that differ by only one phoneme, e.g., pat vs. bat. Minimal pairs are how linguists confirm that two phones are separate phonemes rather than allophones of the same phoneme.
Consonant: a speech sound that is articulated with complete or partial closure of the vocal tract.
Vowel: a speech sound that is produced without significant constriction of the vocal tract.
Diphthong: a single vowel sound formed by gliding from one vowel quality to another within the same syllable, e.g., the vowel in face, eɪ.
Syllable: a unit of organization for a sequence of speech sounds, typically consisting of a vowel sound with optional surrounding consonants.
Mora: a unit of syllable weight, or length, used to measure the timing of syllables in some languages.
Phonotactics: the rules governing which phoneme sequences are allowed in a language, e.g., English allows words to start with str- but not with rtk-.
Suprasegmentals: features such as stress, tone, duration, juncture, and intonation that apply to syllables or larger units rather than to individual phonemes.
One more....
Telephone: tele- (far) + -phone (sound) or “far sound,” a device for transmitting speech sounds over a distance.
Classifying Sounds
There are quite a few dimensions on which sounds can be classified. For consonants, we have place, manner, and voicing. The place of articulation refers to where in the vocal tract the sound is produced; the manner of articulation refers to how the sound is produced; voicing refers whether the vocal cords vibrate during the production of a sound. For example, the sound p is a voiceless bilabial plosive, meaning it is produced by closing both lips and releasing a burst of air without vibrating the vocal cords. Vowels are classified according to height, backness, and roundness. You’ll find that there’s more to sounds than just vowels and consonants!
The IPA
The International Phonetic Association (IPA) is the major organization for phoneticians. Their aim is “to promote the scientific study of phonetics and the various practical applications of that science.” They are best known for creating and maintaining the International Phonetic Alphabet (IPA), a standardized system for representing the sounds of spoken language.
The alphabet gets very minor revisions from time to time. Here is a screenshot of the IPA Chart, taken in 2026. Visit the official IPA Interactive Chart for an interactive version where you can click on each sound for more info and to hear each sound.
BTW, there’s another interactive IPA chart in case you’d like to compare. (This one allows you to click around more quickly.)
Here’s something fun:
Place of Articulation
Here’s a little diagram showing where sounds can be made (taken from Eric Dunn):
This allows us to classify consonants:
Bilabial: Made with both lips together, e.g., p, b, m.
Labiodental: Made by touching the lower lip to the upper teeth, e.g., f, v.
Dental: Made with the tongue against the upper teeth, e.g., t̪, d̪.
Interdental: Made with the tongue against or between the teeth, e.g., θ, ð.
Alveolar: Made with the tongue at the alveolar ridge just behind the upper teeth, e.g., t,d, s, n, l, and the trilled r of Spanish, r).
Postalveolar: Made just behind the alveolar ridge, e.g., ʃ, ʒ, tʃ.
Palatal: Made by raising the tongue body to the hard palate, e.g., j.
Retroflex: Made with the tongue curled back toward the palate, e.g., ʈ, ɖ.
Velar: Made by raising the back of the tongue to the soft palate/velum, e.g., k, g, ŋ.
Uvular: Made by raising the back of the tongue to the uvula, e.g., q, ʁ.
Pharyngeal: Made by constricting the pharynx, e.g., ħ, ʕ.
Glottal: Made with constriction at the glottis or vocal cords, e.g., h.
Manner of Articulation
While place tells us where a consonant is made, manner of articulation tells us how, specifically, how much the airflow through the vocal tract is obstructed:
Plosive (Stop): complete closure of the vocal tract, then a released burst of air, e.g., p, b, t, d, k, g.
Affricate: a stop immediately released into a fricative at the same place, e.g., tʃ (church), dʒ (judge).
Fricative: a narrow constriction forcing air through with audible friction, e.g., f, v, s, z, θ, ʃ.
Nasal: complete mouth, but air escapes through the nose instead, e.g., m, n, ŋ.
Trill: an articulator vibrates rapidly against the place of articulation, e.g., Spanish rrr.
Tap/Flap: a single, quick contact between articulator and place, e.g., American English tt in butter, ɾ, or the single r of Spanish pero.
Approximant: articulators come close together but not close enough to create turbulence, e.g., w, j, English ɹ. In a Lateral approximant, air flows around the side(s) of the tongue rather than over the center, e.g., l.
You may also hear the traditional cover term liquid, which lumps together the laterals and the rhotics (l- and r-type sounds)—it’s a useful phonological grouping, but not an “official” manner.
Exercise: (Classic) Hold your hand in front of your mouth and say pin, then spin. Can you feel the difference in airflow on the p?
Exercise: Give examples of pairs of words in Hindi or Tamil showing that aspiration variants are true phonemes rather than allophones.
Voiced Consonants
Voicing is the third dimension in addition to manner and place. Combining all three dimensions is what gives us fully specific labels like “voiceless bilabial plosive” for p, or “voiced alveolar fricative” for z.
Here are some voiceless and voiced pairs for comparison:
p vs b (pit vs. bit)
t vs d (tip vs. dip)
k vs g (cap vs. gap)
f vs v (fan vs. van)
s vs z (sip vs. zip)
θ vs ð (thin vs. then)
ʃ vs ʒ (assure vs. azure)
Exercise: For each of the following, name the voicing, place, and manner: m, ʃ, d, ŋ, f.
Exercise: English used to have the letter þ (thorn), which represented the voiceless dental fricative θ as in thin, and ð (eth), which represented the voiced dental fricative ð as in this. (Icelandic still has these letters.) When did these disappear from English orthography? And how do we know now, whether th is supposed to be voiced or not? Are there any rules?
Vowels
Vowels are classified along three independent dimensions, based on the position and shape of the tongue and lips:
Height
Close (High): tongue raised close to the roof of the mouth, e.g., i (beet), u (boot)
Mid: tongue in a middle position, e.g., e (French été), o (Spanish solo)
Open (Low): tongue low, mouth relatively open, e.g., a (Spanish casa), ɑ (father)
Backness
Front: tongue pushed forward, e.g., i (beet), ɛ (bet)
Central: tongue in the middle, e.g., ə (final sound in sofa)
Back: tongue pulled back, e.g., u (boot), ɔ (bought)
Roundedness
Rounded: lips pushed forward and rounded, e.g., u, o, ɔ
Unrounded: lips relaxed or spread, e.g., i, ɛ, æ, ɑ
Note that this classifies individual vowel qualities. A diphthong is a quick glide between two of these vowel qualities within a single syllable. Here are some common diphthongs in English:
aɪ as in my, high, fly, light, pie, price
eɪ as in say, face, day, play, break, rain
ɔɪ as in boy, coin, toy, choice
aʊ as in now, loud, town, cow, mouth
oʊ as in go, no, show, home, loan, though
ɪə as in here, beer, deer, fear
eə as in air, hair, bear, chair, square
ʊə as in poor, sure, cure, tour
Exercise: How many dipthongs are officially recognized in different dialects of English?
Suprasegmentals
Not every meaningful sound feature belongs to an individual segment. Suprasegmentalsare (prosodic) features that extend over units larger than a single phoneme, such as syllables, words, or phrases, rather than acting as isolated consonant or vowel segments. Common suprasegmental features include stress, tone, intonation, length, and juncture:
Stress: extra emphasis (loudness, pitch, and length) on one syllable of a word. In English, stress placement can distinguish a noun from a verb: RE-cord (noun) vs. re-CORD (verb); PRO-duce (noun) vs. pro-DUCE (verb).
Exercise: Think of others. Say them aloud, both ways.
Tone: pitch used to distinguish word meaning, not just to convey emphasis or emotion. In Mandarin, for example, ma can mean mother, hemp, horse, or scold depending entirely on the pitch contour used.
Intonation: pitch patterns that apply across a whole phrase or sentence rather than a single word to convey attitudes, emotions, or grammatical structures, e.g., the rising pitch that often signals a yes/no question in English.
Length (duration): the relative duration of a sound or syllable, which can differentiate words in certain languages. Common in Hawaiian, e.g, pakū (to burst open) vs. paku (a curtain or screen).
Juncture: the pausing or transition features that distinguish between word boundaries (such as ice cream vs. I scream).
Exercise: Say the English sentence You’re going to the store once as a statement and once as a question (by changing only your intonation, not the word order). What changes, physically, about how you say it? Make questions in which stress is placed on a single word, but a different word for each sentence question.
The UCLA inventory has 451 languages covered so far, and has identified over 500 consonants and over 200 vowels.
One fun fact: no vowel is common across all languages.
Exercise: Which vowels appear most frequently across different languages?
Exercise: List some vowels or consonants that seem to appear in only one or in a very small number of languages.
Exercise: Which languages have the fewest phonemes? Which have the most?
Exercise: Find some consonant pairs easy to distinguish by speakers of one language but difficult for speakers of another language.
You’ll notice IPA symbols written between either / / or [ ], depending on the language context:
Slashes, / / are for phonemic transcriptions. For example, the English p is written the same way whether aspirated or not, so we write p.
Square brackets, [ ] are for phonetic transcriptions that capture the actual, physical sound in as much detail as is useful, including allophones. So in English, we need to write aspirated pʰ and unaspirated p in square brackets.
So the same English word can be transcribed two ways depending on what you want to show: pin is phonemically pɪn, but phonetically pʰɪn.
Exercise: For the words top and stop, give both a phonemic and a phonetic transcription of the initial consonant, showing whether aspiration is present. Use // for phonemic and [] for phonetic transcriptions.
Evolution of Spoken Language
Let’s look at the origins and evolution of spoken language.
Origin Theories
Roughly speaking, we have origin theories that are (1) Gesture-First, (2) Vocal-First, and (3) Multi-Modal.
Gesture-First Theory
Maybe human language did not start off spoken. The Gesture-First Theory (also called the gestural origin hypothesis) argues that language first emerged as a system of manual and bodily gestures, with vocal speech added on later. Evidence and arguments offered in favor include:
In nonhuman apes, gesture is way more flexible than vocalization. Chimpanzees and bonobos use manual gestures intentionally, flexibly, and toward specific goals (e.g., reaching, pointing at food they want another individual to fetch). Their gestures show far more voluntary control than their vocal calls, which tend to be fixed, largely involuntary, and tied to emotional states (like an alarm call). If our closest living relatives already show voluntary, flexible communication mainly through gesture, this suggests that a shared ancestor with humans may have relied on gesture, too.
Mirror neurons. Neurons in the Macaque’s premotor cortex (area F5) fire both when a monkey performs a hand action and when it merely watches another individual perform the same action. Area F5 is considered a homolog of Broca’s area in humans, a region central to human language production and planning and observations of hand and arm movements. This overlap between an action-observation system and the brain region most associated with language suggests these systems may share a deep evolutionary history.
Sign languages are full languages. Natural sign languages (like ASL) have the full expressive and grammatical complexity of any spoken language, which proves that the gestural/visual channel is entirely capable of carrying complete linguistic structure, not just a simplified stand-in for speech.
Gesture comes early developmentally. Human infants typically point, wave, and use other communicative gestures for months before they produce their first words, and even congenitally blind people (who have never seen anyone gesture) spontaneously gesture while speaking to other blind people. This suggests a tight, perhaps ancient, link between gesture and the language faculty, independent of vision.
The theory posits that the shift from gesture to speech may have happened as hominins increasingly needed their hands free (for tool use, carrying food or infants, etc.) and that the advantages of vocal communication (works in the dark, around corners, over longer distances, and while the hands and eyes are busy with something else) offered selective advantages.
Vocal-First Theory
The Vocal-First Theory (or continuity-based theory) posits that human language evolved primarily from primate vocalizations, rather than gestures. Proponents argue that the vocal-auditory channel was the primary medium for early communication, with gestures playing a secondary role.
Evidence and arguments in favor of this theory include:
Primate vocalizations are flexible. While less flexible than gestures, many primate calls show some degree of voluntary control and can convey specific meanings.
Infant vocal development. Human infants produce a wide range of vocalizations (cooing, babbling) before they develop meaningful gestures, indicating a predisposition for vocal communication.
Neuroanatomical support. Brain regions involved in vocal production and perception are highly developed in humans, supporting the idea that vocal communication has deep evolutionary roots.
Sound Changes
Vowel shifts, consonant changes, and stress shifts occur as a natural drive for efficiency and for reasons of identity marking. The clearest examples for English speakers come from comparing the Romance branch of Indo-European (descended from Latin) with the Germanic branch (which includes English), since both diverged from the shared parent language, Proto-Indo-European (PIE).
Grimm’s Law
Voiceless stops → Voiceless fricatives
p → f: Latin pater → English father
t → θ: Latin tres → English three
k → h: Latin centum → English hundred
Dyēus Ph₂tḗr, Zeus, and Jupiter
Reconstructed PIE had a chief sky-god, *Dyḗus Ph₂tḗr, literally “Sky Father.” In Greek, it survived almost unchanged as Zeus Patḗr (Ζεῦ πάτερ), which contracted over time into simply Zeus. In Latin, Dyēu-pater shifted sound-by-sound into Iuppiter, or Jupiter. The same PIE root, *dyew- (sky, to shine), shows up all over the Indo-European family without the “father” part attached: Latin deus (god) and dies (day), Sanskrit deva (the Vedic sky-father), and even English Tuesday, named for the Germanic god Tiw (Old Norse Týr).
Voiced stops → Voiceless stops
b → p: Latin labium (lip) → English lip
d → t: Latin decem → English ten (also duo → two)
g → k: Latin ager (field) → English acre
PIE voiced aspirated stops → Germanic voiced stops but Latin voiceless fricatives
bʰ → b in Germanic, bʰ → f in Latin: PIE *bʰrā́tēr → Latin frāter, English brother
dʰ → d in Germanic, dʰ → f in Latin: PIE *dʰwer- → Latin foris (doorway), English door
gʰ → g in Germanic, gʰ → h in Latin: PIE *ghóstis → Latin hostis (stranger, enemy), English guest
Voiced aspirated stops
In PIE, voiced aspirated stops were distinct phonemes. Sanskrit and other Indo-European languages that preserve this feature, while English and others do not. Here are some Sanskrit/English pairs: bʰrā́tēr / brother, dʰātr / father, agha / egg.
Exercise: Languages like Hindi have to mark aspiration on consonants explicitly in writing, unlike English where aspiration is allophonic and not indicated in the orthography. Is the reason for this that aspiration is predictable in English but not in Hindi? Research whether this is true, and if so, what is the basic rule for aspiration predictability in English and whether there are exceptions.
Consonant cluster reduction
A 16th–17th century English development (16th–17th century). The spelling did not change, even though the pronunciation did.
Loss of k before n: knight, knee, know, knife
Loss of w before r: write, wrist, wrong, wrap
Loss of l before certain consonants: walk, talk, half, calm
Loss of final b after m: comb, climb, limb, thumb, dumb, numb, bomb
Lenition (intervocalic voicing)
Common in the development of Spanish, French, and other Western Romance languages from Latin; a voiceless stop between vowels becomes voiced.
p → b: Latin sapere → Spanish saber
t → d: Latin vita → Spanish vida; Latin pater → Spanish padre
k → g: Latin amica → Spanish amiga
Diphthongization
Common in Spanish: certain short, stressed Latin vowels split into two-vowel sequences.
Latin short e → ie: terra → tierra, tempus → tiempo
Latin short o → ue: porta → puerta, bonus → bueno
Vowel change at word endings
Common in Italian, where Latin’s word-final case endings collapsed into a small, regular set of vowels.
Latin -us → Italian -o: lupus → lupo
Latin -um → Italian -o: amicum → amico
End-of-word erosion
Common in French, which lost most of Latin’s word-final sounds over time.
Loss of the final unstressed -e in pronunciation, though it’s often kept in spelling: table is pronounced tablə in Old French but tabl in Modern French.
Loss of most final consonants in pronunciation: Latin digitus → French doigt, pronounced dwa, with the final -t silent.
Vowel fronting
An early, shared change in the ancestors of English and Frisian: Germanic a fronted to æ.
Proto-Germanic *fadēr → Old English fæder → Modern English father
Proto-Germanic *that → Old English þæt → Modern English that
Vowel diphthongization (Great Vowel Shift and after)
Some English long vowels raised and eventually broke into diphthongs; still ongoing in most varieties of English today.
aː → eː → eɪ: Middle English namenaːmə → Modern English nameneɪm
oː → oʊ: Middle English goɡoː → Modern English goɡoʊ
Umlaut (i-mutation)
A vowel is fronted or raised because of an i or j sound in the following syllable. Fun fact: in Proto-Germanic, many plural nouns were formed by adding a suffix containing i—singular *fōts (foot), plural *fōtiz (feet). That i in the suffix dragged the stem vowel ō forward in the mouth to ē, giving *fōtiz. Once the -iz ending itself eroded away and stopped being pronounced, the fronted vowel was all that was left to signal the plural, which is why foot/feet don’t just add -s like a regular English plural. The same process, on adjectives instead of nouns, produced some irregular comparatives.
foot / feet (Old English fōt / fēt, from Proto-Germanic *fōts / *fōtiz)
mouse / mice
goose / geese
man / men
old / elder (comparative, from Proto-Germanic *althiz)
full / fill (the verb fill comes from Proto-Germanic *fulljan, meaning to make full, with the -j- triggering umlaut)
Metathesis
Two adjacent sounds swap places.
Old English brid → Modern English bird
Old English wæps → Modern English wasp
Old English þridda → Modern English third (compare three)
Assimilation
A sound becomes more like a neighboring sound.
Latin in- + possibilis → English impossible (n → m before p)
Latin in- + legalis → English illegal (n → l before l)
Latin in- + regularis → English irregular (n → r before r)
Exercise: Look up Verner’s Law. Upon discovery, it looked like an exception to Grimm’s Law, but once figured out, it is now considered a rule of its own. What’s the story behind this?
Here’s short showing the evolution of English:
Mapping Sounds to Writing
English is an interesting case study. Today, it is often said to have about 44 phonemes: roughly 24 consonants and 20 vowels (including diphthongs), though as we saw above, the exact count varies by dialect and analysis.
Let’s start with the ancestor language, Old English (OE). It had seven short vowels: i, e, æ, a, o, u,y, and seven long vowels, paired exactly with the short ones: iː, eː, æː, aː, oː, uː, yː. There were four diphthongs that could each be long or short: ea, eːa, eo, eːo, ie, iːe, io, iːo. The consonants were p, b, t, d, k, g, f, v, θ, ð, s, z, h, x, m, n, ŋ, l, r, w, j, ɣ, tʃ.
The Futhorc
OE has been written with symbols from the Anglo-Frisian Futhorc (named after its first six letters, F-U-Þ-O-R-C), which was descended from the Elder Futhark, brought over from the Germanic mainland by the Angles, Saxons, and Jutes. The mapping from phoneme to rune was pretty good! Here is the 28-rune version from Wikipedia:
ᚠ
feoh
fv
ᚢ
ūr
u
ᚦ
þorn
θð
ᚩ
ōs
o
ᚱ
rād
r
ᚳ
ċēn
ktʃ
ᚷ
ġiefu
gɣj
ᚹ
wynn
w
ᚻ
hæġl
hxç
ᚾ
nēod
n
ᛁ
īs
iɪ
ᛄ
ġear
j
ᛇ
īw
ixç
ᛈ
peorþ
p
ᛉ
eolhx
ks
ᛋ
siġel
sz
ᛏ
Tīw
t
ᛒ
beorc
b
ᛖ
eoh
e
ᛗ
mann
m
ᛚ
lagu
l
ᛝ
Ing
ŋ
ᛟ
ēþel
ø
ᛞ
dæġ
d
ᚪ
āc
ɑ
ᚫ
æsċ
æ
ᛠ
ēar
æːa
ᚣ
ȳr
y
Some runes mapped to multiple phonemes because of the phonotactics of OE: for feoh, þorn, and siġel, the first sound is the main one and the second happens between two vowels or between a vowel and a voiced consonant. For hæġl we have h at the start of the word and x elsewhere. For ċēn it’s k before back vowels and tʃ before front vowels. The rules for ġyfu and the others are more complex—look them up.
By the way, some Futhorcs added new runes to make things a little less ambiguous.
The Latin Alphabet arrives in Britain 😮
Later, Christian missionaries brought the Latin alphabet (ABCDEFGHIKLMNOPQRSTVXYZ), which did not have enough letters for Old English, so folks had to add a few. They added:
/æ for æ
J/j for j and dʒ
U/u for u
W/w for w
Þ/þ for θ
Ð/ð for ð (though þ often did double duty)
Œ/œ for œ
ċ for tʃ
ġ for dʒ (in some places, otherwise pronounced j)
As with runes, vowel length usually wasn’t marked in the manuscripts themselves; modern editions mark it with a macron. So god was pronounced god and gōd (“good”) was pronounced goːd.
Middle English
Two runes, þ (thorn) and ƿ (wynn) survived for a while. But the language kept evolving! Then, after the Norman Conquest, English got a massive French influence. Spelling started to get standardized at this time.
By Middle English (ME), the y fused into i, æ dropped into a, and ə appeared in unstressed syllables (no special letter for it), giving 6 short vowel sounds. On the long side, yː merged with iː, and æː merged with eː, but two new long vowels appeared: ɛː and ɔː, so there were seven long vowel sounds in total. (Wikipedia has the full chart). The spelling system did not introduce new letters for these new long vowels; instead, they just added an a, e.g., meat and boat (mɛːt and bɔːt).
In Middle English spelling, scribes began doubling a vowel letter to show length, a convention that in some words survives into Modern English spelling: good, feet, moon, and boot all still show a doubled vowel letter from a historically long vowel, even though the vowel qualities have since shifted.
Exercise: Read the entire Wikipedia article on Middle English Phonology. Then read the one on the Great Vowel Shift. It should help you realize why English spelling seems so weird. It was once fairly phonetic, right? At least as good as it could be with the foreign Latin alphabet.
Wait, æ disappeared in ME?
It did, but it came back. Linguistics is fun, right?
The Great Vowel Shift
Roughly between 1400 and 1700 (with most of the action happening in the 15th and 16th centuries), English underwent the Great Vowel Shift: nearly every long vowel in the language shifted upward or, for the two vowels already at the top, broke into a diphthong. Spelling had already mostly settled down by the time this happened (largely thanks to the printing press, introduced to England in 1476), so English ended up with spellings that reflect Middle English pronunciation while the words themselves are now said very differently.
Here’s a table showing the changes in some representative words, taken from the Wikipedia. The sound files are from user Erutuon.
Word
Pronunciation
Sound Evolution
1400
1500
1600
by 1900
bite
biːt
beit
bɛit
baɪt
out
uːt
out
ɔut
aʊt
meet
meːt
miːt
boot
boːt
buːt
meat
mɛːt
meːt
miːt
boat
bɔːt
boːt
boʊt
mate
maːt
meːt
mɛit
meɪt
Here are how the seven long vowels of Middle English (ME) changed during the Great Vowel Shift (to what we have in Modern English (ModE)):
ME iː became ModE aɪ (time, mice)
ME eː rose to become ModE iː (meet, feet)
ME ɛː rose to eː then either:
to eɪ (great, break, steak)
up to merge with the iː above (meat, sea)
ME uː became ModE aʊ (mouse, house)
ME oː rose to become ModE uː (moon, food)
ME ɔː rose to become ModE oʊ (boat, road)
ME aː rose to become ModE eɪ (name, make)
Notice how meet and meat end up pronounced identically today, even though they started from two different Middle English vowels and are still spelled differently.
Exercise: Keep practicing your IPA skills. Then read aloud the IPA transcription for Middle English.
Exercise: Make a list of words that did not shift. Here is one to get you started: been.
Exercise: What kind of shift happened in took and foot? (You’ll need to do some research)
Like boot and moon, these words had the Middle English long vowel oː, which the Great Vowel Shift regularly raised to uː. But afterward, a separate, sporadic (lexically irregular) shortening lowered that new uː to ʊ in a subset of words — foot, took, book, look, good, wood — while others, like boot and moon, escaped it and kept uː. So it isn’t part of the regular GVS pattern itself, but a later shortening layered on top of it.
If you like hearing sounds before, during, and after the Great Vowel Shift, you might like this page.
The Latin Alphabet has become Popular
Like English, a lot of languages (over 3000 worldwide!) have adopted the Latin alphabet, even when the sounds don’t match the letters perfectly. Watch the following video to see if the Latin alphabet is a good fit for Tlingit.
Exercise: How badly does the Latin alphabet fit English? How well does it fit Spanish? Why the difference? What other alphabets are used for English?
Exercise: Research the introduction of Latin Letters into Hawaiian. How did that work?
Exercise: Research the click sounds of Zulu. How are they represented in the Latin alphabet? How are they represented in the IPA?
Recall Practice
Here are some questions useful for your spaced repetition learning. Many of the answers are not found on this page. Some will have popped up in lecture. Others will require you to do your own research.
What is a phoneme?
The smallest unit of sound that can distinguish meaning.
What is phonology?
The study of how phonemes are organized and used in language.
What is a phone?
Any distinct speech sound, regardless of whether it distinguishes meaning in a given language.
What is an allophone?
One of the different phones that can realize the same phoneme without changing meaning.
What is a minimal pair?
A pair of words that differ by only one phoneme, e.g., pat vs. bat.
What is phonotactics?
The rules governing which phoneme sequences are allowed in a language.
What are suprasegmentals?
Prosodic features, such as stress, tone, intonation, length, and juncture, that extend over units larger than a single phoneme.
What three dimensions classify consonants?
Place, manner, and voicing.
What does place of articulation refer to?
Where in the vocal tract a sound is produced.
What does manner of articulation refer to?
How a sound is produced, specifically how much the airflow is obstructed.
What does voicing refer to?
Whether the vocal cords vibrate during the production of a sound.
What is a plosive (stop)?
A consonant made with complete closure of the vocal tract, then a released burst of air.
What is a fricative?
A consonant made by forcing air through a narrow constriction, producing audible friction.
What is an affricate?
A stop immediately released into a fricative at the same place of articulation.
What is the traditional cover term for laterals and rhotics?
Liquid.
What is aspiration?
An extra puff of breath released after a stop, as in the p in pin.
Is aspiration phonemic or allophonic in English?
Allophonic — it never changes meaning in English.
What three dimensions classify vowels?
Height, backness, and roundedness.
What is a diphthong?
A single vowel sound formed by gliding from one vowel quality to another within the same syllable.
What’s the difference between tone and intonation?
Tone uses pitch to distinguish the meaning of individual words; intonation uses pitch patterns over a whole phrase or sentence to convey attitude or grammatical structure.
What suprasegmental feature is illustrated by the difference between ice cream and I scream?
Juncture.
In IPA notation, what does a transcription between slashes, / /, indicate?
A phonemic transcription — sounds at the level of the phoneme.
In IPA notation, what does a transcription between square brackets, [ ], indicate?
A phonetic transcription — the actual, physical sound, including allophones.
About how many phonemes does English have?
About 44 (roughly 24 consonants and 20 vowels, including diphthongs).
What are the three major origin theories of spoken language?
Gesture-First, Vocal-First, and Multi-Modal.
What does the Gesture-First Theory claim?
That language first emerged as a system of manual and bodily gestures, with vocal speech added on later.
What mirror neuron evidence is offered in support of the Gesture-First Theory?
Neurons in the macaque premotor cortex (area F5, a homolog of Broca’s area) fire both when performing and when observing a hand action, linking gesture to the brain region most associated with language.
What does the Vocal-First Theory claim?
That human language evolved primarily from primate vocalizations, with gestures playing a secondary role.
What is Grimm’s Law?
A sound change describing how certain Proto-Indo-European consonants shifted in the Germanic branch, e.g., voiceless stops became voiceless fricatives.
Give a word pair illustrating Grimm’s Law.
Latin pater vs. English father (p → f).
What is Verner’s Law?
A refinement of Grimm’s Law: PIE voiceless consonants became voiced, rather than staying as voiceless fricatives, when they weren’t immediately preceded by the PIE stress accent.
What is lenition (intervocalic voicing)?
A voiceless stop becoming voiced between vowels, e.g., Latin sapere → Spanish saber.
What is metathesis?
Two adjacent sounds swapping places, e.g., Old English brid → Modern English bird.
What is assimilation?
A sound becoming more like a neighboring sound, e.g., Latin in- + possibilis → English impossible.
What is umlaut (i-mutation)?
A vowel being fronted or raised because of an i or j sound in the following syllable.
Why do foot and feet have different vowels?
Umlaut: an i in the Proto-Germanic plural suffix fronted the stem vowel before the suffix itself eroded away.
What is the Great Vowel Shift?
A change, mostly in the 15th and 16th centuries, in which nearly every long vowel in English rose or, for the two already at the top, broke into a diphthong.
Why does English spelling seem so mismatched with pronunciation today?
Spelling was largely standardized, thanks to the printing press, before the Great Vowel Shift finished, so spellings still reflect Middle English pronunciation.
Give an example of two English words that sound identical today but were pronounced differently in Middle English and are still spelled differently.