Sounds in Language

Communication in human language can be signed, spoken, or written. Let’s take a look at the spoken language.

Phonology

Spoken language is made up of sounds, which are organized into patterns that convey meaning.

Phonemes are the smallest units of sound that can distinguish meaning. Phonology is the study of how phonemes are organized and used in language.

The human anatomy is such that we can distinguish several hundred phonemes. No other animal has anywhere near this capability.

Exercise: But what about parrots?
Exercise: Research how the evolution of the human vocal tract and brain enabled the production and perception of such as massive number of phonemes. (A good place to start is Mithen’s The Language Puzzle, Chapter 5.) What were the selective pressures that led to this capability? How does this compare to other animals, such as birds, that can produce a wide variety of sounds?

Some spoken languages have only a few dozen phonemes. Some have over a hundred. They are impossible to count exactly, given that some counters may not distinguish short and long vowels, tones, morae, broadness, or other distinctions. Additionally, some dialects introduce new phonemes. Therefore most reports of phoneme inventory have a range rather than a precise number:

LanguageApprox. Phoneme CountNotes
Rotokas~11–12 
Hawaiian~13–18Depends on whether ā, ē, ī, ō, ū are counted as separate phonemes from a, e, i, o, u
Spanish~2520 consonants, 5 vowels
Italian~3023 consonants, 7 vowels
Japanese~20–26Depending on mora/long vowel count and whether palatalized consonants are counted separately
German~36–48 
English~36–46Depending on dialect and vowel analysis
Russian~40–45Large consonant count comes from palatalized/non-palatalized pairs, e.g., tʲ vs. t
Arabic (Standard)~34 
Irish~49–72High count due to broad/slender consonant distinctions
Archi~100+Lots of consonants

See Wikipedia for a huge list.

Vocabulary Time

Here’s some vocabulary for talking about sounds, roughly ordered from the smallest units up to larger structures:

One more....

Classifying Sounds

There are quite a few dimensions on which sounds can be classified. For consonants, we have place, manner, and voicing. The place of articulation refers to where in the vocal tract the sound is produced; the manner of articulation refers to how the sound is produced; voicing refers whether the vocal cords vibrate during the production of a sound. For example, the sound p is a voiceless bilabial plosive, meaning it is produced by closing both lips and releasing a burst of air without vibrating the vocal cords. Vowels are classified according to height, backness, and roundness. You’ll find that there’s more to sounds than just vowels and consonants!

The IPA

The International Phonetic Association (IPA) is the major organization for phoneticians. Their aim is “to promote the scientific study of phonetics and the various practical applications of that science.” They are best known for creating and maintaining the International Phonetic Alphabet (IPA), a standardized system for representing the sounds of spoken language.

The alphabet gets very minor revisions from time to time. Here is a screenshot of the IPA Chart, taken in 2026. Visit the official IPA Interactive Chart for an interactive version where you can click on each sound for more info and to hear each sound.

IPA Chart

International Phonetic Alphabet in IPA Kiel, available under a Creative Commons Attribution-Sharealike 4.0 International Deed. Copyright © 2026 International Phonetic Association. Date of access: 15 September 2026.

Before exploring the entire site, here’s a short video showing the sounds of English in IPA:

After you watch the video, it’s time to explore!

CLASSWORK
Let‘s tour the IPA chart online. Turn your device’s volume up.

BTW, there’s another interactive IPA chart in case you’d like to compare. (This one allows you to click around more quickly.)

Here’s something fun:

Place of Articulation

Here’s a little diagram showing where sounds can be made (taken from Eric Dunn):

articulation-places.png

This allows us to classify consonants:

Manner of Articulation

While place tells us where a consonant is made, manner of articulation tells us how, specifically, how much the airflow through the vocal tract is obstructed:

You may also hear the traditional cover term liquid, which lumps together the laterals and the rhotics (l- and r-type sounds)—it’s a useful phonological grouping, but not an “official” manner.

Exercise: (Classic) Hold your hand in front of your mouth and say pin, then spin. Can you feel the difference in airflow on the p?
Exercise: Give examples of pairs of words in Hindi or Tamil showing that aspiration variants are true phonemes rather than allophones.

Voiced Consonants

Voicing is the third dimension in addition to manner and place. Combining all three dimensions is what gives us fully specific labels like “voiceless bilabial plosive” for p, or “voiced alveolar fricative” for z.

Here are some voiceless and voiced pairs for comparison:

Exercise: For each of the following, name the voicing, place, and manner: m, ʃ, d, ŋ, f.
Exercise: English used to have the letter þ (thorn), which represented the voiceless dental fricative θ as in thin, and ð (eth), which represented the voiced dental fricative ð as in this. (Icelandic still has these letters.) When did these disappear from English orthography? And how do we know now, whether th is supposed to be voiced or not? Are there any rules?

Vowels

vowels.png

Vowels are classified along three independent dimensions, based on the position and shape of the tongue and lips:

Note that this classifies individual vowel qualities. A diphthong is a quick glide between two of these vowel qualities within a single syllable. Here are some common diphthongs in English:

Exercise: How many dipthongs are officially recognized in different dialects of English?

Suprasegmentals

Not every meaningful sound feature belongs to an individual segment. Suprasegmentalsare (prosodic) features that extend over units larger than a single phoneme, such as syllables, words, or phrases, rather than acting as isolated consonant or vowel segments. Common suprasegmental features include stress, tone, intonation, length, and juncture:

Exercise: Say the English sentence You’re going to the store once as a statement and once as a question (by changing only your intonation, not the word order). What changes, physically, about how you say it? Make questions in which stress is placed on a single word, but a different word for each sentence question.

Sounds Across Different Languages

There are several online databases that catalog the phonemes of different languages, such as the PHOIBLE database, the UCLA Phonological Segment Inventory Database and the Sound Comparisons project.

The UCLA inventory has 451 languages covered so far, and has identified over 500 consonants and over 200 vowels.

One fun fact: no vowel is common across all languages.

Exercise: Which vowels appear most frequently across different languages?
Exercise: List some vowels or consonants that seem to appear in only one or in a very small number of languages.
Exercise: Which languages have the fewest phonemes? Which have the most?
Exercise: Find some consonant pairs easy to distinguish by speakers of one language but difficult for speakers of another language.

You’ll notice IPA symbols written between either / / or [ ], depending on the language context:

So the same English word can be transcribed two ways depending on what you want to show: pin is phonemically pɪn, but phonetically pʰɪn.

Exercise: For the words top and stop, give both a phonemic and a phonetic transcription of the initial consonant, showing whether aspiration is present. Use // for phonemic and [] for phonetic transcriptions.

Evolution of Spoken Language

Let’s look at the origins and evolution of spoken language.

Origin Theories

Roughly speaking, we have origin theories that are (1) Gesture-First, (2) Vocal-First, and (3) Multi-Modal.

Gesture-First Theory

Maybe human language did not start off spoken. The Gesture-First Theory (also called the gestural origin hypothesis) argues that language first emerged as a system of manual and bodily gestures, with vocal speech added on later. Evidence and arguments offered in favor include:

The theory posits that the shift from gesture to speech may have happened as hominins increasingly needed their hands free (for tool use, carrying food or infants, etc.) and that the advantages of vocal communication (works in the dark, around corners, over longer distances, and while the hands and eyes are busy with something else) offered selective advantages.

Vocal-First Theory

The Vocal-First Theory (or continuity-based theory) posits that human language evolved primarily from primate vocalizations, rather than gestures. Proponents argue that the vocal-auditory channel was the primary medium for early communication, with gestures playing a secondary role.

Evidence and arguments in favor of this theory include:

Sound Changes

Vowel shifts, consonant changes, and stress shifts occur as a natural drive for efficiency and for reasons of identity marking. The clearest examples for English speakers come from comparing the Romance branch of Indo-European (descended from Latin) with the Germanic branch (which includes English), since both diverged from the shared parent language, Proto-Indo-European (PIE).

Grimm’s Law
Voiceless stops → Voiceless fricatives
  • p → f: Latin pater → English father
  • t → θ: Latin tres → English three
  • k → h: Latin centum → English hundred
Dyēus Ph₂tḗr, Zeus, and Jupiter

Reconstructed PIE had a chief sky-god, *Dyḗus Ph₂tḗr, literally “Sky Father.” In Greek, it survived almost unchanged as Zeus Patḗr (Ζεῦ πάτερ), which contracted over time into simply Zeus. In Latin, Dyēu-pater shifted sound-by-sound into Iuppiter, or Jupiter. The same PIE root, *dyew- (sky, to shine), shows up all over the Indo-European family without the “father” part attached: Latin deus (god) and dies (day), Sanskrit deva (the Vedic sky-father), and even English Tuesday, named for the Germanic god Tiw (Old Norse Týr).

Voiced stops → Voiceless stops

  • b → p: Latin labium (lip) → English lip
  • d → t: Latin decem → English ten (also duo → two)
  • g → k: Latin ager (field) → English acre

PIE voiced aspirated stops → Germanic voiced stops but Latin voiceless fricatives

  • bʰ → b in Germanic, bʰ → f in Latin: PIE *bʰrā́tēr → Latin frāter, English brother
  • dʰ → d in Germanic, dʰ → f in Latin: PIE *dʰwer- → Latin foris (doorway), English door
  • gʰ → g in Germanic, gʰ → h in Latin: PIE *ghóstis → Latin hostis (stranger, enemy), English guest
Voiced aspirated stops

In PIE, voiced aspirated stops were distinct phonemes. Sanskrit and other Indo-European languages that preserve this feature, while English and others do not. Here are some Sanskrit/English pairs: bʰrā́tēr / brother, dʰātr / father, agha / egg.

Exercise: Languages like Hindi have to mark aspiration on consonants explicitly in writing, unlike English where aspiration is allophonic and not indicated in the orthography. Is the reason for this that aspiration is predictable in English but not in Hindi? Research whether this is true, and if so, what is the basic rule for aspiration predictability in English and whether there are exceptions.
Consonant cluster reduction
A 16th–17th century English development (16th–17th century). The spelling did not change, even though the pronunciation did.
  • Loss of k before n: knight, knee, know, knife
  • Loss of w before r: write, wrist, wrong, wrap
  • Loss of l before certain consonants: walk, talk, half, calm
  • Loss of final b after m: comb, climb, limb, thumb, dumb, numb, bomb
Lenition (intervocalic voicing)
Common in the development of Spanish, French, and other Western Romance languages from Latin; a voiceless stop between vowels becomes voiced.
  • p → b: Latin sapere → Spanish saber
  • t → d: Latin vita → Spanish vida; Latin pater → Spanish padre
  • k → g: Latin amica → Spanish amiga
Diphthongization
Common in Spanish: certain short, stressed Latin vowels split into two-vowel sequences.
  • Latin short e → ie: terra → tierra, tempus → tiempo
  • Latin short o → ue: porta → puerta, bonus → bueno
Vowel change at word endings
Common in Italian, where Latin’s word-final case endings collapsed into a small, regular set of vowels.
  • Latin -us → Italian -o: lupus → lupo
  • Latin -um → Italian -o: amicum → amico
End-of-word erosion
Common in French, which lost most of Latin’s word-final sounds over time.
  • Loss of the final unstressed -e in pronunciation, though it’s often kept in spelling: table is pronounced tablə in Old French but tabl in Modern French.
  • Loss of most final consonants in pronunciation: Latin digitus → French doigt, pronounced dwa, with the final -t silent.
Vowel fronting
An early, shared change in the ancestors of English and Frisian: Germanic a fronted to æ.
  • Proto-Germanic *fadēr → Old English fæder → Modern English father
  • Proto-Germanic *that → Old English þæt → Modern English that
Vowel diphthongization (Great Vowel Shift and after)
Some English long vowels raised and eventually broke into diphthongs; still ongoing in most varieties of English today.
  • aː → eː → eɪ: Middle English name naːmə → Modern English name neɪm
  • oː → oʊ: Middle English go ɡoː → Modern English go ɡoʊ
Umlaut (i-mutation)
A vowel is fronted or raised because of an i or j sound in the following syllable. Fun fact: in Proto-Germanic, many plural nouns were formed by adding a suffix containing i—singular *fōts (foot), plural *fōtiz (feet). That i in the suffix dragged the stem vowel ō forward in the mouth to ē, giving *fōtiz. Once the -iz ending itself eroded away and stopped being pronounced, the fronted vowel was all that was left to signal the plural, which is why foot/feet don’t just add -s like a regular English plural. The same process, on adjectives instead of nouns, produced some irregular comparatives.
  • foot / feet (Old English fōt / fēt, from Proto-Germanic *fōts / *fōtiz)
  • mouse / mice
  • goose / geese
  • man / men
  • old / elder (comparative, from Proto-Germanic *althiz)
  • full / fill (the verb fill comes from Proto-Germanic *fulljan, meaning to make full, with the -j- triggering umlaut)
Metathesis
Two adjacent sounds swap places.
  • Old English brid → Modern English bird
  • Old English wæps → Modern English wasp
  • Old English þridda → Modern English third (compare three)
Assimilation
A sound becomes more like a neighboring sound.
  • Latin in- + possibilis → English impossible (n → m before p)
  • Latin in- + legalis → English illegal (n → l before l)
  • Latin in- + regularis → English irregular (n → r before r)
Exercise: Look up Verner’s Law. Upon discovery, it looked like an exception to Grimm’s Law, but once figured out, it is now considered a rule of its own. What’s the story behind this?

Here’s short showing the evolution of English:

Mapping Sounds to Writing

English is an interesting case study. Today, it is often said to have about 44 phonemes: roughly 24 consonants and 20 vowels (including diphthongs), though as we saw above, the exact count varies by dialect and analysis.

Let’s start with the ancestor language, Old English (OE). It had seven short vowels: i, e, æ, a, o, u,y, and seven long vowels, paired exactly with the short ones: iː, eː, æː, aː, oː, uː, yː. There were four diphthongs that could each be long or short: ea, eːa, eo, eːo, ie, iːe, io, iːo. The consonants were p, b, t, d, k, g, f, v, θ, ð, s, z, h, x, m, n, ŋ, l, r, w, j, ɣ, tʃ.

The Futhorc

OE has been written with symbols from the Anglo-Frisian Futhorc (named after its first six letters, F-U-Þ-O-R-C), which was descended from the Elder Futhark, brought over from the Germanic mainland by the Angles, Saxons, and Jutes. The mapping from phoneme to rune was pretty good! Here is the 28-rune version from Wikipedia:

ᚠ
feoh
f v
ᚢ
ūr
u
ᚦ
þorn
θ ð
ᚩ
ōs
o
ᚱ
rād
r
ᚳ
ċēn
k tʃ
ᚷ
ġiefu
g ɣ j
ᚹ
wynn
w
ᚻ
hæġl
h x ç
ᚾ
nēod
n
ᛁ
īs
i ɪ
ᛄ
ġear
j
ᛇ
īw
i x ç
ᛈ
peorþ
p
ᛉ
eolhx
ks
ᛋ
siġel
s z
ᛏ
Tīw
t
ᛒ
beorc
b
ᛖ
eoh
e
ᛗ
mann
m
ᛚ
lagu
l
ᛝ
Ing
ŋ
ᛟ
ēþel
ø
ᛞ
dæġ
d
ᚪ
āc
ɑ
ᚫ
æsċ
æ
ᛠ
ēar
æːa
ᚣ
ȳr
y

Some runes mapped to multiple phonemes because of the phonotactics of OE: for feoh, þorn, and siġel, the first sound is the main one and the second happens between two vowels or between a vowel and a voiced consonant. For hæġl we have h at the start of the word and x elsewhere. For ċēn it’s k before back vowels and tʃ before front vowels. The rules for ġyfu and the others are more complex—look them up.

By the way, some Futhorcs added new runes to make things a little less ambiguous.

The Latin Alphabet arrives in Britain 😮

Later, Christian missionaries brought the Latin alphabet (ABCDEFGHIKLMNOPQRSTVXYZ), which did not have enough letters for Old English, so folks had to add a few. They added:

As with runes, vowel length usually wasn’t marked in the manuscripts themselves; modern editions mark it with a macron. So god was pronounced god and gōd (“good”) was pronounced goːd.

Middle English

Two runes, þ (thorn) and ƿ (wynn) survived for a while. But the language kept evolving! Then, after the Norman Conquest, English got a massive French influence. Spelling started to get standardized at this time.

By Middle English (ME), the y fused into i, æ dropped into a, and ə appeared in unstressed syllables (no special letter for it), giving 6 short vowel sounds. On the long side, yː merged with iː, and æː merged with eː, but two new long vowels appeared: ɛː and ɔː, so there were seven long vowel sounds in total. (Wikipedia has the full chart). The spelling system did not introduce new letters for these new long vowels; instead, they just added an a, e.g., meat and boat (mɛːt and bɔːt).

In Middle English spelling, scribes began doubling a vowel letter to show length, a convention that in some words survives into Modern English spelling: good, feet, moon, and boot all still show a doubled vowel letter from a historically long vowel, even though the vowel qualities have since shifted.

Exercise: Read the entire Wikipedia article on Middle English Phonology. Then read the one on the Great Vowel Shift. It should help you realize why English spelling seems so weird. It was once fairly phonetic, right? At least as good as it could be with the foreign Latin alphabet.
Wait, æ disappeared in ME?

It did, but it came back. Linguistics is fun, right?

The Great Vowel Shift

Roughly between 1400 and 1700 (with most of the action happening in the 15th and 16th centuries), English underwent the Great Vowel Shift: nearly every long vowel in the language shifted upward or, for the two vowels already at the top, broke into a diphthong. Spelling had already mostly settled down by the time this happened (largely thanks to the printing press, introduced to England in 1476), so English ended up with spellings that reflect Middle English pronunciation while the words themselves are now said very differently.

gvs.jpg

Here’s a table showing the changes in some representative words, taken from the Wikipedia. The sound files are from user Erutuon.

Word Pronunciation Sound Evolution
1400 1500 1600 by 1900
bite biːt beit bɛit baɪt
out uːt out ɔut aʊt
meet meːt miːt
boot boːt buːt
meat mɛːt meːt miːt
boat bɔːt boːt boʊt
mate maːt meːt mɛit meɪt

Here are how the seven long vowels of Middle English (ME) changed during the Great Vowel Shift (to what we have in Modern English (ModE)):

Notice how meet and meat end up pronounced identically today, even though they started from two different Middle English vowels and are still spelled differently.

Exercise: Keep practicing your IPA skills. Then read aloud the IPA transcription for Middle English.
Exercise: Make a list of words that did not shift. Here is one to get you started: been.
Exercise: What kind of shift happened in took and foot? (You’ll need to do some research)
Like boot and moon, these words had the Middle English long vowel oː, which the Great Vowel Shift regularly raised to uː. But afterward, a separate, sporadic (lexically irregular) shortening lowered that new uː to ʊ in a subset of words — foot, took, book, look, good, wood — while others, like boot and moon, escaped it and kept uː. So it isn’t part of the regular GVS pattern itself, but a later shortening layered on top of it.

If you like hearing sounds before, during, and after the Great Vowel Shift, you might like this page.

The Latin Alphabet has become Popular

Like English, a lot of languages (over 3000 worldwide!) have adopted the Latin alphabet, even when the sounds don’t match the letters perfectly. Watch the following video to see if the Latin alphabet is a good fit for Tlingit.

Exercise: How badly does the Latin alphabet fit English? How well does it fit Spanish? Why the difference? What other alphabets are used for English?
Exercise: Research the introduction of Latin Letters into Hawaiian. How did that work?
Exercise: Research the click sounds of Zulu. How are they represented in the Latin alphabet? How are they represented in the IPA?

Recall Practice

Here are some questions useful for your spaced repetition learning. Many of the answers are not found on this page. Some will have popped up in lecture. Others will require you to do your own research.

  1. What is a phoneme?
    The smallest unit of sound that can distinguish meaning.
  2. What is phonology?
    The study of how phonemes are organized and used in language.
  3. What is a phone?
    Any distinct speech sound, regardless of whether it distinguishes meaning in a given language.
  4. What is an allophone?
    One of the different phones that can realize the same phoneme without changing meaning.
  5. What is a minimal pair?
    A pair of words that differ by only one phoneme, e.g., pat vs. bat.
  6. What is phonotactics?
    The rules governing which phoneme sequences are allowed in a language.
  7. What are suprasegmentals?
    Prosodic features, such as stress, tone, intonation, length, and juncture, that extend over units larger than a single phoneme.
  8. What three dimensions classify consonants?
    Place, manner, and voicing.
  9. What does place of articulation refer to?
    Where in the vocal tract a sound is produced.
  10. What does manner of articulation refer to?
    How a sound is produced, specifically how much the airflow is obstructed.
  11. What does voicing refer to?
    Whether the vocal cords vibrate during the production of a sound.
  12. What is a plosive (stop)?
    A consonant made with complete closure of the vocal tract, then a released burst of air.
  13. What is a fricative?
    A consonant made by forcing air through a narrow constriction, producing audible friction.
  14. What is an affricate?
    A stop immediately released into a fricative at the same place of articulation.
  15. What is the traditional cover term for laterals and rhotics?
    Liquid.
  16. What is aspiration?
    An extra puff of breath released after a stop, as in the p in pin.
  17. Is aspiration phonemic or allophonic in English?
    Allophonic — it never changes meaning in English.
  18. What three dimensions classify vowels?
    Height, backness, and roundedness.
  19. What is a diphthong?
    A single vowel sound formed by gliding from one vowel quality to another within the same syllable.
  20. What’s the difference between tone and intonation?
    Tone uses pitch to distinguish the meaning of individual words; intonation uses pitch patterns over a whole phrase or sentence to convey attitude or grammatical structure.
  21. What suprasegmental feature is illustrated by the difference between ice cream and I scream?
    Juncture.
  22. In IPA notation, what does a transcription between slashes, / /, indicate?
    A phonemic transcription — sounds at the level of the phoneme.
  23. In IPA notation, what does a transcription between square brackets, [ ], indicate?
    A phonetic transcription — the actual, physical sound, including allophones.
  24. About how many phonemes does English have?
    About 44 (roughly 24 consonants and 20 vowels, including diphthongs).
  25. What are the three major origin theories of spoken language?
    Gesture-First, Vocal-First, and Multi-Modal.
  26. What does the Gesture-First Theory claim?
    That language first emerged as a system of manual and bodily gestures, with vocal speech added on later.
  27. What mirror neuron evidence is offered in support of the Gesture-First Theory?
    Neurons in the macaque premotor cortex (area F5, a homolog of Broca’s area) fire both when performing and when observing a hand action, linking gesture to the brain region most associated with language.
  28. What does the Vocal-First Theory claim?
    That human language evolved primarily from primate vocalizations, with gestures playing a secondary role.
  29. What is Grimm’s Law?
    A sound change describing how certain Proto-Indo-European consonants shifted in the Germanic branch, e.g., voiceless stops became voiceless fricatives.
  30. Give a word pair illustrating Grimm’s Law.
    Latin pater vs. English father (p → f).
  31. What is Verner’s Law?
    A refinement of Grimm’s Law: PIE voiceless consonants became voiced, rather than staying as voiceless fricatives, when they weren’t immediately preceded by the PIE stress accent.
  32. What is lenition (intervocalic voicing)?
    A voiceless stop becoming voiced between vowels, e.g., Latin sapere → Spanish saber.
  33. What is metathesis?
    Two adjacent sounds swapping places, e.g., Old English brid → Modern English bird.
  34. What is assimilation?
    A sound becoming more like a neighboring sound, e.g., Latin in- + possibilis → English impossible.
  35. What is umlaut (i-mutation)?
    A vowel being fronted or raised because of an i or j sound in the following syllable.
  36. Why do foot and feet have different vowels?
    Umlaut: an i in the Proto-Germanic plural suffix fronted the stem vowel before the suffix itself eroded away.
  37. What is the Great Vowel Shift?
    A change, mostly in the 15th and 16th centuries, in which nearly every long vowel in English rose or, for the two already at the top, broke into a diphthong.
  38. Why does English spelling seem so mismatched with pronunciation today?
    Spelling was largely standardized, thanks to the printing press, before the Great Vowel Shift finished, so spellings still reflect Middle English pronunciation.
  39. Give an example of two English words that sound identical today but were pronounced differently in Middle English and are still spelled differently.
    Meet and meat.

Summary

We’ve covered:

  • Phonemes and Phonology
  • Vocabulary
  • Sound Classification
  • The International Phonetic Alphabet (IPA)
  • Suprasegmentals
  • Sounds across different languages
  • Historical Sound Changes
  • Mapping sounds to spelling