02 Phonetics
Understand how speech sounds are produced, shaped into acoustic signals, perceived by listeners, and represented with the International Phonetic Alphabet.
Three perspectives on speech sounds
studies speech sounds from three connected perspectives: articulatory examines how speakers produce sounds, acoustic examines the sound signal, and auditory examines hearing and perception. Together, these perspectives describe speech as production, sound, and perception rather than as spelling alone.
How speech is produced
Most speech sounds are made with air pushed out from the lungs. The airflow passes through the larynx, where the vocal folds may vibrate, and then through the vocal tract: the throat, mouth, and sometimes the nasal cavity. The tongue, lips, jaw, and soft palate shape the airflow and the spaces through which it passes.
When the vocal folds vibrate, a sound is ; when they do not, it is . One way to feel the contrast is to touch the throat while alternating the initial sounds of zoo and Sue: [z] is and [s] is . Voicing is caused by vocal-fold vibration, not simply by a sound being loud.
Speech movements overlap in connected speech. For example, the lips may begin rounding in preparation for a rounded vowel while the preceding consonant is still being produced. This overlap is called and is one reason the same sound can vary slightly across contexts.
Describing consonants and vowels
Consonants are commonly described by voicing, place of articulation, and manner of articulation. Voicing indicates whether the vocal folds vibrate; place identifies where airflow is narrowed or blocked; and manner describes how airflow is shaped or restricted.
The examples show how these properties combine:
[p] is a bilabial stop: both lips close, briefly blocking airflow, and then release it.
[b] is also bilabial and a stop, but is typically .
[f] is a labiodental fricative: the lower lip approaches the upper teeth, creating a narrow gap that produces turbulent airflow.
[m] is a bilabial nasal: the lips close while the soft palate lowers, allowing air to escape through the nose.
Vowels are produced with a relatively open vocal tract. They are described mainly by tongue height, from high to low; tongue backness, from front to back; and lip rounding. In [i], as in the vowel of machine, the tongue is high and front and the lips are unrounded. In [u], as in many pronunciations of goose, the tongue is high and back and the lips are rounded. Vowel categories and exact pronunciations vary across languages and speakers.
Acoustic properties of speech
Speech travels as changing air pressure. A waveform shows changes in amplitude over time, while a spectrogram shows how energy at different frequencies changes over time.
Several acoustic properties help describe speech:
Frequency, measured in hertz (Hz), relates to the repetition rate of a sound wave. For sounds, the rate of vocal-fold vibration is the , written as , and is an important cue to perceived pitch.
Amplitude relates to the strength of pressure variations and is associated with perceived loudness. Loudness also depends on frequency and hearing conditions.
Duration is the length of a sound or speech event.
are resonant frequency regions shaped by the vocal tract. The first two, and , are especially useful for describing vowel quality. generally rises as tongue position is lower, while generally tends to be higher for front vowels than for back vowels. These are useful patterns, not perfect one-to-one measurements of tongue position.
The offers a simplified account of speech: vocal-fold vibration supplies an acoustic source, and the vocal tract filters it by strengthening some frequency regions and weakening others. Changing vocal-tract shape changes the resulting spectrum, helping distinguish vowels. Consonants create other characteristic patterns: a stop may show a brief closure followed by a release burst, while a fricative often has sustained turbulent noise.
How listeners perceive speech
Sound is conducted through the outer and middle ear to the inner ear. The cochlea converts sound vibrations into neural signals, which the auditory system analyzes over time and frequency. Listeners use acoustic cues to identify sound categories and interpret them in sequence.
There is no simple one-to-one relationship between a speech sound and a single acoustic feature. Listeners combine cues, and a cue can be influenced by neighboring sounds, speaking rate, voice characteristics, and context. For example, listeners use timing cues as well as other information to distinguish consonants such as [b] and [p].
Perception is not just a passive copy of the signal: expectations and linguistic context can also help listeners resolve unclear or noisy speech.
Representing speech with the IPA
The represents speech sounds with standardized symbols. Unlike ordinary spelling, IPA transcription aims to record pronunciation rather than written letters. The International Phonetic Association maintains the official chart, which organizes consonant symbols by place and manner of articulation and displays vowels according to tongue position and lip rounding.
Square brackets mark a phonetic transcription, as in [ʃ], the first sound in ship. Slashes are commonly used for phonological representations. A symbol in square brackets represents a sound, not necessarily a letter.
Diacritics add phonetic detail. The small raised [ʰ] marks aspiration, a puff of air after a stop; [pʰ] is an aspirated [p]. A broad transcription records selected sound details, while a narrow transcription records more precise phonetic details. The amount of detail depends on the purpose.
In many English pronunciations, the initial [p] in pin is aspirated, while the [p] in spin is not. A relatively narrow transcription can show the contrast as [pʰɪn] versus [spɪn]. The IPA represents this pronunciation difference consistently even though English spelling uses the same letter p.