Head-Related Transfer Function

What is a head-related transfer function (HRTF)? This article explains how humans perceive the direction of sound and how HRTFs are applied to binaural and spatial audio.

A head-related transfer function (HRTF) is a function that expresses mathematically the changes in sound that occur as it travels to the two ears, as a result of the effects of the head, pinnae, shoulders, torso, and other parts of the body.

Although humans have only two ears, we can perceive where a sound is coming from in three dimensions, including not only front-back and left-right directions but also elevation. This is because, as sound travels to the two ears, reflections and diffraction caused by the shape of the head and ears produce differences in arrival time, sound pressure, and frequency response between the two ears.

When a sound arrives from the right, for example, the sound reaching the left ear is delayed and slightly lower in level than the sound reaching the right ear because it must travel around the head. The complex shape of the pinna also produces reflections and resonances that vary with frequency. As a result, sounds arriving from above or below, or from the front or rear, have different frequency characteristics. The brain uses these characteristics, learned from early childhood, as cues to estimate the direction and distance of the sound source.

To obtain an HRTF, miniature microphones are placed near the ears of an actual human subject, and test signals are reproduced from various directions. The changes in sound caused by the shape of the pinnae, head, chest, and other parts of the body are measured for each direction. Based on these measurements, a function is derived that represents changes in the frequency characteristics of sound as parameters of the sound source’s direction and distance.

In recent years, technology known as binaural rendering has become widely used to convert multichannel immersive audio sources produced for loudspeaker playback into binaural audio for headphones, using HRTFs. Spatial audio streaming services such as Apple Music and Amazon Music also use technologies based on this approach, including Dolby Atmos.

Because HRTFs differ from person to person according to the shape of the head and pinnae, general-purpose systems typically use a generic HRTF representing an average listener. As a result, for some people, aspects such as sound localization, the sense of height, and front-back perception may not be reproduced sufficiently.

The most effective way to address this issue is to use an HRTF specific to each individual. However, measuring an individual’s HRTF directly is a large-scale and demanding process.

Research and practical applications of personalization technologies are therefore advancing. These technologies estimate an individual’s HRTF from photographs of the ears, 3D scans, and other information. Apple’s “Personalized Spatial Audio” is one example, and such technologies are making immersive audio reproduction over headphones increasingly natural and accurate.