Audio Compression

Data compression that reduces the size of audio files. General-purpose compression algorithms are poorly suited to audio data and achieve little reduction. Compression algorithms intended for audio were therefore devised: reversible, lossless algorithms and irreversible, lossy ones.

Lossless

Lossless compression reduces the quantity of data without losing any of the original audio information. (Note 1) It compresses by exploiting regularities in the audio data, and the original audio is restored perfectly on playback. Its greatest advantage is therefore that the sound quality captured at the recording stage is preserved. In general, however, lossless compression can reduce the data to around 50% of the original size at best.

Lossy

Lossy compression achieves a far larger reduction by omitting sound information held to be difficult for humans to hear. (Note 2) It keeps data traffic down, but information once discarded cannot be recovered. The music remains enjoyable, yet fine detail — the lustre and texture of the sound, the decay, the sense of air — is easily lost. As a result, the data can be reduced to between 5% and 20% of the original.

Compression and Immersive Audio

The immersive audio recording technique promoted by this laboratory is a time-difference method that extends AB recording.

The minute differences in time, phase and correlation between the channels are themselves the spatial information that forms the spread and depth of the space. However, some lossy compression methods (Note 3) can simplify this phase and time information, potentially resulting in a loss of envelopment and spaciousness.
Lossless compression, which loses no information, is therefore recommended for immersive audio.
Combining it with high-resolution sources (Note 4), which hold more information than a CD, makes a more natural and richer spatial expression possible. The Irimajiri Immersive Audio Laboratory considers that it is immersive audio above all that requires high resolution if a deeper sense of presence is to be achieved. In practice, the difference between a compressed source and a lossless high-resolution source is more readily perceived in immersive audio than in stereo.

Note 1: An example of the mathematics behind lossless compression

Sound is like a wave travelling across the surface of water, so there are no extreme jumps between the heights of adjacent points of the waveform. Transmitting the differences between adjacent values, rather than the height values themselves, therefore reduces the quantity of information while still allowing the original data to be reconstructed completely.

Note 2: How lossy compression works

Also known as perceptual coding or psychoacoustic coding. It is an audio compression technique that uses the characteristics of the human ear (psychoacoustics) to remove from the data sounds that people cannot hear or are unlikely to notice. MP3 and AAC are representative examples.

  1. The minimum audible level varies with frequency (this is called the absolute threshold of hearing, or ATH).
  2. When a loud sound and a quiet sound occur at the same time, the quieter sound becomes more difficult to hear (this is called masking).

In other words, components quieter than the absolute threshold of hearing, and components drowned out by louder sounds, can be discarded without anyone noticing.
This processing reduces the data to about one tenth of the original without greatly harming the overall impression of the music.

Note 3: Lossy compression of multichannel audio

Lossy coding schemes such as the Dolby Digital family and MPEG-H may use joint coding, channel coupling or the parameterisation of spatial information — all of which exploit correlation between channels — in order to reduce the bit rate. At high frequencies in particular, instead of holding several channels as entirely independent waveforms, they may treat them as a common component together with per-channel level and envelope information. In a conventional, panning-based mix this processing rarely produces any marked incongruity. In immersive recording founded on time-difference AB techniques, however, it is the minute delays between channels, the phase differences, the fluctuations in correlation and the uncorrelated components of reverberation and reflections that sustain the spatial density, the sense of distance, the depth and the envelopment. Even where the timbre and the broad sense of direction are preserved, therefore, lossy compression may smooth out and merge the spatial information, diminishing the natural sense of space captured at the recording.

Note 4: High resolution

An abbreviation for high-resolution audio. High-resolution audio refers to audio sources with a higher sampling rate and bit depth than CD audio, and therefore a greater amount of audio information. On the basis of research finding that humans hear only up to 20 kHz, CD discards information above 20 kHz. It is true that very few people can hear a pure tone above 20 kHz, yet many people can distinguish a difference in sound quality between high-resolution audio and CD audio, and research has also reported differences in brain-wave responses depending on whether information above 20 kHz is present. Analogue records, incidentally, can carry information up to around 50 kHz.