Skip to content

Loudness Metering, An Overview

The objective measurement and quantification of an audio signal’s perceived loudness has recently seen growing interest. Broadcast and music/video streaming services, in particular, depend on it to standardize and automate music delivery and audio production.

Out of the earliest attempts grew a seemingly confusing maze of metering methods and terminology. This article attempts to put the most popular approaches into historical, technical, and perceptual context. Specifically from a modern music production engineering perspective.

Decibel, Phon and Sone

These three foundational terms represent an important chronological evolution in sound measurement.

Introduced in 1920, the decibel measures the physical intensity or pressure of sound. It is a logarithmic scale that compresses the massive physical range of sound that our ears can tolerate (from a whisper to a jet engine) into a manageable scale.

Introduced in 1925 by German physicist Heinrich Barkhausen and further established by researchers Harvey Fletcher and W. A. Munson (creating the famous Fletcher-Munson equal-loudness contours), the Phon measures the perceived loudness level.
Because human ears do not perceive all frequencies equally (we are terrible at hearing low or ultra-high frequencies), the phon matches the loudness of any sound to a pure 1000 Hz tone, which is then expressed in phons.

In 1936, American psychologist Stanley Smith Stevens introduced the concept of Sone. He recognized that while the Phon tracked sine levels at different frequencies, it still wasn’t a linear scale (40 phons is not “twice as loud” as 20 phons).
One sone is universally defined as the loudness of a 1kHz tone at 40 phons. If a sound is measured at 2 sones, it sounds exactly twice as loud as a 1 sone sound.

To summarize: the decibel (dB) provides physical measurement, the phon introduces subjective frequency adjustments, and the sone translates this into a linear perception of volume.

Equal Loudness Contours

As we’ve seen, mapping from the Decibel (sound pressure) to the Phon (perceived, contextual, subjective loudness) involves some form of psychoacoustic model. The initial works involved the famous Fletcher/Munson tonal measurements, which were later revised and refined, finding their modern equivalent in the ISO 226 equal loudness contours. Note that both sets of curves describe purely tonal signals.


Interestingly, noise-like signals such as a noise floor, bursts or clicks (i.e. transients, percussive signals) are perceived quite differently. ITU-R 468 describes the noise-specific equal loudness weighting:

This poses a challenge to music content loudness metering, because music signals consist both of tonal and noise-like content, in more or less equal amounts.

In addition, acoustics and many other factors can also have significant influence on the effective spectral weighting, acoustics in particular.

Advanced Perceptual Phenomena

Ideally, this psychoacoustic model also takes into account the individual critical bands and their interaction. This involves phenomena such as Auditory Masking, Time Envelope Effects, sound-field considerations and other contextual phenomena.

This animation shows Auditory Masking effect of a single tone at various loudnesses:


Analogue circuit design quickly hit a complexity issue in this context, unable to cope with the physical/computational requirements. Hence, and out of realistic options, most modern and historical standards still tend to ignore these seeming “subtleties”.
Modern digital systems do not suffer under these costs and restrictions and will likely develop further.

Metering

The earliest and still the most reliable form of loudness estimation is the questionnaire. While it is somewhat hopeless to ask a group of people to directly assign their sensation of loudness to a number on a scale, it is possible to observe their sensations in relative terms.

In the context of music production, the audio engineer will intuitively balance out the loudness of different sources in a creatively meaningful way, and with a bit of experience, even on the fly. A DJ is a great example, literally acting as a human loudness controller.

VU Meter

One of the first attempts at measuring and representing the loudness of a signal electro-mechanically was the VU meter (VU = Volume Unit). Standardized in 1942 (ANSI C16.5-1942), it employs a rather crude approach to average out the audio signal. This was initially done by the physical inertia of the needle’s own mass and later by electronic/digital filters. The standard suggests raw rectification, flipping negative values to positive, followed by some form of filtering that imposes a 300 ms rise/fall time. Mathematically, this process yields the average rectified value (ARV).

AC Input -> Rectifier (Absolute Value) -> Low Pass Filtering -> DC Output

While this method is cheap and easy to implement, its physical basis somewhat shaky. It estimates signal energy with error, and generally doesn’t have much in common with the loudness a listener perceives when subject to an arbitrary signal.
Given any sine signal, the average level will be 3.92 dB below its peak. Mathematically, its energy is really 3.01 dB below its peak, imposing significant error.
VU meters ignore the strongly frequency-dependent nature of loudness perception, as well as many of its more complex mechanisms.

The VU meter further has a linear scale. This doesn’t match the way the auditory system perceives audio and thus greatly restricts the range of operation to a very small range around the headroom limits of the system. In the case of VU meters, decibels are not spread evenly across the range, making them impractical for low or moderate signal levels, especially those that modern digital workflows demand.

This makes VU meters a questionable choice for loudness metering.

It is worth noting, however, that the VU meter has turned out to be useful for representing how analogue processors work against the available headroom. VU meters are foremost headroom meters. Their lack of frequency weighting has turned into a feature in this context.

Root-Mean-Square (RMS) Meters

The physically correct method to detect a signal’s average energy content is the Root-Mean-Square operation. The way it operates is very similar to the ARV described above, except that it first squares, Applies a low-pass filter, and then scales it back to the original via square-root operation.

AC Input -> Square -> Low Pass Filtering -> Square Root -> DC Output

Analogue implementation of the RMS detector remained challenging until David E. Blackmer invented the “true RMS detector” in 1971, offering a reliable and precise solution to the broader market.

This development suddenly made loudness metering far more practical, preparing ground for a wide range of weighted RMS metering approaches and standards that followed. These include various SPL measurement standards (typically classified by their spectral weighting), the K-System, and the basis for the ITU-R BS.1770/EBU R128 broadcast standard.

While weighted RMS offers an approximation to perceived loudness, it does not reliably match the perceived loudness of arbitrary signals.

CBS Loudness Meter

In 1967, CBS Laboratories published a paper describing an elaborate loudness meter with improved frequency weighting and a mechanism that took into account the auditory system’s critical bands and their individual time-envelope effects.

Interestingly, CBS Meter Frequency Weighting leans more toward the ISO 468 curve:


CBS Meter structure:


The CBS loudness meter is a broadcast hardware/software system developed for television and radio that targets dialogue intelligibility and helps prevent listener fatigue. Popularized by Orban broadcast processors, the CBS algorithm has proven its effectiveness not only in the form of visual indication, but also by directly processing millions of hours of on-air programming.

CBS meters measure levels directly related to broadcast standards, such as LKFS (Loudness, K-weighted, relative to Full Scale) and Loudness Units (LU).

Zwicker Loudness

Building on the work of Karl Eberhard Zwicker, the Zwicker method first appeared in an international standard in 1975, as Method B in the original ISO 532. Similar to the CBS method, it calculates perceived loudness based on auditory excitation patterns, but with greater depth. It mimics how the human ear processes sound by incorporating auditory masking, a psychoacoustic phenomenon where the presence of a louder sound (the masker) renders a softer, nearby partial inaudible.

Approximation of auditory filters, i.e., the critical bands:

The Zwicker models are primarily used to analyze the loudness of noise-like signals, including complex audio signals and music. They require SPL-calibrated monitoring.
A distinction is made between a free sound field (anechoic, with no reflecting surfaces) and a diffuse sound field (reverberant, with sound arriving from all directions).
For non-stationary signals a percentile mechanism (similar to median calculation) is used to derive the overall loudness.

Zwicker requires analyzing wide-band audio to observe specific patterns and how the ear receives all frequencies at once. This is computationally heavy but excellent for testing psychoacoustic harshness.

Zwicker produces physical loudness values measured in Sone or Phon.

ITU-R BS1770/EBU R128

Broadcasters use Peak Programme Meters to comply with broadcast regulations. These meters measure the highest electrical peaks of an audio signal, which fail to reflect human perception of “how loud” a program is.

The ITU-R BS.1770 standard and EBU R128 recommendation revolutionized audio by moving the broadcast industry’s attention away from “peak normalization” to loudness normalization. This transition, in theory, ended the jarring volume jumps—often referred to as the “loudness war”—between television shows and commercials. Twofold, though, as the technology can also be used to secure premium loudness specifically to advertisers.

In 2006, the International Telecommunication Union (ITU) published the ITU-R BS.1770 standard, which suggested a universal algorithm to measure and quantify perceived audio loudness. It introduced:

  • K-weighting: A specific frequency-weighting filter that mimics how human ears perceive different frequencies.
  • LUFS (Loudness Units relative to Full Scale): An absolute unit of measurement where 1 LU = 1 dB.

While the ITU specified the method (frequency-weighted RMS), it did not define a target level or an operational practice. In 2010 the EBU published EBU R128, introducing:

  • Target Level: An integrated loudness target of -23 LUFS.
  • Loudness Range (LRA): A metric used to describe the variation of loudness within a program to ensure aesthetic compatibility.
  • True Peak: A maximum ceiling of -1 dBTP (decibels true peak) to prevent digital clipping.
  • EBU Mode Metering: Defined a specific gating method (the “G8” gate) that ignores quiet passages to ensure the measurement reflects foreground loudness.

The ITU and EBU worked collaboratively to continually refine the measurement parameters. in 2011 they further introduced:

  • ITU-R BS.1770-2: Added the relative gating method pioneered by the EBU, meaning that both the ITU standard and EBU R128 would calculate foreground loudness in the same way.
  • ITU-R BS.1770-3 (2012) & 1770-4 (2015): Introduced further technical refinements to multi-channel audio handling and filter coefficients.
  • EBU R128 updates: The EBU continually reviewed and expanded the R128 guidelines, with Version 4.0 published in 2020, formalizing stricter tolerances.

EBU R128 became the foundational standard for broadcasting across Europe—with stations adopting it as early as 2012 in Germany.

Today, these foundational principles have been adopted by streaming giants like Spotify, Apple Music, and YouTube to keep content consistently leveled. However, music services usually target far greater loudness levels around -14 LUFS.

It is worth noting that the primary objective of these initiatives is standardization and implementation simplicity to enable wide adoption in broadcast circles. Contrary to the CBS Loudness algorithm and Zwicker method, they do not seek to build a realistic model of human auditory perception or assist music production.

Conclusion

Loudness metering is a rather recent appearance and is likely subject to further developments. While modern digital environments make it increasingly easy to adopt more elaborate models of our now reasonably well understood auditory system, any added complexity also hinders widespread acceptance.

In music production specifically, workflows and aesthetic preferences have already adapted to the concept of signal loudness, moving away from peak-level-centric decision making.

The weighted RMS approach, as imperfect as it is, has proven sufficient for many musical applications. It remains to be seen whether and how far more elaborate metering methods offer practical benefits in the music production context. But it is worth considering the revolutionary effect that modern psychoacoustic models have had on the success of lossy audio compression schemes. These are exactly the aspects “modern” weighted RMS loudness metering methods use to ignore.

Published inTech Articles
Copyright © Tokyo Dawn Records