Audio measurements & fidelity:
the ten basic rules
What is the connection between measurements and hearing? What are measurements for and what is listening to music for? What is the problem with 'blind testing'? - Principles, methods, misconceptions.
May 4, 2025 (Last edited: 2026.07.28.)
The validity of audio measurements are a subject of ongoing debate in audiophile blogs and forums. To outsiders and newbies, measurements seems quite chaotic, as it's difficult to see the connections between different disciplines and their different approaches. Without a compass it's easy to get lost in the noisy jungle of misconceptions, false analogies, unfounded claims and marketing frauds. The following article provides a brief overview of the most important ground rules covering the science and logic behind sound reproduction and audio measurement. I hope these principles not only dispels many misconceptions, but also shed new light on old knowledge.
In audio magazines and forums, it's very common to blur the line between subjective and objective, that is to claim that a characteristics is subjective, when in fact it is not. If we try to characterize the components of the audio chain with the emotional reaction to music, we will certainly fool ourselves. The musical experience, the enjoyment of music is influenced by many factors besides audio fidelity - audio fidelity is only a part of the musical experience, although an important and determining component. (Problems usually start when the audible differences become minor, belief mess up judgment and the wire drama begins...)
1. Audio science is an applied science
Audio science (science of sound reproduction) is an applied science and not a fundamental science. However, as other applied fields it has roots in fundamental sciences. - Any claim about amplifiers, DACs, audio formats or loudspeakers is a claim about hearing, physics or signal behavior.
2. How much can we trust our hearing?
We can easily recall the sound of a piano or guitar without much difficulty, but comparing subtle differences requires special methods. This also applies to vision, with the difference that comparing images is a somewhat more straightforward process, since the analytical and focusing ability of vision is more refined. Separating one part of an image from the rest is a simple task, but separating one part of the music from the rest - for example, the sound of an instrument in an orchestral music - can be extremely difficult. It's not surprising that comparing minor audible differences requires careful attention and special conditions. We need a method that eliminates the possibility of making a wrong judgment.
However, even if we can definitely hear a difference, that's not the end of the story. By testing with music we can only make claims about the sound and not about the technology or the real cause. Revealing the cause-effect ("mechanism") requires additional test conditions. For example, we need to rule out the possibility that the difference is caused by an unknown (hidden) variable (the test is not false positive or false negative).
Key steps to do a proper listening test:
- audio samples should be short (maximum five seconds long),
- switching between the samples should be fast and silent (maximum two seconds),
- only one parameter should be changed at a time, otherwise the test is invalid,
- in addition to musical excerpts, test signals, solo instrumental recordings may be needed,
- the selection of test material is critical, as test tones affect the sensitivity of the test and the variance of the results,
- in some tests, digital filters can also be applied to improve the sensitivity of the test (e.g. testing audio formats),
- the tests should be designed in such a way that they not only provide results, but also to infer some kind of causal relationship,
- one test is not a test, test should be repeated, preferably on another day.
3. Blind test is not a "golden hammer"!
Blind test, or its most commonly used form, the ABX test, is a universal tool for eliminating self-deception arising from our expectations. Blind test doesn't guarantee that the test is free of flaws, all variables are known (to avoid false positive or false negative results) and if the test is free of errors, correlation still doesn't imply causation.
Blind testing methodology is the most misunderstood topic in audio. Rejection of blind testing is absurd, on the other hand, searching for the truth with blind tests and statistical methods is a hopeless and completely pointless endeavour. Blind testing with music is not a verification method, it shouldn't be used as a primary tool, only as a complementary method to support measurements. Testing with music only makes sense if we already have a well-founded knowledge of the system to be tested. Furthermore, the constant forcing of blind testing as a golden hammer can be a barrier in a deeper understanding of hearing, audio components, formats... The exclusive use of tests can lead not only to misconceptions but also to superficial knowledge! (Note: a listening test is a verification method only if we can prove that the selected audio sample has special qualities and these special qualities turn the test into a conclusive test. Otherwise we just collect data...)
The idea that non-blind testing ("sighted test") is inherently flawed is also wrong. People have different susceptibility to confirmation bias which depends on experience and mental state, as people have different susceptibility to marketing or political propaganda. People believe in nonsense because they don't understand a scientific theory well, or they don't know how to brake down a complex problem into smaller, easily verifiable tasks, and not due to the lack of blind tests.
Key to science is the distinction between knowledge without comprehension (~test results) and knowledge with comprehension (~mechanism). Numerous scientific tools (surveys, correlation analysis, data science in general) can only contribute to the knowledge without comprehension at most. Understanding causal relationships requires special experiments. (The biggest mistake in science is using statistical analysis without deductive logic, reducing scientific knowledge to "correlations" (correlation between data) and ignoring causality.)
Instead of ABX testing, a more reasonable goal is to find test methods that can be considered a conclusive test ("experimentum crucis"). Fortunately, there are plenty of them.
Two uncommon "artifacts" that may lead to false positive results (as false cause)
Polarity inversion (or polarity reversal). A "hidden" polarity inversion may generate false positive test results. Though the ear is insensitive to polarity inversion, we can detect polarity indirectly by distortion byproducts. If we send an asymmetrical signal through a nonlinear stage with asymmetrical nonlinearity, the generated distortion products change according to the polarity of the signal. Cone displacement in loudspeakers is rarely symmetrical, which result in 2nd order distortion at low frequencies, ear has its own asymmetrical nonlinearity (increasing 2nd order distortion from 80 dBSPL). Since music is a good masker, polarity inversion can only be detected with solo instruments (e.g. trumpet) and special tones (asymmetrical periodic signals).
It's also worth checking the quality of (re)sampling at least with a high-frequency tone before making hasty statements. Non-transparent resampling is rare today, but who knows...
4. Key to measurements - the six questions
The logic behind audio measurements isn't that complicated: audio measurements are signal measurements, measuring only those characteristics that can affect what we hear. Audio fidelity of a component is determined objectively by measuring the change(s) in the signal as it passes through the component and comparing with the corresponding threshold(s).
Audio measurements are only useful if they correlate with perception. Unfortunately, some old-school measurement methods taken directly from electrical measurements with minor modifications correlate poorly with hearing (SNR, SINAD). However, understanding how we hear can be very helpful in interpreting less accurate measurements. If we know the limitations of SINAD and SNR, we will not fool ourselves with SINAD and SNR.
All problems related to audio measurements can be grouped into six categories (or six "key points"):
- What is the physical process behind the technology used? What does the system do with the signal? (This is the level of physical explanation, or in other words, the physical model. In digital systems, the question is what the mathematical calculations do with the signal.)
- How can we measure signal changes, signal attributes, model variables? (I mean completely general signal measurements, not specific audio measurement.)
- What is the auditory process behind the perception of the "artifact" or characteristic, and how can we create meaningful audio measurements? What is the nature of the signal change that we want to measure? (three main auditory process: auditory masking, absolute threshold of hearing, loudness sensation)
- With which type of signal the error is the most audible? (we look for the signal representing the "worst-case")
- What are the individual differences between the thresholds? (mean, P90, P75, P10... )
- How do measurements, measurement signals and thresholds relate to the sound sources around us? (musical instruments, speech, dog barking, thunder, crickets chirping, etc.)
We can use these six-point framework for anything: quantization, resampling, lossy audio compression, nonlinear distortion, resonances, speaker cables, loudspeaker feet / spikes, etc...
Measurements can't describe the full performance of an audio component...
Measurements can't describe the full performance of an audio component is not only a myth, but in order to prove that measurements fail, one has to perform all existing relevant measurements and only if they fail we can say that there is something wrong with the measurements. Comparing DACs without measurements, picking basic measurements (frequency response, distortion at 1 kHz) or picking irrelevant measurements don't prove anything...
5. A bit more about the accuracy of measurements
A huge advantage of measurements is that they are more accurate than our hearing. But what does this really mean? Unfortunately, life is full of ambiguous expressions... and accuracy is just one of them.
We must distinguish numerical accuracy from modeling accuracy (validity of prediction). All modern hardware and software provide exceptional numerical accuracy, but not all measurement method provide modeling accuracy. For example, SINAD/THD and SNR are based on an oversimplified model of human hearing. As a result, SINAD/THD measurements and comparisons done with SINAD/THD can be misleading if the measurement result is lower than ~70 dB (when distortion can be audible). SNR measurements are also problematic when noise spectra differ significantly. Without a hearing model, FFT measurements (frequency spectrum analysis) provide low modeling accuracy for noise and transients.
An audio measurement can be only considered accurate, if it has numerical accuracy and the model behind the measurement is also accurate.
It's worth mentioning that "measurements are reliable and hearing is unreliable" is not exactly true, because by choosing the right test signal and method hearing can become a very accurate and reliable tool. Just think about psychoacoustic experiments: the whole point of psychoacoustic experiments is that they are more reliable and accurate than ordinary tests.
A huge benefit of measurements
Measurements are a way to create verifiable models by providing non-ambiguous results. A model with ambiguous output or trivial predictions can't be verified (actually rejected) and improved...
6. What we can hear...
Psychoacoustics studies the human hearing with special test tones, creates hearing models and determines various thresholds. What we can hear is determined by the Absolute Thresholds of Hearing (ATH) and auditory masking. Masking means that in the neighborhood of a (loud) tone the hearing threshold is raised. Signals below the threshold are inaudible.
Not only the audibility of pure tones or compex tones, but even the audibility of nonlinear distortion, noise or resonances is related to masking and ATH. (In fact, there is a third mechanism: adaptation or compression, a shift in the non-masked threshold.)
7. "Purity" of the signal is irrelevant
It is not the purity of the signal that matters... In an audio system, the goal is to preserve and transmit the signal in such a way that accumulated errors cannot be heard or low enough to not affect playback fidelity.
The shape of the signal itself is also irrelevant (square wave response, impulse response). We can't assess fidelity by looking at the waveform, only by applying the appropriate "perceptual rules".
This also means that the aversion to software resampling and the hype around bitperfect playback that is typical today is just another nonsense.
High-end audio is not necessarily about high fidelity
Many people confuse high fidelity with high-end audio. High fidelity means lack of audible coloration ("transparency"), whereas high-end audio is an unfair business based mainly on people's insecurity and gullibility.
8. Audio measurements & fidelity - the main categories
(A short overview.)
Audio fidelity of a component is determined by frequency response, nonlinear distortion curves and noise (necessary for expressing dynamic range). Time domain measurements (impulse response, phase shift, group delay) are secondary, as DACs and amplifiers have negligible phase distortion in the audible range. In multi-way loudspeakers the phase distortion is orders of magnitude higher, but even this is not audible in the impulse response. Time domain measurements are essential only in room acoustics and differ significantly from phase/group delay measurements. (It makes little sense to argue about phase-shift and "time resolution", since time-domain behavior can be easily tested with a pulse or pulse series. More about in this article.)
In amplifiers and DACs, crosstalk between channels can also be considered an important parameter, although it's very rare to find a system with audible crosstalk. Jitter (fluctuation of the clock signal) doesn't require separate measurements, as it manifests itself as nonlinear distortion.
Audio fidelity measurements describe to what extent an audio system can reproduce the original performance (live, electronic). Poor frequency response can change the timbre (certain harmonics are emphasized while others are attenuated). Distortion can also change the timbre (new harmonics are generated) and high level of distortion can hide low-level details. Noise can also hide low-level signals. Usually a system or component with poor measured performance will not sound good. Very poor audio fidelity can even be painful to listen to...
Audiophoolery from Ethan Winer is a great introduction to this topic.
9. Thresholds, signals and "worst-case" thresholds
Just-noticeable distortion, just-noticeable noise, just-noticeable level difference, just-noticeable change in group delay varies according to the harmonic and temporal properties of the audio signal. There always exists a particular signal with the lowest threshold, which can be identified as "worst-case".
For example, nonlinear distortion is best heard with pure tones and two-tones. Detection is more difficult with complex time-varying signals and pulses. In contrast, time-related errors are more audible in pulses or pulse series. (Yes, some measurement signals are worst-case signals.)
The definition of "transparency" and transparent system is closely related to the concept of 'worst-case'. For example, transparency is reached for nonlinearity when distortion is not audible with worst-case type test signals. The only exception is the frequency response of loudspeakers and headphones, where the “worst-case” (just-noticeable level difference with pure tone) is a too strict criterion.
Audibility threshold and the relationship between threshold and waveform are perhaps the most important principles in audio. Any question related to fidelity is about threshold and the relationship between threshold and waveform.
10. What is a trained ear?
"Trained ear" is also a controversial concept. Maybe, the controversy partially can be attributed to the word "trained", because the expression suggests that we can train our ears as we can train our muscles. Though we can improve our discrimination skills, this doesn't mean that listening to a lot of music some day we will be able to hear something that can't be measured. This is just wishful thinking. The other problem is that experience with music is no substitute for expertise, and training our hearing is not an effective way to solve scientific problems, otherwise all audiophiles would be audio experts.
Hearing is partially a learning process, which means that some characteristics of our hearing can be improved by learning, but the basic characteristics cannot. The skills of a musically trained person are similar to the skills of a photographer, painter or a sports expert. A sports expert can see things that an untrained person can't, because he can categorize events very fast and knows where to see, and not because he has sharper eyes.
On the flip side, our discrimination skills and accuracy of a listening test can be improved much more by finding the right test signal(s) than by "training". This is the most important rule. A main benefit of psychoacoustic experiments and specialized conclusive listening tests is that they don't require excessive training. Anyone can be trained quickly to determine the audibility of a resonance or pre-echo with pulses. Also, the results can be interpreted and communicated in a simple way.
And another one:
11. Music is for enjoyment and not for comprehensive testing (+1)
Maybe I should have started with this one because of its importance, but it's more logical to discuss at the end.
I always thought that we should listen to music for enjoying it and not for comprehensive testing. Detection of performance issues is usually harder with music and easier with special test tones, not to mention measurements. As a rule of thumb, more complex the signal (spectrum), higher the masking threshold and harder to detect nonlinear distortion and noise. The other problem is that if we mix different signals, we have to reduce their level so that the resulting signal fits in the dynamic range. For example, if we mix a 50 Hz tone with a pulse, the level of both signals should be reduced and the 50 Hz tone will put less stress on the amplifier and the pulse will also put less stress on the amplifier (DAC, loudspeaker...).
There are some areas where it's impossible to avoid music as a test signal, for example testing the quality of lossy audio compression, or studying the percpetion of spacioussness and envelopment in rooms and recordings. All in all, I'm happy that I don't have to compare DACs and amps with music, and I can skip boring (and rather dumb) ABX tests, and measurements can do that job.
Notes: Testing loudspeakers and headphones with music can reveal a lot because there are large differences between them. Besides, there are test signals for reliable detection of certain types of nonlinear distortion and resonance, but there is no good test signal for reliable detection of frequency response flaws.
Conclusive test: a simple way to test audio components by ear
I mentioned conclusive tests previously. You can find more about these tests in this article.
Csaba Horváth
Related articles:
Lossy audio compression: principles, methods, misconceptions 🔊 🎧
Sampling rate controversy: simple and conclusive test methods 🔊 🎧
Audibility thresholds for SINAD / THD+N measurements
Noise perception, detection threshold & dynamic range
Speaker Driver Simulation With Room Response (online simulator)



