REVIEW 5 major objections 6 minor 66 references
Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that adaptively weighting PEAQ's distortion metrics by cognitive effect measures—perceptual streaming, informational masking, and speech probability—generalizes better to unseen codecs than standard tools and…
desk verdict A solid PEAQ extension with genuine gains on unseen databases, but the interaction-selection procedure needs a stability check before I'd fully trust the mechanism. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Cognitive Salience Model (CSM), a weighted-sum architecture that replaces PEAQ's neural-network mapping stage. Its load-bearing parts are the salience measure $S_m(j)$, computed as the Pearson correlation between one DM's basis-function output and the mean subjective scores across the treatments of signal $j$; the interaction metric $C_m$ that scores each candidate cognitive-effect/distortion pair; and the Detection Probability Weights (DPWs), sigmoid functions of the cognitive effect metrics that scale each DM's contribution. The CSM's defining move is that cognitive effects enter only as multipliers on distortion terms, never as direct predictors of the final score, which restricts the model's learnable interaction space and is the claimed source of its generalization.
What would settle it
Re-run the interaction selection using a calibration set in which each content item has many more codec/bitrate conditions—say ten or more—so that Eq. (1)'s per-signal correlations are computed from more than three or four points; if the selected interactions change materially or the validation correlations on the seven databases drop below the reported average of 0.80, then the salience-by-correlation assumption is doing the work and the model as published is not stable.
Extended reading notes
Core claim
The central claim is that prediction generalization in objective audio quality assessment can be improved by an explicit model of how cognitive effects modulate distortion salience. The proposed architecture keeps PEAQ's psychoacoustic front end and its six model output values as distortion metrics (DMs), but replaces the fixed or generally-learned mapping from metrics to score with a Cognitive Salience Model: each DM is first mapped to a quality scale by a basis function trained on isolated-artifact listening tests, and then weighted by Detection Probability Weights derived from three cognitive effect metrics (EPN for perceptual streaming, PDEV for informational masking, and a speech-music probability). The interaction structure is chosen by a two-stage data-driven procedure—an interaction metric that measures how well each transformed cognitive effect predicts each DM's salience, followed by step-wise regression to select the final terms. Validated on seven unseen databases (MUSHRA, BS.1116, and blind source separation), the resulting PEAQ-CSM+ achieves an average Pearson correlation of about 0.80 and specifically improves prediction on parametrically coded, non-waveform-preserving audio, where the authors report other methods fail.
Load-bearing premise
The method's interaction selection rests on estimating how salient each distortion type is from the correlation between one distortion metric and the mean listener scores across only a few codec/bitrate versions of each signal; with so few data points per signal, that estimate is noisy, and if it misranks distortion salience the chosen cognitive weights—and with them the claimed generalization—lose their support.
Editorial extensions
If this is right
- If the reported generalization holds, quality prediction for parametric and low-bitrate codecs improves enough that the method can be used in codec development and selection where standard PEAQ, ViSQOL, and PEMO-Q give weak correlations.
- The separation of basis-function calibration (isolated artifacts) from interaction calibration (USAC VT1) means new distortion types or new cognitive effects can be added without retraining the entire mapping stage.
- The explicit interaction terms align with psychoacoustic results, for example speech probability raising the salience of noise loudness and lowering the salience of linear distortions, so the model doubles as a quantitative statement about which distortion types matter in which listening contexts.
- Because the reduced parameter count keeps inference time near 0.5 ms, the approach can be deployed in large-scale codec evaluation sweeps with negligible added cost.
Reading between the lines
- The interaction-selection step is the fragile point: replacing the calibration database (USAC VT1) with another multi-distortion database could change which CEM/DM pairs survive, so a leave-one-database-out rerun would isolate whether the architectural constraint (cognitive effects only as multipliers) or the specific interaction table drives the reported generalization.
- The same salience-weighting scheme could transfer to non-intrusive, reference-free assessment and to multidimensional attributes such as timbre, loudness, or spatial quality, because the CSM separates distortion measurement from cognitive weighting; a P.800 speech-quality variant is one explicit direction the paper's future work points to.
- Because the salience estimate needs several treatments per signal, the benefit of the method should grow as subjective databases include more codec/bitrate conditions per content item; databases with only one treatment per signal cannot support the interaction analysis at all.
- The USAC VT2 result (correlation 0.63 versus 0.30 for standard PEAQ) shows parametric stereo distortions remain largely unexplained; adding the planned binaural model is the natural test of whether cognitive weighting extends to spatial quality or marks the method's scope boundary.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PEAQ-CSM+, an extension of the PEAQ perceptual audio quality assessment system. Individual PEAQ distortion metrics are mapped to quality scores via MARS basis functions trained on an isolated-artifacts listening database. Cognitive effect measures (speech probability, perceptual streaming, informational masking) are then used as adaptive weights on the distortion metrics through sigmoidal detection-probability weights (DPWs), with the CEM/DM interaction graph selected by an interaction metric and stepwise regression on the USAC VT1 database. The resulting model is validated on seven previously unseen subjective-quality databases, reporting a mean Pearson correlation of about 0.80 and outperforming several general-purpose ML mapping stages and established objective metrics. The authors argue that the architecture's explicit cognitive-interaction structure explains its generalization advantage.
Significance. If the reported generalization holds, this is a valuable contribution to objective audio quality assessment: it offers an interpretable, low-parameter-count alternative to black-box ML mappings, with consistent performance across diverse codecs, bitrates, and even an out-of-domain blind-source-separation database (SASSEC). The paper's strengths include validation on multiple large independent databases, comparison against a wide set of state-of-the-art metrics (PEAQ DI, ViSQOL, PEMO-Q, CombAQ, GPSMq), a transparent architecture, and an unusually honest limitations section. The promise of extending the framework to spatial and multi-dimensional quality measurement without full retraining is also substantively appealing. However, the statistical evidence for the central superiority claim and the stability of the data-driven interaction selection currently require additional support.
major comments (5)
- [Section IV-C, Figures 6 and 7] The statement 'CI95%R ≤ ±0.01 for all estimates' is not credible for the sample sizes involved (N between 144 and 280 for the validation databases). For example, a correlation of 0.80 with N=200 has an approximate 95% confidence interval of about ±0.06, not ±0.01. Because the central claim that PEAQ-CSM+ outperforms other systems depends on differences such as 0.84 vs 0.88 on ELD VT(A) (Figure 7), the authors must report proper confidence intervals, significance tests, or a bootstrap analysis; otherwise the observed differences may be within sampling noise on several databases.
- [Section II-C and Section III-B] The interaction selection and DPW parameter optimization are performed on the same USAC VT1 data, with salience estimates S_m(j) in Eq. (1) computed as per-signal Pearson correlations. With N=216 and a plausible split of J≈24 signals and I≈9 treatments per signal, each S_m(j) is based on roughly 7 residual degrees of freedom, and the subsequent C_m in Eq. (2) is a correlation across only about 24 signals. The exhaustive search over two sigmoid parameters to maximize |C_m| on this same data is therefore a selection-on-noise procedure. The paper provides no bootstrap, cross-validation, or perturbation analysis to show that the selected interactions in Table V and the coefficients in Table VI are stable. This is load-bearing because the paper's identity as a reproducible cognitive model, rather than a favorable draw from a noisy selection procedure, depends on such stability.
- [Section IV-B, Table V] The exact DPW sigmoid parameters (steepness and crossover midpoint) are not reported anywhere, and no implementation or code is provided. As a result, the proposed PEAQ-CSM+ model is not fully specified and cannot be independently implemented or compared. The authors should either report the optimized sigmoid parameters for each DPW in Table V or make the implementation publicly available.
- [Section III-B, Table III] The manuscript does not state the number of signals J and the number of treatments I per signal in the USAC VT1 calibration database. These values are necessary to assess the reliability of the salience measure in Eq. (1) and the interaction metric in Eq. (2). Please report them explicitly and discuss the consequences of small I and J for the stability of the selected interactions.
- [Section II-C3] The stepwise linear regression that selects the final quality terms and reports 'p < 0.05' (Table VI) is applied to the same USAC VT1 data that was used for interaction selection and DPW optimization. Post-selection p-values and R=0.91 on the calibration data do not provide evidence of generalizability. The authors should clarify that these statistics are descriptive and should not be interpreted as inferential evidence for the selected model, or should provide a properly separated validation of the selection procedure.
minor comments (6)
- [Section IV-C] In Figure 7, the 'MEAN' row is an unweighted average of Pearson correlations across databases of different sizes and quality ranges. Please state whether Fisher z-transformation was considered, or justify the simple average as a summary measure.
- [Section II-C2] Equation (2) takes the absolute value of the correlation, discarding sign information. The subsequent discussion of positive and negative interactions is clear, but the sign convention should be stated more explicitly in the text near Eq. (2).
- [Figure 4] The labels 'propForSpeech' in Figure 4 are inconsistent with the CEM name 'probSpeech' used in Table II and elsewhere; please unify the nomenclature.
- [Section II-A] Footnote 1 states that the MATLAB implementation is not publicly available but that 'similar results could be achieved' with the open-source implementation in [28]. Please specify whether the reported validation results were obtained with the authors' private implementation, which limits reproducibility.
- [Section III-G] The description of [28] as 'The C implementation of PEAQ' is imprecise; [28] is an open-source implementation (GstPEAQ) but not necessarily a C implementation in the sense used. Please rephrase.
- [Section V-A] The discussion of Q6 states that its inverse polarity 'might be reacting to border effects' and 'needs further investigation.' This post-hoc interpretation is speculative; it would be helpful to present it explicitly as a hypothesis rather than a conclusion.
Circularity Check
No significant circularity: the model is calibrated on USAC VT1 and the central claim rests on validation with unseen databases.
full rationale
The paper's derivation chain is not circular. The basis functions are estimated on the isolated-artifacts database [43]; the salience measure (Eq. 1), interaction metric (Eq. 2), DPW optimization, interaction selection, and final regression are calibration operations on USAC VT1. The paper does not present these as predictions, and Section III-D explicitly states: 'With the exception of USAC VT1, none of the validation databases were used for calibrating any of the parameters of the proposed quality measurement system.' The headline result (Figure 7) is computed on ITU DB4, ELD VT (A/T), MPEG-H, USAC VT2/VT3, and SASSEC, all unseen during calibration. No equation reduces to its input by construction: Eq. 1 and Eq. 2 are fit criteria, and Eq. 3 is the resulting weighted sum whose coefficients are estimated on the calibration set. The self-citations to [18], [24], and [63] are contextual; [18] introduced the CSM architecture, but the present paper provides its own detailed derivation and independent empirical validation. The acknowledged limitations (Section V-G: low-rated speech signals, HE-AACv2 mono overestimation, and the lack of a spatial model for USAC VT2) are honestly stated performance limitations, not evidence of circularity. The interaction-selection procedure involves estimating per-signal salience correlations from small treatment counts, which is a robustness/generalization risk, but the external validation on several independent databases is exactly the test of that risk, so it is not a circularity.
Assumptions & free parameters
free parameters (4)
- Regression coefficients Q0-Q7 =
Q0=58.3, Q1=2.69, Q2=2.61, Q3=1.97, Q4=1.54, Q5=1.64, Q6=-1.89, Q7=8.49 (Table VI)
- DPW sigmoid parameters (steepness and crossover midpoint) =
Not reported
- Basis function (MARS) parameters =
Shown as curves in Figure 3
- MARS maxFuncs hyperparameter =
3
assumptions (5)
- domain assumption PEAQ advanced perceptual model produces valid excitation patterns and distortion metrics for audio quality assessment
- domain assumption Cognitive effects (perceptual streaming, informational masking, speech probability) modulate the salience of distortions multiplicatively via DPWs
- domain assumption Per-signal DM salience S_m(j) can be estimated by Pearson correlation between BF and mean subjective scores across the treatments per signal
- domain assumption The interaction metric C_m in Eq. (2) identifies meaningful cognitive-distortion interactions
- domain assumption The isolated artifacts and USAC VT1 databases are representative for calibrating BFs, DPWs, and regression coefficients
Cite this review
Pith. "Pith review of Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model." pith.science (2026). https://pith.science/paper/35DAHMIF
@misc{pith2026241118222,
author = {Pith},
title = {Pith review of: Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/35DAHMIF}},
note = {Machine review of arXiv:2411.18222}
}
read the original abstract
Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for quality judgment. Many of these tools use perceptual auditory models to extract audio features that are mapped to a basic audio quality score prediction using machine learning algorithms and subjective scores as training data. However, existing tools struggle with generalization in quality prediction, especially when faced with unknown signal and distortion types. This is particularly evident in the presence of signals coded using non-waveform-preserving parametric techniques. Addressing these challenges, this two-part work proposes extensions to the Perceptual Evaluation of Audio Quality (PEAQ - ITU-R BS.1387-1) recommendation. Part 1 focuses on increasing generalization, while Part 2 targets accurate spatial audio quality measurement in audio coding. To enhance prediction generalization, this paper (Part 1) introduces a novel machine learning approach that uses subjective data to model cognitive aspects of audio quality perception. The proposed method models the perceived severity of audible distortions by adaptively weighting different distortion metrics. The weights are determined using an interaction cost function that captures relationships between distortion salience and cognitive effects. Compared to other machine learning methods and established tools, the proposed architecture achieves higher prediction accuracy on large databases of previously unseen subjective quality scores. The perceptually-motivated model offers a more manageable alternative to general-purpose machine learning algorithms, allowing potential extensions and improvements to multi-dimensional quality measurement without complete retraining.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[18]
A data-driven cognitive salience model for objective perceptual audio quality assessment,
P. M. Delgado and J. Herre, “A data-driven cognitive salience model for objective perceptual audio quality assessment,” in ICASSP 2022- 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 986–990
work page 2022
-
[1]
ISO/MPEG-1 audio: A generic standard for coding of high-quality digital audio,
K. Brandenburg and G. Stoll, “ISO/MPEG-1 audio: A generic standard for coding of high-quality digital audio,” Journal of the Audio Engineer- ing Society, vol. 42, no. 10, pp. 780–792, 1994
work page 1994
-
[2]
C. Timmerer, M. Wien, L. Yu, and A. Reibman, “Special issue on open media compression: Overview, design criteria, and outlook on emerging standards,” Proceedings of the IEEE , vol. 109, no. 9, pp. 1423–1434, 2021
work page 2021
-
[3]
Objective assessment of speech and audio quality—technology and applications,
A. W. Rix, J. G. Beerends, D. . Kim, P. Kroon, and O. Ghitza, “Objective assessment of speech and audio quality—technology and applications,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 14, no. 6, pp. 1890–1901, 2006
work page 1901
-
[4]
BS.1387, Method for objective measurements of perceived audio quality, Geneva, Switzerland, 2001
ITU-R Rec. BS.1387, Method for objective measurements of perceived audio quality, Geneva, Switzerland, 2001
work page 2001
-
[5]
P.863, Perceptual Objective Listening Quality Assessment , Geneva, Switzerland, 2014
ITU-T Rec. P.863, Perceptual Objective Listening Quality Assessment , Geneva, Switzerland, 2014
work page 2014
-
[6]
Objective assessment of perceptual audio quality using ViSQOLAudio,
C. Sloan, N. Harte, D. Kelly, A. C. Kokaram, and A. Hines, “Objective assessment of perceptual audio quality using ViSQOLAudio,” IEEE Transactions on Broadcasting, vol. PP, no. 99, pp. 1–13, 2017
work page 2017
-
[7]
PEMO-Q—a new method for objective audio quality assessment using a model of auditory perception,
R. Huber and B. Kollmeier, “PEMO-Q—a new method for objective audio quality assessment using a model of auditory perception,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 14, no. 6, pp. 1902–1911, Nov 2006
work page 1902
Show all 66 references
-
[8]
High fidelity neural audio compression,
A. D ´efossez, J. Copet, G. Synnaeve, and Y . Adi, “High fidelity neural audio compression,” 2022
2022
-
[9]
Deep noise suppression maxi- mizing non-differentiable PESQ mediated by a non-intrusive PESQNet,
Z. Xu, M. Strake, and T. Fingscheidt, “Deep noise suppression maxi- mizing non-differentiable PESQ mediated by a non-intrusive PESQNet,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1572–1585, 2022
2022
-
[10]
Objective measures of perceptual audio quality reviewed: An evaluation of their application domain dependence,
M. Torcoli, T. Kastner, and J. Herre, “Objective measures of perceptual audio quality reviewed: An evaluation of their application domain dependence,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1530–1541, 2021
2021
-
[11]
P.862, Perceptual evaluation of speech quality (PESQ) , Geneva, Switzerland, Dec
ITU-R Rec. P.862, Perceptual evaluation of speech quality (PESQ) , Geneva, Switzerland, Dec. 2001
2001
-
[12]
PEAQ - the ITU standard for objective measurement of perceived audio quality,
T. Thiede, W. C. Treurniet, R. Bitto, C. Schmidmer, T. Sporer, J. G. Beerends, and C. Colomes, “PEAQ - the ITU standard for objective measurement of perceived audio quality,” J. Audio Eng. Soc. , vol. 48, no. 1/2, pp. 3–29, January/February 2000
2000
-
[13]
Perceptual techniques in audio quality assessment,
A. Rix, “Perceptual techniques in audio quality assessment,” Ph.D. dissertation, University of Edinburgh, 2003
2003
-
[14]
Novel deep autoencoder features for non- intrusive speech quality assessment,
M. H. Soni and H. A. Patil, “Novel deep autoencoder features for non- intrusive speech quality assessment,” in 2016 24th European Signal Processing Conference (EUSIPCO), 2016, pp. 2315–2319
2016
-
[15]
A multitask teacher- student framework for perceptual audio quality assessment,
C.-W. Wu, P. A. Williams, and W. Wolcott, “A multitask teacher- student framework for perceptual audio quality assessment,” in 2021 29th European Signal Processing Conference (EUSIPCO) , 2021, pp. 396–400
2021
-
[16]
A differentiable perceptual audio metric learned from just noticeable differences,
P. Manocha, A. Finkelstein, R. Zhang, N. J. Bryan, G. J. Mysore, and Z. Jin, “A differentiable perceptual audio metric learned from just noticeable differences,” in Interspeech, Oct. 2020
2020
-
[17]
Generative machine listener,
G. Jiang, L. Villemoes, and A. Biswas, “Generative machine listener,” in Audio Engineering Society Convention 155 , Oct 2023
2023
-
[19]
A. S. Bregman, Auditory scene analysis: The perceptual organization of sound. MIT press, 1994
1994
-
[20]
BS.1534, Method for the subjective assessment of interme- diate quality levels of coding systems , Geneva, Switzerland, 2015
ITU-R Rec. BS.1534, Method for the subjective assessment of interme- diate quality levels of coding systems , Geneva, Switzerland, 2015
2015
-
[21]
BS.1116, Methods for the subjective assessment of small impairments in audio systems , Geneva, Switzerland, 2015
ITU-R Rec. BS.1116, Methods for the subjective assessment of small impairments in audio systems , Geneva, Switzerland, 2015
2015
-
[22]
Sound quality assessment: Concepts and criteria,
T. Letowski, “Sound quality assessment: Concepts and criteria,” in Audio Engineering Society Convention 87 , New York, Oct 1989. [Online]. Available: http://www.aes.org/e-lib/browse.cfm?elib=5869
1989
-
[23]
Assessment and prediction of binaural aspects of audio quality,
J.-H. Fleßner, R. Huber, and S. D. Ewert, “Assessment and prediction of binaural aspects of audio quality,” J. Audio Eng. Soc, vol. 65, no. 11, pp. 929–942, 2017. [Online]. Available: http://www.aes.org/e-lib/browse.cfm?elib=19361
2017
-
[24]
Can we still use PEAQ? A performance analysis of the ITU standard for the objective assessment of perceived audio quality,
P. M. Delgado and J. Herre, “Can we still use PEAQ? A performance analysis of the ITU standard for the objective assessment of perceived audio quality,” in 2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX) , 2020, pp. 1–6
2020
-
[25]
A new cognitive model for objective assessment of audio quality,
J. G. A. Barbedo and A. Lopes, “A new cognitive model for objective assessment of audio quality,” J. Audio Eng. Soc , vol. 53, no. 1/2, pp. 22–31, 2005. [Online]. Available: http: //www.aes.org/e-lib/browse.cfm?elib=13386
2005
-
[26]
Objective evaluation of speech signal quality by the prediction of multiple foreground diagnostic acceptability measure attributes,
D. Sen and W. Lu, “Objective evaluation of speech signal quality by the prediction of multiple foreground diagnostic acceptability measure attributes,” The Journal of the Acoustical Society of America, vol. 131, no. 5, pp. 4087–4103, 2012. [Online]. Available: https://doi.org/...
2012 doi
-
[27]
Natick, Massachusetts: The MathWorks Inc., 2023
The MathWorks, version 9.14.0 (R2023a). Natick, Massachusetts: The MathWorks Inc., 2023
2023
-
[28]
GstPEAQ–an open source implementation of the PEAQ algorithm,
M. Holters and U. Z ¨olzer, “GstPEAQ–an open source implementation of the PEAQ algorithm,” in Proc. 18th Int. Conf. Digital Audio Effects (DAFx), Trondheim, Norway, 2015
2015
-
[29]
A perceptual audio quality measure based on a psychoacoustic sound representation,
J. G. Beerends and J. A. Stemerdink, “A perceptual audio quality measure based on a psychoacoustic sound representation,” J. Audio Eng. Soc , vol. 40, no. 12, pp. 963–978, 1992. [Online]. Available: http://www.aes.org/e-lib/browse.cfm?elib=7019
1992
-
[30]
Zwicker and H
E. Zwicker and H. Fastl, Psychoacoustics. Facts and Models. Springer, 1999
1999
-
[31]
Perceptual audio quality assessment using a non-linear filter bank,
T. Thiede, “Perceptual audio quality assessment using a non-linear filter bank,” Ph.D. dissertation, Fachbereich Elektrotechnik, Technische Universit¨at Berlin, 1999
1999
-
[32]
A robust speech/music discriminator for switched audio coding,
G. Fuchs, “A robust speech/music discriminator for switched audio coding,” in 23rd European Signal Processing Conference, EUSIPCO 2015, Nice, France, August 31 - September 4, 2015 . IEEE, 2015, pp. 569–573. [Online]. Available: https://doi.org/10.1109/EUSIPCO.2015. 7362447
2015 doi
-
[33]
Perceptual Audio Codecs - What to Listen For, Web Edition,
S. Dick and AES Technical Committee on Coding of Audio Signals (TC- CAS), “Perceptual Audio Codecs - What to Listen For, Web Edition,” in Audio Engineering Society , May 2021. [Online]. Available: https:// aes2.org/resources/audio-topics/audio coding/perceptual-audio-codecs/
2021
-
[34]
The role of informational masking and perceptual streaming in the measurement of music codec quality,
J. G. Beerends, W. A. C. van den Brink, and B. Rodger, “The role of informational masking and perceptual streaming in the measurement of music codec quality,” in Audio Engineering Society Convention 100 , Copenhagen, May 1996. [Online]. Available: http://www.aes.org/e-lib/brow...
1996
-
[35]
Influence of working memory and attention on sound-quality ratings,
R. Huber, S. R ¨ahlmann, T. Bisitz, M. Meis, S. Steinhauser, and H. Meister, “Influence of working memory and attention on sound-quality ratings,” The Journal of the Acoustical Society of America, vol. 145, no. 3, pp. 1283–1292, 2019. [Online]. Available: https://doi.org/10.11...
2019 doi
-
[36]
ARESLab: Adaptive regression splines toolbox for MATLAB. http://www.cs.rtu. lv/jekabsons/regression.html,
G. Jekabsons, “ARESLab: Adaptive regression splines toolbox for MATLAB. http://www.cs.rtu. lv/jekabsons/regression.html,” 2019
2019
-
[37]
Psychometric functions for level discrimination,
S. Buus, C. R. Mason, and M. Florentine, “Psychometric functions for level discrimination,” The Journal of the Acoustical Society of America, vol. 82, no. S1, pp. S25–S25, 1987. [Online]. Available: https://doi.org/10.1121/1.2024720
1987 doi
-
[38]
Hastie, R
T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction , ser. Springer series in statistics. Springer, 2009
2009
-
[39]
Submission and evaluation procedures for 3D audio N13633,
ISO/IEC JTC1/SC29/WG11, “Submission and evaluation procedures for 3D audio N13633,” International Organisation for Standardisation, Tech. Rep., 2013
2013
-
[40]
Progress report towards revision of recommenda- tion ITU-R BS.1387-1 (Annex 13 to Document 6C/415-E),
ITU-R BS.1387-1, “Progress report towards revision of recommenda- tion ITU-R BS.1387-1 (Annex 13 to Document 6C/415-E),” Geneva, Switzerland, Nov. 2010
2010
-
[41]
Report on the verification test of MPEG- 4 Enhanced Low Delay AAC N10032,
ISO/IEC JTC1/SC29/WG11, “Report on the verification test of MPEG- 4 Enhanced Low Delay AAC N10032,” International Organisation for Standardisation, Hannover, Germany, Tech. Rep., 2008
2008
-
[42]
USAC verification test report N12232,
——, “USAC verification test report N12232,” International Organisa- tion for Standardisation, Tech. Rep., 2011
2011
-
[43]
Generation and evaluation of isolated audio coding artifacts,
S. Dick, N. Schinkel-Bielefeld, and S. Disch, “Generation and evaluation of isolated audio coding artifacts,” in Audio Engineering Society Convention 143 , New York, Oct 2017. [Online]. Available: http://www.aes.org/e-lib/browse.cfm?elib=19206
2017
-
[44]
The sebass-db: A consolidated public data base of listening test results for perceptual evaluation of bss quality mea- sures,
T. Kastner and J. Herre, “The sebass-db: A consolidated public data base of listening test results for perceptual evaluation of bss quality mea- sures,” in 2022 International Workshop on Acoustic Signal Enhancement (IWAENC), 2022, pp. 1–5
2022
-
[45]
First stereo audio source separation evaluation campaign: data, algorithms and results,
E. Vincent, H. Sawada, P. Bofill, S. Makino, and J. P. Rosca, “First stereo audio source separation evaluation campaign: data, algorithms and results,” in International Conference on Independent Component Analysis and Signal Separation . Springer, 2007, pp. 552–559
2007
-
[46]
P.1401, Methods, metrics and procedures for statistical evaluation, qualification and comparison of objective quality prediction models, Geneva, Switzerland, 2012
ITU-T Rec. P.1401, Methods, metrics and procedures for statistical evaluation, qualification and comparison of objective quality prediction models, Geneva, Switzerland, 2012. IEEE/ACM TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING, VOL. XX, 2023 15
2012
-
[47]
Workplan towards draft revision of recommendation ITU- R BS.1387-1,
ITU-R, “Workplan towards draft revision of recommendation ITU- R BS.1387-1,” Annex 12 to Working Party 6C Chairman’s Report. International Telecommunication Union, Tech. Rep., 2009
2009
-
[48]
Quality assessment of multi-channel audio processing schemes based on a binaural auditory model,
J. H. Flessner, S. D. Ewert, B. Kollmeier, and R. Huber, “Quality assessment of multi-channel audio processing schemes based on a binaural auditory model,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2014, pp. 1340–1344
2014
-
[49]
Vision and technique behind the new studios and listening rooms of the Fraunhofer IIS audio laboratory,
A. Silzle, S. Geyersberger, G. Brohasga, D. Weninger, and M. Leistner, “Vision and technique behind the new studios and listening rooms of the Fraunhofer IIS audio laboratory,” in Audio Engineering Society Convention 126 , Munich, May 2009. [Online]. Available: http://www.aes....
2009
-
[50]
MATLAB implementation of the Com- bined Audio Quality Model,
T. Biberger and J. Fleßner, “MATLAB implementation of the Com- bined Audio Quality Model,” https://gitlab.uni-oldenburg.de/kuxo2262/ combinedaudioqualitymodel, 2019, accessed: 2022-07-07
2019
-
[51]
Subjective and objective as- sessment of monaural and binaural aspects of audio quality,
J. Fleßner, T. Biberger, and S. D. Ewert, “Subjective and objective as- sessment of monaural and binaural aspects of audio quality,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, no. 7, pp. 1112–1125, 2019
2019
-
[52]
An objective audio quality measure based on power and envelope power cues,
T. Biberger, J.-H. Fleßner, R. Huber, and S. D. Ewert, “An objective audio quality measure based on power and envelope power cues,” Journal of the Audio Engineering Society , vol. 66, no. 7/8, pp. 578– 593, 2018
2018
-
[53]
The PEASS toolkit-perceptual evaluation methods for audio source separation,
V . Emiya, E. Vincent, N. Harlander, and V . Hohmann, “The PEASS toolkit-perceptual evaluation methods for audio source separation,” in 9th Int. Conf. on Latent Variable Analysis and Signal Separation , 2010
2010
-
[54]
ViSQOL Audio MATLAB implementation. http://www.sigmedia.tv/tools,
A. Hines, E. Gillen, D. Kelly, J. Skoglund, A. Kokaram, and N. Harte, “ViSQOL Audio MATLAB implementation. http://www.sigmedia.tv/tools,” Accessed 2019
2019
-
[55]
Compression artifacts in perceptual audio coding,
C.-M. Liu, H.-W. Hsu, and W.-C. Lee, “Compression artifacts in perceptual audio coding,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 16, no. 4, pp. 681–695, 2008
2008
-
[56]
Spectral pattern, harmonic relations, and the perceptual grouping of low-numbered components,
B. Roberts and J. M. Brunstrom, “Spectral pattern, harmonic relations, and the perceptual grouping of low-numbered components,” The Journal of the Acoustical Society of America , vol. 114, no. 4, pp. 2118–2134, 2003
2003
-
[57]
Enhancing the performance of perceptual audio coders by using Temporal Noise Shaping (TNS),
J. Herre and D. Johnston, “Enhancing the performance of perceptual audio coders by using Temporal Noise Shaping (TNS),” in 101st AES Convention, Los Angeles, 1996, preprint 4384
1996
-
[58]
The ISO/MPEG unified speech and audio coding standard—consistent high quality for all content types and at all bit rates,
M. Neuendorf, M. Multrus, N. Rettelbach, G. Fuchs, J. Robilliard, J. Lecomte, S. Wilde, S. Bayer, S. Disch, C. Helmrich, R. Lefebvre, P. Gournay, B. Bessette, J. Lapierre, K. Kj ¨orling, H. Purnhagen, L. Villemoes, W. Oomen, E. Schuijers, K. Kikuiri, T. Chinen, T. Norimatsu, K...
2013
-
[59]
Objective assessment of spatial audio quality using directional loudness maps,
P. M. Delgado and J. Herre, “Objective assessment of spatial audio quality using directional loudness maps,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019
2019
-
[60]
Perceptual objective quality evaluation method for high quality multichannel audio codecs,
J.-H. Seo, S. B. Chon, K.-M. Sung, and I. Choi, “Perceptual objective quality evaluation method for high quality multichannel audio codecs,” J. Audio Eng. Soc , vol. 61, no. 7/8, pp. 535–545, 2013. [Online]. Available: http://www.aes.org/e-lib/browse.cfm?elib=16869
2013
-
[61]
Temporal envelope-based psychoacoustic modelling for evaluating non-waveform preserving audio codecs,
S. van de Par, S. Disch, A. Niedermeier, E. Burdiel P ´erez, and B. Edler, “Temporal envelope-based psychoacoustic modelling for evaluating non-waveform preserving audio codecs,” in AES Convention, New York, 2019, p. 10314. [Online]. Available: http: //www.aes.org/e-lib/browse...
2019
-
[62]
An efficient model for estimating subjective quality of separated audio source signals,
T. Kastner and J. Herre, “An efficient model for estimating subjective quality of separated audio source signals,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2019, pp. 95–99
2019
-
[63]
An improved metric of informational masking for perceptual audio quality measurement,
P. M. Delgado and J. Herre, “An improved metric of informational masking for perceptual audio quality measurement,” in 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2023, pp. 1–5
2023
-
[64]
P.800, Methods for subjective determination of transmission quality, Geneva, Switzerland, Aug
ITU-T Rec. P.800, Methods for subjective determination of transmission quality, Geneva, Switzerland, Aug. 1996
1996
-
[65]
NISQA: A deep CNN- self-attention model for multidimensional speech quality prediction with crowdsourced datasets,
G. Mittag, B. Naderi, A. Chehadi, and S. M ¨oller, “NISQA: A deep CNN- self-attention model for multidimensional speech quality prediction with crowdsourced datasets,” arXiv preprint arXiv:2104.09494 , 2021. Pablo M. Delgado is a member of the scientific staff at the Advanced ...
2021 arXiv
-
[1989]
In 1995, Dr
Since then he has been involved in the de- velopment of perceptual coding algorithms for high quality audio, including the well-known ISO/MPEG- Audio Layer III coder (aka “MP3”). In 1995, Dr. Herre joined Bell Laboratories for a Post-Doctoral term working on the development of...
1995
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.