{"id":"0e939a4e-a6f4-4e4c-ab16-16ba588cd460","arxiv_id":"2411.18222","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A cognitive weighting layer on top of PEAQ's distortion metrics improves prediction of unseen subjective audio quality scores across codec and source-separation databases.","lead":"This paper presents a new way to predict how listeners rate audio quality by making the quality metric adapt its focus based on the type of signal and the kind of distortion. If the results hold, audio codec developers could rely more on this automated metric and less on expensive listening tests.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The interaction graph in Table VI is selected from per-signal salience correlations (Eq.","rationale":"The strongest claim is empirical: PEAQ-CSM+ predicts previously unseen subjective scores better than the alternatives. The mechanism claimed to deliver this is the explicit cognitive interaction model, so the load-bearing condition is that the interaction graph in Table VI is a real, reproducible property of the data rather than an artifact of the calibration procedure. The reader's weakest assumption identifies exactly this point: Eq. 1's per-signal salience correlations are noisy, and the DPW optimization and stepwise selection amplify that noise by fitting on the same USAC VT1 data. If the selected interactions are unstable, the model structure is not identifiable, and the external validation, while encouraging, is only one draw from a noisy selection process. The proposed bootstrap or leave-one-signal-out check directly tests this: stable selected terms and stable external R would resolve the concern, whereas instability would require substantial qualification of the generalization claim. I therefore agree with the reader's identification of the weakest assumption and see no reason to shift the conditional verdict; the concern is addressable and the paper should provide the stability evidence and the missing parameter/code details.","tokens_in":22819,"tokens_out":15274,"duration_ms":155873,"concrete_test":"Run a bootstrap or leave-one-signal-out resampling of the USAC VT1 calibration set and re-run the full pipeline on each resample: fix the BFs from the isolated-artifact database, perform the exhaustive DPW sigmoid optimization, carry out stepwise interaction selection, and re-estimate regression coefficients. Then measure (a) how often each of the seven Q terms in Table VI is selected and (b) the distribution of validation R over the seven external databases. If term selection flips frequently (e.g., a term is selected in fewer than 80% of resamples) or the standard deviation of mean external R exceeds roughly 0.03, the current single calibration cannot support the claimed generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that PEAQ-CSM+ generalizes because cognitive effects modulate the salience of different distortion types. The selection of those modulations (Table V) rests on Eq. 1: S_m(j) is the Pearson correlation between one DM's basis function and the mean subjective scores across the I treatments available for signal j. In USAC VT1 (N=216), a plausible split of J≈24 signals implies only I≈9 treatments per signal, so each S_m is estimated from roughly 7 residual degrees of freedom; the standard error of a correlation near 0.7 is then about 0.2. The interaction metric C_m in Eq. 2 correlates these noisy S_m values with thresholded CEM outputs across J≈24 signals, and the two sigmoid parameters of each DPW are chosen by exhaustive search to maximize |C_m| on the same data. This is a selection-on-noise procedure: even near-random salience estimates can be made to track a CEM by fitting the sigmoid. The subsequent stepwise regression and coefficient estimation again use the same USAC VT1 subjective scores. If the selected interaction graph in Table VI is not stable under small perturbations of the calibration database, the architecture has no demonstrated identity: the good external R values could be one favorable draw from a noisy selection process rather than evidence of a reproducible cognitive mechanism. The paper reports no calibration stability analysis, no internal cross-validation, and no release of the exact DPW parameters or code to assess this directly.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PEAQ-CSM+, an extension of the PEAQ perceptual audio quality assessment system. Individual PEAQ distortion metrics are mapped to quality scores via MARS basis functions trained on an isolated-artifacts listening database. Cognitive effect measures (speech probability, perceptual streaming, informational masking) are then used as adaptive weights on the distortion metrics through sigmoidal detection-probability weights (DPWs), with the CEM/DM interaction graph selected by an interaction metric and stepwise regression on the USAC VT1 database. The resulting model is validated on seven previously unseen subjective-quality databases, reporting a mean Pearson correlation of about 0.80 and outperforming several general-purpose ML mapping stages and established objective metrics. The authors argue that the architecture's explicit cognitive-interaction structure explains its generalization advantage.","tokens_in":23080,"tokens_out":6386,"duration_ms":59394,"significance":"If the reported generalization holds, this is a valuable contribution to objective audio quality assessment: it offers an interpretable, low-parameter-count alternative to black-box ML mappings, with consistent performance across diverse codecs, bitrates, and even an out-of-domain blind-source-separation database (SASSEC). The paper's strengths include validation on multiple large independent databases, comparison against a wide set of state-of-the-art metrics (PEAQ DI, ViSQOL, PEMO-Q, CombAQ, GPSMq), a transparent architecture, and an unusually honest limitations section. The promise of extending the framework to spatial and multi-dimensional quality measurement without full retraining is also substantively appealing. However, the statistical evidence for the central superiority claim and the stability of the data-driven interaction selection currently require additional support.","major_comments":[{"comment":"The statement 'CI95%R ≤ ±0.01 for all estimates' is not credible for the sample sizes involved (N between 144 and 280 for the validation databases). For example, a correlation of 0.80 with N=200 has an approximate 95% confidence interval of about ±0.06, not ±0.01. Because the central claim that PEAQ-CSM+ outperforms other systems depends on differences such as 0.84 vs 0.88 on ELD VT(A) (Figure 7), the authors must report proper confidence intervals, significance tests, or a bootstrap analysis; otherwise the observed differences may be within sampling noise on several databases.","section":"Section IV-C, Figures 6 and 7"},{"comment":"The interaction selection and DPW parameter optimization are performed on the same USAC VT1 data, with salience estimates S_m(j) in Eq. (1) computed as per-signal Pearson correlations. With N=216 and a plausible split of J≈24 signals and I≈9 treatments per signal, each S_m(j) is based on roughly 7 residual degrees of freedom, and the subsequent C_m in Eq. (2) is a correlation across only about 24 signals. The exhaustive search over two sigmoid parameters to maximize |C_m| on this same data is therefore a selection-on-noise procedure. The paper provides no bootstrap, cross-validation, or perturbation analysis to show that the selected interactions in Table V and the coefficients in Table VI are stable. This is load-bearing because the paper's identity as a reproducible cognitive model, rather than a favorable draw from a noisy selection procedure, depends on such stability.","section":"Section II-C and Section III-B"},{"comment":"The exact DPW sigmoid parameters (steepness and crossover midpoint) are not reported anywhere, and no implementation or code is provided. As a result, the proposed PEAQ-CSM+ model is not fully specified and cannot be independently implemented or compared. The authors should either report the optimized sigmoid parameters for each DPW in Table V or make the implementation publicly available.","section":"Section IV-B, Table V"},{"comment":"The manuscript does not state the number of signals J and the number of treatments I per signal in the USAC VT1 calibration database. These values are necessary to assess the reliability of the salience measure in Eq. (1) and the interaction metric in Eq. (2). Please report them explicitly and discuss the consequences of small I and J for the stability of the selected interactions.","section":"Section III-B, Table III"},{"comment":"The stepwise linear regression that selects the final quality terms and reports 'p < 0.05' (Table VI) is applied to the same USAC VT1 data that was used for interaction selection and DPW optimization. Post-selection p-values and R=0.91 on the calibration data do not provide evidence of generalizability. The authors should clarify that these statistics are descriptive and should not be interpreted as inferential evidence for the selected model, or should provide a properly separated validation of the selection procedure.","section":"Section II-C3"}],"minor_comments":[{"comment":"In Figure 7, the 'MEAN' row is an unweighted average of Pearson correlations across databases of different sizes and quality ranges. Please state whether Fisher z-transformation was considered, or justify the simple average as a summary measure.","section":"Section IV-C"},{"comment":"Equation (2) takes the absolute value of the correlation, discarding sign information. The subsequent discussion of positive and negative interactions is clear, but the sign convention should be stated more explicitly in the text near Eq. (2).","section":"Section II-C2"},{"comment":"The labels 'propForSpeech' in Figure 4 are inconsistent with the CEM name 'probSpeech' used in Table II and elsewhere; please unify the nomenclature.","section":"Figure 4"},{"comment":"Footnote 1 states that the MATLAB implementation is not publicly available but that 'similar results could be achieved' with the open-source implementation in [28]. Please specify whether the reported validation results were obtained with the authors' private implementation, which limits reproducibility.","section":"Section II-A"},{"comment":"The description of [28] as 'The C implementation of PEAQ' is imprecise; [28] is an open-source implementation (GstPEAQ) but not necessarily a C implementation in the sense used. Please rephrase.","section":"Section III-G"},{"comment":"The discussion of Q6 states that its inverse polarity 'might be reacting to border effects' and 'needs further investigation.' This post-hoc interpretation is speculative; it would be helpful to present it explicitly as a hypothesis rather than a conclusion.","section":"Section V-A"}],"recommendation":"major_revision","confidential_remarks":"This is a technically interesting paper with a strong external validation dataset setup and an honest limitations section. The main weaknesses are statistical reporting (the impossible 95% CI claim), the lack of stability analysis for the interaction-selection procedure, and missing DPW parameters/code. These are all addressable within the scope of a revision and do not, in my view, require rejection. I would encourage the editor to send the paper back for major revision with emphasis on bootstrap/cross-validation stability of the selected cognitive interactions and rigorous reporting of confidence intervals."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a credible, practical extension of PEAQ, not a revolution. The genuinely new pieces are the detection probability weights, the interaction cost function for deciding which cognitive effects modulate which distortion metrics, and the broad validation across seven unseen databases, including parametric codecs and blind source separation. The consistent advantage over PEAQ DI, ViSQOL, PEMO-Q, and CombAQ is the strongest evidence in the paper. Mean correlation around 0.80, with large gains on the ELD VT sets and SASSEC, is a real result. The architecture is also cheap to run and much simpler than the ANN and KSOM alternatives. That alone earns peer review.\n\nWhat worried me is the interaction selection. The salience measure in Eq. 1 is a Pearson correlation between one DM basis function and mean subjective scores across the treatments of a signal. In USAC VT1, that likely means something like 7-9 residual degrees of freedom per signal, so each S_m is noisy. Eq. 2 then correlates those noisy estimates with thresholded CEM outputs, and the sigmoid parameters are chosen by exhaustive search on the same data. The stepwise regression that fixes the final coefficients and drops candidate terms also runs on USAC VT1. So Table VI is selected on the calibration set, and the paper gives no bootstrap, split-half, or internal cross-validation to show the interaction graph is stable. Without released code or the exact DPW parameters, a reader cannot tell whether the selected interactions would survive perturbation of the calibration database. That is a real gap.\n\nBut it is a gap in demonstrated stability of the mechanism, not in the empirical claim. The validation databases are genuinely unseen, and the gains are broad and consistent. The paper is also candid about its limits: USAC VT2 spatial distortions, low-rated speech, and HE-AACv2 outliers are discussed openly. Some of the post-hoc psychoacoustic interpretations in Section V are speculative, but they are flagged as such. The self-citation to [18] is appropriate because this is an incremental extension of their own published cognitive salience model.\n\nWho should read this: audio quality researchers, codec developers, and people in standardization. For a serious referee, my asks would be: release the DPW parameters and code or detailed pseudocode, add a stability analysis of the interaction selection, and report the number of signals and treatments behind Eq. 1 so the noise level is visible. I would send it to peer review with those revision requirements.","headline":"A solid PEAQ extension with genuine gains on unseen databases, but the interaction-selection procedure needs a stability check before I'd fully trust the mechanism.","tokens_in":23635,"tokens_out":2043,"would_cite":true,"duration_ms":21628,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adaptively weighting PEAQ's distortion metrics by cognitive effect measures—perceptual streaming, informational masking, and speech probability—generalizes better to unseen codecs than standard tools and…","keywords":["objective audio quality assessment","PEAQ","cognitive salience model","perceptual streaming","informational masking","speech-music classification","subjective quality prediction","parametric audio codecs"],"falsifier":"Re-run the interaction selection using a calibration set in which each content item has many more codec/bitrate conditions—say ten or more—so that Eq. (1)'s per-signal correlations are computed from more than three or four points; if the selected interactions change materially or the validation correlations on the seven databases drop below the reported average of 0.80, then the salience-by-correlation assumption is doing the work and the model as published is not stable.","tokens_in":22580,"feed_emoji":"🎧","tokens_out":8732,"duration_ms":70777,"temperature":0.7,"pith_summary":"Objective audio quality assessment tries to predict how listeners would rate a coded signal, but standard tools lose accuracy when they meet unseen codecs and signal types, especially parametric codecs that do not preserve the waveform. This paper argues that the missing ingredient is cognitive: the same physical distortion can be more or less noticeable depending on whether it forms a separate percept, is masked by signal complexity, or appears in speech versus music. To capture this, the authors extend the PEAQ standard's perceptual front end with a Cognitive Salience Model (CSM) that turns two existing cognitive measures—perceptual streaming and informational masking, plus a speech probability estimate—into adaptive weights on PEAQ's six distortion metrics. They train the weights on subjective scores, and report that the resulting PEAQ-CSM+ system reaches an average correlation near 0.80 with subjective scores across seven validation databases, outperforming other mapping stages and established metrics. The point is that thoughtful, perceptually motivated structure beats unrestricted machine learning when training data is small.","feed_headline":"Audio quality metric hits 0.80 correlation on unseen codec tests","feed_subtitle":"Weighting each distortion by perceived salience beats PEAQ, ViSQOL, PEMO-Q on seven unseen databases.","key_machinery":"The central object is the Cognitive Salience Model (CSM), a weighted-sum architecture that replaces PEAQ's neural-network mapping stage. Its load-bearing parts are the salience measure $S_m(j)$, computed as the Pearson correlation between one DM's basis-function output and the mean subjective scores across the treatments of signal $j$; the interaction metric $C_m$ that scores each candidate cognitive-effect/distortion pair; and the Detection Probability Weights (DPWs), sigmoid functions of the cognitive effect metrics that scale each DM's contribution. The CSM's defining move is that cognitive effects enter only as multipliers on distortion terms, never as direct predictors of the final score, which restricts the model's learnable interaction space and is the claimed source of its generalization.","core_discovery":"The central claim is that prediction generalization in objective audio quality assessment can be improved by an explicit model of how cognitive effects modulate distortion salience. The proposed architecture keeps PEAQ's psychoacoustic front end and its six model output values as distortion metrics (DMs), but replaces the fixed or generally-learned mapping from metrics to score with a Cognitive Salience Model: each DM is first mapped to a quality scale by a basis function trained on isolated-artifact listening tests, and then weighted by Detection Probability Weights derived from three cognitive effect metrics (EPN for perceptual streaming, PDEV for informational masking, and a speech-music probability). The interaction structure is chosen by a two-stage data-driven procedure—an interaction metric that measures how well each transformed cognitive effect predicts each DM's salience, followed by step-wise regression to select the final terms. Validated on seven unseen databases (MUSHRA, BS.1116, and blind source separation), the resulting PEAQ-CSM+ achieves an average Pearson correlation of about 0.80 and specifically improves prediction on parametrically coded, non-waveform-preserving audio, where the authors report other methods fail.","pith_inferences":["The interaction-selection step is the fragile point: replacing the calibration database (USAC VT1) with another multi-distortion database could change which CEM/DM pairs survive, so a leave-one-database-out rerun would isolate whether the architectural constraint (cognitive effects only as multipliers) or the specific interaction table drives the reported generalization.","The same salience-weighting scheme could transfer to non-intrusive, reference-free assessment and to multidimensional attributes such as timbre, loudness, or spatial quality, because the CSM separates distortion measurement from cognitive weighting; a P.800 speech-quality variant is one explicit direction the paper's future work points to.","Because the salience estimate needs several treatments per signal, the benefit of the method should grow as subjective databases include more codec/bitrate conditions per content item; databases with only one treatment per signal cannot support the interaction analysis at all.","The USAC VT2 result (correlation 0.63 versus 0.30 for standard PEAQ) shows parametric stereo distortions remain largely unexplained; adding the planned binaural model is the natural test of whether cognitive weighting extends to spatial quality or marks the method's scope boundary."],"forward_implications":["If the reported generalization holds, quality prediction for parametric and low-bitrate codecs improves enough that the method can be used in codec development and selection where standard PEAQ, ViSQOL, and PEMO-Q give weak correlations.","The separation of basis-function calibration (isolated artifacts) from interaction calibration (USAC VT1) means new distortion types or new cognitive effects can be added without retraining the entire mapping stage.","The explicit interaction terms align with psychoacoustic results, for example speech probability raising the salience of noise loudness and lowering the salience of linear distortions, so the model doubles as a quantitative statement about which distortion types matter in which listening contexts.","Because the reduced parameter count keeps inference time near 0.5 ms, the approach can be deployed in large-scale codec evaluation sweeps with negligible added cost."],"supporting_citations":[{"why":"Introduces the Cognitive Salience Model idea and the data-driven weighting of distortion metrics by cognitive effects; this paper extends and validates that architecture.","marker":"[18]"},{"why":"PEAQ ITU-R BS.1387-1 standard: supplies the perceptual model, the six distortion metrics (MOVs), and the baseline quality mapping this work replaces.","marker":"[4]"},{"why":"Barbedo and Lopes cognitive model: source of the perceptual-streaming (EPN) and informational-masking (PDEV) cognitive effect metrics and of the cognitive-model mapping approach.","marker":"[25]"},{"why":"Isolated-audio-artifacts listening-test database used to estimate each distortion metric's basis function to the MUSHRA quality scale.","marker":"[43]"},{"why":"USAC verification test 1 database used to select and optimize the cognitive-effect/distortion interactions and detection-probability weights.","marker":"[42]"},{"why":"ITU-R workplan for PEAQ revision: provides the ANN mapping configuration used as the BASELINE DM(ANN) system.","marker":"[47]"},{"why":"The authors' MATLAB implementation of PEAQ's advanced mode, used as the perceptual front end and as the basis for the complexity/performance validation and the PEAQ DI comparison.","marker":"[24]"},{"why":"ViSQOL Audio metric: state-of-the-art baseline the proposed model must beat on the validation databases.","marker":"[6]"},{"why":"PEMO-Q metric: state-of-the-art auditory-model baseline used in the comparison.","marker":"[7]"},{"why":"CombAQ binaural/timbre quality metric: strongest prior multi-dimensional baseline and the one comparison system with a binaural model relevant to the USAC VT2 limitation discussion.","marker":"[51]"}],"fun_headline_variants":["Audio metric beats PEAQ on blind tests","Cognitive weighting lifts audio quality prediction","Audio metric hits 0.80 on unseen codec tests","Salience-weighted metrics outdo classic PEAQ","Audio quality model generalizes to new distortions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's interaction selection rests on estimating how salient each distortion type is from the correlation between one distortion metric and the mean listener scores across only a few codec/bitrate versions of each signal; with so few data points per signal, that estimate is noisy, and if it misranks distortion salience the chosen cognitive weights—and with them the claimed generalization—lose their support.","fun_headline_variants_meta":{"raw":{"variants":["Audio metric beats PEAQ on blind tests","Cognitive weighting lifts audio quality prediction","Audio metric hits 0.80 on unseen codec tests","Salience-weighted metrics outdo classic PEAQ","Audio quality model generalizes to new distortions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1686,"prompt_tokens":1055,"completion_tokens":631,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":560}},"tokens_in":671,"tokens_out":631,"duration_ms":6349,"temperature":1.0,"reasoning_tokens":560,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:24:03.005131+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the interaction selection using a calibration set in which each content item has many more codec/bitrate conditions—say ten or more—so that Eq. (1)'s per-signal correlations are computed from more than three or four points; if the selected interactions change materially or the validation correlations on the seven databases drop below the reported average of 0.80, then the salience-by-correlation assumption is doing the work and the model as published is not stable.","supporting_citations":[{"cited_title":"A data-driven cognitive salience model for objective perceptual audio quality assessment,","cited_arxiv_id":null,"evidence_quote":"Introduces the Cognitive Salience Model idea and the data-driven weighting of distortion metrics by cognitive effects; this paper extends and validates that architecture."},{"cited_title":"BS.1387, Method for objective measurements of perceived audio quality, Geneva, Switzerland, 2001","cited_arxiv_id":null,"evidence_quote":"PEAQ ITU-R BS.1387-1 standard: supplies the perceptual model, the six distortion metrics (MOVs), and the baseline quality mapping this work replaces."},{"cited_title":"A new cognitive model for objective assessment of audio quality,","cited_arxiv_id":null,"evidence_quote":"Barbedo and Lopes cognitive model: source of the perceptual-streaming (EPN) and informational-masking (PDEV) cognitive effect metrics and of the cognitive-model mapping approach."},{"cited_title":"Generation and evaluation of isolated audio coding artifacts,","cited_arxiv_id":null,"evidence_quote":"Isolated-audio-artifacts listening-test database used to estimate each distortion metric's basis function to the MUSHRA quality scale."},{"cited_title":"USAC verification test report N12232,","cited_arxiv_id":null,"evidence_quote":"USAC verification test 1 database used to select and optimize the cognitive-effect/distortion interactions and detection-probability weights."},{"cited_title":"Workplan towards draft revision of recommendation ITU- R BS.1387-1,","cited_arxiv_id":null,"evidence_quote":"ITU-R workplan for PEAQ revision: provides the ANN mapping configuration used as the BASELINE DM(ANN) system."},{"cited_title":"Can we still use PEAQ? A performance analysis of the ITU standard for the objective assessment of perceived audio quality,","cited_arxiv_id":null,"evidence_quote":"The authors' MATLAB implementation of PEAQ's advanced mode, used as the perceptual front end and as the basis for the complexity/performance validation and the PEAQ DI comparison."},{"cited_title":"Objective assessment of perceptual audio quality using ViSQOLAudio,","cited_arxiv_id":null,"evidence_quote":"ViSQOL Audio metric: state-of-the-art baseline the proposed model must beat on the validation databases."},{"cited_title":"PEMO-Q—a new method for objective audio quality assessment using a model of auditory perception,","cited_arxiv_id":null,"evidence_quote":"PEMO-Q metric: state-of-the-art auditory-model baseline used in the comparison."},{"cited_title":"Subjective and objective as- sessment of monaural and binaural aspects of audio quality,","cited_arxiv_id":null,"evidence_quote":"CombAQ binaural/timbre quality metric: strongest prior multi-dimensional baseline and the one comparison system with a binaural model relevant to the USAC VT2 limitation discussion."}],"review_version":1}