REVIEW 4 major objections 5 minor 23 references
Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit Rates
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Parametric side information about just two dominant objects per frequency band lets a stereo downmix carry three or four audio objects at 24.4-32 kbit/s while preserving the spatial image.
desk verdict IVAS ParamISM is a real engineering validation with a coherent mechanism and an honest, if under-reported, listening test; the main caveat is that the 'faithfully reconstructing' claim overreaches for overlapping talkers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The two-dominant-object side-information model: for each of 11 psychoacoustic bands, the encoder ranks the objects by power, selects the two loudest, and transmits their indices with a 3-bit quantized power ratio $r_1 = P_1/(P_1 + P_2)$, while the azimuth and elevation of every object are sent once per frame. A stereo downmix is formed by cardioid gains derived from object azimuths, and the decoder reconstructs the loudspeaker signals with covariance synthesis, using Edge Fading Amplitude Panning gains for the two dominant objects and a prototype matrix determined by the target loudspeaker layout. The encoder and decoder use low-delay filterbanks shared with the rest of IVAS, which is what keeps the mode within the codec's 40 ms delay and complexity budgets.
What would settle it
Take a four-talker scene at 32 kbit/s, vary the fraction of frames in which more than two talkers have comparable energy within the same band, and run the same MUSHRA test; if scores fall to the anchor level at moderate overlap values rather than only at the 50 to 100 percent overlap the paper flags, the claim that ParamISM faithfully reconstructs realistic conferencing scenes is refuted.
Extended reading notes
Core claim
The central claim is that for scenes of three or four uncorrelated objects, the spatial information needed for faithful replay reduces to, per psychoacoustic band, the identities of the two loudest objects and the power ratio between them, plus per-object direction metadata, with the audio itself carried by a two-channel downmix. Because only two object indices and one ratio are sent per band, the side-information rate stays essentially constant as the object count grows, unlike earlier parametric object-coding schemes whose side information scales with the number of objects. The paper reports that in listening tests this configuration outperforms independently EVS-coded objects by 15 to 20 MUSHRA points at the same total bit rate, that subtracting metadata leaves ParamISM using roughly 61 to 63 percent of the EVS bit rate for the objects themselves, and that the mode also has lower WMOPS complexity.
Load-bearing premise
The whole scheme assumes that every time-frequency tile can be described by the two loudest objects, so the energy of all remaining objects can be dropped from the side information; the paper's own results show quality drops when heavy overlap makes that assumption false.
Editorial extensions
If this is right
- At the same total bit rate, ParamISM is reported to beat multi-mono EVS by 15 to 20 MUSHRA points and to support super-wideband audio where EVS at 8 kbit/s per object only supports wideband.
- After subtracting the 1.6 kbit/s per object spent on metadata, ParamISM uses about 61 to 63 percent of the bit rate that multi-mono EVS needs for the object signals themselves.
- The mode is a single coding-and-rendering chain, whereas the EVS comparison requires external rendering with unquantized metadata, so ParamISM removes a post-processing step.
- Measured WMOPS complexity for three and four objects is lower for ParamISM than for the corresponding set of EVS instances.
- Quality drops in scenes with 50 to 100 percent talker overlap and in choir content, which the paper attributes to the limitations of the two-dominant-object model rather than to typical conferencing scenarios.
Reading between the lines
- A natural extension is adaptive mode switching: detect sustained overlap in a band and hand those frames to a higher-rate non-parametric IVAS mode, reserving ParamISM for polite-conversation content where its model holds.
- Because the side-information cost depends only on the number of bands and not on the object count, the same scheme may scale beyond four objects with only per-object metadata growth, though the downmix and two-dominant assumptions would need new listening tests.
- The stereo downmix uses only azimuth for the cardioid gains, so elevation is carried entirely by metadata and panning; the scheme's robustness to large elevation spreads or moving objects is a question the current test set only partially covers.
- The metadata bit rate treated as part of IVAS's total means the real network-level rate advantage over EVS is larger than the raw 24.4 versus 32 kbit/s comparison suggests; billing or capacity planning should count ParamISM's object-coding rate after metadata.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper describes ParamISM, the parametric object coding mode of the 3GPP IVAS codec, which is designed to efficiently transmit three or four audio objects at 24.4 or 32 kbit/s. The encoder computes a stereo downmix from cardioid-based gains and transmits, per parameter band, the indices of the two most dominant objects and a quantized power ratio between them, together with once-per-frame quantized direction metadata. The decoder reconstructs the loudspeaker signals by covariance synthesis, using EFAP panning gains for the two dominant objects to build a target covariance matrix. The paper reports MUSHRA listening tests on twelve scenes, comparing ParamISM against multi-mono EVS at 8 and 9.6 kbit/s per object, and WMOPS complexity measurements. The central claim is that ParamISM offers a comparable immersive experience to independent EVS coding at lower total bit rate and with lower complexity, with quality described as good for speech and fair for music.
Significance. If the claims are validated, ParamISM is a noteworthy achievement: it brings parametric object coding to a standardized, low-delay mobile codec at bit rates compatible with current networks, and it does so with a complexity profile that is competitive with (or better than) running multiple mono EVS instances. The paper is valuable as a public algorithmic description of a 3GPP standard component, and the evaluation is appropriately anchored to an external codec (EVS) rather than to an internal variant of the proposed method, so the result is not circular. The algorithm contains no free parameters fitted to the test results; the design choices (number of bands, dominant-object count, quantization) are fixed by the standard. At the same time, the evidence as presented is not yet fully convincing: the listening-test analysis omits confidence intervals and significance testing, the complexity data contain a configuration that contradicts the unqualified 'lower complexity' claim, and the two-dominant-object assumption, which becomes a rank-2 covariance limit at the decoder, is not discussed as a structural limitation despite the observed quality drops in overlap-heavy scenes.
major comments (4)
- [IV.B, Fig. 2] The MUSHRA results are presented only as bar charts with no confidence intervals, raw scores, per-listener variance, or significance tests. With 13 listeners, the claimed 15–20 point advantage over multi-mono EVS at the same total bit rate (Section IV.B) cannot be assessed for statistical reliability. Please report per-condition mean scores with 95% confidence intervals (as recommended by ITU-R BS.1534-3) and results of a paired significance test, or at least clearly state that the differences are not statistically tested.
- [II.A / III.B, Eqs. (10)-(11)] The target covariance Cy in the decoder is constructed from the two dominant objects' EFAP panning vectors and a 2x2 diagonal power matrix, so Cy has rank at most 2 in every time-frequency tile. When three or four objects are simultaneously active in a parameter band, the ideal rendered covariance has higher rank, and the synthesis cannot reproduce the true spatial image; energy from non-dominant objects is necessarily reassigned to the two dominant directions. The paper observes quality drops in scenes i1 and i5 (50–100% temporal overlap) and i11 (choir, full spectral overlap) but does not connect these to this structural limitation. The abstract's phrase 'faithfully reconstructing the spatial image of the original audio scene' and 'arbitrarily placed objects' overstates the capability for exactly the multi-party overlap conditions that the use case implies. Please either qualify the claim or provide objective/perceptual evidence that the rank-2 approximation is adequate for realistic conferencing scenes.
- [IV.C, Table II] The complexity numbers in Table II do not support the unqualified 'lower complexity' claim in the abstract. For three objects, ParamISM at 24.4 kbit/s has a total of 241.14 WMOPS, which is higher than 3x EVS at 8 kbit/s (237.57 WMOPS). The lower-complexity statement is true only when comparing against EVS at 9.6 kbit/s per object (313.71 WMOPS) or for the four-object case. Please specify the reference EVS bit rate in the abstract and conclusions, or qualify the claim accordingly.
- [IV.A, Table I] The test set is strongly biased toward speech scenes with limited overlap; only scenes i1 and i5 contain 50–100% temporal overlap and only i11 has full spectral overlap. Given the abstract's claim about 'arbitrarily placed objects' and 'multiple audio objects' without content-type restriction, the MUSHRA evidence does not support generalization to music and mixed content with substantial object overlap. Please narrow the claim to the tested conditions or extend the test set.
minor comments (5)
- [Title] The title and running header render 'IVAS' as 'IV AS' with a space; the codec name should be written consistently as 'IVAS'.
- [III.B] The phrase 'mixing mixing matrix' should be 'mixing matrix'.
- [IV.B] The sentence 'In the three-object test, the fourth objects of the scenes are not used' should read 'the fourth object'.
- [III.B, Eq. (11)] The definition of PDM X(k) uses Xi(k,n) to denote the downmix channels, but earlier Xi denotes the input objects; please clarify the notation.
- [II.A] Equation (3) maps r1 in [0.5,1] to indices 0..7; the decoder reconstruction in Eq. (7) is correct, but it may be helpful to state explicitly that r1 cannot be below 0.5 because P1 is the larger power.
Circularity Check
No significant circularity; the derivation chain is self-contained and the performance claim rests on an external listening test.
full rationale
The paper's derivation chain is self-contained and externally grounded. The parametric side information is computed directly from the input objects (Eqs. 1-3): the dominant-object indices and the power ratio r1 = P1/(P1+P2) are measured quantities, not fitted values. The decoder's rendering (Eqs. 8-11) is a standard covariance-synthesis step, citing the independent Vilkamo et al. framework, with the target covariance built from transmitted ratios and the decoded downmix power. No equation reduces to a fitted parameter renamed as a prediction, and no uniqueness theorem or ansatz is smuggled in via self-citation. The central performance claim - comparable immersion at lower bit rate and complexity - is established by an external MUSHRA comparison against multi-mono EVS and by ITU-T WMOPS measurements, with EVS given the advantage of free (unquantized) metadata while IVAS's metadata bit rate is counted in its total. The two-dominant-object assumption (Sec. II.A) is a disclosed design model, and the paper explicitly reports quality drops on heavy-overlap scenes (i1, i5, i11) and choir content (i11), which is a limitation statement rather than a circular step. The self-citations [2], [3] (the IVAS standard specification and the authors' own overview) point to the algorithmic specification inside the standard, but they are not load-bearing for the performance claim, which rests on the independent listening test and complexity measurements. The rank-2 covariance limitation of the rendering (a consequence of the two-dominant model) is a technical correctness concern about the 'faithfully reconstructing' phrasing, not a circularity in the derivation.
Assumptions & free parameters
free parameters (4)
- Dominant object count per band =
2
- Number of parameter bands and borders =
11 bands, borders B(l) not listed
- Side-information bit allocation =
3 bits per power ratio, 2 bits per object index, 7+6 bits per direction per frame
- Smoothing and energy compensation constants =
not specified
assumptions (5)
- domain assumption Two dominant objects per frequency band are sufficient to represent the spatial image.
- domain assumption Input objects are mutually uncorrelated, making the downmix covariance matrix diagonal.
- domain assumption MUSHRA with 13 expert listeners on a 7.1.4 layout generalizes to the intended use cases.
- standard math The IVAS core coder and filterbanks from TS 26.253 behave as specified and add no unmodeled error.
- domain assumption EFAP panning produces correct target covariance for arbitrary object directions on the output layout.
Cite this review
Pith. "Pith review of Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit Rates." pith.science (2026). https://pith.science/paper/UT5VFQJW
@misc{pith2026250705409,
author = {Pith},
title = {Pith review of: Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit Rates},
year = {2026},
howpublished = {\url{https://pith.science/paper/UT5VFQJW}},
note = {Machine review of arXiv:2507.05409}
}
read the original abstract
The recently standardized 3GPP codec for Immersive Voice and Audio Services (IVAS) includes a parametric mode for efficiently coding multiple audio objects at low bit rates. In this mode, parametric side information is obtained from both the object metadata and the input audio objects. The side information comprises directional information, indices of two dominant objects, and the power ratio between these two dominant objects. It is transmitted to the decoder along with a stereo downmix. In IVAS, parametric object coding allows for transmitting three or four arbitrarily placed objects at bit rates of 24.4 or 32 kbit/s and faithfully reconstructing the spatial image of the original audio scene. Subjective listening tests confirm that IVAS provides a comparable immersive experience at lower bit rate and complexity compared to coding the audio objects independently using Enhanced Voice Services (EVS).
Figures
Reference graph
Works this paper leans on
-
[1]
Codec for Immersive V oice and Audio Services (IV AS); General overview,
“Codec for Immersive V oice and Audio Services (IV AS); General overview,” 3GPP, TS 26.250, July 2024. [Online]. Available: https://www.3gpp.org/DynaReport/26250.htm
work page 2024
-
[2]
“Codec for Immersive V oice and Audio Services (IV AS); Detailed Algorithmic Description including RTP payload format and SDP parameter definitions,” 3GPP, TS 26.253, July 2024. [Online]. Available: https://www.3gpp.org/DynaReport/26253.htm
work page 2024
-
[3]
M. Multrus et al., “Immersive V oice and Audio Services (IV AS) Codec – The New 3GPP Standard for Immersive Communication,” in AES Convention 157, no. 10188, New York, NY , USA, October 2024
work page 2024
-
[4]
Codec for Enhanced V oice Services (EVS); Detailed algorithmic description,
“Codec for Enhanced V oice Services (EVS); Detailed algorithmic description,” 3GPP, TS 26.445, May 2024. [Online]. Available: https://www.3gpp.org/DynaReport/26445.htm
work page 2024
-
[5]
Overview of the EVS codec architecture,
M. Dietz et al., “Overview of the EVS codec architecture,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia, April 2015, pp. 5698–5702
work page 2015
-
[6]
E. Fotopoulou et al., “Use-Cases of the new 3GPP Immersive V oice and Audio Services (IV AS) Codec and a Web Demo Implementation,” in IEEE 5th International Symposium on the Internet of Sounds (IS2) , Erlangen, Germany, September 2024
work page 2024
-
[7]
Binaural Cue Coding: A Novel and Efficient Representation of Spatial Audio,
C. Faller and F. Baumgarte, “Binaural Cue Coding: A Novel and Efficient Representation of Spatial Audio,” in IEEE International Con- ference on Acoustics, Speech, and Signal Processing (ICASSP) , Orlando, FL, USA, May 2002, pp. II–1841–II–1844
work page 2002
-
[8]
Binaural Cue Coding - Part II: Schemes and Applications,
——, “Binaural Cue Coding - Part II: Schemes and Applications,” IEEE Transactions on Speech and Audio Processing , vol. 11, no. 6, pp. 520– 531, January 2004
work page 2004
Show all 23 references
-
[9]
Binaural Cue Coding - Part I: Psychoacous- tic Fundamentals and Design Principles,
F. Baumgarte and C. Faller, “Binaural Cue Coding - Part I: Psychoacous- tic Fundamentals and Design Principles,” IEEE Transactions on Speech and Audio Processing , vol. 11, no. 6, pp. 509–519, January 2004
2004
-
[10]
Spatial Audio Object Coding (SAOC) - The Upcoming MPEG Standard on Parametric Object Based Audio Coding,
J. Breebaart et al., “Spatial Audio Object Coding (SAOC) - The Upcoming MPEG Standard on Parametric Object Based Audio Coding,” in AES Convention 124 , no. 7377, Amsterdam, The Netherlands, May 2008
2008
-
[11]
MPEG Spatial Audio Object Coding — The ISO/MPEG Standard for Efficient Coding of Interactive Audio Scenes,
J. Herre et al., “MPEG Spatial Audio Object Coding — The ISO/MPEG Standard for Efficient Coding of Interactive Audio Scenes,” Journal of the Audio Engineering Society , vol. 60, pp. 655–673, October 2012
2012
-
[12]
ISO/MPEG-H 3D Audio: SAOC-3D Decoding and Rendering,
A. Murtaza et al., “ISO/MPEG-H 3D Audio: SAOC-3D Decoding and Rendering,” in AES Convention 139 , no. 9434, New York, NY , USA, October 2015
2015
-
[13]
Immersive Audio Delivery Using Joint Object Coding,
H. Purnhagen et al., “Immersive Audio Delivery Using Joint Object Coding,” in AES Convention 140 , no. 9587, Paris, France, May 2016
2016
-
[14]
AC-4 — The Next Generation Audio Codec,
K. Kj ¨orling et al., “AC-4 — The Next Generation Audio Codec,” in AES Convention 140, no. 9491, Paris, France, May 2016
2016
-
[15]
Efficient Parametric Audio Coding for Interactive Rendering: The Upcoming ISO/MPEG Standard on Spatial Audio Object Coding (SAOC),
L. Terentiev et al., “Efficient Parametric Audio Coding for Interactive Rendering: The Upcoming ISO/MPEG Standard on Spatial Audio Object Coding (SAOC),” in International Conference on Acoustics . Rotterdam, The Netherlands: DAGA, March 2009
2009
-
[16]
Software tools for speech and audio coding standardization,
“Software tools for speech and audio coding standardization,” ITU-T, Tech. Rep. G.191, May 2024. [Online]. Available: https://handle.itu.int/11.1002/1000/15940
2024
-
[17]
Directional audio coding – perception-based reproduc- tion of spatial sound,
V . Pulkki et al., “Directional audio coding – perception-based reproduc- tion of spatial sound,” in International Workshop on the Principles and Applications of Spatial Hearing , Zao, Miyagi, Japan, November 2009
2009
-
[18]
Low Delay Filterbanks for Enhanced Low Delay Audio Coding,
M. Schnell et al., “Low Delay Filterbanks for Enhanced Low Delay Audio Coding,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , New Paltz, NY , USA, October 2007, pp. 235– 238
2007
-
[19]
Information technology — Dynamic adaptive streaming over HTTP (DASH) — Part 3: Media presentation description and segment format,
“Information technology — Dynamic adaptive streaming over HTTP (DASH) — Part 3: Media presentation description and segment format,” ISO/IEC, Tech. Rep. 23091-3, 2018
2018
-
[20]
Optimized Covariance Domain Framework for Time-Frequency Processing of Spatial Audio,
J. Vilkamo, T. B ¨ackstr¨om, and A. Kuntz, “Optimized Covariance Domain Framework for Time-Frequency Processing of Spatial Audio,” Journal of the Audio Engineering Society , vol. 61, no. 6, pp. 403–411, July 2013
2013
-
[21]
A Polygon-Based Panning Method for 3D Loudspeaker Setups,
C. Borß, “A Polygon-Based Panning Method for 3D Loudspeaker Setups,” in AES Convention 137 , no. 9106, Los Angeles, CA, USA, October 2014
2014
-
[22]
Virtual Sound Source Positioning Using Vector Base Ampli- tude Panning,
V . Pulkki, “Virtual Sound Source Positioning Using Vector Base Ampli- tude Panning,” Journal of the Audio Engineering Society , vol. 45, no. 6, pp. 456–466, June 1997
1997
-
[23]
Method for the subjective assessment of intermediate quality level of audio systems,
“Method for the subjective assessment of intermediate quality level of audio systems,” Recommendation BS.1534-3, 2015
2015
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.