Pith. sign in

REVIEW 3 major objections 64 references

Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats

T0 review · 3 major / 0 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read A parametric method transcodes spatial audio from any capture format to any playback format by estimating source and ambience metadata that fits observed covariances.

desk verdict The paper gives a single parametric pipeline that fits variable primaries plus ambience angular power to capture covariances then builds target covariances for any playback format plus independent rotations, but the abstract supplies no numbers or derivations to check whether the model actually extrapolates reliably. read the letter →

arxiv 2606.18480 v1 pith:IYAXBJNG submitted 2026-06-16 eess.AS cs.SD

classification eess.AScs.SD
keywords spatialaudiotranscodingparametricrenderingAmbisonicsmicrophonearrayscovariancemixingmatricesreproduction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces a framework that processes spatial sound captured as Ambisonic signals or raw microphone array signals. It extracts time-frequency metadata describing a variable number of primary sources along with an ambience component whose angular power distribution is fitted to match the captured signals' spatial covariances. This metadata is then used to build the spatial covariances needed for any chosen target playback format and to compute the optimal mixing matrices that convert the scene for reproduction. The method also accommodates separate rotations of the capture and playback arrangements. Listening tests on simulated scenes from different array types indicate perceptual gains over prior parametric approaches, especially when using lower-order or constrained arrays.

What carries the argument

Time-frequency-dependent spatial metadata consisting of primary source components and an ambience angular power distribution whose parameters are chosen to match observed spatial covariances of the input signals.

What would settle it

A controlled listening test in which the proposed transcoder produces audible spatial distortions or loses detail relative to direct non-parametric methods when the input contains many simultaneous overlapping sources that cannot be well approximated by the primary-plus-ambience model.

Watch

Extended reading notes

Core claim

The central claim is that a single analysis stage estimating time-frequency-dependent spatial metadata for a variable number of primary source components plus an ambience component with its own fitted angular power distribution can characterise the captured scene sufficiently to construct the spatial covariances of arbitrary target playback formats and thereby derive optimal mixing matrices for transcoding, while also handling independent rotations of capture and playback setups.

Load-bearing premise

A scene model built from a variable number of primary sources plus one ambience component whose angular power distribution is fitted to the observed covariances is accurate enough to support high-quality transcoding to any playback format.

Editorial extensions

If this is right

  • The same metadata can be reused to derive mixing matrices for any combination of capture and playback formats without re-analysis.
  • Independent rotation of capture and playback coordinate systems is supported without additional processing stages.
  • Perceptual quality remains high even when the capture array is limited to low-order Ambisonics or geometrically constrained microphone placements.
  • Real-time implementations can be compared directly against existing parametric renderers on the same simulated scenes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The covariance-fitting step could be made fully causal to support live capture and rendering of moving sources.
  • The framework may allow consumer devices with only a few microphones to deliver spatial audio for arbitrary loudspeaker or headphone layouts.
  • Similar covariance-matching logic might be applied to other array-processing tasks such as source separation or noise reduction in spatial scenes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper introduces a unified parametric framework for transcoding spatial audio scenes captured as Ambisonic signals or raw microphone array signals to arbitrary playback formats. It estimates time-frequency-dependent spatial metadata consisting of a variable number of primary source components plus an ambience component whose angular power distribution parameters are fitted to the observed capture covariances; this metadata is then used to construct target-format covariances from which optimal mixing matrices are derived. The method also supports independent rotations of capture and playback setups. Real-time implementations are compared against existing state-of-the-art parametric renderers via listening tests on simulated scenes from Ambisonic, spherical, and head-worn arrays, with claims of perceptual benefits especially for lower-order and geometrically constrained arrays.

Significance. If the parametric model proves sufficient for accurate covariance extrapolation, the framework would offer a general solution for arbitrary capture-to-playback transcoding that improves upon existing methods for constrained arrays and diverse content. The inclusion of listening tests across multiple array types provides direct perceptual evidence, which is a positive aspect of the evaluation design.

major comments (3)
  1. [Abstract] Abstract: The description of the covariance-fitting procedure states that parameters are fitted to observed capture covariances and then used to construct target covariances for deriving mixing matrices, but provides no equations or derivation showing that the resulting matrices are independent of the fit rather than reducing to it by construction; this is load-bearing for the claim of general transcoding to arbitrary formats.
  2. [Abstract] Abstract: The listening-test comparison claims perceptual benefits over state-of-the-art renderers, yet the abstract (and available description) supplies no quantitative results such as mean scores, confidence intervals, or statistical tests; without these, the magnitude and reliability of the claimed improvements cannot be assessed.
  3. [Abstract] Abstract: The central assumption that a model of variable primary sources plus an ambience component with fitted angular power distribution is expressive enough to characterise scenes for covariance extrapolation to arbitrary targets is not supported by any analysis of model mismatch (e.g., correlated primaries or non-diffuse ambience); this directly affects the validity of the transcoding claim for real scenes.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for the constructive feedback. We address each major comment below, clarifying the manuscript content and indicating planned revisions to improve clarity and completeness.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The description of the covariance-fitting procedure states that parameters are fitted to observed capture covariances and then used to construct target covariances for deriving mixing matrices, but provides no equations or derivation showing that the resulting matrices are independent of the fit rather than reducing to it by construction; this is load-bearing for the claim of general transcoding to arbitrary formats.

    Authors: The abstract summarises the approach at a high level. The full manuscript (Section 3) derives the mixing matrices via an optimisation that minimises the Frobenius distance between the parametrically constructed target-format covariances and the rendered covariances; the target covariances are formed from the fitted primary-source directions/intensities and the fitted angular power distribution of the ambience, which are distinct from the observed capture covariances. This construction is what enables transcoding to arbitrary formats. We will revise the abstract to include a concise clause referencing this independence. revision: yes

  2. Referee: [Abstract] Abstract: The listening-test comparison claims perceptual benefits over state-of-the-art renderers, yet the abstract (and available description) supplies no quantitative results such as mean scores, confidence intervals, or statistical tests; without these, the magnitude and reliability of the claimed improvements cannot be assessed.

    Authors: We agree that quantitative indicators would strengthen the abstract. Detailed results (mean scores, confidence intervals, and statistical tests) appear in Section 5. We will add a short quantitative summary of the key perceptual improvements to the abstract within length constraints. revision: yes

  3. Referee: [Abstract] Abstract: The central assumption that a model of variable primary sources plus an ambience component with fitted angular power distribution is expressive enough to characterise scenes for covariance extrapolation to arbitrary targets is not supported by any analysis of model mismatch (e.g., correlated primaries or non-diffuse ambience); this directly affects the validity of the transcoding claim for real scenes.

    Authors: The listening-test scenes were generated with a range of source counts, correlations, and reverberation conditions to probe the model. While an explicit mismatch analysis is absent from the current version, the perceptual outcomes provide supporting evidence. We will add a dedicated paragraph in the discussion section addressing model assumptions and potential mismatch cases. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: parametric model fit to capture data used for independent target-format extrapolation

full rationale

The paper describes a standard parametric approach: fit time-frequency metadata (variable primaries + ambience angular power distribution) to observed capture covariances, then use the fitted parameters to synthesize covariances for arbitrary target playback formats and derive mixing matrices. This is model-based extrapolation between distinct capture and playback geometries, not a self-definitional loop or a fitted input renamed as prediction. No equations reduce the target construction to the capture fit by algebraic identity, no self-citation chain is load-bearing for the central claim, and the derivation remains self-contained against external array geometries. The reader's concern about model sufficiency is a question of representational power, not circularity.

Assumptions & free parameters 2 free parameters · 1 assumptions · 0 invented entities

Only the abstract is available, so the ledger is limited to elements explicitly stated or implied therein.

free parameters (2)
  • number of primary sources
    Described as variable and estimated per time-frequency tile; no detail on how the number is chosen or regularised.
  • ambience angular power distribution parameters
    Fitted to match observed covariances; exact functional form and number of parameters not stated.
assumptions (1)
  • domain assumption Observed spatial covariances of captured signals can be accurately represented by a sum of primary source components plus an ambience component with its own angular power distribution
    The entire metadata estimation and subsequent transcoding step rests on this model fitting the data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats." pith.science (2026). https://pith.science/paper/IYAXBJNG

@misc{pith2026260618480,
  author       = {Pith},
  title        = {Pith review of: Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IYAXBJNG}},
  note         = {Machine review of arXiv:2606.18480}
}
read the original abstract

This article introduces a unified framework for the parametric analysis and reproduction of spatial sound scenes captured either as Ambisonic signals or as raw microphone array signals. The proposed method estimates time-frequency-dependent spatial metadata that characterises a variable number of primary source components and an ambience component with its own angular power distribution, whose parameters fit the observed spatial covariances of the captured signals. This metadata is used to construct spatial covariances of the target playback formats, which are then used to derive optimal mixing matrices for transcoding the scene for playback over the target reproduction system. The method additionally handles independent rotations of both capture and playback setups. Real-time implementations of the method and other existing state-of-the-art parametric renderers are compared in a listening test using simulated scenes from Ambisonic, spherical, and head-worn arrays. The results highlight perceptual benefits of the proposed framework across a diverse range of content and receiver configurations, particularly for lower-order and geometrically constrained microphone arrays.

Figures

Figures reproduced from arXiv: 2606.18480 by the authors.

Figure 1
Figure 1. The sound field model. The blue line indicates the ambience power distribution, while the orange and purple lines [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Block diagram of the proposed framework. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Medians and 95% confidence intervals for listening test parts 1, 2 and 3. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 3 canonical work pages

  1. [1]

    Sound externalization: A review of recent research,

    V . Best, R. Baumgartner, M. Lavandier, P. Majdak, and N. Kop ˇco, “Sound externalization: A review of recent research,”Trends in hearing, vol. 24, p. 2331216520948390, 2020

  2. [2]

    Multichannel 3D microphone arrays: A review,

    H. Lee, “Multichannel 3D microphone arrays: A review,”J. Audio Eng. Soc., vol. 69, no. 1/2, pp. 5–26, 2021

  3. [3]

    Microphone array beam forming for multichannel record- ing,

    J. Backman, “Microphone array beam forming for multichannel record- ing,” inAES Convention 114, 2003

  4. [4]

    Portable spherical microphone for Super Hi-Vision 22.2 multichannel audio,

    K. Ono, T. Nishiguchi, K. Matsui, and K. Hamasaki, “Portable spherical microphone for Super Hi-Vision 22.2 multichannel audio,” inAES Convention 135, 2013

  5. [5]

    External ear transfer function modeling: A beamforming approach,

    J. Chen, B. D. Van Veen, and K. E. Hecox, “External ear transfer function modeling: A beamforming approach,”J. Acoust. Soc. Am., vol. 92, no. 4, pp. 1933–1944, 1992

  6. [6]

    Headphone-based reproduction of 3D audi- tory scenes captured by spherical/hemispherical microphone arrays,

    Z. Li and R. Duraiswami, “Headphone-based reproduction of 3D audi- tory scenes captured by spherical/hemispherical microphone arrays,” in IEEE ICASSP, 2006

  7. [7]

    Periphony: With-height sound reproduction,

    M. A. Gerzon, “Periphony: With-height sound reproduction,”J. Audio Eng. Soc., vol. 21, no. 1, pp. 2–10, 1973

  8. [8]

    Binaural rendering of Ambisonic signals via magnitude least squares,

    C. Sch ¨orkhuber, M. Zaunschirm, and R. H ¨oldrich, “Binaural rendering of Ambisonic signals via magnitude least squares,” inDAGA, 2018

Show all 64 references
  1. [9]

    All-round ambisonic panning and decoding,

    F. Zotter and M. Frank, “All-round ambisonic panning and decoding,” J. Audio Eng. Soc., vol. 60, no. 10, pp. 807–820, 2012

  2. [10]

    Rafaely,Fundamentals of spherical array processing

    B. Rafaely,Fundamentals of spherical array processing. Springer, 2015, vol. 8

  3. [11]

    Investigation on localisation accuracy for first and higher order ambisonics reproduced sound sources,

    S. Bertet, J. Daniel, E. Parizet, and O. Warusfel, “Investigation on localisation accuracy for first and higher order ambisonics reproduced sound sources,”Acta Acustica united with Acustica, vol. 99, no. 4, pp. 642–657, 2013

  4. [12]

    Spatial perception of sound fields recorded by spherical microphone arrays with varying spatial resolution,

    A. Avni, J. Ahrens, M. Geier, S. Spors, H. Wierstorf, and B. Rafaely, “Spatial perception of sound fields recorded by spherical microphone arrays with varying spatial resolution,”J. Acoust. Soc. Am., vol. 133, no. 5, pp. 2711–2721, 2013

  5. [13]

    Investigating sound-field reproduction methods as perceived by bilateral hearing aid users and normal-hearing listeners,

    J. Fernandez, L. McCormack, P. Hyv ¨arinen, and A. A. Kressner, “Investigating sound-field reproduction methods as perceived by bilateral hearing aid users and normal-hearing listeners,”J. Acoust. Soc. Am., vol. 155, no. 2, pp. 1492–1502, 2024

  6. [14]

    Comparing ambisonic microphones–part 1,

    E. Bates, M. Gorzel, L. Ferguson, H. O’Dwyer, and F. M. Boland, “Comparing ambisonic microphones–part 1,” inAES Int. Conf. on Sound Field Control, 2016

  7. [15]

    3D sound field recording with higher order ambisonics–objective measurements and validation of a 4th order spherical microphone,

    S. Moreau, J. Daniel, and S. Bertet, “3D sound field recording with higher order ambisonics–objective measurements and validation of a 4th order spherical microphone,” inAES Convention 120, 2006

  8. [16]

    First- order directional audio coding (DirAC),

    V . Pulkki, A. Politis, M.-V . Laitinen, J. Vilkamo, and J. Ahonen, “First- order directional audio coding (DirAC),” inParametric Time-Frequency Domain Spatial Audio, V . Pulkki, S. Delikaris-Manias, and A. Politis, Eds. John Wiley & Sons, 2017, pp. 89–138. PREPRINT. UNDER REV...

  9. [17]

    High angular resolution planewave expansion,

    S. Berge and N. Barrett, “High angular resolution planewave expansion,” in2nd Int. Symp. on Ambisonics and Spherical Acoustics, 2010

  10. [18]

    COMPASS: Coding and multidirec- tional parameterization of ambisonic sound scenes,

    A. Politis, S. Tervo, and V . Pulkki, “COMPASS: Coding and multidirec- tional parameterization of ambisonic sound scenes,” inIEEE ICASSP, 2018

  11. [19]

    Parametric first-order am- bisonic decoding for headphones utilising the cross-pattern coherence algorithm,

    L. McCormack and S. Delikaris-Manias, “Parametric first-order am- bisonic decoding for headphones utilising the cross-pattern coherence algorithm,” inEAA Spatial Audio Sig. Proc. Symp., 2019

  12. [20]

    Linearly and quadratically constrained least-squares decoder for signal-dependent binaural rendering of am- bisonic signals,

    C. Sch ¨orkhuber and R. H¨oldrich, “Linearly and quadratically constrained least-squares decoder for signal-dependent binaural rendering of am- bisonic signals,” inAES Int. Conf. on Immersive and Interactive Audio, 2019

  13. [21]

    Parametric spatial audio effects based on the multi-directional decomposition of ambisonic sound scenes,

    L. McCormack, A. Politis, and V . Pulkki, “Parametric spatial audio effects based on the multi-directional decomposition of ambisonic sound scenes,” in24th Int. Conf. on Digital Audio Effects (DAFx20in21), 2021

  14. [22]

    A spatial enhancement approach for binaural rendering of head-worn microphone arrays,

    J. Fernandez, L. McCormack, P. Hyv ¨arinen, A. Politis, and V . Pulkki, “A spatial enhancement approach for binaural rendering of head-worn microphone arrays,” inInt. Congress on Acoustics (ICA), 2022

  15. [23]

    Upscaling ambisonic sound scenes using compressed sensing techniques,

    A. Wabnitz, N. Epain, A. McEwan, and C. Jin, “Upscaling ambisonic sound scenes using compressed sensing techniques,” inIEEE Work. on Applications of Sig. Proc. to Audio and Acoustics (WASPAA), 2011

  16. [24]

    Mixed source sound field translation for virtual binaural application with perceptual validation,

    L. Birnie, T. Abhayapala, V . Tourbabin, and P. Samarasinghe, “Mixed source sound field translation for virtual binaural application with perceptual validation,”IEEE/ACM Trans. Audio, Speech, and Lang. Proc., vol. 29, pp. 1188–1203, 2021

  17. [25]

    Deep-sound field analysis for upscaling ambisonic signals,

    G. Routray, S. Basu, P. Baldev, and R. M. Hegde, “Deep-sound field analysis for upscaling ambisonic signals,” inEAA Spatial Audio Sig. Proc. Symp., 2019

  18. [26]

    UP-WGAN: Upscaling ambisonic sound scenes using wasserstein generative adversarial networks,

    Y . Wang, X. Wu, and T. Qu, “UP-WGAN: Upscaling ambisonic sound scenes using wasserstein generative adversarial networks,” inAES Con- vention 152, 2022

  19. [27]

    Ambisonics super-resolution using a waveform-domain neural network,

    I. Nawfal, S. D. Manias, M. Souden, J. Merimaa, J. Atkins, E. McMullin, S. Pirhosseinloo, and D. Phillips, “Ambisonics super-resolution using a waveform-domain neural network,” inAES Int. Conf. on Audio for Virtual and Augmented Reality, 2024

  20. [28]

    Para- metric ambisonic encoding of arbitrary microphone arrays,

    L. McCormack, A. Politis, R. Gonzalez, T. Lokki, and V . Pulkki, “Para- metric ambisonic encoding of arbitrary microphone arrays,”IEEE/ACM Trans. Audio, Speech, and Lang. Proc., vol. 30, pp. 2062–2075, 2022

  21. [29]

    Ambisonics capture using microphones on head-worn device of arbitrary geometry,

    A. Bastine, L. Birnie, T. D. Abhayapala, P. Samarasinghe, and V . Tourbabin, “Ambisonics capture using microphones on head-worn device of arbitrary geometry,” in30th European Sig. Proc. Conf. (EUSIPCO), 2022, pp. 309–313

  22. [30]

    Neural ambisonics encoding for compact irregular microphone arrays,

    M. Heikkinen, A. Politis, and T. Virtanen, “Neural ambisonics encoding for compact irregular microphone arrays,” inIEEE ICASSP, 2024

  23. [31]

    End-to-end magnitude least squares binaural rendering of spherical microphone array signals,

    T. Deppisch, H. Helmholz, and J. Ahrens, “End-to-end magnitude least squares binaural rendering of spherical microphone array signals,” in Immersive and 3D Audio Conf. (I3DA), 2021

  24. [32]

    Beamforming- based binaural reproduction by matching of binaural signals,

    L. Madmoni, J. Donley, V . Tourbabin, and B. Rafaely, “Beamforming- based binaural reproduction by matching of binaural signals,” inAES Int. Conf. on Audio for Virtual and Augmented Reality, 2020

  25. [33]

    Feasibility of iMagLS-BSM-ILD informed binaural signal matching with arbitrary microphone arrays,

    O. Berebi, Z. Ben-Hur, D. L. Alon, and B. Rafaely, “Feasibility of iMagLS-BSM-ILD informed binaural signal matching with arbitrary microphone arrays,” in18th Int. Work. on Acoustic Sig. Enhancement (IWAENC), 2024

  26. [34]

    Enhancing binaural rendering of head-worn microphone arrays through the use of adaptive spatial covariance matching,

    J. Fernandez, L. McCormack, P. Hyv ¨arinen, A. Politis, and V . Pulkki, “Enhancing binaural rendering of head-worn microphone arrays through the use of adaptive spatial covariance matching,”J. Acoust. Soc. Am., vol. 151, no. 4, pp. 2624–2635, 2022

  27. [35]

    Six-degrees-of-freedom binaural reproduction of head-worn microphone array capture,

    L. McCormack, N. Meyer-Kahlen, D. L. Alon, Z. Ben-Hur, S. V . A. Gari, and P. Robinson, “Six-degrees-of-freedom binaural reproduction of head-worn microphone array capture,”J. Audio Eng. Soc., vol. 71, no. 10, pp. 638–649, 2023

  28. [36]

    Perceptual comparison of dynamic binaural reproduction methods for sparse head-mounted microphone arrays,

    S. Benjamin and R. Stefan, “Perceptual comparison of dynamic binaural reproduction methods for sparse head-mounted microphone arrays,”J. Audio Eng. Soc., vol. 72, pp. 442–454, july 2024

  29. [37]

    Performance and robustness of signal-dependent vs. signal-independent binaural signal matching with wearable microphone arrays,

    A. Berger, M. Yifrach, O. Berebi, V . Tourbabin, J. Donley, Z. Ben-Hur, and B. Rafaely, “Performance and robustness of signal-dependent vs. signal-independent binaural signal matching with wearable microphone arrays,”J. Audio, Speech, and Mus. Proc., 2026

  30. [38]

    Enhancement of ambisonic binaural reproduction using directional audio coding with optimal adap- tive mixing,

    A. Politis, L. McCormack, and V . Pulkki, “Enhancement of ambisonic binaural reproduction using directional audio coding with optimal adap- tive mixing,” inIEEE Work. on Applications of Sig. Proc. to Audio and Acoustics (WASPAA), 2017

  31. [39]

    SPARTA & COMPASS: Real-time implementations of linear and parametric spatial audio reproduction and processing methods,

    L. McCormack and A. Politis, “SPARTA & COMPASS: Real-time implementations of linear and parametric spatial audio reproduction and processing methods,” inAES Int. Conf. on Immersive and Interactive Audio, 2019

  32. [40]

    Computing Fourier transforms and convolutions on the 2-sphere,

    J. R. Driscoll and D. M. Healy, “Computing Fourier transforms and convolutions on the 2-sphere,”Advances Appl. Math., vol. 15, no. 2, pp. 202–250, 1994

  33. [41]

    Rotation matrices for real spherical harmonics. Direct determination by recursion,

    J. Ivanic and K. Ruedenberg, “Rotation matrices for real spherical harmonics. Direct determination by recursion,”J. Phys. Chem. A, vol. 102, no. 45, pp. 9099–9100, 1998

  34. [42]

    Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and ambisonics,

    A. Politis, “Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and ambisonics,”arXiv preprint arXiv:2407.06847, 2024

  35. [43]

    Parametric spatial audio processing utilising com- pact microphone arrays,

    S. Delikaris-Manias, “Parametric spatial audio processing utilising com- pact microphone arrays,” Ph.D. dissertation, Aalto University, 2017

  36. [44]

    Easycom: An augmented reality dataset to support algorithms for easy communication in noisy environments,

    J. Donley, V . Tourbabin, J.-S. Lee, M. Broyles, H. Jiang, J. Shen, M. Pantic, V . K. Ithapu, and R. Mehra, “Easycom: An augmented reality dataset to support algorithms for easy communication in noisy environments,”arXiv preprint arXiv:2107.04174, 2021

  37. [45]

    Virtual sound source positioning using vector base amplitude panning,

    V . Pulkki, “Virtual sound source positioning using vector base amplitude panning,”J. Audio Eng. Soc., vol. 45, no. 6, pp. 456–466, 1997

  38. [46]

    Zotter and M

    F. Zotter and M. Frank,Ambisonics: A practical 3D audio theory for recording, studio production, sound reinforcement, and virtual reality. Springer, 2019

  39. [47]

    A super-resolution ambisonics-to- binaural rendering plug-in,

    P. M. Giller and C. Sch ¨orkhuber, “A super-resolution ambisonics-to- binaural rendering plug-in,” inDAGA, 2019

  40. [48]

    Sector-based parametric sound field reproduction in the spherical harmonic domain,

    A. Politis, J. Vilkamo, and V . Pulkki, “Sector-based parametric sound field reproduction in the spherical harmonic domain,”IEEE J. Sel. Topics in Sig. Proc., vol. 9, no. 5, pp. 852–866, 2015

  41. [49]

    Parametric spatial audio processing of spaced microphone array recordings for multichannel reproduction,

    A. Politis, M.-V . Laitinen, J. Ahonen, and V . Pulkki, “Parametric spatial audio processing of spaced microphone array recordings for multichannel reproduction,”J. Audio Eng. Soc, vol. 63, no. 4, pp. 216– 227, 2015

  42. [50]

    Parametric spatial sound processing using linear microphone arrays,

    O. Thiergart, M. Kallinger, G. D. Galdo, and F. Kuech, “Parametric spatial sound processing using linear microphone arrays,” inMicroelec- tronic Systems. Springer, 2011, pp. 321–329

  43. [51]

    MPEG surround- the ISO/MPEG standard for efficient and compatible multichannel audio coding,

    J. Herre, K. Kj ¨orling, J. Breebaart, C. Faller, S. Disch, H. Purnhagen, J. Koppens, J. Hilpert, J. R ¨od´en, W. Oomenet al., “MPEG surround- the ISO/MPEG standard for efficient and compatible multichannel audio coding,”J. Audio Eng. Soc., vol. 56, no. 11, pp. 932–955, 2008

  44. [52]

    Diffuse-field coherence of sensors with arbitrary directional responses,

    A. Politis, “Diffuse-field coherence of sensors with arbitrary directional responses,”arXiv preprint arXiv:1608.07713, 2016

  45. [53]

    Estimating and reproducing ambience in ambisonic recordings,

    L. McCormack and A. Politis, “Estimating and reproducing ambience in ambisonic recordings,” in30th European Sig. Proc. Conf. (EUSIPCO), 2022

  46. [54]

    Optimized covariance domain framework for time–frequency processing of spatial audio,

    J. Vilkamo, T. B ¨ackstr¨om, and A. Kuntz, “Optimized covariance domain framework for time–frequency processing of spatial audio,”J. Audio Eng. Soc., vol. 61, no. 6, pp. 403–411, 2013

  47. [55]

    Direction finding using noise covariance modeling,

    B. Friedlander and A. J. Weiss, “Direction finding using noise covariance modeling,”IEEE Trans. Sig. Proc., vol. 43, no. 7, pp. 1557–1567, 1995

  48. [56]

    A perceptual evaluation of individual and non-individual HRTFs: A case study of the SADIE II database,

    C. Armstrong, L. Thresh, D. Murphy, and G. Kearney, “A perceptual evaluation of individual and non-individual HRTFs: A case study of the SADIE II database,”Applied Sciences, vol. 8, no. 11, p. 2029, 2018

  49. [57]

    Time-frequency processing: Methods and tools,

    J. Vilkamo and T. B ¨ackstrom, “Time-frequency processing: Methods and tools,” inParametric Time-Frequency Domain Spatial Audio, V . Pulkki, S. Delikaris-Manias, and A. Politis, Eds. John Wiley & Sons, 2017, pp. 1–24

  50. [58]

    Improved source number detection and direction estimation with nested arrays and ULAs using jackknifing,

    K. Han and A. Nehorai, “Improved source number detection and direction estimation with nested arrays and ULAs using jackknifing,” IEEE Trans. Sig. Proc., vol. 61, no. 23, pp. 6118–6128, 2013

  51. [59]

    Multiple emitter location and signal parameter estimation,

    R. Schmidt, “Multiple emitter location and signal parameter estimation,” IEEE Trans. Antennas Propag., vol. 34, no. 3, pp. 276–280, 1986

  52. [60]

    Spatial reconstruction-based rendering of microphone array room impulse responses,

    L. McCormack, N. Meyer-Kahlen, and A. Politis, “Spatial reconstruction-based rendering of microphone array room impulse responses,”J. Audio Eng. Soc., vol. 71, no. 5, pp. 267–280, 2023

  53. [61]

    Image method for efficiently simulating small-room acoustics,

    J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics,”J. Acoust. Soc. Am., vol. 65, no. 4, pp. 943–950, 1979

  54. [62]

    Anechoic recording system for symphony orchestra,

    J. P ¨atynen, V . Pulkki, and T. Lokki, “Anechoic recording system for symphony orchestra,”Acta Acustica united with Acustica, vol. 94, no. 6, pp. 856–865, 2008

  55. [63]

    Spatially oriented format for acoustics 2.1: Introduction and recent advances,

    P. Majdak, F. Brinkmann, J. De Muynke, M. Mihocic, and M. Noisternig, “Spatially oriented format for acoustics 2.1: Introduction and recent advances,”J. Audio Eng. Soc., vol. 70, pp. 565–584, 2022

  56. [64]

    Perceptual evaluation of approaches for binaural reproduction of non-spherical microphone array signals,

    T. L ¨ubeck, S. V . Amengual Gar ´ı, P. Calamia, D. L. Alon, J. Crukley, and Z. Ben-Hur, “Perceptual evaluation of approaches for binaural reproduction of non-spherical microphone array signals,”Frontiers in Sig. Proc., vol. 2, p. 883696, 2022

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.