Pith. sign in

REVIEW 3 major objections 5 minor 16 references

Interactive Multimodal Integral Field Spectroscopy

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that a binaural sonification tool for galaxy datacubes lets listeners estimate a spectrum's position, distance from the center, and stellar-age class at above-chance rates after brief training.

desk verdict A genuine, open-source sonification tool with an honest but statistically thin user study; the location and distance results hold, while the spectral-type generalization claim rests on an untested cross-galaxy latent-space assumption. read the letter →

arxiv 2412.00185 v1 pith:IEYRDMS4 submitted 2024-11-29 astro-ph.IM astro-ph.GA

classification astro-ph.IMastro-ph.GA
keywords sonificationintegralfieldspectroscopydatacubevisualizationautoencoderlatentspacebinauralsoundscapeaccessibilityCALIFAsurveygalaxyspectra
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes that integral field spectroscopy datacubes—grids of thousands of spectra covering a galaxy's face—can be explored with sound as well as sight, and reports a working tool that does it. The central claim is that listeners who hear a short binaural rendering of a spectrum can estimate where that spectrum sits in the galaxy, whether it is near or far from the center, and whether it comes from a star-forming, intermediate-age, or retired galaxy, at rates above random choice. The supporting evidence is an online user study with 67 participants: the mean success rate on simple questions was 0.516, above the per-question random reference, and 79 percent of participants rated the application useful. The authors read these results as indicating that multimodal IFS analysis is feasible and that it could make datacube exploration more accessible to blind and low-vision researchers.

What carries the argument

The load-bearing object is a six-layer, six-dimensional sparse autoencoder that compresses each 1901-wavelength spectrum into a six-component latent vector; the sonification module multiplies those components by 10,000 to produce the fundamental frequencies of a six-oscillator additive synthesizer. Spatial information is added by encoding the spaxel's azimuth as a head-related-transfer-function binaural position and its distance from the galaxy center as the direct-to-reverberant ratio of a reverberation effect. The autoencoder's reconstruction quality is high for most datacubes—coefficients of determination range from 0.882 to 0.998, with 49 percent of datacubes above 0.96—so the six latent coordinates serve as a compact auditory fingerprint of spectral shape, including strong emission lines and continuum.

What would settle it

A listening experiment in which participants train only on star-forming, intermediate-age, and retired examples from one set of galaxies and are then tested on galaxies never heard in the training phase; if test accuracy does not exceed random choice, the claim of a generalizable auditory vocabulary fails. A complementary quantitative check is to compute the six-dimensional latent vectors for the same spectral class across many CALIFA datacubes and test whether the clusters overlap; large between-galaxy scatter would show the sounds are per-cube rather than per-class.

Watch

Extended reading notes

Core claim

The core discovery is that a six-dimensional latent vector produced by a sparse autoencoder trained on a single CALIFA datacube can be turned into six simultaneous oscillator tones and still deliver perceptible astronomical information. In the user study, participants identified left/right/front/rear sound position, close/intermediate/far distance from the galaxy center, and star-forming/intermediate/retired spectral class at rates above chance, with a global mean success of 0.516 against random-choice references. The authors also find that experience with sound events predicted success better than prior knowledge of the data, that the two blind or low-vision professional astronomers scored on par with or slightly above comparable sighted peers on simple questions, and that combined three-attribute questions were markedly harder than single-attribute ones.

Load-bearing premise

The load-bearing premise is that the sonification vocabulary is consistent across galaxies: because the autoencoder is retrained on each datacube, a star-forming spectrum could produce one set of pitches in one galaxy and a different set in another, and the study never tests whether the six latent coordinates are comparable between training and test galaxies.

Editorial extensions

If this is right

  • A researcher inspecting an unfamiliar datacube can hear a quick qualitative overview of each spaxel's spectrum, complementing visual 2D maps and spectra.
  • The same architecture can be pointed at datacubes from other instruments and surveys once they pass through the flexible FITS reader, so the sonification palette is not tied to CALIFA.
  • Blind and low-vision astronomers can perform basic IFS orientation and spectral-class discrimination tasks after short training, based on the two BLV participants' performance.
  • Training and attention are decisive: participants with sound-analysis experience or focused non-expert learning performed as well as or better than professional astronomers, so practical deployment should invest in guided listening exercises.
  • Combined spatial-and-spectral identification is still unreliable, implying the tool should initially support single-attribute tasks rather than full simultaneous classification.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper never tests whether the six latent coordinates mean the same thing across galaxies, because each datacube's autoencoder is trained independently; a direct check would be to compare latent vectors for the same spectral classes in many galaxies and see whether they cluster consistently.
  • If cross-cube alignment fails, the apparent success could reflect within-datacube timbre differences rather than a universal sonification vocabulary; aligning the latent space across galaxies would turn the current per-cube timbre language into a stable one.
  • The two BLV participants are suggestive but not enough to establish the accessibility claim; a larger study designed with BLV researchers as co-designers would be a natural next step.
  • The six-tone additive sound is probably perceived as a single timbre rather than six separate pitches, so future sonification palettes could exploit timbre dimensions directly instead of mapping each latent coordinate to a frequency.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents ViewCube/SoniCube, a Python application for interactive multimodal (visual and auditory) exploration of integral field spectroscopy datacubes, with sonification based on a six-dimensional sparse autoencoder trained per datacube. The authors report a user study with 67 participants (including two BLV astronomers) who completed an online questionnaire with training videos and in-person demonstrations. Quantitative results show above-chance performance in estimating spaxel location, distance from center, and spectral type/age in simple questions, with a global success rate of 0.516. Qualitative feedback was largely positive. The paper concludes that multimodal IFS can convey useful information and may improve accessibility for BLV researchers.

Significance. If the central claims are reliable, the paper makes a useful contribution to astronomical sonification: it provides an open-source tool and an evaluated sonification scheme that could complement visual analysis of datacubes and support inclusive access. The study includes multiple galaxies, participant groups, and quantitative and qualitative measures, and the data and analysis notebooks are publicly available. However, the type/age result depends on an untested cross-galaxy consistency of the autoencoder latent spaces, and the BLV accessibility claim rests on two participants; both points materially temper the conclusions as currently stated.

major comments (3)
  1. [Section 2.4, Section 3.1, Section 3.3] The autoencoder is trained independently on each datacube, so the six latent coordinates are not aligned across galaxies. The training videos for the age/type task use spectra from NGC 3395, NGC 2347, and NGC 6125, while all test questions use sonifications from NGC 5784, NGC 5732, NGC 5682, NGC 6060, NGC 7562, NGC 7671, NGC 7800, NGC 2638, and UGC 00148. For the type/age task (Section 3.3, 40.7% success versus the 33% random reference) to support the claim that listeners can identify spectral classes, the latent representation of a star-forming region must be sonically similar across galaxies. Because nothing in the autoencoder architecture or loss constrains the latent axes across training runs, this assumption is unverified. The observed above-chance performance could instead reflect within-question low-level cues (for example, overall flux, number of strong emission lines, or time/frequency envelope) that would not generalize. The authors should either demonstrate cross-galaxy latent consistency (e.g., by comparing latent encodings of matched spectral classes from the Zenodo data) or restrict the type/age conclusion to within-cube training.
  2. [Section 3.2 and Conclusions] The claim that the tool "can improve the access of BLV astronomers to IFS analysis" is based on two self-declared BLV professional astronomers. The paper correctly notes the sample is too small for statistical significance, but the conclusion still presents the suggestive comparison (8% higher success than a random two-person subgroup) as support. With n=2, the comparison is not informative; a single response change flips the sign. The authors should either remove or explicitly label this as anecdotal and not as evidence of accessibility, or provide additional evidence (e.g., qualitative feedback from the BLV participants or accessibility evaluation with a larger sample).
  3. [Section 3.3, Fig. 9] The combined-question success rate (0.157) is only modestly above the random baseline of 0.125 (all three two-alternative sub-questions correct by chance), and no confidence interval or statistical test is reported for the combined block. The text states that "all participants were able to retrieve information from the sonifications" based on the global simple-question rate, but the combined task is the more realistic test of integrated use. The authors should report uncertainty on the combined success rate and discuss whether it actually exceeds chance.
minor comments (5)
  1. [Section 3.4] One participant (1.49%) is reported to have considered the application "Usefulness"; this likely should read "Useless" or "Doubtfully useful".
  2. [Abstract and Introduction] The description of the deep learning module as generating "binaural unsupervised auditory representations" conflates the autoencoder-based sound synthesis with the HRTF-based binaural rendering described in Appendix B; consider clarifying that binaural encoding is separate from the deep learning component.
  3. [Appendix A, Eq. (A1)] The sum in Equation A1 runs from i=0 to 6 but the text states there are six oscillators; please adjust the index bounds so the equation matches the described number of oscillators.
  4. [References] The reference to "PHOENIX, 1990" is listed with only a URL and is incomplete relative to the other references; please provide full bibliographic details.
  5. [Figure 7] The caption and text refer to panels as "up-left", "left-down", etc.; use standard panel labels (a)-(d) for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the user study is an external perceptual test, and the autoencoder's reconstruction objective does not encode the survey labels.

full rationale

The paper's central claim is that listeners can recover spaxel position, radial distance, and spectral type from ViewCube/SoniCube sonifications. That claim is tested by an online user study (Section 3) with training videos and held-out question galaxies, not derived from the autoencoder's fitted parameters. The autoencoder (Section 2.4) is trained only to reconstruct spectra, with reconstruction quality reported as R^2, and the labels used to score the Type/Age questions are external physical classifications (star-forming, intermediate, retired), not outputs of the network. Position and distance cues are generated by direct azimuth-to-HRTF and distance-to-reverberation mappings (Appendix A, Eq. A4; Appendix B), and the perceptual task is whether listeners can hear those cues; that is an external behavioral test. The only author self-citation is for the choice of six latent dimensions, in the sentence 'This dimension also worked effectively for the Calcium II Triplet library (CaT) reduced with sparse autoencoders (García Riber & Serradilla June, 2024), which also agrees with our preliminary tests on the CALIFA survey galaxies.' This citation is not the sole load-bearing justification, because the same passage also cites Portillo et al. (2020) and the authors' own CALIFA preliminary tests. The per-datacube independent training is a genuine generalization limitation: latent axes may not align across galaxies, so the training examples from NGC 3395, NGC 2347, and NGC 6125 may not map to the same sonic categories as the question galaxies. However, this is a threat to external validity and to the interpretation of the Type/Age results, not a circular reduction: the study's success rates are not defined by the fitted parameters, and the autoencoder never sees the Type/Age labels during training.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim rests mainly on the assumption that per-cube autoencoder latent spaces are consistent enough for listeners to generalize, on unverified participant compliance with headphone instructions, and on a volunteer sample. No new physical entities are introduced.

free parameters (6)
  • latent_dimension = 6
    The number of autoencoder latent dimensions sets the number of oscillators and was chosen from preliminary CALIFA tests and prior work (Section 2.4), not derived from first principles.
  • latent_to_frequency_scale = 10,000
    The latent vector components are multiplied by 10,000 to reach audible frequencies (Section 2.3). This hand-chosen scaling is central to all sonifications.
  • reverberation_feedback_gain_g = 0.9
    The feedback gain in Appendix A is set to 0.9 to provide a long reverberation effect, which encodes the distance cue.
  • reverberation_delay_tau = not specified in text
    The fixed delay time per line in Appendix A is not given a value, which limits exact reproduction of the distance encoding.
  • compressor_stage_slope_r = not fully specified
    The slope of the two-stage limiter and compressor appears in Equation A1 and Figure A1, but the numerical slope is not stated, affecting loudness mapping.
  • autoencoder_architecture = six layers, 100 epochs
    Network depth, sparsity level, and training duration were chosen from tests in Section 2.4 and are not justified from first principles.
assumptions (5)
  • domain assumption Per-datacube autoencoder latent spaces are comparable across galaxies
    Section 2.4 trains each autoencoder independently, so latent axes are not aligned. The evaluation assumes listeners can transfer class cues from training galaxies to test galaxies.
  • domain assumption Self-reported expertise in astronomy and music is a valid grouping variable
    Subgroup analyses in Section 3.3 rely on participants' own declarations about their expertise rather than objective tests.
  • domain assumption Volunteer sample represents the intended user population
    The 67 participants were self-selected, mostly from Spain, with only two BLV astronomers. Generalizability to the broader astronomy community is assumed.
  • domain assumption Headphones were worn and left and right channels correctly placed
    The survey instructed participants to use headphones and check placement (Section 3.1), but compliance was not verified.
  • standard math HRTF-based binaural rendering provides reliable azimuth cues
    The hrtfmove opcode and HRTF formalism are referenced from prior literature (Brown and Duda 1997; Carty 2008) and treated as a standard audio technique.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Interactive Multimodal Integral Field Spectroscopy." pith.science (2026). https://pith.science/paper/IEYRDMS4

@misc{pith2026241200185,
  author       = {Pith},
  title        = {Pith review of: Interactive Multimodal Integral Field Spectroscopy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IEYRDMS4}},
  note         = {Machine review of arXiv:2412.00185}
}
read the original abstract

Using sonification on scientific data analysis provides additional dimensions to visualization, potentially increasing researchers' analytical capabilities and fostering inclusion and accessibility. This research explores the potential of multimodal Integral Field Spectroscopy (IFS) applied to galaxy analysis through the development and evaluation of a tool that complements the visualization of datacubes with sound. The proposed application, \textit{ViewCube}, provides interactive visualizations and sonifications of spectral information across a two-dimensional field-of-view, and its architecture is designed to incorporate future sonification approaches. The first sonification implementation described in this article uses a deep learning module to generate binaural unsupervised auditory representations. The work includes a qualitative and quantitative user study based on an online questionnaire, aimed at both specialized and non-specialized participants, focusing on the case study of datacubes of galaxies from the Calar Alto Integral Field Spectroscopy Area (CALIFA) survey. Out of 67 participants who completed the questionnaire, 42 had the opportunity to test the application in person prior to filling out the online survey. 81\% of these 42 participants expressed the good interactive response of the tool, 79.1\% of the complete sample found the application "Useful", and 58.2\% rated its aesthetics as "Good". The quantitative results suggest that all participants were able to retrieve information from the sonifications, pointing to previous experience in the analysis of sound events as more helpful than previous knowledge of the data for the proposed tasks, and highlighting the importance of training and attention to detail for the understanding of complex auditory information.

Figures

Figures reproduced from arXiv: 2412.00185 by the authors.

Figure 1
Figure 1. ViewCube UI displaying the datacube of the spiral (Sbc) galaxy NGC 5732. 2D image window (left) and multimodal representation –spectral window and sonification– of the spaxel (35,35) (right). 2.2 Standalone application ViewCube1 is a lightweight, standalone application written entirely in Python, designed for the efficient browsing of datacubes. Originally developed for the quick assessment of the quality and physic… view at source ↗
Figure 2
Figure 2. SoniCube block diagram. Pre-processing and real time calculations including data, OSC, and audio signal flows. 2.3 Sonification module This section describes the sound module implemented within ViewCube to allow the sonification of the spectra associated with each spatial element of a datacube. The module, named SoniCube, pro￾vides an open, comprehensive, and general-purpose multimodal tool for IFS analysis, support… view at source ↗
Figure 3
Figure 3. Autoencoder comparative for the spiral (Scd) galaxy NGC 5406. Six-layer six-dimensional autoencoder (black, continuous) vs four-layer six dimensional VAE (orange, dotted). Reconstructed spectra and residual error from the original spectrum for spaxel (34,34). Sparse autoencoder: R2 = 0.99 (spectrum), R2 = 0.98 (datacube), 39.12% of the spectra with R2 > 0.9, 100 epochs, one hour per cube. VAE: R2 = 0.97 (spectrum), … view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Six-layer six-dimensional autoencoder results. Reconstructed (black dashed line) and original (blue solid line) spectrum with residual error (orange dotted line) of the spaxel (35,35) from the spiral (Sbc) galaxy NGC 5784 (left), and the elliptical (E4) galaxy NGC 5682…
Figure 5
Figure 5. Figure 5: Learning curves showing mean square error vs epoch during the training and validation processes for the spiral (SAb) galaxy NGC7047, the elliptical (E1) galaxy UGC10331, and the elliptical (E5) galaxy UGC03960. These datacubes correspond respectively to the best, mediu…
Figure 6
Figure 6. Figure 6: Age/Galaxy type examples presented in the training videos. Spectrum of a star-forming region in the spiral galaxy (S0) NGC 3395 (left), spectrum close to the center of the intermediate-age spiral galaxy (Sd) NGC 2347 (center), and spectrum of a region close to the cent…
Figure 7
Figure 7. Figure 7: Evaluation results. Up-left: Average success rates for 67 participants on simple questions (left, magenta), for 31 professional astronomers (center, blue) and for 31 non-astronomers (right, green). Up-right: Average success rates on simple questions by field of experti…
Figure 8
Figure 8. Figure 8: Success rate by question and group of expertise (notice the difference in the number of participants). Results for Astronomers vs Non-astronomers (up), and subgroup rates for Astronomers musicians, Astronomers non-musicians, Musicians non-astronomers, and Non-experienc…
Figure 9
Figure 9. Figure 9: Average success rate of the two combined questions by blocks (Location, Distance, and Type/Age). From left to right, success rate for As￾tronomers vs Non-astronomers, and subgroup rates for Astronomers musi￾cians, Astronomers non-musicians, Musicians non-astronomers, a…
Figure 10
Figure 10. Figure 10: Success rate on simple questions for participants trained only with videos (V) vs participants testing the application live (L). From left to right, global results, astronomers vs non-astronomers, and expertise subgroups. Dotted line shows averaged random choice rate …
Figure 11
Figure 11. Figure 11: Qualitative evaluation. Interactivity: feedback from 42 participants who tested the application "in person". Usefulness and aesthetics: full sample, 67 participants. 81% declared that the application had a "good interactivity", 79.1% found it "useful" and 58.2% "good …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 16 canonical work pages

  1. [1]

    {TensorFlow}:asystem for{Large-Scale} machine learning, in 12th USENIX symposium on operating systems design and implementation (OSDI 16), pp

    Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat,S.,Irving,G.,Isard,M.,etal.,2016. {TensorFlow}:asystem for{Large-Scale} machine learning, in 12th USENIX symposium on operating systems design and implementation (OSDI 16), pp. 265–283. Allington-Smith, J.,

  2. [4]

    Analysis of astronomical data through sonification: reaching more inclusion for visual disable scientists

    Analysis of astronomical data through sonification: reaching more inclusion for visual disable scientists,arXiv preprint arXiv:2402.00611. Charbonneau, J., Novak, C., Gaspar, R., & Ule, H.,

  3. [14]

    Introducing STRAUSS: A flexible sonification Python package

    Introducing strauss: A flexible sonification python package,arXiv preprint arXiv:2311.16847. Trayford,J.W.,Harrison,C.,Hinz,R.,KavanaghBlatt,M.,Dougherty,S.,& Girdhar, A.,

  4. [15]

    West,R.,Johnson,V.,Yeh,I.C.,Thomas,Z.,&Tarlton,M.,2018

    Theory of sonification,The sonification handbook,1, 9–39. West,R.,Johnson,V.,Yeh,I.C.,Thomas,Z.,&Tarlton,M.,2018. Experienc- ing a slice of the sky: Immersive rendering and sonification of antarctic astronomy data,Electronic Imaging,30, 1–10. Wright, M.,

  5. [291]

    Hansen, B., Burchett, J

    Goodfellow, I., 2016.Deep Learning, MIT Press. Hansen, B., Burchett, J. N., & Forbes, A. G.,

  6. [1997]

    1–4, IEEE, New Paltz, NY, USA

    An efficient hrtf model for 3-d sound, in Proceedings of 1997 Workshop on Applications of Signal Processing to Audio and Acoustics, pp. 1–4, IEEE, New Paltz, NY, USA. Carty, B.,

  7. [2005]

    Lazzarini,V.&Carty,B.,2008

    Principles of digital dynamic-range compression,Trends RASTI000, 1–12 (2024) Interactive Multimodal IFS 11 in amplification,9(2), 45–76. Lazzarini,V.&Carty,B.,2008. Newcsoundopcodesforbinauralprocessing, inProc.6thLinuxAudioConference ,pp.28–35,Citeseer,Köln,Germany. Lu, Y.-C. & Cooke, M.,

  8. [2010]

    Mas-Buitrago, P., González-Marcos, A., Solano, E., Passegger, V., Cortés- Contreras, M., Ordieres-Meré, J., Bello-García, A., Caballero, J., Schweitzer,A.,Tabernero,H.,etal.,2024

    Binaural estimation of sound source distance via the direct-to-reverberant energy ratio for static and moving sources, IEEE Transactions on Audio, Speech, and Language Processing, 18(7), 1793–1805. Mas-Buitrago, P., González-Marcos, A., Solano, E., Passegger, V., Cortés- Contreras, M., Ordieres-Meré, J., Bello-García, A., Caballero, J., Schweitzer,A.,Tabe...

Show all 16 references
  1. [2011]

    PHOENIX,

    Sparse autoencoder,CS294A Lecture notes, 72(2011), 1–19. PHOENIX,

  2. [2013]

    Kahrs,M.&Brandenburg,K.,1998

    Standard 61672-1: 2013, electroacoustics—sound level me- ters—part 1: Specifications. Kahrs,M.&Brandenburg,K.,1998. Applicationsofdigitalsignalprocessing to audio and acoustics, Springer Science & Business Media. Kates, J. M.,

  3. [2017]

    Dubus,G.&Bresin,R.,2013

    The PyCASSO database: spatially resolved stellar population properties for CALIFA galaxies,MNRAS, 471(3), 3727–3752. Dubus,G.&Bresin,R.,2013. Asystematicreviewofmappingstrategiesfor the sonification of physical quantities,PloS one,8(12), e82491. Enge, K., Elmquist, E., Caiola,...

  4. [2018]

    Trayford, J

    Soniverse, https://soniverse.space/2018/02/15/ listen-to-the-radio-cube-of-the-antennae-galaxies . Trayford, J. W. & Harrison, C. M.,

  5. [2021]

    72–80, Association for Computing Machinery, NY, USA

    Sonification of planetary orbits in asteroid belts, inProceedings of the 16th International Audio Mostly Conference, pp. 72–80, Association for Computing Machinery, NY, USA. Rönnberg,N.&Jimmy,J.,2016. Interactivesonificationforvisualdensedata displays,in ISon2016,5thInteractiv...

  6. [2023]

    The sonified hertzsprung-russell diagram,arXiv preprint arXiv:2401.00488. IEC,

  7. [2024]

    Baldi,P.,2012

    A universe of sound: processing nasa data into sonifications to explore participant response, Frontiers in Communication,9, 1288896. Baldi,P.,2012. Autoencoders,unsupervisedlearning,anddeeparchitectures, in Proceedings of ICML workshop on unsupervised and transfer learn- ing, ...

  8. [4793]

    APPENDIX A: SYNTHESIZER DESCRIPTION The equation of the six-oscillator additive synthesizer implemented in SoniCubecan be expressed for each sonification as: 𝑆(𝑡) = 6∑︁ 𝑖=0 𝐴i𝐹𝑟 sin(2𝜋𝑓 i𝑡+𝜙i) (A1) where Ai is the A-weighting coefficient for each frequency,F is the median of a...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.