Pith. sign in

Paper Citation Record · LEDGER

Visual-based spatial audio generation system for multi-speaker environments

As of 20 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2502.07538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07538 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:30:26.197293Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:04:00.782536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T22:54:55.849359Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01c1e241-7246-4d53-9b27-b2b1450eed7a · outbound

This paper cites Exploring audio- visual information fusion for sound event localization and detection in low-resource realistic scenarios,.

Visual-based spatial audio generation system for multi-speaker environments Exploring audio- visual information fusion for sound event localization and detection in low-resource realistic scenarios,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.536832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.103584Z digest=sha256:e3bc65df29bf17816a73bcd61db23da0ee0f4ef6a4d8ec673a3d8e14147b5f11

Observation 047e4701-6346-4cf6-8f4d-9ca1b7958006 · outbound

This paper cites Aligning audiovisual features for audiovisual speech recognition,.

Visual-based spatial audio generation system for multi-speaker environments Aligning audiovisual features for audiovisual speech recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.522970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.108687Z digest=sha256:d40101db13fef695d92b00c4a0951d6ab91106fc03cd272aedc02b0555b8a8c4

Observation c3174224-52d0-4dab-9458-417abf7c78be · outbound

This paper cites Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360° images,.

Visual-based spatial audio generation system for multi-speaker environments Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360° images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.508310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.113528Z digest=sha256:fc0193f4db92425bcffeb597f7be56577f475a493d1d9bf32827cd34d364f000

Observation 0a28dc3b-e7de-4625-87b6-e6a8ac9ff269 · outbound

This paper cites Scene-aware audio for 360° videos,.

Visual-based spatial audio generation system for multi-speaker environments Scene-aware audio for 360° videos,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.494322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.118296Z digest=sha256:41413184051c09b054c53c7bcf269a3db9e832f20cf75ec8b46a358b441028ba

Observation f115ea40-7d76-4016-b312-4845ec80fcd7 · outbound

This paper cites Analysis of a distributed processing model for spatialized audio conferences,.

Visual-based spatial audio generation system for multi-speaker environments Analysis of a distributed processing model for spatialized audio conferences,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.480615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.122985Z digest=sha256:074a9ae4b841321549cfd3e7393950c8b0358afb1490ded199bc62f5941ee028

Observation 78e8b1d3-6199-452e-a843-782c4e2ff069 · outbound

This paper cites Realistic audio in immersive video conferencing,.

Visual-based spatial audio generation system for multi-speaker environments Realistic audio in immersive video conferencing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.465535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.127638Z digest=sha256:ae757780b8f24930e6e579f4defaefae8069f3896be05d0b5b01a6aa4b61840a

Observation 69c65982-9250-4479-9afb-2e3243dc5188 · outbound

This paper cites Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement,.

Visual-based spatial audio generation system for multi-speaker environments Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.451213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.132615Z digest=sha256:ccbdfdc5ec540b01ded00d210ae281610e66a4e14cfebd845f0bab6744fd6639

Observation 55096907-3dba-4bce-b6ba-1c792c2ef5a4 · outbound

This paper cites 2.5d visual sound,.

Visual-based spatial audio generation system for multi-speaker environments 2.5d visual sound,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.434867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.136964Z digest=sha256:ac77e1eb771071a77eec7fe225cd18192ab65945671020a2c521e7a527d74b06

Observation 118c667c-4c36-4529-bc12-a5fddc806369 · outbound

This paper cites Visually informed binaural audio generation without binaural audios,.

Visual-based spatial audio generation system for multi-speaker environments Visually informed binaural audio generation without binaural audios,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.418915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.141461Z digest=sha256:fd8895e961a2d2c5c087ee367d0d6ea3e96fb095b2b341903b99e30c5ae367fd

Observation ed4ccc89-e858-4317-8a1e-85e3381ececf · outbound

This paper cites A review on yolov8 and its advancements,.

Visual-based spatial audio generation system for multi-speaker environments A review on yolov8 and its advancements,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.403157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.145960Z digest=sha256:2f51c8eaec48f9d3cefae26b53f8b557db9834e1e674c6d70d4145dc68e4a632

Observation 7687ebb2-2f03-4d28-adb5-1085d7bacbe0 · outbound

This paper cites Wider face: A face detection benchmark,.

Visual-based spatial audio generation system for multi-speaker environments Wider face: A face detection benchmark,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.388110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.150357Z digest=sha256:5de5fd4ba54387447969a4a3e039e455ffefd0d7ac200043c0ed9ba9d1e7db0c

Observation b017bd20-49bb-4854-bd06-289fa64f87f5 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data,.

Visual-based spatial audio generation system for multi-speaker environments Depth anything: Unleashing the power of large-scale unlabeled data,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.372085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.154803Z digest=sha256:ce258703b1f4ef0eb2aa01861fa6142fd348906d18359bbdc9bdd966d8017258

Observation 535446c5-1b96-47b9-aefd-f809af2ef142 · outbound

This paper cites You only look once: Unified, real-time object detection,.

Visual-based spatial audio generation system for multi-speaker environments You only look once: Unified, real-time object detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.356023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.160648Z digest=sha256:c65372ffbbc215446cc0ff846015c4894cf627e3d00cfce72d4ff9581e68cb81

Observation e8396a1d-7eb7-4c8d-aedc-f42dbcdd999c · outbound

This paper cites Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Visual-based spatial audio generation system for multi-speaker environments Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.339161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.165241Z digest=sha256:aa21bea03d2c45e404e96503e7083e3360afb714c40cb1abe65173a18ac8ec8a

Observation f09ed45e-7d93-432f-83d1-ddcd5ac4e1e3 · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

Visual-based spatial audio generation system for multi-speaker environments Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.169424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.169424Z digest=sha256:fb02e2d37d01396a091d0014e508fb935b1b7e30efd465800404f4bdd6bdeb47

Observation f87a8c9c-c216-4bbd-b62f-33cba63d9bda · outbound

This paper cites A perceptual evaluation of individual and non-individual hrtfs: A case study of the sadie ii database,.

Visual-based spatial audio generation system for multi-speaker environments A perceptual evaluation of individual and non-individual hrtfs: A case study of the sadie ii database,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.322485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.174917Z digest=sha256:5d3d6b7b318a454b48b2ff0fda912f9d7a0ed627289133cd29b497f6fd789dcc

Observation df9dfaea-879f-4ace-9f8d-7a3361eb6daa · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Visual-based spatial audio generation system for multi-speaker environments LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.180184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.180184Z digest=sha256:c2adec12d767f825a0f45fc02ef968fb63c48988c7231b17d83e290005f7aac9

Observation ca24f598-e728-466c-8ccb-b99e6fe79ba8 · outbound

This paper cites Self- supervised generation of spatial audio for 360 video,.

Visual-based spatial audio generation system for multi-speaker environments Self- supervised generation of spatial audio for 360 video,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.308132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.184894Z digest=sha256:9a3d068de6b4b4e7f696916d506c18011dfe33dbfe7815de7e474c25d81ce7de

Observation 37c35292-c3e0-4301-a298-4f30aa1ebff7 · outbound

This paper cites Peaq-the itu standard for objective measurement of perceived audio quality,.

Visual-based spatial audio generation system for multi-speaker environments Peaq-the itu standard for objective measurement of perceived audio quality,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.293031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.188982Z digest=sha256:efcfd0858e6432265713ec505f1dc16d4f600b0b92cf2f07a3560a9de5f5c0db

Observation 196a87c3-a672-415b-8523-196aa0eb2830 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

Visual-based spatial audio generation system for multi-speaker environments An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.278058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T12:30:26.193102Z digest=sha256:3e38207fea1f0743c689366d1afd3325bd5bf53d3924154482280b5cc4d14ad3

Observation 35500a2e-48ca-4ba1-8b71-f261a0836319 · outbound

This paper cites MOSNet: Deep Learning based Objective Assessment for Voice Conversion.

Visual-based spatial audio generation system for multi-speaker environments MOSNet: Deep Learning based Objective Assessment for Voice Conversion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.197293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.197293Z digest=sha256:0540ef023bd8c2847a5cf505c4fde58c0e22f39b616c6dc39ed2f3a69f57c850

Pith citing papers

Observation e32e88ea-5594-4674-8822-0c9f4b73b63c · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations Visual-based spatial audio generation system for multi-speaker environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.782536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.782536Z digest=sha256:990a79381607fa9ea159b5baedf6af67e773a9dc339be4d32d22086d29fc32a5

Observation 2eafc3f1-4726-4f15-9c63-e70ead68af53 · inbound

ASAudio: A Survey of Advanced Spatial Audio Research cites this paper.

ASAudio: A Survey of Advanced Spatial Audio Research Visual-based spatial audio generation system for multi-speaker environments

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:54:55.853160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-05T22:54:55.101714Z digest=sha256:7dc6ac3175a7ec6b2b9da13729d6d40bacf4669dfa04db5bf5c2e47bb045ce30