Pith. sign in

Paper Citation Record · LEDGER

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.21174.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21174 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:53.842471Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 594845b1-4df0-45ac-9b2f-5fbbad5e9c55 · outbound

This paper cites This complex task requires a system to identify active sound classes (audio tagging) and to isolate their corresponding anechoic source signals accurately.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 This complex task requires a system to identify active sound classes (audio tagging) and to isolate their corresponding anechoic source signals accurately

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:57.666390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:51.984029Z digest=sha256:1e2d4645489784f50289c5584a55157f8db3ba2990fce36de96eaac5b2bafa8d

Observation 6d07eef9-ab98-4f10-b40b-044f7210a1ab · outbound

This paper cites This section outlines the com- position of the official dataset and the specific data curation and augmentation steps to improve the performance of the model.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 This section outlines the com- position of the official dataset and the specific data curation and augmentation steps to improve the performance of the model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:57.489728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:52.076241Z digest=sha256:04c942fe005c386c4cb8801388870516761bdd73027e0dc10da008f2febd4218

Observation 9aef203b-bec4-4fff-b5d5-01e5669690f0 · outbound

This paper cites an unresolved cited work.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:57.034367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:52.269462Z digest=sha256:5269543765ab8447c4c39e34a695e1c39493c9dc1d99a0c9b5549bb64ac7f835

Observation df393d6e-d508-49e4-83be-ef658276eaea · outbound

This paper cites Model training The audio-tagging and source-separation models were trained in- dependently.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Model training The audio-tagging and source-separation models were trained in- dependently

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T22:36:56.632808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:52.564116Z digest=sha256:fce63a34fc7669ae86533431e7d180b1e3f37f23b1cfa392af580442acdf96d0

Observation 7d65e086-be41-4beb-b043-31c4a2c1be09 · outbound

This paper cites The proposed strategy combined additional audio feature input (spectral roll-off and chroma), a dataset refinement process, and an agent-based er- ror correction system.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 The proposed strategy combined additional audio feature input (spectral roll-off and chroma), a dataset refinement process, and an agent-based er- ror correction system

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:56.421479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:52.670206Z digest=sha256:ba4cf57c23f4b4893b0478fac39b0d835a328a3d94fbc350f4828d4ff9b209b6

Observation b8cb2d5b-f9bb-406d-8a2e-200ac2f866fd · outbound

This paper cites SpatialScaper: A library to simulate and augment soundscapes for sound event localization and detection in re- alistic rooms,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 SpatialScaper: A library to simulate and augment soundscapes for sound event localization and detection in re- alistic rooms,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:55.479863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.260325Z digest=sha256:c5a82538dab85b35bc5d4bd9d095c8fd3c790899454a4f59af6daadbe508267e

Observation c64136c5-da74-48c4-865f-95c227379066 · outbound

This paper cites FSD50K: An open dataset of human -labeled sound events,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 FSD50K: An open dataset of human -labeled sound events,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:55.270723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.336148Z digest=sha256:e744e42b3236ce7b66b132d8bf1ae1c0279b743985a25376527190d4190c7de2

Observation 7fe2ea60-7d33-4e04-b7c3-bff7316d9dae · outbound

This paper cites Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:52.750273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:52.750273Z digest=sha256:8ea5018675cd57dba55c4e5c6e8aab9b11e3ab6741b0036e5182d35df86e28bc

Observation 7020296b-b2a9-42cf-b4d9-ecbe937a319f · outbound

This paper cites Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:52.846269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:52.846269Z digest=sha256:5cafb1652453ed0d0e4875880a6f649a935fa3a929ea1312da7e97228e386ed7

Observation a98befa5-111e-4ae3-83a4-b9bf9eb68ec6 · outbound

This paper cites an unresolved cited work.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:56.829497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:52.410628Z digest=sha256:219a6622454bf994d16aebce585fc436096520bc92de8570690d5d64c5870959

Observation 3cfc95c4-4be1-4441-81b5-2830196e37f8 · outbound

This paper cites All audio mixes of 10 seconds long each were generated at a sampling rate of 32 kHz.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 All audio mixes of 10 seconds long each were generated at a sampling rate of 32 kHz

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:57.244612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:52.165942Z digest=sha256:eb67d8e12b21abd4c3dcc55660055eb915edf81fd59e9bafdb2d5eceae4f0c57

Observation ea2b1f06-cf0c-40a7-9cb1-7769dbc84a58 · outbound

This paper cites Classification in the presence of label noise: A survey,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Classification in the presence of label noise: A survey,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:56.231767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:52.958954Z digest=sha256:3610b0013da12553e6534838c14bafd501d8850eae40bea0902de68dd2918c77

Observation 8f659957-e00e-4b53-9744-bb82fbce96c8 · outbound

This paper cites Acoustic classification and seg- mentation using modified spectral roll-off and variance-based features.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Acoustic classification and seg- mentation using modified spectral roll-off and variance-based features

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:55.984683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.071778Z digest=sha256:6a11063a65fefd35fa2c5e7b3bfae3aaf8d638348e07e960c50512e995dcdc62

Observation 9f57c45c-423a-40b3-9400-ff259b0e0f05 · outbound

This paper cites Making chorma features more robust to timber changes.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Making chorma features more robust to timber changes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:55.727487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.153606Z digest=sha256:bc961ec1b801d5b65380a17a7fc9044b38f46fadd9363c4de76a3a3c2006d786

Observation 2b769b2d-e693-43e8-a747-fdbea6bcf71f · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverbera- tion,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverbera- tion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:55.076894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.424366Z digest=sha256:02159bfb65d732936ce1cfc4ef34b3388e179eaa098fda30be461afa28399556

Observation 8c429450-61fb-43f4-bc84-13cd888635b0 · outbound

This paper cites Echo-aware adaptation of sound event localization and detection in unknown envi- ronments,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Echo-aware adaptation of sound event localization and detection in unknown envi- ronments,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:54.896198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.529845Z digest=sha256:818e9e26efcfdb52c0643f9b5b34854818c999f37480136ba9cf70a199b471c9

Observation ad108375-86e0-4399-a657-6c546a9420cc · outbound

This paper cites ESC: Dataset for environmental sound classi- fication,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 ESC: Dataset for environmental sound classi- fication,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:54.800161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.614307Z digest=sha256:d10458be287f4012dcb1cfa91ff0afddab65645b9def7d358779a5132f9572dd

Observation 8be1995d-3557-4b11-9a06-cda8c31233ba · outbound

This paper cites 13, 2025).

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 13, 2025)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:54.613294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.715954Z digest=sha256:6d1607cf146d323c674b67efce199666ef841c34b03f95d12b6b15639c89479f

Observation a93be642-2af9-4dd4-a2f8-3f88dbf1fad2 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Audio set: An ontology and human- labeled dataset for audio events,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:53.792270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:53.792270Z digest=sha256:3ff9da7e739dcac80c8734e3d298454c575a8ad9c28aa5ee6af4fdeacb124f45

Observation 8abe9694-6c7f-4345-a470-a39c478645ad · outbound

This paper cites Masked modeling duo: Towards a universal audio pre-training fram ework,.

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 Masked modeling duo: Towards a universal audio pre-training fram ework,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:54.357859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.838897Z digest=sha256:d25d8c8a9e3cb58db2848c2705d78f3f643becb93c5e40265b62b90fd3844c06

Observation 5b41a0eb-11a9-4b2a-826a-8d0bb5f47fb4 · outbound

This paper cites 13, 2025).

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4 13, 2025)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:54.087029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:53.842471Z digest=sha256:8a3c3ea2f584d5bda7f3602da49cf68b162cf00c009975d2a2681594bee8fc08

Pith citing papers

No inbound Pith citation observations are available.