Pith. sign in

Paper Citation Record · LEDGER

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2601.22504.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22504 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:20.739422Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:17.487287Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df17951e-f422-4499-b041-d7774b497274 · outbound

This paper cites Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:17.487287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:17.487287Z digest=sha256:7498809d45ba65ee9096404678947f697ad0230f8ab1bf915cd35dbe253e5fb8

Observation 65b18697-f10a-4607-ae06-dbccec36f949 · outbound

This paper cites an unresolved cited work.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:17.550821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:17.550821Z digest=sha256:29410d1d861b809451c21bdedec289e10763affa5dd7dad46360afefdad47ef3

Observation 4c2f6663-d5fb-4234-b9c3-161b8a3704ea · outbound

This paper cites Audio tagging model Figure 3 illustrates the modified M2D AT architecture.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Audio tagging model Figure 3 illustrates the modified M2D AT architecture

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:17.664736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:17.664736Z digest=sha256:00b3e8704747a77461435f774b7087bf15980e8c79c630d2dde595831e899501

Observation 7cca14be-a78e-43da-bdbc-82d7d28ab25a · outbound

This paper cites an unresolved cited work.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:17.884405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:17.884405Z digest=sha256:288f9b93c88fa92af7393b9c2ca310912b72d2c671c4387672d9448df41d8d24

Observation 38046b06-c0b4-4a43-865a-448f1ce6c98e · outbound

This paper cites The order of input labels is also used to align the estimated and reference sources to calculate the SDR loss function.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources The order of input labels is also used to align the estimated and reference sources to calculate the SDR loss function

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:17.769959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:17.769959Z digest=sha256:553e9602bfd359c396c2b2f221af3aa2cf1382770f219230867cafbec81aa8e5

Observation b39f801c-a3f1-42b3-9168-194a263df96e · outbound

This paper cites We also propose an evaluation metric to address the confusion caused by duplicated labels.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources We also propose an evaluation metric to address the confusion caused by duplicated labels

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:17.993759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:17.993759Z digest=sha256:da8ee288fe5a14ed6af2fc3a02b677b67026a5f4b86a5863faeb23b9806e1fc8

Observation 70fcb112-d350-4735-a88b-fb96e8e3f3c1 · outbound

This paper cites Immersive voice and audio services (IV AS) codec-the new 3GPP standard for immersive communication,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Immersive voice and audio services (IV AS) codec-the new 3GPP standard for immersive communication,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.134729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.134729Z digest=sha256:940a5c2062c4b0b2446f624749a73ed50b4f4895e30a2e626a594906eb065c34

Observation 4a1ea614-8ca6-4ccd-bff5-6581c8ad4a13 · outbound

This paper cites Metadata-assisted spatial audio coding in IV AS codec,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Metadata-assisted spatial audio coding in IV AS codec,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.248977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.248977Z digest=sha256:cf25f266b276e2128624f2d2f904930810cc83d585311764830d27cd95f40952

Observation b1c2e27f-175e-413e-ad07-280f89522e1c · outbound

This paper cites 3GPP IV AS codec–perspectives on development, testing and standardiza- tion,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources 3GPP IV AS codec–perspectives on development, testing and standardiza- tion,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.372348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.372348Z digest=sha256:6db6e10cbd544526e5392c2ce93c59541544bfc0470bda256d2475273311a6b2

Observation 9eebb139-8628-4781-bb7d-968c0db505bf · outbound

This paper cites De- scription and discussion on DCASE 2025 challenge task 4: Spatial semantic segmentation of sound scenes,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources De- scription and discussion on DCASE 2025 challenge task 4: Spatial semantic segmentation of sound scenes,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.474971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.474971Z digest=sha256:e3ea3f507d79b4a6eac5ec5c9c5de386674ecbe4bd0b5691a9d963e5427f502c

Observation b7fedba4-7ca3-4b1a-a82f-33afe246e73d · outbound

This paper cites Baseline systems and evaluation metrics for spatial semantic segmentation of sound scenes,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Baseline systems and evaluation metrics for spatial semantic segmentation of sound scenes,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.590809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.590809Z digest=sha256:386de6825a944a00a8a6f9a2688c4af4f83aac032af1a27eef80404ff71c9672

Observation 16e735ce-66e2-4034-8922-45df4c29427f · outbound

This paper cites Transformer-aided audio source separation with temporal guidance and iterative refine- ment,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Transformer-aided audio source separation with temporal guidance and iterative refine- ment,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.702728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.702728Z digest=sha256:9b1193df117345dede0e880dc4c7c14eeaf8c8bef99769b639360a3df40c72bc

Observation d73198d3-15e2-4437-a855-5935ef0ca366 · outbound

This paper cites TS-TFGRIDNET: Extend- ing tfgridnet for label-queried target sound extraction via em- bedding concatentaiton,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources TS-TFGRIDNET: Extend- ing tfgridnet for label-queried target sound extraction via em- bedding concatentaiton,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.880711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.880711Z digest=sha256:84586e098c4efe72774833107848316da005fd05f36ebbd8ca3c3bb69c1d44f2

Observation e444fa05-21fc-4e3e-b038-34da565f1760 · outbound

This paper cites Self-guided target sound extraction and classifi- cation through universal sound separation model and multiple clues,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Self-guided target sound extraction and classifi- cation through universal sound separation model and multiple clues,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:18.983090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:18.983090Z digest=sha256:289ae6da9e23eaa1d9a6dd2ae41800f92582b9f9cd1f3a5e9e6db7bed0e3ef96

Observation a6d564e4-2b65-4a51-9f9f-06530cec6192 · outbound

This paper cites Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.098403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.098403Z digest=sha256:a96eca93d8cea4974d645718a9087edafe177a88873fb52509608801aaf8018d

Observation 6dd41bcd-db68-4c09-b3d3-e743f947eb89 · outbound

This paper cites REDUX: An iterative strategy for semantic source separation,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources REDUX: An iterative strategy for semantic source separation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.245518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.245518Z digest=sha256:5c3d58c95dbc334c3c4ecf8e8d0522376526c995baeee9697a08a3aef345c43f

Observation 8619264d-b833-4c9d-b75c-8cb548be150a · outbound

This paper cites SJTU-AUDIOCC system for DCASE 2025 Challenge Task 4: Spatial semantic segmentation of sound scenes,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources SJTU-AUDIOCC system for DCASE 2025 Challenge Task 4: Spatial semantic segmentation of sound scenes,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.364114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.364114Z digest=sha256:7bc772557127cae285f041635165b29895a04189a6450332140b9e6ce2f03308

Observation 4bc805f5-f82a-441c-9040-8ed6c3dc7bad · outbound

This paper cites A hybrid S5 system based on neural blind source separation,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources A hybrid S5 system based on neural blind source separation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.473441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.473441Z digest=sha256:5c7c710d2741c58c1ca8096cc08bef48a4197cacfe0a6fd0c9a7cc2bede12676

Observation 8ffe3599-aa96-4857-b891-baffa8136112 · outbound

This paper cites Permutation invariant training of deep models for speaker- independent multi-talker speech separation,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Permutation invariant training of deep models for speaker- independent multi-talker speech separation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.585308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.585308Z digest=sha256:cc6bd29db2cf39f69fec57870186df3acd21e55d6960c0b377145c10cb2eef83

Observation 99c8fcb4-894b-4ea9-ba6c-b58640434f0b · outbound

This paper cites An improved event- independent network for polyphonic sound event localization and detection,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources An improved event- independent network for polyphonic sound event localization and detection,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.701734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.701734Z digest=sha256:6774170634d4b71079845aad4ced382fe1ea0024fa92d6be17529a9160d2d1cc

Observation aaf2b0e6-0bd9-4906-a07f-e41aa50b4b39 · outbound

This paper cites Zero-and few-shot sound event localization and detection,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Zero-and few-shot sound event localization and detection,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.807464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.807464Z digest=sha256:72f9e5e1200ea0a3e8ea5240fe8b4cc7ab656d5ecac565c0c7c04e0262e5ddde

Observation 1ffdbdc7-66cb-4958-85a0-4d2e491fd173 · outbound

This paper cites Cross-attention inspired selective state space models for tar- get sound extraction,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Cross-attention inspired selective state space models for tar- get sound extraction,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:19.922593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:19.922593Z digest=sha256:02d9932eb3ed4790779cd293eaea492a59c9c852fcd2172381809cd99acb9a12

Observation 8d5ecf20-02aa-4235-851d-9524e696a0ee · outbound

This paper cites Real-time tar- get sound extraction,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Real-time tar- get sound extraction,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:20.027080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:20.027080Z digest=sha256:8b008e902ef6d146190e82bddd78122f56a1586d2ee3bf1b56e583b94b8362f5

Observation b8757b51-e071-44ae-a35d-84c4cbfaf12e · outbound

This paper cites Soundbeam meets M2D: Target sound extrac- tion with audio foundation model,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Soundbeam meets M2D: Target sound extrac- tion with audio foundation model,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:20.136821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:20.136821Z digest=sha256:6d01da0ba317be291d791b8c674175f24212ef77b8feec14b51f7bc97232675c

Observation 590e13b6-74f1-417f-9b70-0e66f52297ec · outbound

This paper cites Universal Source Separation with Weakly Labelled Data.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Universal Source Separation with Weakly Labelled Data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:20.258010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:20.258010Z digest=sha256:39245f387fe91a3f8c587b61b7f6d232fe59ae59f7a95a0d15ca82d1a3584cf4

Observation 19104afa-571e-43d3-838d-987dea9cf1a0 · outbound

This paper cites Masked modeling duo: Towards a universal audio pre-training framework,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Masked modeling duo: Towards a universal audio pre-training framework,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:20.378620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:20.378620Z digest=sha256:138fe00deb756c5a07eead0af2d81e3ed32365d22a75f7961383fd9362a9e548

Observation 891b0184-3843-4204-820c-950a97eecebb · outbound

This paper cites Spatial scaper: a library to simulate and augment soundscapes for sound event localiza- tion and detection in realistic rooms,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Spatial scaper: a library to simulate and augment soundscapes for sound event localiza- tion and detection in realistic rooms,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:20.483496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:20.483496Z digest=sha256:e321ffb7f2ec734f461b1400e076fb232c7699d2f4965f98ef29548017fc1374

Observation 1981802e-c176-4a14-89e5-51a8af57a1ea · outbound

This paper cites Semantic hearing: Programming acoustic scenes with binaural hearables,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Semantic hearing: Programming acoustic scenes with binaural hearables,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:20.620920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:20.620920Z digest=sha256:3974e2312c397b25d9d2f56a6e79175ba216cd4d9c4ae81b77375efab22f0b38

Observation 33f01359-b034-44b3-bcdc-a1c9150edb49 · outbound

This paper cites Echo-aware adaptation of sound event localization and de- tection in unknown environments,.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Echo-aware adaptation of sound event localization and de- tection in unknown environments,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:20.739422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:20.739422Z digest=sha256:a7c04814010e7d3665f75f94ad9b3ded3c92914905eedfb6a1ecc97354106c38

Pith citing papers

Observation df17951e-f422-4499-b041-d7774b497274 · inbound

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources cites this paper.

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:17.487287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:17.487287Z digest=sha256:7498809d45ba65ee9096404678947f697ad0230f8ab1bf915cd35dbe253e5fb8