Pith. sign in

Paper Citation Record · LEDGER

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing

As of 9 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2607.04314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04314 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T20:10:31.150625Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1671982-0e41-4708-9c42-ee79fef06a12 · outbound

This paper cites ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:f0bf370481686268a67de00d4c8c60be8235d63aae8d3c9fbb0f87a5410d6b2e

Observation 338d1d36-6e05-4454-9fbe-7d8361d3d096 · outbound

This paper cites ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:e0b5e79f74f265c93d4b2be08a5c0c5e579735b80a4fb7514be88ea73ac9616f

Observation 66934558-3e34-4119-97f8-dcf13094b027 · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:e69f730a269a5d05999bc9e5a407bd670e69275dbdefa1e464b9f8b98a590403

Observation 17ec56ad-fa4e-4206-ad33-cd2d60c91cba · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:b0c08d89c7ab6da2975b380f99a3b3c9c8d659025ae15d0f704305cf8f0d4bd0

Observation 73b03c90-d1f2-4487-a431-7d5fec2af6dd · outbound

This paper cites AASIST: Audio anti-spoofing using integrated spectro- temporal graph attention networks,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing AASIST: Audio anti-spoofing using integrated spectro- temporal graph attention networks,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:3bcefb56a64fce4797cf7d63ce472b544fe4a125554adb27936ff856fc0a0c33

Observation 3ec2cbbc-704f-45f0-976d-2f682efd0398 · outbound

This paper cites End-to-end anti-spoofing with RawNet2,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing End-to-end anti-spoofing with RawNet2,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:dc3466c5b41f225dbfe61ebd4ae106d8a518b8097b2e609e31bedc3705616dd1

Observation ddc466e6-d6e1-4303-9a19-1961f4cd210c · outbound

This paper cites Attentive merging of hidden embeddings from pre-trained speech model for anti-spoofing detection,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Attentive merging of hidden embeddings from pre-trained speech model for anti-spoofing detection,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:f9bdee88ae63ae8766959f5f5d7d724a7249cf16be6ddc9363c7bad8790fc427

Observation b6ebf520-8efa-48b4-b7dd-528053a32f29 · outbound

This paper cites Audio deepfake detection with self- supervised XLS-R and SLS classifier,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Audio deepfake detection with self- supervised XLS-R and SLS classifier,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:82281dc658e996065fc24be1d4df8f59992c8ccfc5f9c726ff6837159dfbb875

Observation 421bbb64-dbf0-47b4-93ce-f22483319b06 · outbound

This paper cites Two views, one truth: Spectral and self-supervised features fusion for robust speech deepfake detection,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Two views, one truth: Spectral and self-supervised features fusion for robust speech deepfake detection,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:c2a5c232e9bd4c1dacbd8bc9215cc1e9539a6071fb77b68535a7b60fa8539c3d

Observation 1de197a8-1a37-41a6-8e91-a5d50ce7620b · outbound

This paper cites Pitch Imperfect: Detecting Audio Deepfakes Through Acoustic Prosodic Analysis.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Pitch Imperfect: Detecting Audio Deepfakes Through Acoustic Prosodic Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:1aa4c422ba47813f5c87f4800e45dc277dbe57bedab1c29e174937634b0c3f81

Observation d712b68c-2111-4f27-ba96-d047513e3d84 · outbound

This paper cites Contributions of jitter and shimmer in the voice for fake audio detection,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Contributions of jitter and shimmer in the voice for fake audio detection,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:f112deee5d7c49f76f303d05658b6de239ef2ec4b8131f74817eedb0a328bc25

Observation b143e5fc-1669-4f94-b823-a0496130599c · outbound

This paper cites Toward robust replay attack detection in automatic speaker verification: A study of spectrum estima- tion and channel magnitude response modeling,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Toward robust replay attack detection in automatic speaker verification: A study of spectrum estima- tion and channel magnitude response modeling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:008c0622e1200ff5db9a758623fcb45c47f0822f37d68fc4883348e79df0d2ba

Observation e4933feb-17b9-4a24-a782-8793f9a7e2a0 · outbound

This paper cites Domain-adversarial training of neural networks,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Domain-adversarial training of neural networks,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:542d7a1607cc00a665c46eb060a2de5805f8e78ed963eed00409efd8915f40ef

Observation 95bb5885-f758-479f-87ff-b551ff83833d · outbound

This paper cites Focal loss for dense object detection,.

MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing Focal loss for dense object detection,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T20:10:31.150625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:10:31.150625Z digest=sha256:b2e55613991f5ca45b2e73aa408a1a60a38651566c85a27985f04e9be005f9b2

Pith citing papers

No inbound Pith citation observations are available.