Pith. sign in

Paper Citation Record · LEDGER

Scaling up masked audio encoder learning for general audio classification

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.06992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06992 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:56:46.230938Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:39.522452Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b46a13a5-5276-415c-a05e-6a163b4ef094 · inbound

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning cites this paper.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Scaling up masked audio encoder learning for general audio classification

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.230938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.230938Z digest=sha256:bdb5b75da5f1fe3313cd66f32c1e5b15f3b8f9ec5d6a1ee877de64bf4fe4c472

Observation 491785a8-3459-4e14-8b06-35b25e30ec36 · inbound

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification cites this paper.

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification Scaling up masked audio encoder learning for general audio classification

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:44.825761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:44.825761Z digest=sha256:3b57ba995e2fa9dc5b8414f0618bb64b0a3b87b8ed766a1e60509235afa90842

Observation 7c5a6f8e-1d76-45a2-9fad-a1df8ca714bc · inbound

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity cites this paper.

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity Scaling up masked audio encoder learning for general audio classification

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:39.523962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:58:27.702138Z digest=sha256:a53db1170214332fca5db38a9b5a468e15cc78c278f276a1b76813c96ac51251

Observation 93d0354b-deee-47d8-9f72-a4a6b1d755e0 · inbound

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models cites this paper.

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models Scaling up masked audio encoder learning for general audio classification

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:37.228821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:47:53.308909Z digest=sha256:4e5796f45bb3cfb1ef209b982c7ca1558f133f2facfc1648b6244e310de5c85c

Observation faddfb2d-605d-493d-8f29-330f43649b6b · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Scaling up masked audio encoder learning for general audio classification

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:47.112058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:553ffbf1dd3576a9ad02f8fc662e21186e8331199151783216bb57d9b3f3be2a

Observation fe046f72-72e3-4569-a496-4df20a5c0450 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Scaling up masked audio encoder learning for general audio classification

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.845086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:925dad32bf682cc6946bea59d5e94aca86738cd6d7ba5de716260b1a75827979

Observation 3b53fdf3-0866-4d73-9c63-9cb55f6c3304 · inbound

Large Audio Language Models for Spoofing-Aware Speaker Verification cites this paper.

Large Audio Language Models for Spoofing-Aware Speaker Verification Scaling up masked audio encoder learning for general audio classification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:12:30.683133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:12:30.683133Z digest=sha256:ec1c1ea7762d167fa577f92b30fc41ca22cd7b0a0d7b56c927047cd106138f1d

Observation c46ce0a0-409c-4b8c-bad2-cd897e7b0403 · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Scaling up masked audio encoder learning for general audio classification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.218459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.218459Z digest=sha256:9f8191046e009636e06b69af59bac75593e825d119f9ec3291d5d74ef99b3f5e