Pith. sign in

Paper Citation Record · LEDGER

AudioMosaic: Contrastive Masked Audio Representation Learning

As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2605.14231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14231 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T01:52:01.164694Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact8
  • verified fuzzy5
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 961a22ce-e646-4a7a-b214-b8478e94086f · outbound

This paper cites Optimizing Audio Augmentations for Contrastive Learning of Health-Related Acoustic Signals.

AudioMosaic: Contrastive Masked Audio Representation Learning Optimizing Audio Augmentations for Contrastive Learning of Health-Related Acoustic Signals

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.881087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:03bc1916c769a0c50eb95a550640b3e242da595ec049bd50a95e9616155be19c

Observation c5a1aaf1-6513-4c58-8e53-97280eaa76c2 · outbound

This paper cites A-JEPA: Joint-Embedding Predictive Architecture Can Listen.

AudioMosaic: Contrastive Masked Audio Representation Learning A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.868746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:3cdeb4fb93675ab86e53b6bed55a68a557cb53fc5090cf5970c126e3f89a3aca

Observation 5ce69e51-48d6-495d-9ed8-3da424fc931d · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

AudioMosaic: Contrastive Masked Audio Representation Learning Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:42:46.422993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:5022ed9853cbdb57963cdee65c40957a07f1db360782f074ca7825810bd65174

Observation 6ee27fac-4b9b-460a-ac4c-7ef2b50795d1 · outbound

This paper cites Ast: Audio spectro- gram transformer.

AudioMosaic: Contrastive Masked Audio Representation Learning Ast: Audio spectro- gram transformer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T08:20:18.430043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:cac270dcd4838ef597f7790e081df895e39ca02338944cf18c5601f25222531d

Observation 0b14bd4e-3dfa-4ee1-a336-b0e809aa504a · outbound

This paper cites Sheet: A multi- purpose open-source speech human evaluation estimation toolkit.

AudioMosaic: Contrastive Masked Audio Representation Learning Sheet: A multi- purpose open-source speech human evaluation estimation toolkit

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T08:20:18.434392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:34e50e37396b98bb1b91c4a9c82ce3e63ea49feef6658ee754302ff9c9a25015

Observation ec1072ce-d1cb-4a1e-97ae-f072a79f6845 · outbound

This paper cites AHELM: A Holistic Evaluation of Audio-Language Models.

AudioMosaic: Contrastive Masked Audio Representation Learning AHELM: A Holistic Evaluation of Audio-Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:53:28.861492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:fb58e025b700124ef76daf8afef6265e16a27968d6efae18dd3bc79b264ae700

Observation 79cee91e-fdd0-429e-bf36-7ddac2a43092 · outbound

This paper cites Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech.

AudioMosaic: Contrastive Masked Audio Representation Learning Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.877532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:df2c5c65ef52e5cafe9e36a7ee13d17df03fb2689cb33e2fa3180cf7db49ca47

Observation 6cacd83a-3e74-4d9b-8c46-8adfe7ebd0a4 · outbound

This paper cites Acoustic scene classification: an overview of dcase 2017 challenge en- tries.

AudioMosaic: Contrastive Masked Audio Representation Learning Acoustic scene classification: an overview of dcase 2017 challenge en- tries

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T08:20:18.436530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:8062ea958f052b154190b54e67beb83b111f652e6648f2e217b5a660111ea6dc

Observation fafeee34-0fbe-40bf-8a46-52e1c9d03d0d · outbound

This paper cites Masked Spectrogram Modeling using Masked Autoencoders for Learning General-purpose Audio Representation.

AudioMosaic: Contrastive Masked Audio Representation Learning Masked Spectrogram Modeling using Masked Autoencoders for Learning General-purpose Audio Representation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.865087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:b3f9b9167ca4186ad24962c4af33366d6b673fa873272d7d9664a01bb0b88207

Observation f46b6721-cd00-4a59-a4a9-22fb9796d812 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

AudioMosaic: Contrastive Masked Audio Representation Learning Representation Learning with Contrastive Predictive Coding

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:53:28.857982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:f29909300952c75ac36059485f18b313a2d556e18bbdb9b1903cf268bc2f3e4c

Observation 3488ce68-9a17-456d-a8e6-4a18a5351dae · outbound

This paper cites The kaldi speech recognition toolkit.

AudioMosaic: Contrastive Masked Audio Representation Learning The kaldi speech recognition toolkit

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T08:20:18.432079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:46fdbce5676b82b44364623f0941a0280478ef5cd7a8bb1ce3abb6ce500cf850

Observation 05f792ba-a12e-4388-9066-1c2db1efbf8a · outbound

This paper cites Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification.

AudioMosaic: Contrastive Masked Audio Representation Learning Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:52.993424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:a83299a14eb990411bcac0551f9c8408faf42952650980d4f783d83a12c9bd01

Observation 9ddf441b-6201-4d9d-b4e6-7383bdb30365 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AudioMosaic: Contrastive Masked Audio Representation Learning LLaMA: Open and Efficient Foundation Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:53:28.871557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:caf2a3e23cf35ced9345f2c2e719e724dca1c0f78c40f3b429f24c2851e664a1

Observation 0e4ca1b0-23e7-40b3-8f73-2e802a075736 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

AudioMosaic: Contrastive Masked Audio Representation Learning Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:53:28.874553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:dd0b83bc2a3b48ba5296c9680c1d9607905721d39f5c498d609e97ea160fcd11

Observation 4bcd3ac3-e213-4e36-8687-991e32b4b297 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

AudioMosaic: Contrastive Masked Audio Representation Learning SUPERB: Speech processing Universal PERformance Benchmark

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:53:28.854680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:e31b73216e7fc081e7d60ac62b5db3e7231c0e4d9b71da24a225adf86b7e4cb5

Observation f74a1083-df1d-4a18-843b-2412e7f37e7f · outbound

This paper cites Esdd 2026: Environmental sound deepfake detection challenge evalu- ation plan.

AudioMosaic: Contrastive Masked Audio Representation Learning Esdd 2026: Environmental sound deepfake detection challenge evalu- ation plan

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:53:28.851304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:1c57dd98ad1c336f1aceab2169986aa62e0a1fc4c5ec0fa21b121dc8c75b6963

Observation 69e99cf6-99e5-4429-80a3-8aa52bf612cb · outbound

This paper cites Experimental Settings Table 6 summarizes the detailed experimental settings for both pre-training and fine-tuning, while Table 7 reports the settings used for linear probing.

AudioMosaic: Contrastive Masked Audio Representation Learning Experimental Settings Table 6 summarizes the detailed experimental settings for both pre-training and fine-tuning, while Table 7 reports the settings used for linear probing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T08:20:18.425788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:f6df8d5719e86d84d32b26953bc105555ef753ff7613417924fed7fc79f2d07d

Observation 3bd612a9-a124-47e2-a4b3-9ae0a2464653 · outbound

This paper cites an unresolved cited work.

AudioMosaic: Contrastive Masked Audio Representation Learning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-15T08:20:18.427948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:a151912ed925524cf3536a0fc86d3eeb78729e07013212c22ad49169681ebaf9

Pith citing papers

No inbound Pith citation observations are available.