Pith. sign in

Paper Citation Record · LEDGER

BEATs: Audio Pre-Training with Acoustic Tokenizers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2212.09058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.09058 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:05.080883Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.466366Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fd8edf6-6bcf-447c-a786-d6886f7d1814 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:52:31.926052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:d528f1345131ff89a84236baa357864e490f54588ff8ada8cbbcd11a57d248ca

Observation 7beb06ff-c36e-4a0f-a60f-e687ac17af26 · inbound

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance cites this paper.

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:15.772955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:15.772955Z digest=sha256:d253ad5c81656ead06e0059ef74903391166238bee700d07abde73029e6ba482

Observation 44c14ae2-aab0-4e07-b60d-3d12d1b66ed7 · inbound

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English cites this paper.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:05.080883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:05.080883Z digest=sha256:c88d153c5fe1e6fa5816969673186a85315681ca7ec14092f88df830427271f2

Observation 4df0c23c-75af-4104-9e04-af14abb2e49d · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:10.885941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:10.885941Z digest=sha256:5b9adae549bd0de951898ef914d04a734b31220cb65cac31e251445afa4d637c

Observation e0ee79de-2d55-4c4a-9c8b-f63aba10c8d7 · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:48.281447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:48.281447Z digest=sha256:82b4b0a955f222152337d8f96ef4bf8922dea937d2a5eb3dca1108738457ec2c

Observation 1c2b0f6c-cf99-4bb2-b38e-37f4ea4a29c6 · inbound

Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization cites this paper.

Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.632711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.632711Z digest=sha256:04b7589afeb86085a57028708ba48cc7c7896f1eebfa65d551d277943614d1e0

Observation d39d2e9c-6acf-4e67-9491-54518e61f564 · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.677614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.677614Z digest=sha256:0a585368a5fbcd60997731e923c00bbb0ba015bf714706a1335b46b5a3b40be3

Observation 9917f026-aa6d-414f-b692-c025fa093fe7 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:10.747215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:10.747215Z digest=sha256:1b9df269bf0fdb6354a33925d83f827f647d9886a7a78b22b0dfe6327ee21db4

Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.348025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.348025Z digest=sha256:43530af3edc7ef31e9f17730e2bf6ca657b77292c97f5b1ef69ee28f700ac472

Observation 043fba0e-52af-41da-b7cc-98f402ff4421 · inbound

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine cites this paper.

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:34.333552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:47:34.333552Z digest=sha256:e78d32dee4c273a63f4a454b79b2a3c53a5f4c5b638bc2bfa17bf985524b5623

Observation cc072def-b9ec-4f42-bd4c-f70984df4af5 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:50.976241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:217c6603a78b010251d22f260acec03d78a9e5ee3477ad106d96f0d5ec63fe61

Observation d32b479c-3c7f-4f03-9026-72d988877b7b · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.509860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.509860Z digest=sha256:33cc85a5142e216f30eae612ae8c265a211cafd4be38d693a409d5649fc55f13

Observation f7251430-0160-4158-b836-a7496021ba59 · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.678062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.678062Z digest=sha256:b374ae91d624a1246d5430ecc82833a7b2b79dbbf3af896c8359acf4083cb535

Observation 0fc7d1e3-6c86-42ca-a4ee-6686936fc6ae · inbound

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning cites this paper.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:17.525381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:17.525381Z digest=sha256:2b7f6b3db0dd7bdbd559447aa912ec715a01d1e163f2ea585c1fe224b8fb2a08

Observation 22662b78-6232-4296-82c8-475b2160c9ad · inbound

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation cites this paper.

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:54.793915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:54.793915Z digest=sha256:28098c739ae5adadd1a7b9cdf4065bef4f0ea7a441150666411c4e70a9ea3e2c

Observation 9463a595-eb51-4039-85fc-e466eb461fa9 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:18.136571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:18.136571Z digest=sha256:b002206f0d46d1383a2daea7d49285dbdd4286f2d9af61c2d5ec3baf7130f763

Observation 4d40ba64-a9e4-4823-8020-9045121195db · inbound

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation cites this paper.

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:32.271114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:57:32.271114Z digest=sha256:0dd6de199f5d136989cc0f100a853bd870406fd2374e77da00859c4f1170833f

Observation 88a4edce-cff7-4396-89c3-402c0ff4e7df · inbound

VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results cites this paper.

VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:17:54.443943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:17:54.443943Z digest=sha256:82f6be5ff41c0f7700e243f478439697ee1e72f3c6579946f4ab2a1979b7e61f

Observation be8b3450-5551-4e9d-97c3-ecd45d3e549a · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.654795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.654795Z digest=sha256:596b3d0c85fb503c0a6c913c9d620aaf6ab908037c6b526b0573d3947b49244a

Observation dc38e2c6-390a-4c8f-b3a3-9680f939ff1f · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:45.976102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:45.976102Z digest=sha256:c26906f216775a8ed25563092917e9d9fd417dbd3ecaabbec580962fb5e36531

Observation 560dcc23-698d-4244-813f-07ce1600afe7 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:47.142235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:47.142235Z digest=sha256:bdddb2d81fd9c0826d7007f39836dd0ecbe3dd88f5e826ce939b23aaf07dfae1

Observation 3e716b4b-02e4-436e-8cbf-b5b1995e2786 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:36.790900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:36.790900Z digest=sha256:63e6a7b3b1caeeab209da0bed2d76d3bdbdd3bccaa6de547f32795815fc4c5dc

Observation 0ee47b9f-1b44-48b2-b38b-b9af80c234d7 · inbound

Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection cites this paper.

Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T10:53:15.264765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:53:15.264765Z digest=sha256:0c016d2eaf327bc6d982ed9a7d8b5f6c05b58586247ecd02ea17478bdf4dc918

Observation e8c87b1f-9aa7-4e45-826a-1b53a4dda16e · inbound

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals cites this paper.

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:30:50.864486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:30:30.431777Z digest=sha256:d8af81f8869eaa4da8eae261db7407d19ee632cbf033b8f7183a9f1100acdccf

Observation f0e867ad-df14-4bcb-a331-78aa73ff8341 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.946896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:6082190ae639afb7ff141f319a02503b035ae0f7bd7436360a6bedfec5c3728b

Observation bc82c1a8-7dd8-41eb-ae7c-23fcade130d1 · inbound

TinyMU: A Compact Audio-Language Model for Music Understanding cites this paper.

TinyMU: A Compact Audio-Language Model for Music Understanding BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:17:36.935059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:17:23.740979Z digest=sha256:1d33b2c46e33e15434d6c4594df817163f8081d445706a99814f01ac38f1939b

Observation b07ddca6-efdf-41f7-bd4e-e639700fdbb5 · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.191807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:20:22.200577Z digest=sha256:20aaa130b2955f7d04260ad4041e1030844eefdd2464518c1b4d3f8d9aa33fd8

Observation a0570288-70ac-4bfc-8f51-75240f384f19 · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:49:19.547770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T00:45:01.115573Z digest=sha256:24ba4faa53571526e05c04f08dc0232d40c2855663a701fafe5039a97c32b8e1

Observation c413f0d0-e26d-49e0-803b-013725d60ad6 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 291

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.124351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:3a59adf854a01cf5229067b708ed516a7a22e5e48a9ff8957057ed594ef79990

Observation 5721e150-1e14-4978-a682-d84a42bec402 · inbound

Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations cites this paper.

Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:01:18.489162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:53:11.185208Z digest=sha256:dac89104525b7c189f06c4b871ba209d521e2c53f316d218aa92975e3564df86

Observation 137c5f66-8d64-444b-85bf-8641c9512bbe · inbound

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification cites this paper.

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:57.059110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:31:49.239866Z digest=sha256:e4760ea7b575c0e4cac9fd7d534abbc2dc86675e25652dbad933738b977ef67a

Observation bd003016-04bf-488f-a810-4d272627b8f3 · inbound

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation cites this paper.

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.563590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T18:08:50.574960Z digest=sha256:e7a9d4f3d666845e714baa1204e27e8078867450441ab4aff2e4795b1427632a

Observation 87328366-cfcc-4fad-a0ae-6cf828e5d1b3 · inbound

Finding Needles in the Haystack: Transductive Active Labeling in Ecology cites this paper.

Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.405444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:52:23.274642Z digest=sha256:ad72679c86bfe80b6b2de8eab34ebc93211dc61378639af1ddee8e8e565673e7

Observation 2dcc5b3f-0b7e-46cd-947b-6663208e311c · inbound

Finding Needles in the Haystack: Transductive Active Labeling in Ecology cites this paper.

Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:55:31.231791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T07:48:01.274555Z digest=sha256:ee1f1f2d6850b16b21d65835263ed752e22b6072961f7e04282293f4b1587719

Observation f7a930f7-f2ed-43c5-8098-c40fa06e921e · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.917414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:dab265bb63c570aed43be745627be70a58364497d66ae99065d2252a5b0fcc85

Observation e9834dc0-8da9-4b02-8df1-75ca93cd477b · inbound

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types cites this paper.

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-07-12T03:24:13.814557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:24:13.814557Z digest=sha256:dbadf3938ce159b873207baf082e528fe4e9b7362e1ee8f85658171e27699941

Observation 5ea85c91-d981-4e74-9f00-cad0e7b98f56 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 274

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.467634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:ee7578eb19268591386a5f35012e7d191a12943ad7d93287636ea7fbe3486450

Observation fb69bebd-0691-40d1-99ec-d89a0164e2c9 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 274

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:3bc3697ac5a38a86b2de70e09513cf46c7d903d31cda738c35979b40f9e06924

Observation 35666e69-0923-4e17-9b60-173abe7283e8 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:041cf4877d36ec7db83bea1364c4efb81b745ad1bfdff40f94493d39bf312af4

Observation 11f819c1-a7d3-476d-9f32-67771820b54c · inbound

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 cites this paper.

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T02:03:37.761113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:03:37.761113Z digest=sha256:999d26d36286f89edfe87bd84e18d66ff36dcb25b1b6f5a80193fed7e597c475

Observation 9b447bf8-04ab-4d19-b867-a62cfef7b3a3 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T10:35:03.133791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:35:03.133791Z digest=sha256:40cfbd13cb0126e268ed6486598b697d27c4e326885862da7dacb544fb002fad

Observation 3cdad4a4-601b-452a-afb4-4458397ec441 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.418732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.418732Z digest=sha256:44f5854715931b76b4dff3702ad9985de404cf3609b7c4442ee8977769f7249b

Observation 6b1606f0-2d5d-4d48-9dd8-a5d38a722a5e · inbound

Hidden-Domain Routing for All-Type Audio Deepfake Detection cites this paper.

Hidden-Domain Routing for All-Type Audio Deepfake Detection BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:45.434431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:45.434431Z digest=sha256:441ea272b5c1c6a1ac61df3129d87043d51ba0cb6a700dff9b92c669f002f28b