Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2212.09058.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:05.080883Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T00:04:22.466366Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0fd8edf6-6bcf-447c-a786-d6886f7d1814 · inbound
FAST: Efficient Action Tokenization for Vision-Language-Action Models BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7beb06ff-c36e-4a0f-a60f-e687ac17af26 · inbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44c14ae2-aab0-4e07-b60d-3d12d1b66ed7 · inbound
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df0c23c-75af-4104-9e04-af14abb2e49d · inbound
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ee79de-2d55-4c4a-9c8b-f63aba10c8d7 · inbound
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2b0f6c-cf99-4bb2-b38e-37f4ea4a29c6 · inbound
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39d2e9c-6acf-4e67-9491-54518e61f564 · inbound
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9917f026-aa6d-414f-b692-c025fa093fe7 · inbound
Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · inbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043fba0e-52af-41da-b7cc-98f402ff4421 · inbound
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc072def-b9ec-4f42-bd4c-f70984df4af5 · inbound
Step-Audio 2 Technical Report BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d32b479c-3c7f-4f03-9026-72d988877b7b · inbound
Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7251430-0160-4158-b836-a7496021ba59 · inbound
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc7d1e3-6c86-42ca-a4ee-6686936fc6ae · inbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22662b78-6232-4296-82c8-475b2160c9ad · inbound
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9463a595-eb51-4039-85fc-e466eb461fa9 · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 147
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d40ba64-a9e4-4823-8020-9045121195db · inbound
AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a4edce-cff7-4396-89c3-402c0ff4e7df · inbound
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8b3450-5551-4e9d-97c3-ecd45d3e549a · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc38e2c6-390a-4c8f-b3a3-9680f939ff1f · inbound
Assessing Factual Music Comprehension in Large Audio Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 560dcc23-698d-4244-813f-07ce1600afe7 · inbound
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e716b4b-02e4-436e-8cbf-b5b1995e2786 · inbound
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee47b9f-1b44-48b2-b38b-b9af80c234d7 · inbound
Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c87b1f-9aa7-4e45-826a-1b53a4dda16e · inbound
ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0e867ad-df14-4bcb-a331-78aa73ff8341 · inbound
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc82c1a8-7dd8-41eb-ae7c-23fcade130d1 · inbound
TinyMU: A Compact Audio-Language Model for Music Understanding BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b07ddca6-efdf-41f7-bd4e-e639700fdbb5 · inbound
MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0570288-70ac-4bfc-8f51-75240f384f19 · inbound
MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c413f0d0-e26d-49e0-803b-013725d60ad6 · inbound
Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 291
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5721e150-1e14-4978-a682-d84a42bec402 · inbound
Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 137c5f66-8d64-444b-85bf-8641c9512bbe · inbound
SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd003016-04bf-488f-a810-4d272627b8f3 · inbound
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87328366-cfcc-4fad-a0ae-6cf828e5d1b3 · inbound
Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2dcc5b3f-0b7e-46cd-947b-6663208e311c · inbound
Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7a930f7-f2ed-43c5-8098-c40fa06e921e · inbound
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9834dc0-8da9-4b02-8df1-75ca93cd477b · inbound
Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ea85c91-d981-4e74-9f00-cad0e7b98f56 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 274
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb69bebd-0691-40d1-99ec-d89a0164e2c9 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 274
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35666e69-0923-4e17-9b60-173abe7283e8 · inbound
FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11f819c1-a7d3-476d-9f32-67771820b54c · inbound
Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b447bf8-04ab-4d19-b867-a62cfef7b3a3 · inbound
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cdad4a4-601b-452a-afb4-4458397ec441 · inbound
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1606f0-2d5d-4d48-9dd8-a5d38a722a5e · inbound
Hidden-Domain Routing for All-Type Audio Deepfake Detection BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.