Pith. sign in

Paper Citation Record · LEDGER

BEATs: Audio Pre-Training with Acoustic Tokenizers

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2212.09058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.09058 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:51.536713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.466366Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7df371f9-bab8-4f0b-931c-14d381ef70e5 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.096118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.096118Z digest=sha256:d35bc5ea152354cf290b9d927d9831045b04098f4522ceafe42ae6168352b014

Observation c82c3ff4-178a-45ab-a3cd-dad617685b13 · inbound

State-Space Large Audio Language Models cites this paper.

State-Space Large Audio Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:44.590645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:44.590645Z digest=sha256:b872087c6ecc35839ab63689b708da1bb0e26a450f0f9d253273ad75ad1ecfc1

Observation bf084925-529f-451e-b653-15e588c9573e · inbound

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos cites this paper.

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:53:34.381967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:53:34.381967Z digest=sha256:57794465627cdf0a61c0eae936050f9b5a633b42bb6ea0a0b35bf467cd2a0cfe

Observation 54cd0d13-dd9b-435f-bac3-40338c9c69be · inbound

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning cites this paper.

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:00.267227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:00.267227Z digest=sha256:8b96b08ebd11b1d4ee4ee587af2e379fcaf778396142ae37a947539e287e090f

Observation 60e89ec7-58fd-4021-82b5-0f9b3b91db84 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.101143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.101143Z digest=sha256:09cb3a458845eb035d7bfdc9e054106b00069377b809d29a751dcacd2ed54d69

Observation 5ddfea5d-024b-46a4-b01a-fb5db9d6ef34 · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.113036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.113036Z digest=sha256:3c5f8dd808ac05ee7bb245e6e08a499255fcdb5e8c3a879599eab7433e255d5f

Observation 402c1229-3f30-49cb-aea7-8d4a753d7440 · inbound

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch cites this paper.

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:35.812792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:35.812792Z digest=sha256:d719a68a175e9b3d8452a30f7fecde28284c8d96de3922b8232bdfc459f51602

Observation 07018219-5482-4790-8ca2-1dce466a6770 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.369162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.369162Z digest=sha256:8bdfd0adf36ad423d45e3daa1a1b04055eb9b835f950e83b7c1a04200229eca7

Observation e1ca84c6-cd44-4407-8e80-a44132e52880 · inbound

SoundBrush: Sound as a Brush for Visual Scene Editing cites this paper.

SoundBrush: Sound as a Brush for Visual Scene Editing BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:52.130396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:50:52.130396Z digest=sha256:5a515cd9f55c4b383da66c11e82cccd31e8f53bbb1990e77ce60749579390fd4

Observation 0fd8edf6-6bcf-447c-a786-d6886f7d1814 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:52:31.926052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:b00a87da6b1d98e5ad47753580832a85d563731fea52c390504cd93bbc15cfe2

Observation 2d4cb8e7-ccf8-4152-aca4-f65ac25662b3 · inbound

Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning cites this paper.

Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:28.706957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:27:28.706957Z digest=sha256:0d843cec7162f4d6ad33b5f4967eab7e7bc775691afc3747006512f2aa10472e

Observation 21115587-0540-47a4-9271-0a6430d890cd · inbound

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models cites this paper.

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T18:05:52.235239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:05:52.235239Z digest=sha256:94b9ffbb7895d6edbf1d867bd5704c69055f4add089f17ff964ff3dadb2b76a6

Observation 9ba18c47-ad51-489b-a2ad-5cbd56f2e8e3 · inbound

Evaluation of Deep Audio Representations for Hearables cites this paper.

Evaluation of Deep Audio Representations for Hearables BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:47:07.989640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:47:07.989640Z digest=sha256:743e7e1541150324a2f09321fe918dc312cba89aae9e725b2b140b7730b947b0

Observation 8b18d267-df6c-45bf-ae1c-cba8b8fdc570 · inbound

Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing cites this paper.

Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:51.536713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:51:51.536713Z digest=sha256:88d818c6aa15483de3731762623e605cee3f266776bd106c3055b2fc5bffac36

Observation a362eff2-7e00-406b-9f0a-7a568ed037f8 · inbound

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding cites this paper.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.578676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.578676Z digest=sha256:a32dddaa1cdcd0692fdac907c02f6b3abc428b773c22c268ab3ecea08351f663

Observation 7beb06ff-c36e-4a0f-a60f-e687ac17af26 · inbound

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance cites this paper.

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:15.772955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:15.772955Z digest=sha256:be9b8d513c4b9585d6223f860393e5332e222aa68305fc8646d529bc966cd408

Observation 44c14ae2-aab0-4e07-b60d-3d12d1b66ed7 · inbound

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English cites this paper.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:05.080883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:05.080883Z digest=sha256:c0118614df37e2a4ac1aedf30d67edb9cec32447ed7dc436e7a758ec6ae1385f

Observation 4df0c23c-75af-4104-9e04-af14abb2e49d · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:10.885941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:10.885941Z digest=sha256:22bcee128ba5d1c30bbe26fddac3bb3aaf73c0a439dacf5d159839927526ce75

Observation e0ee79de-2d55-4c4a-9c8b-f63aba10c8d7 · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:48.281447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:48.281447Z digest=sha256:0cef4f76dca9ad43c2389c612f951d03ef537bcba2b47793381119a4a73eba6c

Observation 1c2b0f6c-cf99-4bb2-b38e-37f4ea4a29c6 · inbound

Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization cites this paper.

Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.632711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.632711Z digest=sha256:4180ee9230f71948d15b271c8822bde32c429c02842426066f8728f0a0567083

Observation 01302fff-0fc9-4900-8b4b-b2469ddad949 · inbound

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning cites this paper.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.869307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.869307Z digest=sha256:d65c9bacc6ea9fbbd41b849d3829ddf9cdb18c71e271ee1c8ac18809e603f892

Observation f7ec1a1b-d6cf-4c47-b2e6-defab20693bb · inbound

Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training cites this paper.

Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:24:37.105136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:24:37.105136Z digest=sha256:bd32d0b9057b9bdb607ead16f86b3e1221033779d5282c8c2d477cf412e2ec80

Observation d39d2e9c-6acf-4e67-9491-54518e61f564 · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.677614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.677614Z digest=sha256:da37254096d295bb33c6051fe5ccdded9ee1e00032c2d4ae0a4c32ac75298fbd

Observation 9917f026-aa6d-414f-b692-c025fa093fe7 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:10.747215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:10.747215Z digest=sha256:921403b7d05a942dcc2c7c7afbef120081c1f0815c3389f873d321d249538aa6

Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.348025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.348025Z digest=sha256:26155e331096b4c9660860d6bde707acb3ef1085b4944252e666a411cc65bab5

Observation 043fba0e-52af-41da-b7cc-98f402ff4421 · inbound

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine cites this paper.

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:34.333552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:47:34.333552Z digest=sha256:36f0790afa06fb95e00fd460feb7791d02e5de454fb47bcd937dfbd5521eb062

Observation cc072def-b9ec-4f42-bd4c-f70984df4af5 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:50.976241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:e0ca85695a149c92371b41dd59a961a2329613b3c7adc2fc2881e890e6611089

Observation d32b479c-3c7f-4f03-9026-72d988877b7b · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.509860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.509860Z digest=sha256:c37ad0d2f945b1cba2699484a41a2fe84c8f8d0b7eae75006b83092e6baa323b

Observation f7251430-0160-4158-b836-a7496021ba59 · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.678062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.678062Z digest=sha256:f64de7f0d061ef1cfa0ede81dd64aa2b5b54bdac5ebc3bb1d3ef5895059dc80a

Observation 0fc7d1e3-6c86-42ca-a4ee-6686936fc6ae · inbound

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning cites this paper.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:17.525381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:17.525381Z digest=sha256:5e5b1c10805eeb56c0f87e4b681a53db676f1fd564a6f6e7650b09070eaf9ce9

Observation 22662b78-6232-4296-82c8-475b2160c9ad · inbound

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation cites this paper.

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:54.793915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:54.793915Z digest=sha256:1c4202e4ac1cda92d32eb345986b6cc3dfb4b475b79b1ba1cfc6e001f5243dfc

Observation 9463a595-eb51-4039-85fc-e466eb461fa9 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:18.136571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:18.136571Z digest=sha256:42e52f300983ba4d3935d560976b05d9c9902aeb93c198c56eb8928fdfea4f73

Observation 4d40ba64-a9e4-4823-8020-9045121195db · inbound

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation cites this paper.

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:32.271114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:57:32.271114Z digest=sha256:475a75ebe607a6c6be1a00d7fac50424b59ddd3b03e69006bce8252eea016755

Observation 88a4edce-cff7-4396-89c3-402c0ff4e7df · inbound

VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results cites this paper.

VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:17:54.443943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:17:54.443943Z digest=sha256:16240ee9eac60f96a0c5bd40dc92ae6da982d9e9dc4fae79b09723d463dd2c1c

Observation be8b3450-5551-4e9d-97c3-ecd45d3e549a · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.654795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.654795Z digest=sha256:290be3941b2a5793f46ff385aca69e70f9c4ddf0208d4df62cadf9944be8c9dd

Observation dc38e2c6-390a-4c8f-b3a3-9680f939ff1f · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:45.976102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:45.976102Z digest=sha256:df9ce4aa4af37e1eec06890e97d2d1e2b35bcf4171ed90d34701a06e42d44d2e

Observation 560dcc23-698d-4244-813f-07ce1600afe7 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:47.142235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:47.142235Z digest=sha256:3c14e7636f7895dee03812ae6bf1e805600fc86f86601ffaf1a1857b2e0b4a59

Observation 3e716b4b-02e4-436e-8cbf-b5b1995e2786 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:36.790900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:36.790900Z digest=sha256:8778aead98a8a525ca94089c8449964a255f99d8c58d8b5bba8479686c73ed2c

Observation 0ee47b9f-1b44-48b2-b38b-b9af80c234d7 · inbound

Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection cites this paper.

Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T10:53:15.264765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:53:15.264765Z digest=sha256:e9ad64f41dcb7bacb33c1afd7565ac5edb6e563d30bd8dc2655d5699326431b6

Observation e8c87b1f-9aa7-4e45-826a-1b53a4dda16e · inbound

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals cites this paper.

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:30:50.864486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:30:30.431777Z digest=sha256:f9c400dfb040600a46ddd9131df75bea8ab2011bf5384f30b5c10ffb51617555

Observation f0e867ad-df14-4bcb-a331-78aa73ff8341 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.946896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:f3078d3b2b8e95de6bd4bb70d7f699d9a14df0e2765f64d3d210d3144e0fb32c

Observation bc82c1a8-7dd8-41eb-ae7c-23fcade130d1 · inbound

TinyMU: A Compact Audio-Language Model for Music Understanding cites this paper.

TinyMU: A Compact Audio-Language Model for Music Understanding BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:17:36.935059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T08:17:23.740979Z digest=sha256:0ca15b1edb3b3029a48b1116fcb5556c4e87acc3f4113eeb1be97f50e4f85273

Observation b07ddca6-efdf-41f7-bd4e-e639700fdbb5 · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.191807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T09:20:22.200577Z digest=sha256:e07c2a01c6c48709b9cac7e80d47cde80b83fd770b2f704fb77f8400691d3aaf

Observation a0570288-70ac-4bfc-8f51-75240f384f19 · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:49:19.547770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T00:45:01.115573Z digest=sha256:1a3a3788b80011667e9f2560f5ebe69b98f8e2b3f3333a575521ef414aa64ac1

Observation c413f0d0-e26d-49e0-803b-013725d60ad6 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 291

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.124351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:3e80b4d07d6417b1ee33fc1d1cc42483fe1795317714f0534aa610b3b30fcaff

Observation 5721e150-1e14-4978-a682-d84a42bec402 · inbound

Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations cites this paper.

Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:01:18.489162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T12:53:11.185208Z digest=sha256:e0a4359ff1301eaf00dd9b13bcf4595fac4239aa67bac32887c07f96a5f962c6

Observation 137c5f66-8d64-444b-85bf-8641c9512bbe · inbound

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification cites this paper.

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:57.059110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:49.239866Z digest=sha256:5bb9e8d3dbec2fd7ae5b5aca069b03abff7b02537273f7f479d9699d9acffc8a

Observation bd003016-04bf-488f-a810-4d272627b8f3 · inbound

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation cites this paper.

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.563590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T18:08:50.574960Z digest=sha256:1cb1e1e430f932f849e5134ae3e8d1e8fc23d34ed977455318302e4e3b32bf9c

Observation 87328366-cfcc-4fad-a0ae-6cf828e5d1b3 · inbound

Finding Needles in the Haystack: Transductive Active Labeling in Ecology cites this paper.

Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.405444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T10:52:23.274642Z digest=sha256:413f252d08b9fecac3292562b84c03503a62a15b8dd34abd9470aaca1bd939c6

Observation 2dcc5b3f-0b7e-46cd-947b-6663208e311c · inbound

Finding Needles in the Haystack: Transductive Active Labeling in Ecology cites this paper.

Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:55:31.231791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T07:48:01.274555Z digest=sha256:ba940cc94395ddc9df2956b2b7c05d06019fd8a94d58937ee1215d6109df809a

Observation f7a930f7-f2ed-43c5-8098-c40fa06e921e · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.917414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:da25e415d64215f4e79434452c42576dbbb994ece5ec3bc17617a5f2d4792210

Observation e9834dc0-8da9-4b02-8df1-75ca93cd477b · inbound

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types cites this paper.

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-07-12T03:24:13.814557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:24:13.814557Z digest=sha256:57157f2b25010beb5dd8104799b03b621350023e04f57f4f427f7b91adf3bc62

Observation 5ea85c91-d981-4e74-9f00-cad0e7b98f56 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 274

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.467634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:db664c12e5a2eae0d7cd6b0295809c4c45a4b5b7d8875e4b6cf3c4adc94201b7

Observation fb69bebd-0691-40d1-99ec-d89a0164e2c9 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 274

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:d4475dc11c1afc9875c94dc14b5da0bf29f35f2167187a5d039404a87b89b3e8

Observation 35666e69-0923-4e17-9b60-173abe7283e8 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:acb3d685d150ffc9b7139e0ae04ade257175222493243cafcd36f69cd0047b36

Observation 11f819c1-a7d3-476d-9f32-67771820b54c · inbound

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 cites this paper.

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T02:03:37.761113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:03:37.761113Z digest=sha256:de663b926fe40c574830535036860663800a9a00fc7974e251c8b197c5ae64c8

Observation 9b447bf8-04ab-4d19-b867-a62cfef7b3a3 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T10:35:03.133791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:35:03.133791Z digest=sha256:1b76b27491a676845257d87650328b278f9d89a7b748a45cf81e5f960d042b53

Observation 3cdad4a4-601b-452a-afb4-4458397ec441 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.418732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.418732Z digest=sha256:f9bcf90b7604181d333ae237dd610a6a90d8f28ade4505836f9e601c401b99fa

Observation 6b1606f0-2d5d-4d48-9dd8-a5d38a722a5e · inbound

Hidden-Domain Routing for All-Type Audio Deepfake Detection cites this paper.

Hidden-Domain Routing for All-Type Audio Deepfake Detection BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:45.434431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:45.434431Z digest=sha256:bbc6faac29eafc4f1aaec84421d74c0461cedaf73880c7a5e762bb3de208b71a

Observation bd586c6b-3c76-4d1c-b27f-9f6a5deddbe0 · inbound

Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport cites this paper.

Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 1984

Resolution
unresolved
no resolver link, observed 2026-08-15T14:48:39.299052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:48:39.299052Z digest=sha256:271e73590d0f8301f179e321de27351b0ab7063b8a252fe469c771376eac4849

Observation 5320dc2d-b18a-49e3-93ee-81955859e82a · inbound

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment cites this paper.

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T17:17:41.297661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:17:41.297661Z digest=sha256:8fda0a32e2cac35412dbf58242423c0734727dd4606fd507b548e50e87b700e6