Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2507.13264.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T12:02:01.060858Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T20:27:36.606337Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 662da643-4ff2-4908-b0df-431df5ff5a86 · inbound
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Voxtral
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1cd5ced8-1692-4beb-8b08-f654536cbb5a · inbound
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages Voxtral
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0e3317c-5ce5-44cf-9647-e5ce27401e88 · inbound
When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models Voxtral
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0f6ce27a-0879-40da-ba58-40b297daf860 · inbound
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages Voxtral
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aedec695-5e30-49da-8b5f-1e495ca7d192 · inbound
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Voxtral
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f2c17ec-f531-4595-844d-f988b6c02138 · inbound
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Voxtral
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fee6458-3f1a-4934-a8fb-d11077dcf17f · inbound
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus Voxtral
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f18f1dd-31ad-4a52-824c-bc3794f08c4b · inbound
Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Voxtral
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead7a03c-a489-4e31-909e-3f4ffdf8286e · inbound
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 30e63df0-e21b-49e6-987d-3d7423224a24 · inbound
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding Voxtral
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ed46d8e-0b15-47c3-a503-5d2f63050ffb · inbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ce1c60b-9d74-46cb-a146-1d958d5f5268 · inbound
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning Voxtral
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee50c181-d1bb-48e4-8ade-7d783804a987 · inbound
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs Voxtral
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 80077c05-c593-4017-a3fc-d8b955e24824 · inbound
Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference Voxtral
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90b0e157-88b6-4957-bf2c-41ebed10af62 · inbound
A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Voxtral
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd193cf0-bba9-4e8f-8dc3-90a2dbfc38af · inbound
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Voxtral
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b62a8db8-b7fa-4358-8b82-bd05cd3a6147 · inbound
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Voxtral
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb5755be-f0df-4074-9c12-38e2410bcec6 · inbound
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Voxtral
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 256912e3-6864-4baf-9ea3-8de8e581005e · inbound
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps Voxtral
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b0de199-c1d6-4d80-96d5-4d692b1707f5 · inbound
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation Voxtral
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cddb048c-95c5-4ca8-89b1-7f2772e7c382 · inbound
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Voxtral
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 60e21894-abda-4684-9f83-0ce4229d7933 · inbound
CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings Voxtral
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d3e0dafa-6f9e-4827-918a-b1b55040a066 · inbound
PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding Voxtral
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0f92b36-1701-424a-996a-d33d1a2d8f79 · inbound
RealityTest: How People Probe AI Identity and Whether Models Disclose It Voxtral
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bbdcde03-ba1e-4cec-ae6b-f9221912c4b9 · inbound
MURMUR: An Efficient Inference System for Long-Form ASR Voxtral
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7584a4a5-0da4-48d6-8531-a78848bbc4cc · inbound
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0d2211e-f890-4faf-950f-365b3c96cbf4 · inbound
SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech Voxtral
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aafcd460-72a7-47fd-8edf-415cf5d7ba2e · inbound
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios Voxtral
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d44d6b97-d702-4272-b6db-bac2e0c8b265 · inbound
FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition Voxtral
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6a6fc55-6346-4ddd-a7b6-aaa034707fd8 · inbound
GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models Voxtral
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1dd1ca22-5f05-4303-a152-1a539f8dad00 · inbound
Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages Voxtral
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e1e30ead-2568-4317-951a-3ec879112b81 · inbound
Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations Voxtral
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6725bc1e-3f7f-4e9a-9927-0a6805a368f8 · inbound
LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning Voxtral
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 40abd6fa-fb90-4d9c-aa60-a7a9a732fbf9 · inbound
Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning Voxtral
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d77aea3a-6b1f-445a-a6fd-03536a02d0bc · inbound
IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages Voxtral
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bf1f3ad5-679f-4774-ab11-7bdcb3f241b0 · inbound
Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi Voxtral
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8fea7a4-5eef-4f8c-8b54-d05aaf81e144 · inbound
From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models Voxtral
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d2974562-cdae-4ac4-9b9b-b750ba3a5404 · inbound
RedVox: Safety and Fairness Gaps in Speech Models Across Languages Voxtral
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe7dcb65-5b71-4ae1-8988-2dca377b95f5 · inbound
Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Voxtral
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58468181-dd03-488c-b1e6-6a5ecce3a2d3 · inbound
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a04cd9-a983-4ec9-a239-ed53e31d7c02 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Voxtral
Reference 234
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f1cf07ff-4252-4e63-9d6d-39b3b03ac91f · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Voxtral
Reference 234
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db9f10c-f72e-4a63-802f-e476e6864eec · inbound
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Voxtral
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 382579c1-70e4-4e99-9485-5309bbe4cacb · inbound
GigaChat Audio: Time-aware Large Audio Language Model Voxtral
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608bbfa2-db29-48f0-b7a6-465c09ffa510 · inbound
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Voxtral
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee421ae-6f81-4633-954c-6d145337df2e · inbound
Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation Voxtral
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8acb12-a731-4957-b1d3-2163a6c723c3 · inbound
Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge Voxtral
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccec7c0-c125-411f-809c-446281cc4a93 · inbound
ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions Voxtral
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66feb1f-3887-4922-a1bb-388b081b4484 · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Voxtral
Reference 186
Source-reported events for the cited work
Unavailable: canonical work link unavailable.