Pith. sign in

Paper Citation Record · LEDGER

Voxtral

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2507.13264.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13264 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:02:01.060858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:27:36.606337Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 662da643-4ff2-4908-b0df-431df5ff5a86 · inbound

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs cites this paper.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Voxtral

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:11:43.160420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:111954e2aee907dec52247fe86f21404650e0b93b7482b219d8f33616e3cfd86

Observation 1cd5ced8-1692-4beb-8b08-f654536cbb5a · inbound

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages cites this paper.

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages Voxtral

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.256969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T16:27:37.596817Z digest=sha256:e03a059cc8d552ab9a1e0bad77ab89d4ceb4df0d96597c82bcc11f2bd0b23915

Observation a0e3317c-5ce5-44cf-9647-e5ce27401e88 · inbound

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models cites this paper.

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models Voxtral

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:17.928697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T11:08:08.916893Z digest=sha256:adf4f88471bc3b26c99534b034b045d8955bc019090b3444e1a83e95588192bc

Observation 0f6ce27a-0879-40da-ba58-40b297daf860 · inbound

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages cites this paper.

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages Voxtral

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:18:57.194337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:15:04.685150Z digest=sha256:9b5aacb1195c8564b32316c87e442f7e85c25ce5d44c3aa4760c68ada94230ca

Observation aedec695-5e30-49da-8b5f-1e495ca7d192 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Voxtral

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:51:17.744592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:1c7af7d33cc33ed38e3ba64a1860ea2243754b0fe6a2267cb1cc8a364552ec1a

Observation 6f2c17ec-f531-4595-844d-f988b6c02138 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Voxtral

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:01.060858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:01.060858Z digest=sha256:55da1c010704394e07d9d838bfa31c8284b342a3227c8af4eb3f1930826ae62a

Observation 7fee6458-3f1a-4934-a8fb-d11077dcf17f · inbound

MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus cites this paper.

MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus Voxtral

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:31:01.948157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:30:46.664861Z digest=sha256:2d7180f313e4e46b6038dddf834ab0436aa30bba0097006170939e007bd3a09f

Observation 6f18f1dd-31ad-4a52-824c-bc3794f08c4b · inbound

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization cites this paper.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Voxtral

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:17:47.177241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.177241Z digest=sha256:7f0350c106b6f62cab8a470cd1281f465c5508bae71d4c12d49c02cd3517d910

Observation ead7a03c-a489-4e31-909e-3f4ffdf8286e · inbound

Voxtral Realtime cites this paper.

Voxtral Realtime Voxtral

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:17:07.534451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:12:16.230433Z digest=sha256:5e6f8c395b85167754664868d61d85bf3c0eb580339adf419d3ccce758cd54ab

Observation 30e63df0-e21b-49e6-987d-3d7423224a24 · inbound

Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding cites this paper.

Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding Voxtral

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T13:58:09.323985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:58:09.323985Z digest=sha256:37eb46ba19e797c6965c8658be8e1196fde43fd52f4cb9a4273d8fbb4bc4395f

Observation 1ed46d8e-0b15-47c3-a503-5d2f63050ffb · inbound

Voxtral TTS cites this paper.

Voxtral TTS Voxtral

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.904186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:12360021bd00f6de0b475aba0d1fd9f37bbbfeefef2f194e1a19f256bb179d69

Observation 5ce1c60b-9d74-46cb-a146-1d958d5f5268 · inbound

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning cites this paper.

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning Voxtral

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:38:02.912734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T17:34:10.089555Z digest=sha256:0841bdfc0d15a809d7765ee797a06177be07f9286eed3d4ffcf7b8f3f2d790b3

Observation ee50c181-d1bb-48e4-8ade-7d783804a987 · inbound

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs cites this paper.

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs Voxtral

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:30:57.273062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:06:50.408402Z digest=sha256:13e14b16ca13b0c5c46a10b7839316073b01f61d864067794033bac28c2827cf

Observation 80077c05-c593-4017-a3fc-d8b955e24824 · inbound

Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference cites this paper.

Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference Voxtral

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:21.947631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:04:25.418910Z digest=sha256:eea0339e401b1419d3b3aa0938b87c374cea80a509c14219a97fa203daee0a10

Observation 90b0e157-88b6-4957-bf2c-41ebed10af62 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Voxtral

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:0aec20ee27a51ac84ebfa7b6ef2daef6f220dbdfd4ae9f6e3901ea75490c3428

Observation fd193cf0-bba9-4e8f-8dc3-90a2dbfc38af · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Voxtral

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.928382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:afc6231c284df79965ac5398758b1b137fecad151df523d365994d8160988c8f

Observation b62a8db8-b7fa-4358-8b82-bd05cd3a6147 · inbound

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR cites this paper.

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Voxtral

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:21:05.956205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:48:14.211240Z digest=sha256:75a13ff2c29a71c3465cd9d2a063d50395930fae613d7300031a068c787982d2

Observation eb5755be-f0df-4074-9c12-38e2410bcec6 · inbound

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR cites this paper.

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Voxtral

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T13:21:06.288305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-05T13:15:50.794969Z digest=sha256:acdaa837ce4e43379ea68b5beb875f13135eb783a20db0206b81758399d530b9

Observation 256912e3-6864-4baf-9ea3-8de8e581005e · inbound

Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps cites this paper.

Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps Voxtral

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:04.863844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T01:56:44.270045Z digest=sha256:dffa9ec199301d5f890fec166977e5b507da7c82bfecfebacc7d6f0b3a7a92de

Observation 4b0de199-c1d6-4d80-96d5-4d692b1707f5 · inbound

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation cites this paper.

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation Voxtral

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:36.207134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T17:39:38.234052Z digest=sha256:4dc79cbfae945c15defbd668ae261100b6c0a5c32edc3f0dda25b5fda3b5ab49

Observation cddb048c-95c5-4ca8-89b1-7f2772e7c382 · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Voxtral

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:12.313110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:a97ef32f4d6ea4b20c194cf0e01897fc61ee1b8d762b6dcadc040bbc439a1ce7

Observation 60e21894-abda-4684-9f83-0ce4229d7933 · inbound

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings cites this paper.

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings Voxtral

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:28:19.127591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:25:09.024054Z digest=sha256:100f8e7a0945aae20464f850163d9794219cc4babe730455f0423c5ba580097a

Observation d3e0dafa-6f9e-4827-918a-b1b55040a066 · inbound

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding cites this paper.

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding Voxtral

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.390433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:56:38.345335Z digest=sha256:5db463d7002cc67fc6347491408bbb0705c58d2b30d721ae584d4e22ba24b5ed

Observation d0f92b36-1701-424a-996a-d33d1a2d8f79 · inbound

RealityTest: How People Probe AI Identity and Whether Models Disclose It cites this paper.

RealityTest: How People Probe AI Identity and Whether Models Disclose It Voxtral

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:36:08.724825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:22:43.432328Z digest=sha256:a0212480f268a504bc860fa95058508255fda047252613c54fdd8786a3397341

Observation bbdcde03-ba1e-4cec-ae6b-f9221912c4b9 · inbound

MURMUR: An Efficient Inference System for Long-Form ASR cites this paper.

MURMUR: An Efficient Inference System for Long-Form ASR Voxtral

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:12:25.096494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T17:06:02.978205Z digest=sha256:5de237037b1a44225da65baff5b4f81055def6f9f086e9d305de3e854b5976fd

Observation 7584a4a5-0da4-48d6-8531-a78848bbc4cc · inbound

Benchmarking Speech-to-Speech Translation Models cites this paper.

Benchmarking Speech-to-Speech Translation Models Voxtral

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:46:29.444473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T10:34:10.510906Z digest=sha256:d51d4abb44f5a4a2eb8c89e0a72f0eea97fea1fa5622264a6a6ddff228ecc6eb

Observation a0d2211e-f890-4faf-950f-365b3c96cbf4 · inbound

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech cites this paper.

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech Voxtral

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:37:06.724970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T23:44:42.188662Z digest=sha256:ea992fca11b5b69bf2a605b54f4556beae8ca6a477fd05c09e32793786b8a0ae

Observation aafcd460-72a7-47fd-8edf-415cf5d7ba2e · inbound

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios cites this paper.

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios Voxtral

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.045645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T01:11:56.037674Z digest=sha256:2ec95929cad2ab846012a013f9dc58259ca05106323677b3532e832a44cb6c27

Observation d44d6b97-d702-4272-b6db-bac2e0c8b265 · inbound

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition cites this paper.

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition Voxtral

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:57.087582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T01:58:13.871771Z digest=sha256:7e8aaa89e7f7c6286dfcdca676194e4dfb085ceb09bc6a8b6eb7b1de88413107

Observation b6a6fc55-6346-4ddd-a7b6-aaa034707fd8 · inbound

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models cites this paper.

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models Voxtral

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.526077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T19:51:59.265579Z digest=sha256:b67b6b4a3357210534810cd21f9e804059068c954f1e950476ffd78affc8a0b1

Observation 1dd1ca22-5f05-4303-a152-1a539f8dad00 · inbound

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages cites this paper.

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages Voxtral

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.489322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T15:10:03.282212Z digest=sha256:a07178d2c2d09e8df54b29fcbe636611c0b2069e44c71b060a605de26eeb045a

Observation e1e30ead-2568-4317-951a-3ec879112b81 · inbound

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations cites this paper.

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations Voxtral

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-27T07:10:41.472377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T07:06:13.658919Z digest=sha256:7bd2332d69dbdf06fbc373b381c9feb83880b2a3df28e1b061fa2e88dd0bc41a

Observation 6725bc1e-3f7f-4e9a-9927-0a6805a368f8 · inbound

LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning cites this paper.

LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning Voxtral

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:13.111737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T08:13:38.489672Z digest=sha256:c44ec85f748880b38e5548ba0e5a05058cb0d4885e4a3cc69927b51db649c53c

Observation 40abd6fa-fb90-4d9c-aa60-a7a9a732fbf9 · inbound

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning cites this paper.

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning Voxtral

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:03.351616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T22:39:09.967750Z digest=sha256:c924411126d5ce9eb8e9dec5189a4553056b1801e1c56c9fb371b3b997c2d454

Observation d77aea3a-6b1f-445a-a6fd-03536a02d0bc · inbound

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages cites this paper.

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages Voxtral

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.345552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T19:14:26.452851Z digest=sha256:9dfaf5e123ea9f927237e31bae934164c36e8535d971a91a6a50233d84f69a98

Observation bf1f3ad5-679f-4774-ab11-7bdcb3f241b0 · inbound

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi cites this paper.

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi Voxtral

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:39.279437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T13:07:37.187581Z digest=sha256:e56a050b80ed324060a634b6f084056953baa490bdb5164f6a22040c5f823c75

Observation d8fea7a4-5eef-4f8c-8b54-d05aaf81e144 · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models Voxtral

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.978640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:ddda5761bb2808ff1a9dc63d6ec2390b8c575990f6aacb716b7f91e265628c29

Observation d2974562-cdae-4ac4-9b9b-b750ba3a5404 · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages Voxtral

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.910091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:64d7fe7e80af0cefc3e3b0ae30187a873d6772da3d9c81f61f6cc2604af1737e

Observation fe7dcb65-5b71-4ae1-8988-2dca377b95f5 · inbound

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving cites this paper.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Voxtral

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.663262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:54ed931434157f647765c1be8cafa27ac6d1fdaa099430e239b56f2c90af2b6c

Observation 58468181-dd03-488c-b1e6-6a5ecce3a2d3 · inbound

S-DiverSe: Spanish Diverse Speech cites this paper.

S-DiverSe: Spanish Diverse Speech Voxtral

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T04:08:13.798171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:08:13.798171Z digest=sha256:2e3a590654d273d4426b078f6fa22ad13752454ac47ead6e4a9df2ec3594e809

Observation 16a04cd9-a983-4ec9-a239-ed53e31d7c02 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Voxtral

Reference 234

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.545770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:b9579ff29c0a8072c5adb996b1ee7a0445d6d5ab6d0263fd9fe24f4f07bc13b1

Observation f1cf07ff-4252-4e63-9d6d-39b3b03ac91f · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Voxtral

Reference 234

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:18d5879c790a50cfcab0aa46418b0885ba7dd3e676014fc80e67839842db58e1

Observation 3db9f10c-f72e-4a63-802f-e476e6864eec · inbound

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs cites this paper.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Voxtral

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.607654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:0292276a1c35f70fb04b4101a335af9f21720ca65cb3822b1ba3a3a830688db5

Observation 382579c1-70e4-4e99-9485-5309bbe4cacb · inbound

GigaChat Audio: Time-aware Large Audio Language Model cites this paper.

GigaChat Audio: Time-aware Large Audio Language Model Voxtral

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T12:07:39.817307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:07:39.817307Z digest=sha256:2353393544a5098907109ab58289eadfb9ddbc459867014dfde06161b958d2f7

Observation 608bbfa2-db29-48f0-b7a6-465c09ffa510 · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Voxtral

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.370603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.370603Z digest=sha256:0f2010d87c45a67a2fd94f2c83b0c62e1d0a1b1556b663d9f9e74c90106cd162

Observation eee421ae-6f81-4633-954c-6d145337df2e · inbound

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation cites this paper.

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation Voxtral

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:08:03.719567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:08:03.719567Z digest=sha256:32cbf432a9204781d953ce1ed5609cc437f0945ac8ddde9e8f246a39490c5a7f

Observation 8e8acb12-a731-4957-b1d3-2163a6c723c3 · inbound

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge cites this paper.

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge Voxtral

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:42:01.682549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:42:01.682549Z digest=sha256:1dd6126fec51caf7e67e13c76a3ea1c27beb18103b7fd76fd8c0e59e1f377a57

Observation 9ccec7c0-c125-411f-809c-446281cc4a93 · inbound

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions cites this paper.

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions Voxtral

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T16:57:11.564423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:57:11.564423Z digest=sha256:c43c379c5075c69365764d16564c6facdaf663fafd38c3fa99f63131d4b5f1f6

Observation f66feb1f-3887-4922-a1bb-388b081b4484 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Voxtral

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.795578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.795578Z digest=sha256:8db38444135d1c352ca8af86b3205b6e1c646fc5ae2d15b2c7c6a76e22b6780e