Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:32.396611Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 7 inbound Pith citation observations for arXiv:2505.13237.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:32.396611Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:32.201734Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T19:35:32.876839Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 220f6e95-2554-4342-a9f0-7f49053abf4f · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information What is the animal in the sound ?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4c13b9ac-25c0-4f0a-bd63-731527226b1c · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eccb163-3f2e-479e-9430-beed7a69738a · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information determining the age of the speaker
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1596be03-4632-424f-b663-cb13a911e7bf · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information the association of animals and human personality
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82f54d82-16b2-4222-89d6-f98e324e0b4f · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information feeding habits
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f2cfe7c5-c50a-40c3-b6b2-1d20fc7625a1 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Evaluation metrics Since SAKURA comprises multiple-choice questions, accu- racy is a natural metric
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ce3c14d-5bfe-47bf-b6a8-5a63ee1d856e · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6dab4d2d-dd47-4a2f-be37-b6af1cd00e5d · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1a8ee3e0-e0f2-4b28-aa05-d75b5b37716c · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information cor- rect/incorrect
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9304f2e5-aa61-4d2d-96c4-fe957780d27c · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information The animal making the sound is cat
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 08e88832-d58e-4cf7-984d-c1aec4d716a8 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Our findings show that LALMs struggle to recognize certain speech and audio attributes, exhibiting perception blind spots
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69d419af-1c17-42e5-af17-106368558be4 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation df627084-5ec7-434e-a390-38670676510a · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6208b81-601c-478c-b1d4-d26fe7ea1830 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information GPT-4o System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76c6f9ea-5555-49d5-ace0-e900c99a6281 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Vipergpt: Visual inference via python execution for reasoning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 584c2320-3e69-40db-864e-55828a41fcb6 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Audiogpt: Understanding and generating speech, music, sound, and talking head,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ffa26111-18b8-4810-ade7-d37de708f3e3 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Speech-copilot: Leveraging large language models for speech processing via task decomposition, modular- ization, and program generation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 18f55ae7-869c-41bd-987f-b7223f036914 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Visual instruction tuning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1ea6430a-1dc2-4c33-a0ae-dfae28f764c7 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ead504f-3ed8-40ec-b36b-c322eda65934 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Joint audio and speech understanding,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30b44235-586e-4815-9852-dda99f3f0037 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c815fee-4775-487b-aa1b-e3a98d4600b0 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information SALMONN: Towards generic hearing abilities for large language models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 37ec7eed-b848-4949-b7bc-98b3b8f1c52a · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec010cc-24e9-4668-a73d-521798b53e67 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed537ded-5d6c-4063-8574-f63aa56d655c · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Qwen2-Audio Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acea8f48-0190-4efc-a583-0b575169e643 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information A peek into token bias: Large language models are not yet genuine reasoners,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1ca49faf-2d01-4ddc-bd12-72af4a3160fa · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Large language models cannot self-correct rea- soning yet,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ce53e49-17c5-42aa-baec-6e1d331cd2b6 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Premise order matters in reasoning with large language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fd902063-b37b-4816-8200-0e012c9486ff · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Measuring and narrowing the compositionality gap in language models,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f55d04c0-601c-471e-8634-704944054c43 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ebca65-3274-46b8-8230-00a8ca8696a1 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Do large language models latently perform multi- hop reasoning?
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dc7deb39-27c9-4983-a7a3-28c91a14cc9f · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Hopping too late: Exploring the limitations of large language models on multi-hop queries,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 690ccb98-c203-4867-92e7-db237b0abe5d · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Investigating multi-hop factual shortcuts in knowl- edge editing of large language models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4f74e161-5e76-42ca-8120-bf4935db7ab3 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Dynamic-SUPERB phase-2: A collabora- tively expanding benchmark for measuring the capabilities of spo- ken language models with 180 tasks,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ff81b993-a8ad-474c-a1c7-2d8a2c985df9 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information AIR-bench: Benchmarking large audio-language models via generative comprehension,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9152f3d8-54dd-47f5-a434-86ff5ecfad57 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e669ad24-3ff2-44b0-81f0-4f90e32575e1 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ef16883-0dfa-490d-9617-d2c5201cb568 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Listen and speak fairly: a study on semantic gender bias in speech integrated large language models,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3edb1952-9549-45e0-856e-5ed3977b9377 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Spoken stereoset: on eval- uating social bias toward speaker in speech large language mod- els,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 08bf72d3-b129-4616-9cc0-4656b2e22f82 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Compa: Addressing the gap in compositional reasoning in audio-language models,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b35f9a55-78e8-48fb-bec1-304d13c7d394 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information MMAU: A massive multi-task audio understand- ing and reasoning benchmark,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a33989dd-b7a6-4dd9-932a-7e3f33601be1 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f02676a3-5c95-494c-aa50-a5f3ec4c02b0 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Common voice: A massively-multilingual speech corpus,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d3c3a2db-b78c-4f8a-852e-ccb507698131 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Crema-d: Crowd-sourced emotional multimodal actors dataset,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd4425f2-73e1-4248-b773-d9c702ec0d21 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information MELD: A multimodal multi-party dataset for emotion recognition in conversations,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 15e44130-5b1e-46d9-8302-bf8b0ba04df6 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Emotion detection on tv show transcripts with sequence-based convolutional neural networks,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e9d44583-5f75-4195-b726-08e7658b512b · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Esc: Dataset for environmental sound classifica- tion,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1afc1e41-2823-425b-9693-79406bc848b9 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Animal sound classification using a convolutional neural network,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 32077ba6-b845-44b7-95f0-8bb3e4eb089d · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information A Survey on LLM-as-a-Judge
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e60d6e0-1752-4fc7-8427-acc0c7eef80a · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Robust speech recognition via large-scale weak supervision,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3388e39a-3943-45a1-bfc0-4c41da4c7261 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information A Preliminary Exploration with GPT-4o Voice Mode
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be34161-d75b-4350-bbdb-f542090ae388 · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44c08ff7-7763-4a7c-be55-7f4cba09cb1c · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Moshi: a speech-text foundation model for real-time dialogue
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f223fe9a-a743-4c20-8ed4-c5df12d9007d · outbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c13b9ac-25c0-4f0a-bd63-731527226b1c · inbound
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d7972e2-000d-4c3b-9f50-2eb310168e9e · inbound
Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0d912e-77d4-477e-a30b-7c271789c230 · inbound
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8555cec-11ee-4dd0-86a2-6b22ac14f501 · inbound
A Survey of Audio Reasoning in Multimodal Foundation Models SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bd865246-2e24-4a45-a29e-31ce1e8a1e8f · inbound
From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 865f8dc1-fbc9-4c8b-8930-afa1b769520c · inbound
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3439d7d3-c933-4473-a804-f21bf2009162 · inbound
Large Audio Language Models for Spoofing-Aware Speaker Verification SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.