Pith. sign in

Paper Citation Record · LEDGER

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

As of 23 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 7 inbound Pith citation observations for arXiv:2505.13237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13237 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:32.396611Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:32.201734Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T19:35:32.876839Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 220f6e95-2554-4342-a9f0-7f49053abf4f · outbound

This paper cites What is the animal in the sound ?.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information What is the animal in the sound ?

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:33.022387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.197075Z digest=sha256:faaf2665d3d20f60aaad50a41c5e781c876dc6fa3703ce41691b3a25229d4fcb

Observation 4c13b9ac-25c0-4f0a-bd63-731527226b1c · outbound

This paper cites SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.201734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.201734Z digest=sha256:1db0ba264f74c768e9e592f9d4e97487846611365248a64305295f9515376d98

Observation 6eccb163-3f2e-479e-9430-beed7a69738a · outbound

This paper cites determining the age of the speaker.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information determining the age of the speaker

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:33.011252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.206065Z digest=sha256:be1a21a5eb389a0f7a984d93b6ea24f36bd931af582583ffc2c33ffb4058d551

Observation 1596be03-4632-424f-b663-cb13a911e7bf · outbound

This paper cites the association of animals and human personality.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information the association of animals and human personality

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:33.000596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.210053Z digest=sha256:0574d8ad23c7b7ff8c9b4e603dbbd19665cadc006f65b18a96d6f2cefe215a40

Observation 82f54d82-16b2-4222-89d6-f98e324e0b4f · outbound

This paper cites feeding habits.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information feeding habits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.988792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.214388Z digest=sha256:e76c49b89490173fb05c201cd6cca944449895caee5664ba7b38af2f1110a5c0

Observation f2cfe7c5-c50a-40c3-b6b2-1d20fc7625a1 · outbound

This paper cites Evaluation metrics Since SAKURA comprises multiple-choice questions, accu- racy is a natural metric.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Evaluation metrics Since SAKURA comprises multiple-choice questions, accu- racy is a natural metric

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.976646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.218515Z digest=sha256:b087c43a80aa34b5e967903fa55820141f3258820874d2524d22fdaab9e3986d

Observation 5ce3c14d-5bfe-47bf-b6a8-5a63ee1d856e · outbound

This paper cites an unresolved cited work.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:32.965380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.222331Z digest=sha256:3f1b62fe9559561456f8c085286d296eb4f9cbe2284c576326217eccfe399dea

Observation 6dab4d2d-dd47-4a2f-be37-b6af1cd00e5d · outbound

This paper cites an unresolved cited work.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:32.954893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.226358Z digest=sha256:33433af2cf11d6740f4cb198ba3e53eee3442417b6f12863714a4b99b441e6b6

Observation 1a8ee3e0-e0f2-4b28-aa05-d75b5b37716c · outbound

This paper cites cor- rect/incorrect.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information cor- rect/incorrect

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.944431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.230095Z digest=sha256:79ad3b4176851fc58fa8c9419edd6ebaa96e6328d3f66c8b9ca9af17c00ba7ec

Observation 9304f2e5-aa61-4d2d-96c4-fe957780d27c · outbound

This paper cites The animal making the sound is cat.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information The animal making the sound is cat

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.933058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.233913Z digest=sha256:3a6dd64fd2ea516278eb24d5fe6b398e45d0bc105cffe3e337dce7794b15d250

Observation 08e88832-d58e-4cf7-984d-c1aec4d716a8 · outbound

This paper cites Our findings show that LALMs struggle to recognize certain speech and audio attributes, exhibiting perception blind spots.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Our findings show that LALMs struggle to recognize certain speech and audio attributes, exhibiting perception blind spots

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.922456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.238091Z digest=sha256:87c8a32127a473f66a772cafd720647cae99e055dd8b7dad359908811acd0344

Observation 69d419af-1c17-42e5-af17-106368558be4 · outbound

This paper cites an unresolved cited work.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:32.910134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.241681Z digest=sha256:9aab4f60e692841b24e257c21cd52cdcece074ecf7748348fe77b109f2a095f2

Observation df627084-5ec7-434e-a390-38670676510a · outbound

This paper cites The Llama 3 Herd of Models.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.245573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.245573Z digest=sha256:7416c7755e3478a1d7f6e86cc12c4addcad828ab11cd866bea8e10b00f712fca

Observation a6208b81-601c-478c-b1d4-d26fe7ea1830 · outbound

This paper cites GPT-4o System Card.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.249490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.249490Z digest=sha256:d62799f055b30eb42878472ff285d1166843b574b17a41a39b3935c997d8150e

Observation 76c6f9ea-5555-49d5-ace0-e900c99a6281 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Vipergpt: Visual inference via python execution for reasoning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.899629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.253711Z digest=sha256:477ff3dfac291266808bbc593b3bcfea22f791cfc0f6f61094f409e1b3f7b225

Observation 584c2320-3e69-40db-864e-55828a41fcb6 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.889553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.257761Z digest=sha256:73ff93a57832edeea8397ff491252ecb6e5ec38eb3b7771adf7385a84e607add

Observation ffa26111-18b8-4810-ade7-d37de708f3e3 · outbound

This paper cites Speech-copilot: Leveraging large language models for speech processing via task decomposition, modular- ization, and program generation,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Speech-copilot: Leveraging large language models for speech processing via task decomposition, modular- ization, and program generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.879308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.261763Z digest=sha256:1e69d30a4f9a35c8677c7b558ee462e677e7d6af8bb8434e8598c359c82bbc9d

Observation 18f55ae7-869c-41bd-987f-b7223f036914 · outbound

This paper cites Visual instruction tuning,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Visual instruction tuning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.867915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.265698Z digest=sha256:6979ce8cfd603c1e4190ec74846bdd7e2b308c70d1fd0e4e676e70576f86be43

Observation 1ea6430a-1dc2-4c33-a0ae-dfae28f764c7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.269489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.269489Z digest=sha256:4ad0c1b51981f01a749ca121e488b6d9e2b9c09bbce80c283101bcffbecb0452

Observation 0ead504f-3ed8-40ec-b36b-c322eda65934 · outbound

This paper cites Joint audio and speech understanding,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Joint audio and speech understanding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.857566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.273650Z digest=sha256:7b60dd9e8d5fa9b4a5b1ea62d24a6a3e1507b955dcf5e70f194b31541b7c1d0d

Observation 30b44235-586e-4815-9852-dda99f3f0037 · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.277711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.277711Z digest=sha256:bea6b9663cb7755601b23cdab90875aec81dbe6459365dae4ac6c62a964ca1aa

Observation 3c815fee-4775-487b-aa1b-e3a98d4600b0 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information SALMONN: Towards generic hearing abilities for large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.847207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.282098Z digest=sha256:294f28df8dc6320b296df67b8715cde079fbe2aae96ce18614fbd1961f1f8d2f

Observation 37ec7eed-b848-4949-b7bc-98b3b8f1c52a · outbound

This paper cites DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.285531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.285531Z digest=sha256:0eb03506c7415979043c9742dd30d11c840f173f552cfe3c86344dbd85531804

Observation bec010cc-24e9-4668-a73d-521798b53e67 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.289365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.289365Z digest=sha256:74e72e2590b02b1d3a852f723037be0e081c2843f1c7b9295ee27153db81d968

Observation ed537ded-5d6c-4063-8574-f63aa56d655c · outbound

This paper cites Qwen2-Audio Technical Report.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Qwen2-Audio Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.293038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.293038Z digest=sha256:698c95e7aa4fb0d216969992c60205ff6078dc90c8a969b623e28379c5f5a137

Observation acea8f48-0190-4efc-a583-0b575169e643 · outbound

This paper cites A peek into token bias: Large language models are not yet genuine reasoners,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information A peek into token bias: Large language models are not yet genuine reasoners,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.835964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.296768Z digest=sha256:36c460b6e3f4748efbc09f4190ab4106b35024a896e3da3a8f20a27efea838b3

Observation 1ca49faf-2d01-4ddc-bd12-72af4a3160fa · outbound

This paper cites Large language models cannot self-correct rea- soning yet,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Large language models cannot self-correct rea- soning yet,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.824357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.300368Z digest=sha256:46b4a586627fb3044f244a93424a0cb4667dfb95de9ac1099c2c9b36e9492614

Observation 4ce53e49-17c5-42aa-baec-6e1d331cd2b6 · outbound

This paper cites Premise order matters in reasoning with large language models,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Premise order matters in reasoning with large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.813640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.303863Z digest=sha256:8a107fead01bdc041b95e89570143a12f4602952f3a2d13ba89a6377368a9e49

Observation fd902063-b37b-4816-8200-0e012c9486ff · outbound

This paper cites Measuring and narrowing the compositionality gap in language models,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Measuring and narrowing the compositionality gap in language models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.802918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.307710Z digest=sha256:9f2d8c7b5791a5b400a8992f6a6f458af1227120b7cbec28435e2089d10dbcaf

Observation f55d04c0-601c-471e-8634-704944054c43 · outbound

This paper cites Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.311200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.311200Z digest=sha256:dc76f4c30c5d17ea7c362c1766f65c6b49348525bc63b28a4d508dccaaaf7a94

Observation 14ebca65-3274-46b8-8230-00a8ca8696a1 · outbound

This paper cites Do large language models latently perform multi- hop reasoning?.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Do large language models latently perform multi- hop reasoning?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.792689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.315001Z digest=sha256:b78875f3ca46f6311c1c6fef20078f10986cd8c8a531cb3292c37365630bb54c

Observation dc7deb39-27c9-4983-a7a3-28c91a14cc9f · outbound

This paper cites Hopping too late: Exploring the limitations of large language models on multi-hop queries,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Hopping too late: Exploring the limitations of large language models on multi-hop queries,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.782218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.318751Z digest=sha256:97e0b3d9ce6a1737054f2e30f31432e3837b0576c65bc895115a2431904f5029

Observation 690ccb98-c203-4867-92e7-db237b0abe5d · outbound

This paper cites Investigating multi-hop factual shortcuts in knowl- edge editing of large language models,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Investigating multi-hop factual shortcuts in knowl- edge editing of large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.771545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.322290Z digest=sha256:af8af7eecc3ab95167bfdd7047c92a3fc357bc20442fa24df2f53a6333d0d7fe

Observation 4f74e161-5e76-42ca-8120-bf4935db7ab3 · outbound

This paper cites Dynamic-SUPERB phase-2: A collabora- tively expanding benchmark for measuring the capabilities of spo- ken language models with 180 tasks,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Dynamic-SUPERB phase-2: A collabora- tively expanding benchmark for measuring the capabilities of spo- ken language models with 180 tasks,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.760585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.325857Z digest=sha256:b7eac9cc72e8747dc15dbcd9733863c21977336b5ea56a533ec90baf937d52da

Observation ff81b993-a8ad-474c-a1c7-2d8a2c985df9 · outbound

This paper cites AIR-bench: Benchmarking large audio-language models via generative comprehension,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information AIR-bench: Benchmarking large audio-language models via generative comprehension,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.749575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.329530Z digest=sha256:54e653317e7859e153d315a18ba2282d91dd66d3a2e386ef7545ca34162288d6

Observation 9152f3d8-54dd-47f5-a434-86ff5ecfad57 · outbound

This paper cites Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.738002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.333173Z digest=sha256:ff256f0e7500a18223c68a052fd608950a289b998e89bce6e3aa1b3482adab96

Observation e669ad24-3ff2-44b0-81f0-4f90e32575e1 · outbound

This paper cites Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.726406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.337049Z digest=sha256:130ed89a18605c8ffef38226a08586e0716639da97be462d684925882f236b2a

Observation 5ef16883-0dfa-490d-9617-d2c5201cb568 · outbound

This paper cites Listen and speak fairly: a study on semantic gender bias in speech integrated large language models,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Listen and speak fairly: a study on semantic gender bias in speech integrated large language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.713188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.340670Z digest=sha256:20f3ad5c09b23190bff9a67b0cedae29d25fdeed95b90c7ae7fd33f9730b3fd4

Observation 3edb1952-9549-45e0-856e-5ed3977b9377 · outbound

This paper cites Spoken stereoset: on eval- uating social bias toward speaker in speech large language mod- els,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Spoken stereoset: on eval- uating social bias toward speaker in speech large language mod- els,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.700574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.344133Z digest=sha256:7e9195c2a9e8cf27f399f015c46b31d442ae42f9cb5f4b67384e3363b6d4a3c7

Observation 08bf72d3-b129-4616-9cc0-4656b2e22f82 · outbound

This paper cites Compa: Addressing the gap in compositional reasoning in audio-language models,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Compa: Addressing the gap in compositional reasoning in audio-language models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.689090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.347722Z digest=sha256:9652bbb9aa2d61fb35fd89f9036ca7c388268ca3d9d4c6a8af7390a2c5e8af0b

Observation b35f9a55-78e8-48fb-bec1-304d13c7d394 · outbound

This paper cites MMAU: A massive multi-task audio understand- ing and reasoning benchmark,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information MMAU: A massive multi-task audio understand- ing and reasoning benchmark,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.677756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.351070Z digest=sha256:8b7ec931409ada991a04571fffd2416167d8e7017ce6c999a9a42badaf0a5fc9

Observation a33989dd-b7a6-4dd9-932a-7e3f33601be1 · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.664808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.354348Z digest=sha256:cb043dde58f4fffa363d4a42688220b1e1b25aac26807eb066579512136c39f6

Observation f02676a3-5c95-494c-aa50-a5f3ec4c02b0 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Common voice: A massively-multilingual speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.653515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.357698Z digest=sha256:ef2e39da7c9bf5bda312dc48ae8157e7d43a86e9ba7e84b88f3bfffb17f5e3c2

Observation d3c3a2db-b78c-4f8a-852e-ccb507698131 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Crema-d: Crowd-sourced emotional multimodal actors dataset,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.641895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.361081Z digest=sha256:38a6afe7af3e7a66783339ff40e005be9b634b8763e4fff1309269cd773360ce

Observation cd4425f2-73e1-4248-b773-d9c702ec0d21 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversations,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information MELD: A multimodal multi-party dataset for emotion recognition in conversations,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.628960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.364575Z digest=sha256:4c1ed3fac1c3fd634f72ed92e42418e22627bc4faf260645e8fbdaa79e3bea78

Observation 15e44130-5b1e-46d9-8302-bf8b0ba04df6 · outbound

This paper cites Emotion detection on tv show transcripts with sequence-based convolutional neural networks,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Emotion detection on tv show transcripts with sequence-based convolutional neural networks,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.615043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.367947Z digest=sha256:23d69dcc95d764da5b3e7349bc15141160c0b5c901bec4c29525d921db4d35cb

Observation e9d44583-5f75-4195-b726-08e7658b512b · outbound

This paper cites Esc: Dataset for environmental sound classifica- tion,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Esc: Dataset for environmental sound classifica- tion,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.371393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.371393Z digest=sha256:82cf926f3a86f68ae3e895bcb50c5c13138a0148de82415930a819beca4143cb

Observation 1afc1e41-2823-425b-9693-79406bc848b9 · outbound

This paper cites Animal sound classification using a convolutional neural network,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Animal sound classification using a convolutional neural network,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.595626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.374940Z digest=sha256:a9f87f2a48d290da0b2bfc35b1019a3f4d6ee3484267adc9a641fac99dda281b

Observation 32077ba6-b845-44b7-95f0-8bb3e4eb089d · outbound

This paper cites A Survey on LLM-as-a-Judge.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information A Survey on LLM-as-a-Judge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.378192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.378192Z digest=sha256:38d6989c2536d62789046fe15277df23e0b699685af77d514cb8b95a629380f1

Observation 7e60d6e0-1752-4fc7-8427-acc0c7eef80a · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Robust speech recognition via large-scale weak supervision,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:32.582671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:21:32.381882Z digest=sha256:b89ba4500f382269211b274618013bcda4d8f218319e11ce0d0a7f086b229e22

Observation 3388e39a-3943-45a1-bfc0-4c41da4c7261 · outbound

This paper cites A Preliminary Exploration with GPT-4o Voice Mode.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information A Preliminary Exploration with GPT-4o Voice Mode

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.385350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.385350Z digest=sha256:3596402fc634bddac7436f289e93395f98cb69dfa1a2a4a6466f4d8cdeacbdb0

Observation 2be34161-d75b-4350-bbdb-f542090ae388 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.388918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.388918Z digest=sha256:cec7418f1fa78375dae4d17ae952c399f0dade9fae0536ae71e74cee8030c6e4

Observation 44c08ff7-7763-4a7c-be55-7f4cba09cb1c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Moshi: a speech-text foundation model for real-time dialogue

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.393008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.393008Z digest=sha256:09e0f6d535f5c4be9eec6f11822e4018902251a2e19fc7cb7cd00354b5d7dde0

Observation f223fe9a-a743-4c20-8ed4-c5df12d9007d · outbound

This paper cites Building a Taiwanese Mandarin Spoken Language Model: A First Attempt.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Building a Taiwanese Mandarin Spoken Language Model: A First Attempt

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.396611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.396611Z digest=sha256:4a9e8c030f7db82b864f494757ca39e4c88bb0d535fe3636d3095d5695de8f4b

Pith citing papers

Observation 4c13b9ac-25c0-4f0a-bd63-731527226b1c · inbound

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information cites this paper.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.201734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.201734Z digest=sha256:1db0ba264f74c768e9e592f9d4e97487846611365248a64305295f9515376d98

Observation 3d7972e2-000d-4c3b-9f50-2eb310168e9e · inbound

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding cites this paper.

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:06.889837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:06.889837Z digest=sha256:c1ff4851a351681a8e61acb787b0d53b7416874de90f41316de3efc97d1549a3

Observation fe0d912e-77d4-477e-a30b-7c271789c230 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.074599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:c7f463a6d40d9bb13f5f05fc594fd9b195732a0881a5b756f9af3915636865bf

Observation c8555cec-11ee-4dd0-86a2-6b22ac14f501 · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.095207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:5d43bf64afaff4ee96ecd7e93a151598cd3962fb4aa3ebda3005ec7674db30a3

Observation bd865246-2e24-4a45-a29e-31ce1e8a1e8f · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.947534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:52a06865fc9d1b62c173fc1a1dbdc4d8b06545ab379be67c0d309e5e33d2dedf

Observation 865f8dc1-fbc9-4c8b-8930-afa1b769520c · inbound

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models cites this paper.

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:35:32.879177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:31:48.172439Z digest=sha256:bf3ac22179ad8c5197ed04a819de47d42ad1551d5e5a7400c46a34ef95711b17

Observation 3439d7d3-c933-4473-a804-f21bf2009162 · inbound

Large Audio Language Models for Spoofing-Aware Speaker Verification cites this paper.

Large Audio Language Models for Spoofing-Aware Speaker Verification SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:12:30.629444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:12:30.629444Z digest=sha256:49ed0d4b8fa1f033753552b2b7a51ba6493488dffc240068f12b859d1f75a71c