Pith. sign in

Paper Citation Record · LEDGER

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

As of 13 August 2026, this Paper Citation Record lists 100 of 212 outbound references and 3 inbound Pith citation observations for arXiv:2605.20266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20266 v1

Coverage vector

measured 100 of 212 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T07:38:23.099479Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:31:00.822147Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T14:31:01.440724Z

Reference resolution

100 of 212 outbound references displayed

  • verified exact74
  • verified fuzzy23
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 694d4f4d-6c8c-43a8-9e9f-bee119defb1f · outbound

This paper cites Train- ing language models to follow instructions with human feed- back.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Train- ing language models to follow instructions with human feed- back

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.817198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:f2d2f1a09c33c2ca42f112522558c91d263c5a6c2affd06658eb8b536fc2a993

Observation 1a6a01ad-9ee4-43a4-a0ed-bb0edc875175 · outbound

This paper cites GPT-4 Technical Report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.861990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:bab4218d0dc6e8d8da64021d17d88b77029ce826e8faa3802e3c6aeaf922266e

Observation 6d50fc00-bda8-4c79-9149-59500ddbf3dd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.799576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:a9fbee22350ad6e47c72aa2444dc89268c5462f8c97b2b4f0c0b09933ca5ca4c

Observation c4b18f69-8e38-4a43-856e-e7d1ea250db9 · outbound

This paper cites Qwen Technical Report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Qwen Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.802464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:9fe950723dc271af19c66ad2922aff3b947965b3ce1e1f0faa85ba851d12f355

Observation ccfb269b-f7e0-4e32-ad46-b412b6f1b0c0 · outbound

This paper cites DeepSeek-V3 Technical Report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook DeepSeek-V3 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.975191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:db997a82a9f69591c43de7e54b4ea917cce7f76298c39eddfac2c1e53971737d

Observation c88ab41e-9ac0-40ea-b897-df2b065cb616 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.972211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:8671647010e34fb413581b89015c32b54425be8100de1b055f80d64822ee61c8

Observation 71218e0b-0241-4385-9a0d-1001978ec24b · outbound

This paper cites Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.726419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2b5076b4f3cdc72c18a05c8ef52a1999c9a83241f042c99f0df95a5aa5601bb9

Observation 665e8213-d143-4275-8771-2388157dc191 · outbound

This paper cites A survey of mathematical reasoning in the era of multimodal large language model: Benchmark, method & challenges.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook A survey of mathematical reasoning in the era of multimodal large language model: Benchmark, method & challenges

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.759049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:c0dc380350b9f2671c5d20304600c65baffc68a4ebc2551bc475dc226ca65d9e

Observation 4e07d990-05bc-4a69-9622-ced1ccad604f · outbound

This paper cites Qwen3.5-Omni Technical Report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Qwen3.5-Omni Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.774649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:fd583a1b17cbdb4c64c52567701b40c9777175c3fefd4e9b9780469cbf129688

Observation 6c6d05f5-03a7-48d3-b83c-8c7ed597307e · outbound

This paper cites Sparks of Large Audio Models: A Survey and Outlook.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Sparks of Large Audio Models: A Survey and Outlook

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:48.960571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:b5ef33b8b2572fb0bdae9ded4af65b3b6896482a9bbaf15c3fa0bd135ae526d6

Observation d7f102fe-a4ed-49a3-a7ce-567ab1f56662 · outbound

This paper cites Audio-conditioned diffusion llms for asr and deliberation pro- cessing.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio-conditioned diffusion llms for asr and deliberation pro- cessing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.754984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:1225ec556069c378ea45a9e834d2b0455de2815390d5a698f6a545f00ec93237

Observation bc4ed5c6-98d8-4860-bb69-3c269283cbac · outbound

This paper cites Qwen3-ASR Technical Report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Qwen3-ASR Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.719800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:a71a163bf3cf72e7680805bd9d59832eb0bb00b47c1ce716e334ea17c7d8dbbb

Observation 4a6d298b-aaec-4927-a64f-f4e518264932 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio set: An ontology and human-labeled dataset for audio events

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.778557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:9c7490da19ed7ef6e84be3a0f09f4f1c105632d38afa9222c917e95621a62937

Observation 26867c5d-7cdf-48ca-b644-c27dd9cc5ded · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Panns: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.846414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:ef68fef4f5a7b7a35bdfc1ea80a43909d0902e9bb8d17a46aade12a9e7b81479

Observation eb3712d9-5bb2-4d85-8236-08592f036432 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.032020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:32e1728f223e37672c4f36f2fd35ac234b63f3b5b9c171d8ba26eaf9bb445abc

Observation c6e1a597-9785-44ea-9428-e31770419ecb · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.005020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:94a431e4077d739223f2cd3c7719bc8dbcb649cd5b7ec5f49f1f71f35ee83cdd

Observation d6f59265-9359-4906-b01c-3441777c4884 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.121346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:4e28c4d09e1e1b2a89a09d4b33aa26024687b30e77e3a81810e0e4702f8f12d1

Observation 0f2f1cc2-7c84-4d62-ace3-918330bf26f6 · outbound

This paper cites Qwen2-Audio Technical Report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Qwen2-Audio Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.995254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:226d27008adc3b81be95498f2cad10db442cad595df0b2827cd52de445c9c43e

Observation 1127c340-aff0-405a-a49c-8502257f8e5a · outbound

This paper cites Step-Audio 2 Technical Report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Step-Audio 2 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.989142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:935877b9e60bf77528f54e76c949fcd85400202e43a96fce9cf582f2373f4d57

Observation a6c798d3-1c16-4c0b-bb8a-df65629c8dfa · outbound

This paper cites Large Language Model Safety: A Holistic Survey.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Large Language Model Safety: A Holistic Survey

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.118397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:121fef1b7da9aaff7ff83c4b4a36172dd1d4d29ad3dfe2d0decb080bb8e32383

Observation 1ebf339c-e46d-4d65-addf-a2d9d2dc71a7 · outbound

This paper cites A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.124688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:98e08c7c88552ae35830b5b9b85bae8054bf7cd042cac9dcd74cf7af12d6e582

Observation 511d99d4-2429-420a-b5bb-040b1dad7325 · outbound

This paper cites A survey on trustworthy llm agents: Threats and countermeasures.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook A survey on trustworthy llm agents: Threats and countermeasures

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.844469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:f4338aa4e38e01333ef88bd50e27d5933b98f9c759721eaf76d88619dad5795c

Observation 6a8e730f-5ff1-46aa-8424-d485d76795a5 · outbound

This paper cites Safety at scale: A comprehensive survey of large model and agent safety.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Safety at scale: A comprehensive survey of large model and agent safety

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.842520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:f3c56eaa749e255397fa3526646ac2e38cf6c1fb9f6a72272d916785b4e68f79

Observation fcb8bd76-8422-4caf-9721-3cce9a275860 · outbound

This paper cites Hidden in the noise: Unveil- ing backdoors in audio llms alignment through latent acoustic pattern triggers.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Hidden in the noise: Unveil- ing backdoors in audio llms alignment through latent acoustic pattern triggers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.135181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:36067d5fb6f1c3656b6e5728e621d2060e3985bba882861c64987e40fd467cca

Observation 8228f3fe-ee4e-416a-8f14-6321626996dc · outbound

This paper cites Synthetic voices, real threats: Evaluating large text-to-speech models in generating harmful audio.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Synthetic voices, real threats: Evaluating large text-to-speech models in generating harmful audio

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.017035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:60b76e8e898daf7db781a166ff534ec3d33275fa23ad0018d4b197cf7ff553af

Observation 17dd6268-9d23-460d-9af3-3a77e27beb01 · outbound

This paper cites Evalu- ation of audio language models for fairness, safety, and security.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Evalu- ation of audio language models for fairness, safety, and security

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.754127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:ef91de1a43968c2ce9f578b68692caa4195eab8b6cc96760a8b69e770845c358

Observation 4e65afcc-bc2c-441c-b480-3433297dde8d · outbound

This paper cites Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.131246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:e4dcee93106e6b0d2ceb3c0a5564c78f5cc4a8e4a9ce66071d40274e087e8bff

Observation 9fc9cc4a-4674-42f3-80ae-4589d960a83d · outbound

This paper cites Spur: A plug-and-play framework for integrating spatial audio understanding and reasoning into large audio-language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Spur: A plug-and-play framework for integrating spatial audio understanding and reasoning into large audio-language models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.089781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:c75131799d8d29d2605bc5a7180d419d03cce80074c1a0baec339c3370796a20

Observation 03366fba-bcde-4bc0-b917-7b224505f92d · outbound

This paper cites Pal: Probing audio encoders via llms-audio information transfer into llms.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Pal: Probing audio encoders via llms-audio information transfer into llms

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.102787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:0e66330270ce1cb782d112cda8af6f0c486e0b568fe6dc8a399834954b80560a

Observation d0344805-9d67-4ee9-b91d-1601b26a7c17 · outbound

This paper cites The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.108827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:045dd89c98186f18ed5eea100165d01fd494cb1fc3be4bde14aaed9fa60db821

Observation 74c04294-e087-4126-9c6b-56957c546dec · outbound

This paper cites Llamapartialspoof: An llm-driven fake speech dataset simulat- ing disinformation generation.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Llamapartialspoof: An llm-driven fake speech dataset simulat- ing disinformation generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.838733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:5b4b3ea13a1b323ec5d0d02e1bff972fc4145a8fb918f03e918d927ce7baeecb

Observation bbefb739-4bcb-4b92-baf6-d0b16a47bf27 · outbound

This paper cites Dfallm: Achieving generalizable multitask deepfake detection by optimizing audio llm components.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Dfallm: Achieving generalizable multitask deepfake detection by optimizing audio llm components

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.080261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:c16e4b2fcb4511024b18aff3c3b0fadff0bcb4d5f6e7fe8d9c566bdfe959bd60

Observation 2a2b6d12-5372-44e8-a201-512ef6c09873 · outbound

This paper cites SARA: Stress Test Reasoning in Audio Deepfake Detection.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SARA: Stress Test Reasoning in Audio Deepfake Detection

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:04:12.916729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:914087d8460cead34d63e5a14bbc8189a74f54fafb7c9686dfa32d00d35e2e0c

Observation 5c629e03-3ee2-4f93-9b3a-70a5f23446ea · outbound

This paper cites A survey on speech large language models for understanding.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook A survey on speech large language models for understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.835077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:065d23e7f5a31e274cfa8a6eba2fdeb5672cfa42db1ea43aa96ac0356eaab485

Observation 9f4bed0e-2a05-402c-a285-dfbabd16d2bc · outbound

This paper cites Audio-Language Models for Audio-Centric Tasks: A Systematic Survey.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:62bbfe9f54bfdbeddc6ecb4ba2587afab2831de4b710fa224f64fd76d9e2d24d

Observation f35eb555-253a-4471-bc51-49ec37d24d6d · outbound

This paper cites Towards holistic evaluation of large audio-language models: A comprehensive survey.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Towards holistic evaluation of large audio-language models: A comprehensive survey

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.832872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:4027142fe82ccf7d64ab301a8ccdca2143c72b74df049594d0c7450dfbfde6ab

Observation 478d7935-bf90-419f-9588-e7702fdd17e8 · outbound

This paper cites A Review of Speech-centric Trustworthy Machine Learning: Privacy, Safety, and Fairness.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook A Review of Speech-centric Trustworthy Machine Learning: Privacy, Safety, and Fairness

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.127787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:c4625974917619ef08be256fb7488bb1d94896fa851deb07805d72af8826b67a

Observation 7b7963dc-e970-485b-bca8-05ac9c338ed2 · outbound

This paper cites Audio Deepfake Detection: A Survey.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio Deepfake Detection: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.684881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:89eb1b54959f12d736104dd163e7cb67bc07067c839c8489fb764f25387b6247

Observation 7a136817-6ac6-4423-a50d-9ceca150c06e · outbound

This paper cites A survey on speech deepfake detection.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook A survey on speech deepfake detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.824885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:7991ef05b81b8a5e627146916441b9e81226ad1d11faf2f3253e133826585d3d

Observation d5605ed9-cb1f-4f14-bcb4-281927ed6c08 · outbound

This paper cites A comprehensive survey with critical analysis for deepfake speech detection.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook A comprehensive survey with critical analysis for deepfake speech detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.820909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:28b3b6313d59285507b283f8e679ac579b6dfde852606d2aa9ffeea364fdb997

Observation f242cafb-2e46-4341-b95a-a5e247b7d1e9 · outbound

This paper cites Recent advances in speech language models: A survey.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Recent advances in speech language models: A survey

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.836923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:b30bb62ed0c23dac7eb161145332a6054bcd2481e4886025385d7e09e727b964

Observation c756bf35-c122-4e1f-9e74-1b68d9f4756e · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.765291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:1e192a5f79b71c2c44ee134f81e2148f5ef9d0e218bdb6f800c0bd94a97df4dd

Observation 622f736f-9925-43a8-a7ca-b854f80fd733 · outbound

This paper cites The interspeech 2026 audio reasoning chal- lenge: Evaluating reasoning process quality for audio reasoning models and agents.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook The interspeech 2026 audio reasoning chal- lenge: Evaluating reasoning process quality for audio reasoning models and agents

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.954048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:695044d0b2b221715c6ca8d360743371d29b68813f2dc0fc02c9b18be7d4c943

Observation 69bf8a91-a929-441f-8d9f-417e2c0257ba · outbound

This paper cites Sci-phi: A large language model spatial audio descriptor.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Sci-phi: A large language model spatial audio descriptor

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.763060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:4baab66c5577c689756aebbe26e518a561e69f9a741a5050aec8f04fa82450d1

Observation e0ca901e-a703-4296-8371-51c885591bfd · outbound

This paper cites It hears, it sees too: Multi-modal llm for depression detection by integrating visual understanding into audio language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook It hears, it sees too: Multi-modal llm for depression detection by integrating visual understanding into audio language models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.665573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:7144c6372c0e056144667ed002a4f645b641f8ce5809686380ca2c7e4f234feb

Observation bd828b01-e59d-422d-8fff-61692e943548 · outbound

This paper cites Wearvox: An egocentric multi- channel voice assistant benchmark for wearables.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Wearvox: An egocentric multi- channel voice assistant benchmark for wearables

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.944786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2de24fc8a7854d8f29493f89d346dabebe8cb9bd3adc679aaea23b1e74184835

Observation f1ac5a27-310a-41f9-9d21-2adf5733ec1b · outbound

This paper cites How auditory knowledge in llm backbones shapes audio language models: A holistic evaluation.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook How auditory knowledge in llm backbones shapes audio language models: A holistic evaluation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.907955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:3c3673584455e901744657b0a5c4595453c41035f373ec081f621a3d487ccca2

Observation 1f82e285-9cf7-4e9f-9430-b6b4621847c5 · outbound

This paper cites Salm: Spatial audio language model with structured embeddings for understanding and editing.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Salm: Spatial audio language model with structured embeddings for understanding and editing

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.019768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:8e0724048e6cef5cc0cd6dd7c07c7d58faf72d18807acb95a8db4959778b157a

Observation 3a85349f-211d-4f1e-bb04-61872c10a105 · outbound

This paper cites Latent speech- text transformer.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Latent speech- text transformer

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.656510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:9ac8cbc706e8e8bda70f7aa29d4b04010825e8fd2e17829d00ba2e73da974905

Observation a6877885-3c6d-49cc-9ddf-860bece034aa · outbound

This paper cites Uniaudio 2.0: A unified audio language model with text-aligned factorized audio tokenization.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Uniaudio 2.0: A unified audio language model with text-aligned factorized audio tokenization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.897242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:9f69100461d5556e434b237a4899664a43bf468597a9521bd24650e32ef3eb6f

Observation 8134bbdb-59df-4ba9-91de-6beb2d1c3d04 · outbound

This paper cites Towards audio token compression in large audio language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Towards audio token compression in large audio language models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.870971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2b21da6417906772cd45cd26cdf77100ade151329d57112c12b307958e66336d

Observation b6264d41-974c-4c6e-81d7-16bb7a80d931 · outbound

This paper cites Vowelprompt: Hearing speech emo- tions from text via vowel-level prosodic augmentation.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Vowelprompt: Hearing speech emo- tions from text via vowel-level prosodic augmentation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.963339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:e33b62cd412bc38b5c88e72d82b8555538a78b5ba6fe5f586e1035747a8624cc

Observation d1e5b316-d06a-4960-a03c-44daa7b02aeb · outbound

This paper cites Moe adapter for large audio language models: Sparsity, disentanglement, and gradient-conflict-free.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Moe adapter for large audio language models: Sparsity, disentanglement, and gradient-conflict-free

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.811484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:487bda304f02e39f4ec7c198cbf37ba34e89a32f1895af311895ee2994aae82a

Observation 4e00dbb0-2c33-46f8-944d-bdb78dfa175b · outbound

This paper cites Segmentwise pruning in audio-language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Segmentwise pruning in audio-language models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.653295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:e142bbc87b3befea642925f418c794cc8ba086cd29e6f69f4a9705334f8e39fd

Observation 9371d7ec-62d0-4181-a271-12d2e274e222 · outbound

This paper cites Fine-tuning large audio-language mod- els with lora for precise temporal localization of prolonged expo- sure therapy elements.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Fine-tuning large audio-language mod- els with lora for precise temporal localization of prolonged expo- sure therapy elements

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.978533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:1333e597ac58f0a8f335c32d669e5992c66ceed849ccc43131a0c07ab6385339

Observation a0be3ba8-caa5-4852-ab22-dc9166fd8491 · outbound

This paper cites ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-28T02:04:13.437222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:685d4c3c859f351137a0cceb61a6bbffe87276abfd996dde45a80a5d764637b2

Observation 85bfcc72-19f4-413b-8ab3-f5cef501f160 · outbound

This paper cites Extending audio context for long-form understanding in large audio-language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Extending audio context for long-form understanding in large audio-language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.753014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:df71d4e0db4816b43d6088c770b2cd396e119e8541785af86c3cf62bb9171a09

Observation a762a375-8756-4105-97dd-8bc156d1cf14 · outbound

This paper cites Listening between the frames: Bridging temporal gaps in large audio-language mod- els.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Listening between the frames: Bridging temporal gaps in large audio-language mod- els

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.840626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:44fd1f91a1d7fb62ce33d54dfcb8ea35610bd503d2e49ec7648f37852e19e106

Observation 16001722-9dbc-43e2-a94f-cdf2e47fdea3 · outbound

This paper cites End-to-end contrastive language-speech pretraining model for long-form spoken ques- tion answering.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook End-to-end contrastive language-speech pretraining model for long-form spoken ques- tion answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.756949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:f9767bfe8c90e11b8e2d26fd60d7ef80ba670c09e6c82eb2e95f18bd212b059a

Observation fc930028-8e59-4a9e-b904-3101dc7422b5 · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Mimo-audio: Audio language models are few-shot learners

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.876692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2034b38fb53f1f28ac5c5252ab5fb8aa1422775370bd1cbfc235e1961683951a

Observation dc7f0c5b-d58e-4804-a8fd-bde7c8fdd1cd · outbound

This paper cites Pay more attention to audio: Mitigating imbalance of cross- modal attention in large audio language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Pay more attention to audio: Mitigating imbalance of cross- modal attention in large audio language models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.762945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:7223944596955c061c10c794400f1b6df5f5a5f3583b0626de9c6d7fa8ec0b12

Observation b2eeec74-f6b4-4aaa-84f7-2f8d2f25d4b5 · outbound

This paper cites Measuring audio’s impact on correctness: Audio-contribution-aware post-training of large audio language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Measuring audio’s impact on correctness: Audio-contribution-aware post-training of large audio language models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.745260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:b90db62fb5b9a5596d998583a4f7b62aac8d79cb1dece8400c8d6e76f2c7a994

Observation 777b3c29-c517-4f61-993a-314ec0faac2b · outbound

This paper cites ALARM: Audio–language align- ment for reasoning models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook ALARM: Audio–language align- ment for reasoning models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.777599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:837d042e2c5f6951c951ee5715a9d9a5c484ee36e3ddbdcb759698f304c20c45

Observation e2e6d3a6-16dc-4ae0-980c-70cdc4b0fd6c · outbound

This paper cites Sightsound- r1: Cross-modal reasoning distillation from vision to audio lan- guage models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Sightsound- r1: Cross-modal reasoning distillation from vision to audio lan- guage models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.852598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:0d0b65cd937242d0fce99d8f893a37aed76ac64b922007567260161172afb32d

Observation 6bf543d6-bdeb-471e-a7fa-2403dd2e9ef8 · outbound

This paper cites Cord: Bridging the audio-text reasoning gap via weighted on-policy cross-modal distillation.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Cord: Bridging the audio-text reasoning gap via weighted on-policy cross-modal distillation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.643960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:bc1976939b364beef76bd0ea162492a27218d884fb4a53ec89fd53b5213bd94f

Observation 50abe11a-3eb8-4798-bb52-8280d0da6e07 · outbound

This paper cites Attention-weighted centered kernel alignment for knowledge distillation in large audio-language models applied to speech emotion recognition.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Attention-weighted centered kernel alignment for knowledge distillation in large audio-language models applied to speech emotion recognition

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.911456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:b6e0974f14161c04eccc3512b800d8461e5d42e74bf47cd6fc0809adca60f237

Observation 0c16837a-701e-4bb2-b07b-f7e282ccd707 · outbound

This paper cites Feedback-driven retrieval-augmented audio generation with large audio language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Feedback-driven retrieval-augmented audio generation with large audio language models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.059461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:0727d82c114866a1bfb2833bace19dbdb2b0a5785bd5ae37273107e0bae916ad

Observation a834effd-d46b-4a95-ad35-ba7c2ef013da · outbound

This paper cites Emo-tta: Improving test- time adaptation of audio-language models for speech emotion recognition.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Emo-tta: Improving test- time adaptation of audio-language models for speech emotion recognition

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.062382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:551bb09d8a16ee5551f728f0635d13a205f288e5bd3dfdfdeee45cac88fd6ab4

Observation 3e1a5b93-42aa-42d3-b65a-0f686d7bb436 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook WavChat: A Survey of Spoken Dialogue Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.050610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:b84750f17af917ab3153b9f197d1c39ecf4161bede63ac69d7f7b5d2d1d1c8b8

Observation 45edfc70-9523-4b0c-bf8c-15961c8ea932 · outbound

This paper cites From turn-taking to synchronous dialogue: A survey of full-duplex spoken language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook From turn-taking to synchronous dialogue: A survey of full-duplex spoken language models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.710604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:12c28a415b8b83a03ce1aa71e0e0200aa2134125e3c8c78ba7c61aff40036da0

Observation af44c0a8-9c92-4863-be2e-b5a09527af87 · outbound

This paper cites Beyond the turn-based game: Enabling real- time conversations with duplex models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Beyond the turn-based game: Enabling real- time conversations with duplex models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.823005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:de436e2a2896af5555da911f0b8bef9997cef4d28085756313223f06b1f298aa

Observation 61fc04fe-291b-453e-aafe-84f5e13f9255 · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.723184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:d2860c5d4e2167e14206fc49a0b8db5ebc8889ea0d504472c46f6cbfe5595ee7

Observation 38d2eaf1-5dd8-4436-85e2-856c68393d1f · outbound

This paper cites SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.873747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2808481ccf29860abdf4d50255753034af3cc06cceec57491b9aafc6e957e948

Observation e606ea92-8c06-4e21-9526-d6c9c9e1fed2 · outbound

This paper cites X-talk: On the underestimated potential of modular speech-to-speech dialogue system.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook X-talk: On the underestimated potential of modular speech-to-speech dialogue system

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.678800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:1b587ae3760bd3260d3cd0475bbfdec47c0b65997d9189d4db7556ace1bef58f

Observation be54a2f3-04cf-4627-8db8-70f9a5e5092f · outbound

This paper cites Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conver- sation.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conver- sation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.115100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:27ff47f37b25e35c1fd6db5038b929017096d2e4c74c9ae4714fba39ee68927c

Observation bf962174-d5e0-472b-8040-867b5abb5218 · outbound

This paper cites TiCo: Time-Controllable Spoken Dialogue Model.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook TiCo: Time-Controllable Spoken Dialogue Model

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.984447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:d0a374ceb7e278ef6540e087c34cc6e6fb91ab3590e8185562ed9fa8b34683fa

Observation 9e8d5479-f8cb-4d95-9e96-ba3ab84b27cb · outbound

This paper cites ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.650067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:31c20a090cdf5c02c871242c7cbcf13866c256dd05bc37518620887d6d0f7153

Observation a7c577a3-c9a2-4781-971f-36e086d45258 · outbound

This paper cites Flm-audio: Natural monologues improves native full-duplex chatbots via dual training.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Flm-audio: Natural monologues improves native full-duplex chatbots via dual training

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.662521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:9f796e6e274796e19a960229ef72bdbd19ab9e5312f6c458b4948555c8fc1db0

Observation 654553f8-955e-4c16-9ce8-bac3cabd52a1 · outbound

This paper cites MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.783203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:50895e3bac60473b4b74c622e2cf2df6ba5b994a4ef1537e4c3c223b43748343

Observation 034d76da-74bc-4a1f-8528-057f34dcb6a4 · outbound

This paper cites The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:48.786532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:85d1ff25988b8ad7d1c3174385ebbfbdb79f4d616c936bd3e80219bf61b0b5f0

Observation 2d0d38c5-f495-476d-b209-703ffee56fbe · outbound

This paper cites Privacy-preserving end-to-end full-duplex speech dialogue models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Privacy-preserving end-to-end full-duplex speech dialogue models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.859379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:bdf12f08939d96c7c4a226dd681477d5f2ed4869a92724a916439c0fd96a22a7

Observation fee3831a-cbb2-4ee3-a1cf-2e8a07c88764 · outbound

This paper cites Covo-audio technical report.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Covo-audio technical report

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.690963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2e43a05d4f0697bb79dc6efd6510a058cb047bce832a49761f900fc981fd401f

Observation 1ab8df7f-4c69-45d5-af14-c9dd7a7ca169 · outbound

This paper cites Thinking with sound: Audio chain-of-thought enables multimodal reasoning in large audio-language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Thinking with sound: Audio chain-of-thought enables multimodal reasoning in large audio-language models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.922078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:8890f4296af378ae8e563303b812fa46f0ef17941fc35eba3901f373a54c3884

Observation 310f6a7a-68d9-4d62-8292-f52092b95b66 · outbound

This paper cites Echo: Towards advanced audio comprehension via audio-interleaved reasoning.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Echo: Towards advanced audio comprehension via audio-interleaved reasoning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.829044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:e1f9db78dd109ecc1b9795cd1b1d358aabc90bec0733253281c6825fe394ba96

Observation 8b58e20f-7775-4376-8202-19d9b5e928bf · outbound

This paper cites Nudging hidden states: Training-free model steering for chain-of-thought reasoning in large audio-language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Nudging hidden states: Training-free model steering for chain-of-thought reasoning in large audio-language models

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.904441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:717b31e7071b03aff2bde9f61e428137c9288290eaf18100d0eece6d86166458

Observation 05e21043-7cbd-4820-90c5-a3cd5bfec5eb · outbound

This paper cites Can speech llms think while listening?arXiv preprint arXiv:2510.07497.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Can speech llms think while listening?arXiv preprint arXiv:2510.07497

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.879572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:9853fb7b25f13227d837768dbacdbfeed214fc7de351732178e59de128589c94

Observation 693ea396-46a3-4d07-986b-0fc30851d3b8 · outbound

This paper cites Incentivizing consistent, effective and scalable reasoning capability in audio llms via reasoning process rewards.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Incentivizing consistent, effective and scalable reasoning capability in audio llms via reasoning process rewards

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.780542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:a45a9e02f6b9df28de800743509faaf62db9f2a59e36c52c1c413b3e0117447f

Observation dea77beb-84cc-4e0a-b12b-ee11ac5e4ff2 · outbound

This paper cites Emo- rl: Emotion-rule-based reinforcement learning enhanced audio- language model for generalized speech emotion recognition.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Emo- rl: Emotion-rule-based reinforcement learning enhanced audio- language model for generalized speech emotion recognition

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.834849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:cd34346f055a742c813bf15741919d3fbebb31da9de85bc891722a23a412b2c7

Observation 8c77be89-b810-42a5-81ff-0036fd64df1d · outbound

This paper cites Audio-thinker: Guiding audio language model when and how to think via reinforcement learning.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio-thinker: Guiding audio language model when and how to think via reinforcement learning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.640261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:de61c935ea25b7859a894da628597c81e3cb6e032e4273fe0c06c8868340513f

Observation fe4f1a89-8be1-4536-802b-8a6fd2f884ca · outbound

This paper cites Soundmind: Rl-incentivized logic rea- soning for audio-language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Soundmind: Rl-incentivized logic rea- soning for audio-language models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.815137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:fa1bc3cc482466dbf175222299cb2857efab7f30df9de479643622a5c59693f8

Observation 08dd4019-555d-41ef-a045-f69dd010ffec · outbound

This paper cites Think smart, not hard: Difficulty adaptive reasoning for large audio language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Think smart, not hard: Difficulty adaptive reasoning for large audio language models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.637079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:ae91ff1f8614cd8c2e08415a228bcc2a8f76c405159f6c7780d3b41c230221a1

Observation bdbaac45-b97f-4d07-8191-80a589687c39 · outbound

This paper cites Decoding ambiguous emotions with test-time scaling in audio-language models.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Decoding ambiguous emotions with test-time scaling in audio-language models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.630868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:3ec5a9225977693c723a5fe5c9fc887d2601568de23fa9d3ce3daee62c72ea25

Observation 5f2a0e8b-49ba-4742-a0bf-0583f24dfd86 · outbound

This paper cites Au- diotoolagent: An agentic framework for audio-language mod- els.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Au- diotoolagent: An agentic framework for audio-language mod- els

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.938902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:04d0e3925460cf9b257b8a54cc8fe3ada55a85f62c2fdbca88480813d26366fe

Observation 12502bd8-2070-4215-a6c0-8382fcec15f9 · outbound

This paper cites Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-14T02:20:20.750907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:867493f99e2676c3886a28871282a1f004e7b4441bf4782283e3d94c0f7d2261

Observation b66ee77e-b10c-42d0-aea1-7a422be4908d · outbound

This paper cites When tone and words disagree: Towards robust speech emotion recognition un- der acoustic-semantic conflict.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook When tone and words disagree: Towards robust speech emotion recognition un- der acoustic-semantic conflict

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.831950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:3d778a024f203b27faf0f1e52141557b6fd1a8eba10f04a6889a471bf48ec027

Observation b9fd82de-09c3-44a5-b2e9-acfc33543015 · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:48.826169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:e2ab5ab44e2e5759dc0fc13e6a39db2b1ecc55d0e23c8242210cd9b601a3d45a

Observation bd69a664-fcfa-40cb-b51b-6f042b620544 · outbound

This paper cites Generative spoken dialogue language modeling.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Generative spoken dialogue language modeling

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.805530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:b57533e2b78a806ce8b4127f422dbbebcc7e72c3e0f658a9b873a9bf1dc4da53

Observation 9c8dfef1-2b9b-477b-96f5-49dd2f8fc4f9 · outbound

This paper cites Pengi: An audio language model for audio tasks.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Pengi: An audio language model for audio tasks

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.803819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:f805c0589419839379f580f5e3b33950a9512ee144452561767f386acaf79e96

Observation cb5bceb9-4a50-4a71-8364-691e61c04010 · outbound

This paper cites an unresolved cited work.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-05-21T07:39:49.807192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:5e785f28c26de2eb78470f5c1b912e0a395aba551bc742c6bbda785c5c2ad93e

Observation b0af11f8-4b65-497a-9583-98c984c857b7 · outbound

This paper cites Spoken question answering and speech con- tinuation using spectrogram-powered llm.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Spoken question answering and speech con- tinuation using spectrogram-powered llm

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:39:49.809145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:94901330abf11806c475f132fd4d038139def4bfb369b94894cd802a4faf152d

Pith citing papers

Observation 08cf7ec4-5cdd-4f04-b090-0b7f13e7fd5c · inbound

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses cites this paper.

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

Reference 251

Resolution
unresolved
no resolver link, observed 2026-07-13T17:08:58.831798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:08:58.831798Z digest=sha256:cfd9218952dba7df70cf44669be615a89ed8d0d3c636388786302855512b687c

Observation 1e59c1ab-3171-48be-92de-7bddcd9d1d8c · inbound

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding cites this paper.

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T02:46:58.160847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:46:58.160847Z digest=sha256:d89bec98f3d105cc46749bb5626a6b16a4921256b7aa6f8bbd39660b27dbff98

Observation 0f456d3a-1629-4e93-9508-c92ebf808190 · inbound

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection cites this paper.

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:31:01.447278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:31:00.822147Z digest=sha256:515442dcb9f4ad3f7021d22a21eba42c3a5ccd213fd2111c128ab8d73270ac97