Pith. sign in

Paper Citation Record · LEDGER

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2406.11768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11768 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:54:13.519625Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6988f0d0-69cc-4727-896c-f2bdbc257da9 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.293952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:97f44888d5e15a7f5c42eb5f072200eabfa8ac6c119ba5ce40a46dabcd11dc7a

Observation 06d1d233-677c-4da6-a2ad-61236a8de789 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.395741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:41e22e94dcf1aea4a25f58a48e54bb98ff0a8cd5cbdb096b88a92977e9dc3cd6

Observation 9c967c50-3543-44f4-b64c-13e5197827fb · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.519625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.519625Z digest=sha256:3c3630bd2b5f93a1da9089baa13506928e3e7317bdcc0c7cef4be4d44706fe82

Observation d7d42a5d-675f-4ef8-b663-e85edb03d41f · inbound

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses cites this paper.

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:53.686731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:53.686731Z digest=sha256:b42a782daef8ac258dc65613110e1a0280c5669380db0a35f3bca4705d543535

Observation 2e732449-e696-4d5f-ae28-3f39fa1338d7 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.380266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.380266Z digest=sha256:e14888111f8711d8963f4af2ceac8a74813810dc024bcda6e1e54ba2a647ea9b

Observation a4ccc679-865f-47a5-9b35-54313167a882 · inbound

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents cites this paper.

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:45.070152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:45.070152Z digest=sha256:3970ae576e6645e362d1334942f84daeb873c2bc030681fd46b393250bfe1f79

Observation 97846861-2b18-4eba-a1bd-0d568262121a · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.254194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.254194Z digest=sha256:3d44e7b8f4598fb3bdc769256a23a07069e2565f7ca247ae1d6a14a936c8d8e8

Observation b03fa92e-d669-4b4c-b3cc-618d47076bdd · inbound

From Sound to Sight: Towards AI-authored Music Videos cites this paper.

From Sound to Sight: Towards AI-authored Music Videos GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:28.522442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:28.522442Z digest=sha256:d8d290e8424cf9b5a7229a039c57f818363b940944c788eccd1d9c882683c4d6

Observation aaef49ed-b065-46dc-bb5a-1772b68ba670 · inbound

Improving Audio Event Recognition with Consistency Regularization cites this paper.

Improving Audio Event Recognition with Consistency Regularization GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:54:57.124210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:54:57.124210Z digest=sha256:3b1ecf889141c2dc7ebbbc221355a22cf54dca25854fff0d1a32378bc438c143

Observation 9ac6ba19-b231-44c4-a136-9a515157f6f3 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.187373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:f38b5daa75e0dc8c91e172135f1a1497aaf75658fd2de7dadccd31fce1eefcf0

Observation b8a97e48-dde2-4ad5-8b68-ff191afc5bd4 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:59.901861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:59.901861Z digest=sha256:c010fc41c84e5ab3f17192d1b1b1596be47680035c0a70aa82d7aef6cc15352e

Observation 807caf8f-a394-43e0-8e96-8cd188daff0d · inbound

EvA: An Evidence-First Audio Understanding Paradigm for LALMs cites this paper.

EvA: An Evidence-First Audio Understanding Paradigm for LALMs GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:28.996341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:28.996341Z digest=sha256:56b18d7b769a25f904cfb644c55fe9fae3c000541c17ccb2a8568060ed9958b7

Observation 2897d192-ccdc-40cd-ae15-6f9526583675 · inbound

Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering cites this paper.

Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:10:58.086055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:16:59.034480Z digest=sha256:9a9e7e79443044e62eb8d6b56b909cee4902f06e7d4a306fb383e94cfc663f60

Observation e8da2d69-0366-4eef-81b8-3841b3aeff5e · inbound

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding cites this paper.

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:38.897129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T09:24:25.616750Z digest=sha256:0d3fbe3fdf93c847b5fd6511e2045a9f6064221d382e221ccca610b6b190dfe4

Observation d74d148d-6114-4231-a16b-f7a6508132cf · inbound

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models cites this paper.

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:19:20.685275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T10:16:26.526986Z digest=sha256:19924513abb99d265bd4b64f5b3e8857862b5be952849436b9db59d711feab35

Observation 34b70cd5-64b5-4786-8ee0-e3be2ff31c97 · inbound

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech cites this paper.

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:35:14.575239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:34:11.297579Z digest=sha256:e923cac3fb13ed14260a157d26c94f68f328137bd99fdf9f93a9fc4ca60e5bfb

Observation 6fed932d-bdcb-4364-b48d-b905fc9f6aa3 · inbound

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models cites this paper.

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T19:28:32.792370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:28:32.792370Z digest=sha256:07a42a9a68bea217267dcd1b547284334b7ac8b3f2974bd75892ab1e6b29032d

Observation b41efbd5-ea20-4409-a767-f73043369a62 · inbound

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints cites this paper.

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.830453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:18:41.816342Z digest=sha256:44b5b3b0fe3f70d10c1b55f33f75c762dfb25437e05e753a4bec1544da70c463

Observation 442e0517-a4c0-4866-b694-c9b6a67c0ec6 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.181500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:72f452c5c27f63d7db9a2bd91bc64276ced83e20f3a18160c48201cd0866f0a3

Observation 4f5c37bc-1184-4884-bf90-ea0a7a740cfe · inbound

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio cites this paper.

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:45.495210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T09:16:25.450336Z digest=sha256:075b7b1d79203bb6e0417477ad80f8e8f023e9ef2c60b5b652caa5bdb2122b2b

Observation c8e42549-7300-4a39-b422-21309c1f4c80 · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:8879d13ec95777e9bc6ca4c74eb75df51cae34be9c218a99fa75d2c80e264502