Pith. sign in

Paper Citation Record · LEDGER

Listen, Think, and Understand

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2305.10790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10790 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:27:18.704042Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae54b952-a671-44b9-a4dd-944c328f4b6e · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models Listen, Think, and Understand

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:29:46.312500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:c9baa40c95433776860e06a8d049aab82281f2e42dbbd98802de462217e9e162

Observation c0c7a3c8-2e85-41e5-a355-99e85750233b · inbound

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models cites this paper.

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models Listen, Think, and Understand

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:18.704042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:18.704042Z digest=sha256:6614b43209c05783ee9fe11f3d5dc30efdd80bbc250a4f4e3098cd0a0412432a

Observation dbd14a94-5ddc-4a21-832d-45d9ccd45631 · inbound

The Sound of Water: Inferring Physical Properties from Pouring Liquids cites this paper.

The Sound of Water: Inferring Physical Properties from Pouring Liquids Listen, Think, and Understand

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T18:52:43.060646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:52:43.060646Z digest=sha256:82277fdf4880fdff06f08fd9186e9fdd7be6ebde55a6f1e5a270c6e2cc26f3aa

Observation 06bc3aeb-6d83-4f5a-bdd0-4ea1250cd62b · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models Listen, Think, and Understand

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.300573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.300573Z digest=sha256:b95fabd942bf8f99a597851de1a74693a1e81a2a02486e51b309dceceb809c1e

Observation 473a6f79-95d2-4681-8828-1fbcfa87a156 · inbound

De-biased Multimodal Electrocardiogram Analysis cites this paper.

De-biased Multimodal Electrocardiogram Analysis Listen, Think, and Understand

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:09.744085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:58:09.744085Z digest=sha256:315256359e89a22b48d1b1c5c1017b19a1c023f5b5d403f0cc92c3331f23249d

Observation 2782f4f9-4361-4c18-841b-6148e704c862 · inbound

State-Space Large Audio Language Models cites this paper.

State-Space Large Audio Language Models Listen, Think, and Understand

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:44.491203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:44.491203Z digest=sha256:077b306d6833560d23d198d14a16770267731649e3d1149d09fee0c5c49a8a81

Observation e1c8a5a8-7a41-4c04-95c9-83791e443ca1 · inbound

Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding cites this paper.

Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:56:41.873408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:56:41.873408Z digest=sha256:3022d2b1661d9a951b88506be881e5e0cc07930b829a553fa83d90aff31443f5

Observation 7e376fd6-1a0e-4692-82fe-5598a7617139 · inbound

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models cites this paper.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Listen, Think, and Understand

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.804426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.804426Z digest=sha256:0a1a363896ff2646bc3091f9cf1085cf935c0dc29748a2ad5dfd460ef34520c2

Observation 549c0c33-09dc-41fe-9e7a-c3c6c616e85e · inbound

The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion cites this paper.

The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion Listen, Think, and Understand

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:57:13.360850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:57:13.360850Z digest=sha256:09a7069c245a9c1c385a387aead2e54997d8334312f7511d097d4640ff5d88d7

Observation 15418cca-6bf9-4945-ac03-c2774b6029c9 · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Listen, Think, and Understand

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.089824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.089824Z digest=sha256:46c73f94a1d5274f97b98cc1305810216b91c3f9af2b769c1fdbb298e270904b

Observation e4cb841c-e23b-4f24-9b3a-8ebabbb7dd51 · inbound

Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio cites this paper.

Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio Listen, Think, and Understand

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:40:31.444333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:40:31.444333Z digest=sha256:75e9697f470e98935702195456c0aea076d126d3342a9cd54a2aedec1987bdf4

Observation 0596e541-45d3-4386-ab7f-ca00c61bd688 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Listen, Think, and Understand

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.726274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.726274Z digest=sha256:b6641399010ebcf05b5e47c843b7c336ea2e309aa4e354ce184539845a43ecb9

Observation 3bbf07de-fedf-4d76-9bbd-affcaac80e81 · inbound

"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer? cites this paper.

"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer? Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:25.603086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:43:25.603086Z digest=sha256:f2373b1cec9ba15a25c22b5fedc805922fd5edeb0604b15d996865c5c708feb5

Observation cd77a381-8ad5-4801-b6df-7d80726574a8 · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.057432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.057432Z digest=sha256:a58847a0cd379be9cbdb18b1c0e8fa45b0774be9ab5073952c26b04f81d0a9da

Observation 1bcd468c-9691-4cef-ad2d-f78d419acf2b · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models Listen, Think, and Understand

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.487032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.487032Z digest=sha256:5b54542e3f2726a8230d0083ceeef3483160441933b1eb14750de26223083fbc

Observation 2dc77620-fe8f-4b0e-9d6d-95e4b0b6e9f9 · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs Listen, Think, and Understand

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:46.290154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:46.290154Z digest=sha256:ba1c3a0e77258cb4e6de305b76fd5c751ea28b39b783ba476e6352121cbb5a97

Observation 4ee1ee6e-52f7-4975-adb3-778e897e1cef · inbound

ALAS: An Automatic Latent Alignment Score for Audio Language Models cites this paper.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Listen, Think, and Understand

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.512637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.512637Z digest=sha256:5ab6ad1803bafffee775c3b3c3b2dbbe61c13332fdf06262481976dce28308aa

Observation 9f77e224-6bc4-4f47-9b94-1564d3bedd32 · inbound

ZeroSep: Separate Anything in Audio with Zero Training cites this paper.

ZeroSep: Separate Anything in Audio with Zero Training Listen, Think, and Understand

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:43.933695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:43.933695Z digest=sha256:ddfb4c1547f6728780926b4ed48c4d343dc48b0d6c75305510990b28545acd1c

Observation b19e6aaa-7840-46aa-9110-6e641672ca3c · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.670777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.670777Z digest=sha256:9938bc4b4680772e6a86dbab48fd235d45f5190150a83fa041cf5eef4d4da708

Observation 552328a7-1953-4cf0-8c24-8804d68a40a1 · inbound

Teaching Physical Awareness to LLMs through Sounds cites this paper.

Teaching Physical Awareness to LLMs through Sounds Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:15:46.367352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:15:46.367352Z digest=sha256:e8d6d17b73b7dbeb19ed59bd2b1f6a487071fbd4640a182ce73ff32bd73d37e9

Observation ed050362-11eb-4766-9b85-376437709e6e · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.382210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.382210Z digest=sha256:13a617c2c862c010dbf35733895a48db2346169ff191d3046fb29d927ba0da7d

Observation da6c1df5-b888-478a-a255-ece4e2b54297 · inbound

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition cites this paper.

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition Listen, Think, and Understand

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:33.044745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:33.044745Z digest=sha256:6a4078ff4bfd2bdccd5a38c03ba0c5077330128dba8333b03cf91fd1feafa09e

Observation 4862910b-5f00-43ed-aa02-27eb8444c0a2 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Listen, Think, and Understand

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:02.588908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:02.588908Z digest=sha256:187db12ab576e6fea7d6a38ae08842744e5f77bd759caee26843c295b1c57586

Observation 74ecce0c-3ad2-4bb4-8d8d-44359d549937 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Listen, Think, and Understand

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.125480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:a95092c04dc6e9a5cbc070cc42db831c60ece1c4291fee4f6edff25a7377d31c

Observation 003dd5a5-a870-433e-8397-f892f8ba0cb1 · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs Listen, Think, and Understand

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:48.796047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:48.796047Z digest=sha256:1a6f1eb96fb694242bd1c7f88c08e52b50a9703dcc9464c6149125251aeb85fa

Observation bac6bdd2-ca36-427f-a2f2-b3064ac08d97 · inbound

Improving Audio Event Recognition with Consistency Regularization cites this paper.

Improving Audio Event Recognition with Consistency Regularization Listen, Think, and Understand

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:54:57.049705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:54:57.049705Z digest=sha256:e5c4edd51d2707150c1a6c114296549ec4851593ca8c3edfae1b99c90f0d9f85

Observation 589131f7-a8e1-4d51-9804-9259c77714ec · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Listen, Think, and Understand

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.396777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.396777Z digest=sha256:6aa07f60715af4bf599ea33e8910dcf4e910e64a9d0450fbdc9586b542bf53dc

Observation 0c8682ac-6584-4734-bcc3-398ed0e6e7e3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Listen, Think, and Understand

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.361593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:e46d8c29e1ffa7fff0ba73b90a09abf6fcfc84da77f21c1e8628a1d207b6716c

Observation 229861e9-348b-4d92-95c9-344776e2ebe3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Listen, Think, and Understand

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:97bb95b89effa39fde141d1fc016103027c7d420e0332eced26dcdc22cb3fbec

Observation 879e6f45-04f2-405d-a31d-8cc5cf3ef6a5 · inbound

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training cites this paper.

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training Listen, Think, and Understand

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:18:07.291974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T07:13:29.041056Z digest=sha256:d560dc6d3b246b5ef8655618f0e65d4db25d35bea217e1327da6ec953323820d

Observation 812223e5-8250-45f0-81da-584609534252 · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Listen, Think, and Understand

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:44:00.649442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T06:43:52.735211Z digest=sha256:9b0a9fd7b6d8f11af6635605f853b8ed27a88a608d7c15a610dd583bd8bd99db

Observation 680c47f9-08d4-44d4-b32f-47f012750bae · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Listen, Think, and Understand

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.640485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T05:44:46.831360Z digest=sha256:784ecba709c2bd582fe287961564b43a51a7546ca21ff4cc814ff16c87bacb76

Observation f870973a-d165-4d40-a668-bba16ed9cd2c · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Listen, Think, and Understand

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.387037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:6cd4f2db414c22afa6be5af2e8a37f1c09c17ef4d1081d52330b64fe9282f574

Observation 66c1f981-7254-4d5d-899a-28b74ca0c293 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Listen, Think, and Understand

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.497306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:364a0b2f679a01b593703ed848afdedf3d273b72157c3d117f7c231deab34c8a

Observation a30199de-f6a5-4b06-947d-5661731f3b2b · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Listen, Think, and Understand

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:47:17.990037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:f920964228d0557cc7ac1a0750e90a526aebcd863fa693a4b7ceb8af981249b0

Observation 85c02625-1456-4d85-8dbf-f919989e25da · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models Listen, Think, and Understand

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.985482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:cecac1507893a6d0f122d1d1ec5eb33708c836db173b65c5ec314d54c65053c8

Observation ce9e0568-b53f-4f2a-9712-b936fe1e0f0c · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding Listen, Think, and Understand

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.186524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T16:59:50.615875Z digest=sha256:366c2f014d5af1d012238eb89b022fb3359dc0401cdb530d6a03fa73e2d37a11

Observation a95e9065-c56c-4d0a-ac26-03416fe65a3e · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding Listen, Think, and Understand

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:34:50.710186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:34:50.710186Z digest=sha256:7bb5c7e1036afc0482d56e8da1ee45b621cf69d213364d3520d49aae45da4da2

Observation 7250104c-9e57-4d1b-890a-bd66ef6cea78 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Listen, Think, and Understand

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.649038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:0ffc3f907ef5a14583d511b42dfff01b0fdba39e0f6254bba741c773f024d37f

Observation c1c12e22-8a62-4bc5-9aac-0135690ebf68 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Listen, Think, and Understand

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:909a2dc4dc791e22d312216af6f2cb3d07406e54adb55d640f6975feb9994f45

Observation e780d499-6b41-4c88-aca9-5acb343089f9 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Listen, Think, and Understand

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.155465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.155465Z digest=sha256:66b5b7aa723622ef0bb96179c562528a29a1eb11fc490aedbfa2dffef012985b