Pith. sign in

Paper Citation Record · LEDGER

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

As of 13 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 0 inbound Pith citation observations for arXiv:2608.09435.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09435 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:31:16.265521Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 103 outbound references displayed

  • verified exact5
  • verified fuzzy43
  • unresolved49
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04ff485b-0e38-41e3-be6e-6ee24df84986 · outbound

This paper cites Proceedings of the 6th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2021) , pages =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the 6th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2021) , pages =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.781188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.781188Z digest=sha256:a2c008cefce73fad6bf9fcd4406132d037ca44f2dd0c2181509d7176721bf561

Observation 47f0f2ea-d125-414a-9cce-cecf045661f8 · outbound

This paper cites arXiv preprint arXiv:2606.19348 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2606.19348 , year=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.788703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.788703Z digest=sha256:84f0551efe9ded831e10fc4a28ecfa797eb9a0c483d7f36b8f164980c98a17dc

Observation c202388e-43b7-4d69-ac9a-606476f21451 · outbound

This paper cites Kimi-Audio Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Kimi-Audio Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.794034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.794034Z digest=sha256:f6d6773279daa1930efe16d21c13ece525d739f400ea30cb682c01ba792270b3

Observation 843c9c46-6a0b-4418-9461-c2a2d923b5ee · outbound

This paper cites doi:10.21437/Interspeech.2021-698 , issn =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models doi:10.21437/Interspeech.2021-698 , issn =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.799394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.799394Z digest=sha256:c8d36e402071f82db054a73e258990fa867383af7b645572e95e185aa535627d

Observation cc6b64d3-383c-444e-a4af-c9a07548dd9f · outbound

This paper cites Advances in neural information processing systems , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Advances in neural information processing systems , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.804735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.804735Z digest=sha256:2a829009bf3d46141be144a362c192354d36203ee0f789ccea1db403bb38762b

Observation 5fe10506-c759-43ce-9ee9-8cbfb99ce73a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.809739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.809739Z digest=sha256:1a4bd09c6768556692345a8e95627791487c22de9ae5a0dd2ae8222fa35ef3dd

Observation f23b135e-6e3b-467c-b500-eb567cae0515 · outbound

This paper cites ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.815138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.815138Z digest=sha256:dc25f79f94e0d9979ee41899ecf212b0e4456ef5aee9f6566157cea91b171f49

Observation 41031b32-f7dc-4a41-8e01-1c39b633167f · outbound

This paper cites Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.820170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.820170Z digest=sha256:faf534418fa96108b5350a2e019a91f01146a95a0f38e7789a3ecf822fa67367

Observation e1c0333d-ec30-4f56-b9b5-d87866de500d · outbound

This paper cites Classification Problem Solving.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Classification Problem Solving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.825345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.825345Z digest=sha256:399de36eb072ea949fd1382dda8d5ecc6d0e5e1ef5e0b30236c5ad5f6417ecb9

Observation 8ee26178-0066-4839-9811-656a3f829e03 · outbound

This paper cites , title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , title =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.830677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.830677Z digest=sha256:0cfd2185405acd8c3f72bc1a1e318467fbb4a64dd54d1b32f14bf2e2b8631f62

Observation 4fb795f0-557a-4ac6-8f84-560f27a32343 · outbound

This paper cites New Ways to Make Microcircuits Smaller---Duplicate Entry.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models New Ways to Make Microcircuits Smaller---Duplicate Entry

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.835603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.835603Z digest=sha256:eb3a4ea826f2d3bb6b6d7cd35f0279ea6b0be6b578492e1d636a6dba0231efbf

Observation 82c3f7f9-b256-4717-9621-46724c1695ed · outbound

This paper cites Clancey and Glenn Rennels , abstract =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Clancey and Glenn Rennels , abstract =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.840692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.840692Z digest=sha256:54ffee110022eb2cfae10b6d2213ceede8f3adb7323833a40dad4d4dadf09d13

Observation a3068a52-6aca-47e8-ae3e-28e220fc6615 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.845666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.845666Z digest=sha256:baffc3b05fecdd20ff5950e43903f81072f8ff80d9a925bfde509ab88a88e018

Observation f0b5e523-5afd-4076-a883-f2d74cf9bf09 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen2.5-Omni Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.850604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.850604Z digest=sha256:9b24443cb7b7d6818a61ec0d99ef247fb130120ad5cd80f038466f57c8672870

Observation 3cbf6da2-c105-44da-adf0-9723ae850916 · outbound

This paper cites Qwen2-Audio Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen2-Audio Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.855588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.855588Z digest=sha256:4956771fdcf38c30d1a63389b1871889e8bc83f1d0f54722ecf0f80b583d9937

Observation 2264e4cb-49df-4d39-810c-01d03a6e82a7 · outbound

This paper cites Qwen3-VL Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen3-VL Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.860490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.860490Z digest=sha256:767cde2221f1b7bdb92c7c8a724ef10be41eabca568739c0fa34c01bf39a7c94

Observation 3a5cc7ef-fb8f-46d9-b5b3-cd28baa34d01 · outbound

This paper cites Qwen2.5-VL Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen2.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.865581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.865581Z digest=sha256:71b3d4eda241a379e58e08660cbaf53fd438ab7246dc5ee9c1dda5c42f550cf4

Observation 5ec19816-a098-41b9-a12b-8c4baeaa26d0 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.870460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.870460Z digest=sha256:7c8f3d4f9b7d270b79e772d5e429057e4e279564c156aedaf5104931d0f347eb

Observation 3023b98b-7fdd-425f-b48a-c0b9cc93c7bb · outbound

This paper cites GPT-4 Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.875213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.875213Z digest=sha256:fcf9ed7fb61556871e14a72378366307a8963dced400bd2f3df5a7ab52a81a3c

Observation 5ee4db9a-a080-4102-8f84-6c395e45cc25 · outbound

This paper cites and Rennels, Glenn R.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Rennels, Glenn R

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.879683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.879683Z digest=sha256:1867110514842faa16ca6db2818264140c910865467fd3fddb7126bd7e697d82

Observation c9cb7848-35a8-43c8-a96c-5626202d39d9 · outbound

This paper cites Poligon: A System for Parallel Problem Solving.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Poligon: A System for Parallel Problem Solving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.884011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.884011Z digest=sha256:4d2b616cc40dfecd45708c31bdc60984027a271d6134ad25fd257a8c4de23dd5

Observation 24a9b67a-1139-43c2-ad0f-9ccf2e4259aa · outbound

This paper cites Transfer of Rule-Based Expertise through a Tutorial Dialogue.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Transfer of Rule-Based Expertise through a Tutorial Dialogue

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.888398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.888398Z digest=sha256:a3e17a2b88ded1be2ac4d83ed0e18e48bf49135c9d1c26ba333abc094fc3f07c

Observation 26d93712-6eb8-4a18-8b51-98d4f49a2f53 · outbound

This paper cites The Engineering of Qualitative Models.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models The Engineering of Qualitative Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.892716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.892716Z digest=sha256:dca4d6341d671bbe8c64ed5c9612f74fb5110d0c9b333f343fb6e7b683985f26

Observation 66219162-e6bf-408a-aa00-6a97cb268638 · outbound

This paper cites 2023 , eprint=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.897094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.897094Z digest=sha256:e445a9af105231292eb9d376e201265af1e61f84d900e867e0422f254c5236ce

Observation 7d1c5648-1dc1-4ccf-97e4-8dd99b8375b1 · outbound

This paper cites Pluto: The 'Other' Red Planet.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Pluto: The 'Other' Red Planet

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.902110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.902110Z digest=sha256:a769c67f3d3889424ce679873669e0ae9d9fba344fc77ea0f48acdf035421856

Observation 676bcba8-5172-4f13-9719-69bdeb104726 · outbound

This paper cites 2023 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , volume =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.906340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.906340Z digest=sha256:327fe595746bfc2e15edf92fc5de795bbc91de838fe7244ef46f9ba71ab43683

Observation a012fd13-8076-4cee-b958-643ac5f54f41 · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.910427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.910427Z digest=sha256:d1dd0b6a141a6978f67e3724c9f51a76f5b76ba92f83fb1e4dcad850eef4a0ea

Observation 7d331210-c885-43b0-85cc-074cacf94faa · outbound

This paper cites 2023 , pages =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , pages =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.914764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.914764Z digest=sha256:067b1264c58c97c7e4238c06ba36ac31ca75b66ffbfffa0feb37b2760f46dcc2

Observation acfb5a09-47b8-4b30-a999-6862a9059b15 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.919326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.919326Z digest=sha256:1efa11b925b851e099e7d4ece6a5784667449a9f8f7041922b5733a3133cf204

Observation 870c1633-b832-4c63-a6a0-fa84da4488c4 · outbound

This paper cites Evidential Deep Learning to Quantify Classification Uncertainty , year =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Evidential Deep Learning to Quantify Classification Uncertainty , year =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.924051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.924051Z digest=sha256:3e1c3c6330290fedb829e62329f1031c887fa45d3643ed4cc11f6e9d3a1d3914

Observation a916ee69-cc71-4b6d-afc6-ce3adca97a07 · outbound

This paper cites 2023 , publisher =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , publisher =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.928813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.928813Z digest=sha256:4b4422667e3228794f5b4213ad158ee8df716deac3ba4ee56a5d0b3c82e18e55

Observation 48550468-28ea-4b59-9425-70adc020939f · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.933913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.933913Z digest=sha256:92f09ca43fc6700a2ad2be52c44967424e28d5897758371599eea9eba36b8308

Observation 7acb0b4c-2c08-4884-b71f-aceedb5a25b8 · outbound

This paper cites 2023 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , volume =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.938937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.938937Z digest=sha256:095e6593e82607551f2bc7f095ed641264eed65ff9c88d4f52867654072a2fc5

Observation 3bd848ad-8b01-4187-9a44-6dcc03eb8959 · outbound

This paper cites 2021 , pages =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , pages =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.943748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.943748Z digest=sha256:434b4de30aac40e6b8d3052dd19bd972a22ed18c92720b4a1237db0ea633c822

Observation 03255404-7148-4dec-a1ec-1120b607f241 · outbound

This paper cites 2015 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2015 , volume =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.948542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.948542Z digest=sha256:b2700d610c2480343c291247beaf3c8556e0ef61a844fc75f3c23ba2b17e2a85

Observation 30225d71-b9c5-45ae-8afc-27b06bfa9847 · outbound

This paper cites NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems , journal =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems , journal =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.071321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.953311Z digest=sha256:5e8540f07cf47f7fec2355fd7c981d85d46869ec98ee716d6309311e16d1d1a4

Observation 515d9a57-cb1d-40d1-ae16-4ae765ba57d5 · outbound

This paper cites 2019 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2019 , volume =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.055527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.958040Z digest=sha256:d5adb1bf7b472bd7f9dbd27dcd339dcc44f6cfad020b1b6334390c3516b504de

Observation 192b7d0f-712a-4257-85ab-9af64b5ac28c · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.039365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.962751Z digest=sha256:4cb02d5a7c6f0362c4793dec99d1dff0075377049c24a927c5c138adc8b95313

Observation 15a6608f-3cf8-48c9-aac5-445bc2d82a56 · outbound

This paper cites 2023 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , volume =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.023376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.967490Z digest=sha256:155851a113e2eaa56a04f9a01437d3368e26b736a98472209a931f54f51b5930

Observation 1757f49d-7cae-4980-bc0c-d850b15210a2 · outbound

This paper cites Industrial fluid pipeline leak detection and localization based on a multiscale Mann-Whitney test and acoustic emission event tracking , journal = MSSP, volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Industrial fluid pipeline leak detection and localization based on a multiscale Mann-Whitney test and acoustic emission event tracking , journal = MSSP, volume =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.007665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.972296Z digest=sha256:95b569b3f1f45fb96755c94376c2166e37259714f112e3199c7173ad397e9150

Observation b3098bc6-a8cb-497f-a78f-da50902bd1c7 · outbound

This paper cites and Carter, G.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Carter, G

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.991751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.977365Z digest=sha256:89d36890216d1ffcb76fbab931974d9802f50f5b8bc9fec5b5a7a278bc073307

Observation bd3c126a-443a-47d2-adee-80981f8ab2f4 · outbound

This paper cites , journal = IEEE_J_AP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_AP, title =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.976536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.982123Z digest=sha256:d60386690f90cc878dc9a3b9322531f593ea38dc0c5db0c58a60fb8ca104b37d

Observation e2263b15-eb2b-4b95-894d-aa33879e4823 · outbound

This paper cites 2010 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2010 , volume =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.961456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.986594Z digest=sha256:7ce79bfcf419964ed036faeb8fefeab0c8a7b5cc8eb177a106b4d597af8cb3d2

Observation eebb7571-31ab-42b3-afd8-238127412615 · outbound

This paper cites 1968 , publisher =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 1968 , publisher =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.946982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.991371Z digest=sha256:02e5dba65b35679189c9a3289a46de60e93e27d34838f47845d69263dbc26127

Observation dfa5e560-54a6-4751-b9e6-f1967cc3f292 · outbound

This paper cites 2018 , publisher =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2018 , publisher =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.932133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.996346Z digest=sha256:676324ced1ef41b5c192bdf57e544c4dbf7a45e42cf9bff1e403d99e25f4e906

Observation b9d10312-a36c-4d15-90b9-7f01114468d8 · outbound

This paper cites Attention is All you Need , year =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Attention is All you Need , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.916281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.001126Z digest=sha256:c5faf4a0ab2df45b05710f77b0eee5073d25c055bf7cd5da50d9c2906b5f4297

Observation f2ef5355-69ed-457f-9d20-ef62f0526384 · outbound

This paper cites , journal = IEEE_J_ASLP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_ASLP, title =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.900260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.005812Z digest=sha256:eeaadcc553a8965df5bb133c1ee7f1ebff17bc5bda4d84556a63990807f59544

Observation 2f783be5-6288-4755-8f07-b8610c13ef83 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 48

Resolution
verified exact
doi, observed 2026-08-11T17:31:16.421301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.010450Z digest=sha256:05a3a5e3299f7525eaec609bb9298b74f1b9c30166fa5555a714d3d9364615f5

Observation 37d938c7-9752-40c4-9bb4-80522f824220 · outbound

This paper cites , journal = IEEE_J_ASLP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_ASLP, title =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.883711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.015719Z digest=sha256:1f2e6313464088566ac831b4de4b7b80f76fc1c5ac84ece11f206c8a8d728c7f

Observation d0b6a5ad-8249-4b2f-82b5-e63cf19c180d · outbound

This paper cites and Evers, Christine and Schmidt, Alexander and Mellmann, Heinrich and Barfuss, Hendrik and Naylor, Patrick A.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Evers, Christine and Schmidt, Alexander and Mellmann, Heinrich and Barfuss, Hendrik and Naylor, Patrick A

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.867098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.020881Z digest=sha256:49cbdc72d868cdebefaf6272cc3b92edf2708efd5e41b2b12126e05f2c660c56

Observation 945f2e39-8ae9-4e48-bb6b-f7798d538c91 · outbound

This paper cites and Moore, AlastairH.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Moore, AlastairH

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.850673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.025748Z digest=sha256:e83eb715a72549b0f6fadd26211f056f1f122b48d2401270697bd49dd7eb298c

Observation 9d615152-aac8-49ea-aa11-751f429677de · outbound

This paper cites 2024 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2024 , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.834112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.030585Z digest=sha256:853425f844b516af57e750d45331d2cd9662729fb755bd3d8b26c23e3baf31e4

Observation 8a5ad485-b3b4-4bce-9aeb-6ebd93e74e0a · outbound

This paper cites 2019 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2019 , volume =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.817280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.035726Z digest=sha256:d3cd2cbe774b406bcdbafa2452fabea41844989b654d33188f2387cd6cf1dc59

Observation dba446fb-85c1-4279-8aa5-3e65257aceba · outbound

This paper cites 2020 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2020 , volume =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.800163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.040610Z digest=sha256:fab07ee4fbd8908475804249a298d6d2b375e5b134db9f39acd3e4f034b4d9f4

Observation e84f8e48-1c60-4126-86f3-bad2a707b77f · outbound

This paper cites 2020 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2020 , volume =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.782696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.045542Z digest=sha256:40580ad3bae88186629286cc60b5ecb3745eab133b958a5231bdab7c8563f6c0

Observation 55391902-fe70-4fcc-8d84-461d3a42c3e7 · outbound

This paper cites , journal = IEEE_J_ASLP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_ASLP, title =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.766501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.050351Z digest=sha256:78c33b2d1b61d9a43140c2023794b69a7f39faeec314f7f61e39ce795fd51b91

Observation 885be313-bb18-4072-bbed-5d7aedea0aab · outbound

This paper cites 2020 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2020 , volume =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.750211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.055548Z digest=sha256:84f38ceb53b8c1c0f4dad067738f7814d22ab624b488456fdfc531be11583dd3

Observation 64ef3442-540c-453f-90ed-a146b1ac9062 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.060192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.060192Z digest=sha256:9117122e0dc759a752696af73921da67d0fc10af303e42e224855c461d828d36

Observation 03caf2b4-19f2-46fe-a543-265d7ac08aaf · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.065047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.065047Z digest=sha256:06f3a1c292318330096eedb34a50fa707eb81628abc93ecd9d410b3818632b90

Observation 1a2bf955-bf87-49f7-bb6f-d63481c66ebb · outbound

This paper cites 2025 , eprint=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2025 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.732636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.070059Z digest=sha256:eecfe4715a313d6095ec73c75bb8ee0e0b2a3a6aaf19faf2a94d8bc93beda87e

Observation d09b27b7-6870-41c5-af54-d4856ad65704 · outbound

This paper cites 2021 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , volume =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.715319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.074823Z digest=sha256:8b3fd018b51f5b1e16d76c809a8aea2880882c3986be757f9973a66df1736604

Observation ce133236-f5d8-41bf-be1d-972fcb397e15 · outbound

This paper cites and Nguyen, Ngoc Khanh and Jones, Douglas L.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Nguyen, Ngoc Khanh and Jones, Douglas L

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.698438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.079100Z digest=sha256:8f29e8852e40497657df84fb8fe617f2ea189ddcd501016e068b731793ccd878

Observation aa264c92-641b-460b-bfed-09de2526bb4d · outbound

This paper cites 2021 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , volume =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.680165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.083925Z digest=sha256:1f6ca28c056e7df1bf2a2d8b3418f0b585f4ed4103e6c9d284478c8878fb524b

Observation 53e4556f-ab24-4dc4-861a-ddcb5221bd10 · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.663645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.088669Z digest=sha256:81d7c654ebf57ed291dc998049f292d3010992b0ef21d42880ec9545666f9833

Observation e314c95b-b1e3-40e8-ab84-95377e68389e · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.647492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.093407Z digest=sha256:7316a0c92879b1b716cf72a488d61981e7e5792b4b9a2e178a83b4906dc3df14

Observation 4fa6ccac-46c1-489f-967e-fb898138dc58 · outbound

This paper cites 2021 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , volume =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.630962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.097963Z digest=sha256:b0fe8de9691e36af03152b0dbfc76249808bd9f45c755c91d358e6eac7941738

Observation 2b25fa94-79b4-4089-85b6-848642f1403a · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.614301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.102583Z digest=sha256:7c326d866d98dcf401a295aed6384daeaf63edb9b0b007dc6561a410fdc0fef4

Observation 1c3709ea-1b4e-4b02-9b31-b8d3a521dac6 · outbound

This paper cites , title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , title =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.107231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.107231Z digest=sha256:83eac1ce7e34351e693559499ea3fd7e743045e06be6bd6f951a5a51ae596f7d

Observation baabb004-dea4-46c7-8c91-4d27febe5c94 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:31:17.585715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.111552Z digest=sha256:394a8713122f3a999b420458fee3d02d75231dfab36a4b5b261dcc8fd1b95ec9

Observation d0af885d-6ab4-4530-b753-52a2ce741c15 · outbound

This paper cites Nature reviews neuroscience , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Nature reviews neuroscience , volume=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.116295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.116295Z digest=sha256:0f216b9929f38308795ef4911c8a7efde7bc2e129d0a0b062976269811fc69ea

Observation 5f5522ad-c541-4fdf-a78f-f57f054a1d70 · outbound

This paper cites Current Biology , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Current Biology , volume=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.558808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.121137Z digest=sha256:50b1115b7d7430f298c2a1dff88204d12d88c7bf9c96114722be2c4a949e688a

Observation 5fa8f8db-6e3b-4feb-8ab2-820d1ee32078 · outbound

This paper cites Current biology , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Current biology , volume=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.542840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.125920Z digest=sha256:1ff189aa92ff3a94ad49b9dd8a59b706852a333c0954d7254674fa5ea7663cca

Observation 1ecfad52-f551-45ee-81c4-0c3397c548bd · outbound

This paper cites Current Biology , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Current Biology , volume=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.526321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.131121Z digest=sha256:aecd799820643a405ce6f8963447532cf98dcab264cb31f1d81b7d87d1e6f929

Observation 401fc6ee-a3d2-4695-bd44-8c719a952e6b · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.136485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.136485Z digest=sha256:7914a6d3203781e82d32ca0c88cc13f21cab8e4637c4639620bd57f1ae55615b

Observation fa536399-949a-4324-ba3e-e496755c37b0 · outbound

This paper cites Proceedings of the European conference on computer vision (ECCV) , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the European conference on computer vision (ECCV) , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.141393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.141393Z digest=sha256:8d5e3d2f8513bbfe30409a1bb04eae62a7cc1cc9c2d36f15496645663b54a82c

Observation 5694338d-cd5b-4870-bc28-bf0f9193b8dd · outbound

This paper cites European conference on computer vision , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models European conference on computer vision , pages=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.488137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.146558Z digest=sha256:32db519eab6eb0eee9cb99de9644e052b7c2a6dcac3a34d739169e9786b078a2

Observation a5ef01d4-88df-4a6d-8004-4bda707ad268 · outbound

This paper cites IEEE Transactions on Audio, Speech and Language Processing , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models IEEE Transactions on Audio, Speech and Language Processing , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.472556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.151687Z digest=sha256:7c17d70bf487bff0ef3773cd3e3fa6c14f96146a390561c3d339027bcf4e18c6

Observation da8538ca-4c86-4a08-acf4-25631d099290 · outbound

This paper cites arXiv preprint arXiv:2606.14141 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2606.14141 , year=

Reference 78

Resolution
verified exact
raw_fallback, observed 2026-08-11T17:31:16.964041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.156836Z digest=sha256:03d2dffc16860285895dfc34bc1c5d2392e65d9991aec6b8ab4d11c9dd67f98d

Observation 760509ae-286c-4763-a90a-dc6d4d4a010b · outbound

This paper cites arXiv preprint arXiv:2509.26140 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2509.26140 , year=

Reference 79

Resolution
verified exact
raw_fallback, observed 2026-08-11T17:31:16.884501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.161674Z digest=sha256:704ea09925c02f41b44f2aad283217fcbb974ee82e65e044b86db7373c8d1077

Observation e00a1f8c-f964-464d-bcdd-df7e295b997d · outbound

This paper cites International Conference on Machine Learning , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Machine Learning , pages=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.455182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.166733Z digest=sha256:0b6aa3cb55e46d02707b489fe667d3826e0bda3ce243178a27d795c76d03afcf

Observation 9d069cdf-dc81-4eb7-a033-e99107dfc131 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Advances in Neural Information Processing Systems , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.172126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.172126Z digest=sha256:65d0706c8e1f23b02a96b50efb1d6fafb301dc76387ae49da994734028eca242

Observation ec917b3b-e32a-403b-8d1f-1e6ace311813 · outbound

This paper cites International Conference on Learning Representations , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Learning Representations , volume=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.427303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.176825Z digest=sha256:b8f44539a95417769dfb5266aa9bdfdbde0e21868ad69d8b6c24b9aa004a8321

Observation 789c37cc-b439-494f-93e0-2a155973fb7f · outbound

This paper cites International Conference on Learning Representations , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Learning Representations , volume=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.181629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.181629Z digest=sha256:daef145874ca293075e6f9595cee1d696d9ab416cc137d4a24195b84f9eb0ecf

Observation 3ea4af13-03f4-4333-a39c-de8504d6a983 · outbound

This paper cites International Conference on Machine Learning , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Machine Learning , pages=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.186425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.186425Z digest=sha256:e349397dc16a0f591f8943e83bfc03b5d6b73fb27d7c8013703e52d9a7b33077

Observation a263c929-b1fe-42a4-b24f-f8b2cd1b23b0 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.389626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.191233Z digest=sha256:7a395fce8ca6d549a97806fcd1219f6a052f931fa4341de9e5831daa392388d9

Observation 8c27b1c2-6684-4867-9a64-b64b0118bf6a · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 86

Resolution
parse uncertain
no resolver link, observed 2026-08-11T17:31:16.196054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.196054Z digest=sha256:193aba10a65209905fd94d82aaa9148e958aa066a3d652b0d15bc2880e9ad424

Observation eb136918-b109-49a4-a57b-a5c4d1f3a405 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 87

Resolution
parse uncertain
no resolver link, observed 2026-08-11T17:31:16.201448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.201448Z digest=sha256:75ec17d16eed6e8e240cbf15c0a4db758c9add42124ad0deedf25a0d1a62ffd4

Observation 0f613bc5-5873-40d7-910e-3aeb8e7868c2 · outbound

This paper cites ArXiv , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ArXiv , year=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.206573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.206573Z digest=sha256:cdc933f25f8b964a56604918913592330d3ef119c1b22e332d550d7971afec79

Observation 9b24d4f4-de91-415f-9130-7d68eeed03b9 · outbound

This paper cites ArXiv , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ArXiv , year=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.341575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.212314Z digest=sha256:b15769711939678beafcb4a7f4e767714f0a846892fedeca69f5f70f1b426048

Observation cffd81c7-e1b4-4ba9-b390-487c6b57013e · outbound

This paper cites ArXiv , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ArXiv , year=

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.325566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.217183Z digest=sha256:03b70018b86e28430513d5ce717edc0bf6e5fef43ec21c2540941036ec2cf633

Observation 463bb80b-cd7c-4bc4-9f91-0df69a51585a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Advances in Neural Information Processing Systems , volume=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.309294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.222076Z digest=sha256:0aa741b4396a197e21c9e13506efab7e86637eda2c31394eb3480f6556395c29

Observation 93e195d4-2c9d-42ac-b53c-72d4a407703a · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.292778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.227036Z digest=sha256:b906a718104ed995e64a14fa963d8d98df277e27859f9b84678c5c10cd8ad612

Observation 0d5efc58-5f2e-42eb-ab42-a1ba8a67e65d · outbound

This paper cites European Conference on Computer Vision , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models European Conference on Computer Vision , pages=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.275683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.231979Z digest=sha256:201c70c5a1084fe32cfcae343fa89ef36e452f7cde706184a4a5e28526eacfd4

Observation 1501dedc-00bf-41b6-b42b-c90e9153d020 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:31:17.259108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.236830Z digest=sha256:20a863700e35a6945897a9bba12deabe12414cf9a604ec29f830e0fd7d213ac6

Observation 891d5b95-18a3-41d9-a149-f6b227265870 · outbound

This paper cites arXiv preprint arXiv:2511.06606 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2511.06606 , year=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.241435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.241435Z digest=sha256:20da5393c2262a184f57db3d60700f3ae528dc5ab455380532bf497826b22e59

Observation db1ed259-4629-4dba-a18f-9765b5dfeab3 · outbound

This paper cites IEEE Open Journal of Signal Processing , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models IEEE Open Journal of Signal Processing , year=

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.243530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.246221Z digest=sha256:a25c8c64e56dac7b637e9358ee5d35151712c370bb527af64a78a6744d20c584

Observation 3e14a13b-8343-46f3-b9a5-78d6240aa579 · outbound

This paper cites Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

Reference 97

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T17:31:16.710107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.251711Z digest=sha256:8301da76c8355369b456980886106200cdffad24eedc66b1bad4d78af1db05ed

Observation e787f4e8-d21d-44bd-b531-256bf1a06540 · outbound

This paper cites arXiv preprint arXiv:2509.14666 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2509.14666 , year=

Reference 98

Resolution
verified exact
raw_fallback, observed 2026-08-11T17:31:16.684681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.256772Z digest=sha256:c60a26227efdf934547407c67ea764749889625c876f6bcdbdcc7080be4c09ef

Observation e2007e8f-a7e2-4323-99fe-a9ad24592ef1 · outbound

This paper cites arXiv preprint arXiv:2602.16334 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2602.16334 , year=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.261198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.261198Z digest=sha256:f4b5edd4d1a180033c9c4e64ab88382ddd9085792f4f7d6aad5dcfd2d43e218d

Observation 8086fb12-512e-4aba-81be-63248a32391e · outbound

This paper cites arXiv preprint arXiv:2606.14141 , year =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2606.14141 , year =

Reference 100

Resolution
verified exact
doi, observed 2026-08-11T17:31:16.404582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.265521Z digest=sha256:a238063dbd50fc02cd74575ef8ee127f3cbfede716f3ec1fef404640ff617010

Pith citing papers

No inbound Pith citation observations are available.