Pith. sign in

Paper Citation Record · LEDGER

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

As of 13 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 0 inbound Pith citation observations for arXiv:2608.09435.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09435 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:31:16.265521Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 103 outbound references displayed

  • verified exact5
  • verified fuzzy43
  • unresolved49
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04ff485b-0e38-41e3-be6e-6ee24df84986 · outbound

This paper cites Proceedings of the 6th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2021) , pages =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the 6th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2021) , pages =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.781188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.781188Z digest=sha256:36d5e1607a21d6a0fc6c86c303aa9f862b0cf4d2aa9d6cb96e2e4f1e757a0fda

Observation 47f0f2ea-d125-414a-9cce-cecf045661f8 · outbound

This paper cites arXiv preprint arXiv:2606.19348 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2606.19348 , year=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.788703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.788703Z digest=sha256:abcaeaa1b1f6fa3a7d15c40e94722ca48f5326b0b6aa2a02dc62e069d8b7f9e2

Observation c202388e-43b7-4d69-ac9a-606476f21451 · outbound

This paper cites Kimi-Audio Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Kimi-Audio Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.794034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.794034Z digest=sha256:16b782b4c934a1dc2a2f870bafc11bf68618610a37fffab315ec611d545a3a10

Observation 843c9c46-6a0b-4418-9461-c2a2d923b5ee · outbound

This paper cites doi:10.21437/Interspeech.2021-698 , issn =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models doi:10.21437/Interspeech.2021-698 , issn =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.799394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.799394Z digest=sha256:4c04d5c6a560379e40ccfd9d92ed548c9084abf3af3b8ae7a520223bcae7646b

Observation cc6b64d3-383c-444e-a4af-c9a07548dd9f · outbound

This paper cites Advances in neural information processing systems , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Advances in neural information processing systems , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.804735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.804735Z digest=sha256:ed070fa2e9cc3e767fe220322930e4e00f7d8c02993c10e653ec61c117314717

Observation 5fe10506-c759-43ce-9ee9-8cbfb99ce73a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.809739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.809739Z digest=sha256:67e6f82f64190d1358a6657ea7f7bea1f7f2b442c7b13d74191e69c03ead300c

Observation f23b135e-6e3b-467c-b500-eb567cae0515 · outbound

This paper cites ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.815138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.815138Z digest=sha256:0668e48924853fbe063768809ba7c9a3f7e6caac909dcf7e556b6689de8dce9f

Observation 41031b32-f7dc-4a41-8e01-1c39b633167f · outbound

This paper cites Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.820170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.820170Z digest=sha256:97779d0dbef542ea76231faa58fa5f7765d3e649ad846d3805c19a823e485976

Observation e1c0333d-ec30-4f56-b9b5-d87866de500d · outbound

This paper cites Classification Problem Solving.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Classification Problem Solving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.825345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.825345Z digest=sha256:b702f118cbf911948bf57d056dc8ad6de7fb88e693f73fecb050092a431d634e

Observation 8ee26178-0066-4839-9811-656a3f829e03 · outbound

This paper cites , title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , title =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.830677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.830677Z digest=sha256:a65df433a65d8094d7d549b2e2e0c57bfb1dc28ccbd3f7b33a08f338062d6c20

Observation 4fb795f0-557a-4ac6-8f84-560f27a32343 · outbound

This paper cites New Ways to Make Microcircuits Smaller---Duplicate Entry.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models New Ways to Make Microcircuits Smaller---Duplicate Entry

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.835603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.835603Z digest=sha256:9a47ce187ee7246bbbb8ce6d9952420270cfa427a0c748542f4e327a1c07e821

Observation 82c3f7f9-b256-4717-9621-46724c1695ed · outbound

This paper cites Clancey and Glenn Rennels , abstract =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Clancey and Glenn Rennels , abstract =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.840692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.840692Z digest=sha256:9271a347990c9a401fe4b9712fb5cc0a14187da11b41de3f0ed19d88ebd3a935

Observation a3068a52-6aca-47e8-ae3e-28e220fc6615 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.845666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.845666Z digest=sha256:3cb67877b189992201232ecff040586f3e4d3a3d3133467a0969196d010f0fe9

Observation f0b5e523-5afd-4076-a883-f2d74cf9bf09 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen2.5-Omni Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.850604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.850604Z digest=sha256:2c8d05b8635b052da6068bc5073d1e265203b301fd17c00107f50a48d4cee7e9

Observation 3cbf6da2-c105-44da-adf0-9723ae850916 · outbound

This paper cites Qwen2-Audio Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen2-Audio Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.855588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.855588Z digest=sha256:d0b14c5f1ceada6d99a5860df9d356d4c0b191fadb53f21db7359b4727a69dd0

Observation 2264e4cb-49df-4d39-810c-01d03a6e82a7 · outbound

This paper cites Qwen3-VL Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen3-VL Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.860490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.860490Z digest=sha256:94a71d2003c727700c68100f348b426167ec175216277a917097a83d52bc5b87

Observation 3a5cc7ef-fb8f-46d9-b5b3-cd28baa34d01 · outbound

This paper cites Qwen2.5-VL Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Qwen2.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.865581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.865581Z digest=sha256:cddc17f2170189d591bea45c2503f581d187d60a4fef2792b5729106daa54094

Observation 5ec19816-a098-41b9-a12b-8c4baeaa26d0 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.870460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.870460Z digest=sha256:a025c7bcdb058aa3836e8f7417dda838e62ae92532cc48fa6729b28bb4d71ed9

Observation 3023b98b-7fdd-425f-b48a-c0b9cc93c7bb · outbound

This paper cites GPT-4 Technical Report.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.875213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.875213Z digest=sha256:f80cf092141a9be47bfd73204fb75cb7d5e895538194dff9b5aaf06f10b0969d

Observation 5ee4db9a-a080-4102-8f84-6c395e45cc25 · outbound

This paper cites and Rennels, Glenn R.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Rennels, Glenn R

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.879683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.879683Z digest=sha256:63866c022c761994c03e1c90f87dbb3b21b91e65df12830261f1de392d1ec91a

Observation c9cb7848-35a8-43c8-a96c-5626202d39d9 · outbound

This paper cites Poligon: A System for Parallel Problem Solving.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Poligon: A System for Parallel Problem Solving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.884011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.884011Z digest=sha256:8296007a9f035ac5be2fcdf937927b17e4b5a90275de9b817cf2066b6b546acd

Observation 24a9b67a-1139-43c2-ad0f-9ccf2e4259aa · outbound

This paper cites Transfer of Rule-Based Expertise through a Tutorial Dialogue.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Transfer of Rule-Based Expertise through a Tutorial Dialogue

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.888398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.888398Z digest=sha256:03af728c8f6c81d18836169bf14791ac0e901080cbc4a5656b7440b826bf02ec

Observation 26d93712-6eb8-4a18-8b51-98d4f49a2f53 · outbound

This paper cites The Engineering of Qualitative Models.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models The Engineering of Qualitative Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.892716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.892716Z digest=sha256:f17a22b18d1a5ae0f6e361db67cc2a0a02fbcb0fe5d21deceed3515189fd9f9e

Observation 66219162-e6bf-408a-aa00-6a97cb268638 · outbound

This paper cites 2023 , eprint=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.897094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.897094Z digest=sha256:797276085a5d2fad4de258456ac32a66ad3e418eb2ff58b1a23e2e72e8bd3443

Observation 7d1c5648-1dc1-4ccf-97e4-8dd99b8375b1 · outbound

This paper cites Pluto: The 'Other' Red Planet.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Pluto: The 'Other' Red Planet

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.902110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.902110Z digest=sha256:897a7ff13f60b69f21a4e85867cda57007e804773fcb5920a3488c7f0ddbb16e

Observation 676bcba8-5172-4f13-9719-69bdeb104726 · outbound

This paper cites 2023 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , volume =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.906340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.906340Z digest=sha256:1cfd05c3cf621d1c40c57230d8dc9894101bc4d98bc843547871f4a31daca78d

Observation a012fd13-8076-4cee-b958-643ac5f54f41 · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.910427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.910427Z digest=sha256:1a3138986783f9933d0eec97e1334662eac66582345a1f136cba6d807b9e8eec

Observation 7d331210-c885-43b0-85cc-074cacf94faa · outbound

This paper cites 2023 , pages =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , pages =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.914764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.914764Z digest=sha256:bbf9fe4aa9ac62d015a90fbd13002cffc9493b5fc62ab809010c21f2e5624f24

Observation acfb5a09-47b8-4b30-a999-6862a9059b15 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.919326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.919326Z digest=sha256:c34c4cf6d3a15411ec7f12d7d7086fc5f4dbe9aadb656a7a4a857382114157f1

Observation 870c1633-b832-4c63-a6a0-fa84da4488c4 · outbound

This paper cites Evidential Deep Learning to Quantify Classification Uncertainty , year =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Evidential Deep Learning to Quantify Classification Uncertainty , year =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.924051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.924051Z digest=sha256:62e0f102f1755060d22253dbba6f6905e7e2848fa6aa1db96468d7f0a284c208

Observation a916ee69-cc71-4b6d-afc6-ce3adca97a07 · outbound

This paper cites 2023 , publisher =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , publisher =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.928813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.928813Z digest=sha256:d189ce2468b491deb2a96134b31caf6c09b4990ea246481ea531c67683e30fea

Observation 48550468-28ea-4b59-9425-70adc020939f · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.933913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.933913Z digest=sha256:3676919c273b0daee3272446da23f3e6a446d101c3d203868845cd23a3b0c58a

Observation 7acb0b4c-2c08-4884-b71f-aceedb5a25b8 · outbound

This paper cites 2023 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , volume =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.938937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.938937Z digest=sha256:326288d285fb5f2100c00026d6837bb0bdcf3fc5af6eb0849de051dff291c711

Observation 3bd848ad-8b01-4187-9a44-6dcc03eb8959 · outbound

This paper cites 2021 , pages =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , pages =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.943748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.943748Z digest=sha256:3ef0cc7151f1d560dd6c32c2e2078ba33c38a57d0da5a8fcc5d2b73b84dc3705

Observation 03255404-7148-4dec-a1ec-1120b607f241 · outbound

This paper cites 2015 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2015 , volume =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:15.948542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:15.948542Z digest=sha256:e6a2aba609464fb40de590b93d4d384d44eb8b86c4b265df56e5eb289597cbb3

Observation 30225d71-b9c5-45ae-8afc-27b06bfa9847 · outbound

This paper cites NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems , journal =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems , journal =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.071321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.953311Z digest=sha256:0bb7432f3fa6ff414f21d1cf602efcd2795e6f55d97911f5284e35ffdaf048db

Observation 515d9a57-cb1d-40d1-ae16-4ae765ba57d5 · outbound

This paper cites 2019 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2019 , volume =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.055527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.958040Z digest=sha256:a4942fcd2a79ef48bb2fb33ebf34476cffa556cd72208ddc952cb9f34009d00c

Observation 192b7d0f-712a-4257-85ab-9af64b5ac28c · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.039365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.962751Z digest=sha256:8c5b66baa8548209aa8ee0b3f9d36bbe8dfc45000b37759fb23c0a357554d4cf

Observation 15a6608f-3cf8-48c9-aac5-445bc2d82a56 · outbound

This paper cites 2023 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2023 , volume =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.023376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.967490Z digest=sha256:bdfe646def3438966bed8bcbcae83b8b517285cb9264950cf3d987970ad9ca63

Observation 1757f49d-7cae-4980-bc0c-d850b15210a2 · outbound

This paper cites Industrial fluid pipeline leak detection and localization based on a multiscale Mann-Whitney test and acoustic emission event tracking , journal = MSSP, volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Industrial fluid pipeline leak detection and localization based on a multiscale Mann-Whitney test and acoustic emission event tracking , journal = MSSP, volume =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:18.007665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.972296Z digest=sha256:869f37d6e63753bd56913bed27bf6ffe37ad696c7d19921a523e74f2155c534e

Observation b3098bc6-a8cb-497f-a78f-da50902bd1c7 · outbound

This paper cites and Carter, G.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Carter, G

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.991751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.977365Z digest=sha256:4712e3e4cfd3d965426df2b7f0b3c521fdab89b0c6395f837aed3d4ed2b6c2ab

Observation bd3c126a-443a-47d2-adee-80981f8ab2f4 · outbound

This paper cites , journal = IEEE_J_AP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_AP, title =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.976536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.982123Z digest=sha256:a946978caa2292af0e999b0e0882f8b7e8ed9293dfc686b7cf160c574f9fe3c9

Observation e2263b15-eb2b-4b95-894d-aa33879e4823 · outbound

This paper cites 2010 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2010 , volume =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.961456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.986594Z digest=sha256:b2dd584308335ab5678339904636afcb53ce4680e47ae133a274f75302c2a24a

Observation eebb7571-31ab-42b3-afd8-238127412615 · outbound

This paper cites 1968 , publisher =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 1968 , publisher =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.946982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.991371Z digest=sha256:b5e813652aeda44089096baa6a88c4e8b07a3be3a06b6292125acbb225f1ac5d

Observation dfa5e560-54a6-4751-b9e6-f1967cc3f292 · outbound

This paper cites 2018 , publisher =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2018 , publisher =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.932133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:15.996346Z digest=sha256:a762f09b71c7c1569be56384d8ea64fe0dc9e8bad7a3c4909004b2b112fa8301

Observation b9d10312-a36c-4d15-90b9-7f01114468d8 · outbound

This paper cites Attention is All you Need , year =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Attention is All you Need , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.916281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.001126Z digest=sha256:69ce805bc4f827ac194e61dbcaa499eea2bf5fe6c84f342fb07fe68f20c7ddbd

Observation f2ef5355-69ed-457f-9d20-ef62f0526384 · outbound

This paper cites , journal = IEEE_J_ASLP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_ASLP, title =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.900260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.005812Z digest=sha256:b6523a73def7e2667a68090eabf86d9663d23ba6d7c626446571fd59d16a9193

Observation 2f783be5-6288-4755-8f07-b8610c13ef83 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 48

Resolution
verified exact
doi, observed 2026-08-11T17:31:16.421301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.010450Z digest=sha256:cc88b1b85101b0f9cd6fa2417a1cc5f3288d92cf3cc01410feb92c575d773773

Observation 37d938c7-9752-40c4-9bb4-80522f824220 · outbound

This paper cites , journal = IEEE_J_ASLP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_ASLP, title =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.883711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.015719Z digest=sha256:1ba09ac5f04845319bea3756bfe170fddf6434b53a4b7da2ab93721a6b1dfeaf

Observation d0b6a5ad-8249-4b2f-82b5-e63cf19c180d · outbound

This paper cites and Evers, Christine and Schmidt, Alexander and Mellmann, Heinrich and Barfuss, Hendrik and Naylor, Patrick A.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Evers, Christine and Schmidt, Alexander and Mellmann, Heinrich and Barfuss, Hendrik and Naylor, Patrick A

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.867098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.020881Z digest=sha256:1a09413fc7f7ec8b2b2a004049b7a535c48315b6829714d1d1fcf459eb24ffcf

Observation 945f2e39-8ae9-4e48-bb6b-f7798d538c91 · outbound

This paper cites and Moore, AlastairH.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Moore, AlastairH

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.850673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.025748Z digest=sha256:58eaf3d33ed29c441fdfe91ee0ceee270b9e766928feca607479fd40d9fed35e

Observation 9d615152-aac8-49ea-aa11-751f429677de · outbound

This paper cites 2024 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2024 , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.834112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.030585Z digest=sha256:d2cd3a38b071d5b5614371e4ba99282502cfe602f102cae5398e9baf7ebb6bcf

Observation 8a5ad485-b3b4-4bce-9aeb-6ebd93e74e0a · outbound

This paper cites 2019 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2019 , volume =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.817280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.035726Z digest=sha256:99cb5a473eabcc52c1bb9d95ec1271b8d890573303d496e2982bfe017f4ee024

Observation dba446fb-85c1-4279-8aa5-3e65257aceba · outbound

This paper cites 2020 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2020 , volume =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.800163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.040610Z digest=sha256:b0e2789f115c6a929ac04e32870b662d1d42e16bd6d0a086fa3bfcfbb000f861

Observation e84f8e48-1c60-4126-86f3-bad2a707b77f · outbound

This paper cites 2020 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2020 , volume =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.782696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.045542Z digest=sha256:048f3740141daf8d230e7331f48d19b1b0172b811867198f910415a78289acaa

Observation 55391902-fe70-4fcc-8d84-461d3a42c3e7 · outbound

This paper cites , journal = IEEE_J_ASLP, title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , journal = IEEE_J_ASLP, title =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.766501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.050351Z digest=sha256:e285be55be73d6b5b4543e0b4b421dacd8ac347cac798ee09222f5c8228a88b3

Observation 885be313-bb18-4072-bbed-5d7aedea0aab · outbound

This paper cites 2020 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2020 , volume =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.750211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.055548Z digest=sha256:80fbbcf74ba3d312c8b56a0bd2202c3e4fec7a8ac3248e8a6e39d39442b8bc8e

Observation 64ef3442-540c-453f-90ed-a146b1ac9062 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.060192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.060192Z digest=sha256:5bea78688b40fac36b5a5904f517aaf778596772dffc9ff9bffd86eb0375e8fd

Observation 03caf2b4-19f2-46fe-a543-265d7ac08aaf · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.065047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.065047Z digest=sha256:1436e71defc3290f9642ed30e6f3f613c148268cb5db2e843b61e241d9b468da

Observation 1a2bf955-bf87-49f7-bb6f-d63481c66ebb · outbound

This paper cites 2025 , eprint=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2025 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.732636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.070059Z digest=sha256:453b3a80bbf193b9dc2525ae3678ac0b6b21356e7ceb80db45d693b6d03ba285

Observation d09b27b7-6870-41c5-af54-d4856ad65704 · outbound

This paper cites 2021 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , volume =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.715319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.074823Z digest=sha256:876c6c3745ac246365e779d881de637fbe03a51c0038f9f5125be826f5ad79c9

Observation ce133236-f5d8-41bf-be1d-972fcb397e15 · outbound

This paper cites and Nguyen, Ngoc Khanh and Jones, Douglas L.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models and Nguyen, Ngoc Khanh and Jones, Douglas L

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.698438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.079100Z digest=sha256:fe7de08a3292624d63f72ba112e5002230c4a00e5e7d0105d969a59688334454

Observation aa264c92-641b-460b-bfed-09de2526bb4d · outbound

This paper cites 2021 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , volume =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.680165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.083925Z digest=sha256:3c956ae3767fbff1999ea92470e73e336a40c6303a78c6edb97dfdf54506b02e

Observation 53e4556f-ab24-4dc4-861a-ddcb5221bd10 · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.663645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.088669Z digest=sha256:e9a1554b4a8d828ebee4a304d4976d7de559d29070fa7a776e1a03bbe667d14a

Observation e314c95b-b1e3-40e8-ab84-95377e68389e · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.647492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.093407Z digest=sha256:ddd822897877589c1bd7a2250a7910b0f235a355233a94d29a31121eb4279d26

Observation 4fa6ccac-46c1-489f-967e-fb898138dc58 · outbound

This paper cites 2021 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2021 , volume =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.630962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.097963Z digest=sha256:847009602da9671d4307433344c65b3c95415ec84dc23777698b5dc4ae0ca765

Observation 2b25fa94-79b4-4089-85b6-848642f1403a · outbound

This paper cites 2022 , volume =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models 2022 , volume =

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.614301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.102583Z digest=sha256:464035455eb8f236692ba19bf5daa2ec2587bfb492130d3944fda51bd5e122d0

Observation 1c3709ea-1b4e-4b02-9b31-b8d3a521dac6 · outbound

This paper cites , title =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models , title =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.107231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.107231Z digest=sha256:ccda1ed437a1bade760d6b20eb027340bbc60d35ba592ccfb8ef85a79922cf23

Observation baabb004-dea4-46c7-8c91-4d27febe5c94 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:31:17.585715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.111552Z digest=sha256:bf36116e2d7532b314210ceee26e727d28a0a2080d220b975650943d973ac5ee

Observation d0af885d-6ab4-4530-b753-52a2ce741c15 · outbound

This paper cites Nature reviews neuroscience , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Nature reviews neuroscience , volume=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.116295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.116295Z digest=sha256:d2c8448e361581337ade390802befea6c5f78a6ab12f55d7b69929ce55ef30f7

Observation 5f5522ad-c541-4fdf-a78f-f57f054a1d70 · outbound

This paper cites Current Biology , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Current Biology , volume=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.558808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.121137Z digest=sha256:6dc2fac98855edfbc6b6597ebf74809c6479364411edc80751bf016912250a20

Observation 5fa8f8db-6e3b-4feb-8ab2-820d1ee32078 · outbound

This paper cites Current biology , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Current biology , volume=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.542840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.125920Z digest=sha256:d29a8d113bff7704f0958b76ac65c2f47bff708b79e0b8484ef507b6778d58d4

Observation 1ecfad52-f551-45ee-81c4-0c3397c548bd · outbound

This paper cites Current Biology , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Current Biology , volume=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.526321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.131121Z digest=sha256:79b9f8df37d22961d5a958d6aad8a41f4c4070b6dbac37022160749b0d356af4

Observation 401fc6ee-a3d2-4695-bd44-8c719a952e6b · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.136485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.136485Z digest=sha256:0985c67c3771cf6e053217e8904b778b0b1d77529f777c1eeb63f8385e07522c

Observation fa536399-949a-4324-ba3e-e496755c37b0 · outbound

This paper cites Proceedings of the European conference on computer vision (ECCV) , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the European conference on computer vision (ECCV) , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.141393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.141393Z digest=sha256:b4a680dc3a2238c6c7c9cb1b0d0c1c61b23b3547ccf0435e89561c13021454f2

Observation 5694338d-cd5b-4870-bc28-bf0f9193b8dd · outbound

This paper cites European conference on computer vision , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models European conference on computer vision , pages=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.488137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.146558Z digest=sha256:742bbf7924711e2d97d2c92ef0bf7a38f64eb156f82afb99c9d1d0f4d4391ef6

Observation a5ef01d4-88df-4a6d-8004-4bda707ad268 · outbound

This paper cites IEEE Transactions on Audio, Speech and Language Processing , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models IEEE Transactions on Audio, Speech and Language Processing , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.472556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.151687Z digest=sha256:b6aed2b02b6467bfc7dc42fb97a6b3a7f5c0c1a4e222b9b2864a166a16abb9b9

Observation da8538ca-4c86-4a08-acf4-25631d099290 · outbound

This paper cites arXiv preprint arXiv:2606.14141 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2606.14141 , year=

Reference 78

Resolution
verified exact
raw_fallback, observed 2026-08-11T17:31:16.964041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.156836Z digest=sha256:2752127606dbade6a35547902e079e50ce76854189d94349251ce105a919f4a4

Observation 760509ae-286c-4763-a90a-dc6d4d4a010b · outbound

This paper cites arXiv preprint arXiv:2509.26140 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2509.26140 , year=

Reference 79

Resolution
verified exact
raw_fallback, observed 2026-08-11T17:31:16.884501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.161674Z digest=sha256:d697f01eb14a2d98a444c8fe7f9eb8c1ad499547b8d03771ec16d28573632c77

Observation e00a1f8c-f964-464d-bcdd-df7e295b997d · outbound

This paper cites International Conference on Machine Learning , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Machine Learning , pages=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.455182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.166733Z digest=sha256:a447a9b5bd7cbad27579afbc9aab56a78aa0b8b2ba8f10da2f4ede9782827331

Observation 9d069cdf-dc81-4eb7-a033-e99107dfc131 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Advances in Neural Information Processing Systems , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.172126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.172126Z digest=sha256:329f7565e50a2853d6145effd97e44ccc99153119f72b54ef113e778d6da5020

Observation ec917b3b-e32a-403b-8d1f-1e6ace311813 · outbound

This paper cites International Conference on Learning Representations , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Learning Representations , volume=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.427303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.176825Z digest=sha256:0169a1fbb29869c00b93322f9f2176c53df1163a0613d3a347ff8430971cb656

Observation 789c37cc-b439-494f-93e0-2a155973fb7f · outbound

This paper cites International Conference on Learning Representations , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Learning Representations , volume=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.181629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.181629Z digest=sha256:08f287bcdacf7d84348aaaf06cffbb34b4954fa7e802ad8c4f8ab310e9ce668c

Observation 3ea4af13-03f4-4333-a39c-de8504d6a983 · outbound

This paper cites International Conference on Machine Learning , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models International Conference on Machine Learning , pages=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.186425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.186425Z digest=sha256:ced2feb28a19626605463fb8504d420cdcff14a59b60cc7f2403aa2457cce9b5

Observation a263c929-b1fe-42a4-b24f-f8b2cd1b23b0 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.389626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.191233Z digest=sha256:e1668768e35a153cc27e89b498f0f4ce4ef0cba52c2610e00c1de05f44d8cdd3

Observation 8c27b1c2-6684-4867-9a64-b64b0118bf6a · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 86

Resolution
parse uncertain
no resolver link, observed 2026-08-11T17:31:16.196054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.196054Z digest=sha256:efe721dd2e24f7b9362a3f508019fe754100904d01d2513a8f867f574944d368

Observation eb136918-b109-49a4-a57b-a5c4d1f3a405 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 87

Resolution
parse uncertain
no resolver link, observed 2026-08-11T17:31:16.201448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.201448Z digest=sha256:a2dfdb089eeca7497e8857c375de3e79353ad033d3737767ce4c8be68dabe3ce

Observation 0f613bc5-5873-40d7-910e-3aeb8e7868c2 · outbound

This paper cites ArXiv , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ArXiv , year=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.206573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.206573Z digest=sha256:ba680dd677b80b71ab4e6748b97924f6b7bd339426f05bd56323e0a25efd20d7

Observation 9b24d4f4-de91-415f-9130-7d68eeed03b9 · outbound

This paper cites ArXiv , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ArXiv , year=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.341575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.212314Z digest=sha256:c23d172c0a39e06a5206748af55792ffd40565fe8cfb2278d8b033e10887a9b5

Observation cffd81c7-e1b4-4ba9-b390-487c6b57013e · outbound

This paper cites ArXiv , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ArXiv , year=

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.325566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.217183Z digest=sha256:e8382a7bff8878579ebec60c5d9e0f9b7a6660f52ee3c3751def14ddba016fd2

Observation 463bb80b-cd7c-4bc4-9f91-0df69a51585a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Advances in Neural Information Processing Systems , volume=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.309294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.222076Z digest=sha256:5d62e1d52f9525d47c86e90bfe8db0ccbaf3dde2184204d28761df006a9c2ab0

Observation 93e195d4-2c9d-42ac-b53c-72d4a407703a · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.292778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.227036Z digest=sha256:d15d47ed44059b329b0e89172a9dd42c85f547c21908f20d4869b3620dfab717

Observation 0d5efc58-5f2e-42eb-ab42-a1ba8a67e65d · outbound

This paper cites European Conference on Computer Vision , pages=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models European Conference on Computer Vision , pages=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.275683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.231979Z digest=sha256:dd0d40facbc0a2510a9a98929a385c03ccce753543ab8aea1fbf2d487a5eb27e

Observation 1501dedc-00bf-41b6-b42b-c90e9153d020 · outbound

This paper cites an unresolved cited work.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:31:17.259108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.236830Z digest=sha256:708d2723797b6c356c217d32f73c375f68f32e461b805a9875d3d0c1d46409a4

Observation 891d5b95-18a3-41d9-a149-f6b227265870 · outbound

This paper cites arXiv preprint arXiv:2511.06606 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2511.06606 , year=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.241435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.241435Z digest=sha256:87b0485b507072c7228062c5ed1dc884689c3beb507eb9e6c9868d75e3478333

Observation db1ed259-4629-4dba-a18f-9765b5dfeab3 · outbound

This paper cites IEEE Open Journal of Signal Processing , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models IEEE Open Journal of Signal Processing , year=

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:31:17.243530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.246221Z digest=sha256:e9f9cfc5b067d46481f972f346ea6ce47d0c79bdb400fed4117ea4f916a4b31f

Observation 3e14a13b-8343-46f3-b9a5-78d6240aa579 · outbound

This paper cites Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

Reference 97

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T17:31:16.710107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.251711Z digest=sha256:3521cf3404972a1ac3394567f8db7f01e9af6cac1454a532f6f8ab96e3177d8f

Observation e787f4e8-d21d-44bd-b531-256bf1a06540 · outbound

This paper cites arXiv preprint arXiv:2509.14666 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2509.14666 , year=

Reference 98

Resolution
verified exact
raw_fallback, observed 2026-08-11T17:31:16.684681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.256772Z digest=sha256:d1ae818b50464430891d1839b32cfe7d745b8def5785e4d8476001b3f44d20c5

Observation e2007e8f-a7e2-4323-99fe-a9ad24592ef1 · outbound

This paper cites arXiv preprint arXiv:2602.16334 , year=.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2602.16334 , year=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T17:31:16.261198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:31:16.261198Z digest=sha256:16afeafe869db8aa0933d1c38bf04fdbc6364c6f6104ed7c2b3ef56af69887e1

Observation 8086fb12-512e-4aba-81be-63248a32391e · outbound

This paper cites arXiv preprint arXiv:2606.14141 , year =.

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models arXiv preprint arXiv:2606.14141 , year =

Reference 100

Resolution
verified exact
doi, observed 2026-08-11T17:31:16.404582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:31:16.265521Z digest=sha256:9e17f66d03f77b75849938505a985d456200f1cb04b21fbba2b036051bbe7c80

Pith citing papers

No inbound Pith citation observations are available.