Pith. sign in

Paper Citation Record · LEDGER

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

As of 21 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2607.10299.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10299 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T12:48:58.688011Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved69
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5122056e-5078-4ae7-a3d7-9d678d82fa28 · outbound

This paper cites In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summa- rization.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summa- rization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:d3f08ddc632578a73c2aa9828b7dc25cafda5d0cb7b5b5fab5b4d6bc525bb154

Observation eec2635a-ca0c-4051-b6d5-35ba18b86735 · outbound

This paper cites arXiv preprint arXiv:2402.12451 (2024).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2402.12451 (2024)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:96ef447e70c1b0d9532769a11450c45271039025b2cfeb272cb1367159937476

Observation fb2f079a-7739-4e4e-b6e2-579857aaa5ea · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:c7d636d3064448db17cc63bd5a0c4f369f3690c2d68a5c6624bf7f2d31d29fc3

Observation f0df176e-058e-478f-8d37-de66312cdbbe · outbound

This paper cites Advances in Neural Information Processing Systems37, 19472–19495 (2024).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Advances in Neural Information Processing Systems37, 19472–19495 (2024)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:1052889e2af472cbb7bde2d3e114267ce91920d6634d4c7fe8b7dba55a672fbb

Observation 018b58cc-c35a-49bc-acd4-6ad7c7afb02a · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:030901909efe02e3badb523fde38fcf5d3ec5990ffba9d60329b1a07dcc3f53b

Observation 5846e4ff-003e-437e-8fc2-de0fed13726c · outbound

This paper cites Advances in Neural Information Processing Systems36, 72842–72866 (2023).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Advances in Neural Information Processing Systems36, 72842–72866 (2023)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:9a590a2524dc7970202021464a8144d60f7fd5b537da84c381f9b3244bb95658

Observation cc933093-9c96-44ed-93a3-20ed34af987f · outbound

This paper cites an unresolved cited work.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:b123ebc0f41108d8555aefc0d20820b71bc9d7db506502c615fee05ffdaa9769

Observation f4a4f653-3e4c-47be-a5af-252d049ee2ee · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:f45550c56ce7b3e44cb306732b368f1a5b9fddb6266755cdadcad0d832fc41af

Observation a6787e69-10f8-45c0-836d-3a92a9d01bcb · outbound

This paper cites MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:0de1e0fd23b704ef2a3dd7f7096db82fc70a5e6b66e89def6db64319d3c55bbc

Observation c1563ff7-ba0d-46bc-be4f-687540325268 · outbound

This paper cites Qwen2-Audio Technical Report.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Qwen2-Audio Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:65dc162fd03fed362fe10c8c35708363790cd51cf68a4999d7d0c1224f5bd498

Observation e34f69ff-9d3f-4cdc-b5b8-c6d410e50366 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:233a69e376e68f19372119d2fd1de90dd3b99757c1154a45ac77b9c4aa89c655

Observation 20cb19fe-baef-466c-9d1c-b8f9a4045965 · outbound

This paper cites Kimi-Audio Technical Report.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Kimi-Audio Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:16a3739f07958a7d9a13105a690a231edefd5ca3ad0f57a9bd892e33672e632b

Observation 8901b41a-3090-472b-a5f4-f6b629133d10 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:a1636e346687224ab472f5d3bc2b541a9de8b5ef2331e80d4b53a8fe2050b9fb

Observation 426ea703-f382-449c-bca5-8ae6abffb455 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:35d3831192542ba94bd5e6843b07c2ef9fbd70f17b54d5ddaabafb43d6c751bd

Observation f22ce0c7-6123-44d8-9e7d-b63d1bbd7f3c · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:a1fed0c163ff04dcdd862824c0738b845b2c3ea65b81e46afa0653c757fc2222

Observation 899adca6-ce50-4d14-adad-bb9e64429028 · outbound

This paper cites In: Proceedings of the IEEE international conference on computer vision.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the IEEE international conference on computer vision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:f403ab6cebeeabd3c8105c37469bbbe8a6082a660930ba1aed7d81e9cf3d8ece

Observation 22214d73-b292-41bf-a7f0-b5230e90e96c · outbound

This paper cites In: 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:f4662c48f9a4c1a50bfe675def1ae05f120f301bcaf9676eca44a50e147107c0

Observation 666ac65c-fc96-4d8d-9b86-cd384491563d · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:fc014b4226480e3c0a9b672ef8e6d189f5f3ea43bb36b10be340193521841103

Observation 8dec227e-335d-4097-a9ca-ee1d075fd9ae · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Con- ference.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the Computer Vision and Pattern Recognition Con- ference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:a10fcddbfafdce0d0812bee1f015e37316a72f663e3edf1257a372570d907602

Observation ea274fe0-9202-4966-b8b5-ad83df298189 · outbound

This paper cites Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:b44ba17db85013ddb582909dbb679a54c51efe3eca967267792ffe217ff6fb59

Observation c8e42549-7300-4a39-b422-21309c1f4c80 · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:6d6ec879b7de77e04f6bc1f3cad9d1e396c65545523e45471c68836466889c34

Observation e192c371-7c04-471a-bb8b-4f75c5959552 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:2c6e767716bde5e8ccd83e5d46bffe138ef78321e61655bc1771b89b7507d8c4

Observation 83f73e97-53f5-4b12-aa23-1f3c8bca789c · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:3c1404e4258f17e99420775fd2151a1b7f6ed5d770d59ae92eac9c60050febf7

Observation 6ea64c05-1253-4716-a1bc-2488ec89a1bf · outbound

This paper cites SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:1e26bef2556dbd6ec99861c2a6abe5e82e7abe3015068bfa3b0f4c6d73f7798e

Observation 0b640bc0-9300-4f94-8396-f551f6d43e7d · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:05e5fad1f6a42078972545cf1ecbe97ff9d3656848737d6b726f8a079d809d7f

Observation 66c8cdc3-85ea-4c68-bbbb-44b54f6a0f15 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:038595fd770db0772ad8c8eefd7f14f16ee69b683194c1e87b5e5ab42806409d

Observation 117905fa-97ff-49ad-8942-0132490d4acb · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:1c1fbdf3c510c0929863a544a0d53d719616825a5067870a10017feeb77adf63

Observation 7d8fdbc0-0e15-4253-8bcf-d3a18c340c6f · outbound

This paper cites In: 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:d9146de3e7d6fd96426d5ef50bc002709178859cdacda79c0eab35f0b0998be7

Observation c4ae2a6e-01ce-45d6-976a-319fac3f1e7f · outbound

This paper cites arXiv preprint arXiv:2501.15368 (2025) 17.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2501.15368 (2025) 17

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:a76386b20668adfea2cef5a0c7b0076c0829144394391502ccde3bdfbae9780a

Observation 8ba5aa49-781b-416c-96b3-5318330c63c5 · outbound

This paper cites arXiv preprint arXiv:2409.15272 (2024).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2409.15272 (2024)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:ab3cd86584f871f5ef01d329880c2a26520e98f59854dced4769c9b729e7b439

Observation 242cdefd-7494-4b76-b68a-d5a9f9f6d434 · outbound

This paper cites In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:983c0d5ccdbae11440374ffd0c6df9cdb744426fb2e50f7871bcba2fb1fe4475

Observation 3e6200c7-c6f1-4dea-a75c-8eb14b16b730 · outbound

This paper cites arXiv e-prints pp.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv e-prints pp

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:11c7b2fb563cf5d9598c696ca58683bedc8a02dac0b85dfb78acacd1bdacb86c

Observation 15ec56f8-173e-4780-b416-249258d469bf · outbound

This paper cites AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:ea7f9fe5dffb785153b240c57e7d77987c39a25d3f781951280844ea7f0ed149

Observation 7186ecf6-6145-494f-9dab-65ba9f74b0a5 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:bb0bf3061143678a86bb707a815ffe1137fb3d54a5fac0f2060fb0775455efd7

Observation f2bbfce8-97a9-4735-a9b9-a3805db6b5e5 · outbound

This paper cites arXiv preprint arXiv:2510.12720 (2025).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2510.12720 (2025)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:6530eaf0b4e465f4030b5729cad850e2833f984d06ca462eb192de65d2f1c8a1

Observation 9a39ee61-66ae-42ff-b61b-8d0f34638da0 · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:92a8127c18779213c69d6cd2763b95d7fc54e660506ba8a7d689ad1c3509e74c

Observation bc24fa89-09ca-4088-b910-5e1e172959bc · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Foundation Models for Video Understanding: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:6892633ed47629fb4f51ef916479059a317030ffa404a3707e8235f64aea6578

Observation c2b19117-48e9-429c-800b-66533200d0a2 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:6ec6a8531afd7f3c71480010a1deb77177b0ae0124aab1f58a03767965fa5726

Observation 45e94a6d-8039-47bf-8763-76a132f391a9 · outbound

This paper cites In: Findings of the Association for Compu- tational Linguistics: ACL 2024.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Findings of the Association for Compu- tational Linguistics: ACL 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:7a4f5dd1d9c950494ecc6a2e0a5813768404e5e88e6ad624ee0088c9e316e7b3

Observation 4be00678-912e-4a64-971e-1bfaa72111f3 · outbound

This paper cites an unresolved cited work.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:1af1167ae71ce826d601bf7c976c528cf12992190a0c3e99bc64ae0a253578f7

Observation 5bf72378-f78f-4eae-8a40-18460747d011 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:75293407a0f46302a3130df6323b6051234e1be7b011b712fe7751ad6697c3c7

Observation d4f64712-ff1a-48f4-8d01-dfbdc52a793f · outbound

This paper cites In: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP)

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:840a625d787bc4bb280baf60681e1c93676c989cb94b729a040bc521c69a32b2

Observation 97a2d668-0f08-4655-b5ac-843dadd5dde6 · outbound

This paper cites IEEE Open Journal of Signal Pro- cessing (2025).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception IEEE Open Journal of Signal Pro- cessing (2025)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:dd7e2660755e7b44b333c5593b41eb660238b98c1225ffaf688f095cd21e5b7f

Observation 25df050b-7fdf-44a3-b143-9b5a5aa79ed2 · outbound

This paper cites arXiv preprint arXiv:2506.15220 (2025).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2506.15220 (2025)

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:f50f9fd21edb9a952937a3a5e981c8a66f10062eb007bb30a0b62ae55e92cfbc

Observation 9de63501-da43-4803-8297-dfea764fabba · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:28082d371e7059bade85f157911b9c1d56d14849d687fb4ff934a698b1951861

Observation 026cc425-c0e4-464b-8452-64bdec5377e4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Gemini: A Family of Highly Capable Multimodal Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:367642cb16eace27cdb8faf9cb885f9193553c726ed0d9d99a0d2cb1a24be3b1

Observation e17b00f3-79d1-48e3-9348-8c5d73cd2cb6 · outbound

This paper cites an unresolved cited work.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:f394ca18065f153dc26ee060e1bc4ebcc30db9315052231c39f9ae55c88955a9

Observation 841482f0-ea6e-4ecb-afcd-d52c1cbe2326 · outbound

This paper cites arXiv preprint arXiv:2505.14142 (2025).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2505.14142 (2025)

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:5d41dce7262a01aa67f7475188852e57858c7981f1d8101e217e6a57eb9da74d

Observation f50cb978-de2a-423d-b438-32d0dbaa3482 · outbound

This paper cites arXiv preprint arXiv:2507.11336 (2025).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2507.11336 (2025)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:b55d9e2ebc2a8940b402ae78b704d67b1d8ea7c3e858c661c064c3b81c2c1363

Observation 297d5e84-f265-43ef-a212-c47902c8ce54 · outbound

This paper cites In: Proc.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proc

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:b57146c98d85cfda4908bb90fda917262bc82c210213852802748dab737ccaa9

Observation 84f6b8d5-4d6c-4a06-b108-520ab69e421e · outbound

This paper cites Qwen2.5-Omni Technical Report.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Qwen2.5-Omni Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:24691df7022b02876a34145fdfb1241c315915fd5aa5945e8778c046f99e106e

Observation 86962074-cfb2-40db-ba8f-bfeb0e247000 · outbound

This paper cites Qwen3-Omni Technical Report.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Qwen3-Omni Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:33c05c848d22aa5f8ebecbe2e2b53732d31a7053ded5fc1872fbe158ad04fb18

Observation fd178ad9-7632-4d54-a204-02d248c10e62 · outbound

This paper cites Qwen3 Technical Report.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Qwen3 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:bb4dea91aa1037b1512dc72ab4b682c01a7a763d94c983cfa6895e48e6506139

Observation 823d461b-d0c3-4667-9928-5656e2fbbc9c · outbound

This paper cites Advances in Neural Information Processing Systems37, 57240–57261 (2024).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Advances in Neural Information Processing Systems37, 57240–57261 (2024)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:88abd72acbb0a756c50a6fb81b31fe0b4d97307a83a815478a4ac32849212730

Observation ebc4918d-32c6-4fef-baf1-6b95a94bc3bf · outbound

This paper cites In: Proceedings of the 30th ACM international conference on multimedia.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the 30th ACM international conference on multimedia

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:c3246a442c3b3d06aed75ea4bcfed3563fd3493803df204dd1e6c97c14ac0a42

Observation 8890fdbd-95f2-49ea-9cb8-00dcbccba527 · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:79e0f663298e0798e345a35feb0e2f570e2588a224f45cf2e6f75bbcb3ed1ae0

Observation 245d2fc8-3b73-4978-bb76-a61bc2991baf · outbound

This paper cites arXiv preprint arXiv:2503.19951 (2025).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2503.19951 (2025)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:9b1f3d356ef8b9e94e8f0d437c60868d81c5abb4024efd639f92c180a5a12ef5

Observation ce02ccee-d154-4f7d-9426-51d53cceb3f5 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:a8951f6c7234ec0ad7d70208e1b27802db8458cfc5d0351705728a0d8b759556

Observation d164e67c-bf6c-4464-b022-82e16c4f643d · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:1b51a4ba88d04f74242c98aad37a1a3cdeb3ae56081ce51d8d72c96587444168

Observation 45946344-9f2b-453a-b9c3-74c76bfe3a06 · outbound

This paper cites arXiv preprint arXiv:2505.17862 (2025).

Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv preprint arXiv:2505.17862 (2025)

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:3782148079167924844e587c0489d82f58a24281df76ec0464d181bee40f679b

Observation 9e709361-2ea4-4002-9f1d-ab988ae2cf75 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:0814f16251b3d1a1f6adefeb0df5fb277a29eeb98e0c84393fb0a907b3166f70

Observation f333eccd-da14-41c8-9aaa-c3ebbfaa84c0 · outbound

This paper cites an unresolved cited work.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:3f99e856e199cd7de182c030376ee1fbd71866c1a50f7127b21869e42deda116

Observation 35bb48cc-c616-4901-8094-f4375117200a · outbound

This paper cites 1/5, 2/5, ... 5/5.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception 1/5, 2/5, ... 5/5

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:481bbdd8c3f08fb824d4c6bd972e65ae71162d0e11cdf8c7161288526dd3ce64

Observation 1ce3a712-2461-4e8f-bdfa-343d1f0aca7f · outbound

This paper cites errors": [ {.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception errors": [ {

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:e5e52f1ee264054d63442c91bd4ad5d4ebe69b7cd98804a59e1ea92afad47aa5

Observation e20afd3b-ac13-472e-a26a-5f4d076f3e12 · outbound

This paper cites an unresolved cited work.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:1c3569a3092d1e081f29986e24ab4a12f30c56c90284de3ca6cb8cbe7372492f

Observation 584c47ec-8c43-499e-8985-cc483d0a0cd4 · outbound

This paper cites Track how events unfold over time and highlight when key actions or changes occur.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Track how events unfold over time and highlight when key actions or changes occur

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:b1f1b00d5db1090421e922d918abcd5b021b41fd2bf4106db7582d66cf67213b

Observation ef7b190d-eee8-4cfb-bb47-b02748e9a6b7 · outbound

This paper cites an unresolved cited work.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:8d4d6edf645a95c9d7c3718dfa12a7bdfecd66c5faa737958a1ae9afca135e0c

Observation 8d00a093-3d3e-4913-a56e-cbeafdbe0214 · outbound

This paper cites an unresolved cited work.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:c43f100cf775153e49a01454fbbad0c5861087ab056d3ba36074342085ec5c77

Observation 90b1728a-5a2d-4474-abc3-216a53a5aa99 · outbound

This paper cites question.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception question

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:dfd9e7581a8f4571c598bed4e9481931f52a5c97034fb18a1614e397bdff3843

Pith citing papers

No inbound Pith citation observations are available.