Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T02:44:52.920816Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 3 inbound Pith citation observations for arXiv:2604.19193.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T02:44:52.920816Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T10:38:22.619277Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T20:57:23.153286Z
100 of 103 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6e4d32ba-24bc-45bd-934e-607289e37cb2 · outbound
How Far Are Video Models from True Multimodal Reasoning? GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation deaab920-39a6-4d1a-baee-42e2dafd5374 · outbound
How Far Are Video Models from True Multimodal Reasoning? Oxford University Press (1984) 6
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e09d2cf-76c8-41de-a044-490ba13a73aa · outbound
How Far Are Video Models from True Multimodal Reasoning? VideoPhy: Evaluating Physical Commonsense for Video Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9864ebcd-f30e-4cc2-9c58-0e545b7f4650 · outbound
How Far Are Video Models from True Multimodal Reasoning? VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec496d56-a06b-4ec3-840b-117c60f06c02 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d42086c-960a-483a-9a37-1dccf399e2d2 · outbound
How Far Are Video Models from True Multimodal Reasoning? Advances in neural information processing systems33, 1877–1901 (2020) 1
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ce4eb6d-a50e-488d-8029-e814b9663fc4 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2015) 4
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8520187-7f24-407e-9afb-5b38b685c028 · outbound
How Far Are Video Models from True Multimodal Reasoning? Training-free group relative policy optimization, October 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b9ddc2a-5337-4393-a4c5-beb2b6385b24 · outbound
How Far Are Video Models from True Multimodal Reasoning? Emerging Properties in Self-Supervised Vision Transformers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b591f5e2-e6ce-48ce-ad74-ac81762f24d8 · outbound
How Far Are Video Models from True Multimodal Reasoning? Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1822d7c-af3c-46ce-9b39-9104241dc763 · outbound
How Far Are Video Models from True Multimodal Reasoning? Emerging Properties in Unified Multimodal Pretraining
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e99af556-8697-44f9-9c12-c6262988f3fb · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE international conference on computer vision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2445045-e6d9-4182-b38a-59f9449287e6 · outbound
How Far Are Video Models from True Multimodal Reasoning? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 973a294d-f044-40c4-a8dd-3fa76d11c722 · outbound
How Far Are Video Models from True Multimodal Reasoning? VILA$^2$: VILA Augmented VILA
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb159a9e-4f83-41e7-80b6-8c522d38ac45 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8fc3120-7dfe-4d30-9945-c2666a8df059 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cd44c22-9197-40fc-b21d-a9360fd3201a · outbound
How Far Are Video Models from True Multimodal Reasoning? RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4c9c31a-6721-4d31-9c71-30050e497e9c · outbound
How Far Are Video Models from True Multimodal Reasoning? LTX-2: Efficient Joint Audio-Visual Foundation Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f54c3ac-48db-49f1-a83b-c091f267e172 · outbound
How Far Are Video Models from True Multimodal Reasoning? Video-Bench: Human-Aligned Video Generation Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09227181-e68c-485b-b41f-72da79208551 · outbound
How Far Are Video Models from True Multimodal Reasoning? arXiv preprint arXiv:2512.07826 , year=
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e4307a2-df9b-4922-a401-24adf4841b6e · outbound
How Far Are Video Models from True Multimodal Reasoning? Huynh-Thu, Q
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9665ed3-4b2b-4ad7-9ce5-ef91a69b4ee7 · outbound
How Far Are Video Models from True Multimodal Reasoning? VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5af0c13-da3b-42c4-adea-d3bae9dc93c3 · outbound
How Far Are Video Models from True Multimodal Reasoning? Measuring Massive Multitask Language Understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d17476e-fcf0-4501-88e5-6652d083fae3 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8c6461c-8aed-4a14-a096-0e57e14776d0 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a2e8c76-5c9c-48f5-9488-ea7bd9579f38 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024) 2, 4
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bede773c-2c44-4bbc-934d-2cbadb9f6fa4 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ac7acfb-08be-49ff-b396-a1aeca5256d3 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1aa385aa-bef8-4eb9-a418-1eeca0bfcf08 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84eab3ba-4a50-4a12-b338-cbb29f745f3e · outbound
How Far Are Video Models from True Multimodal Reasoning? Jiang, Y
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98a7fe84-2002-485f-a39e-436583c22303 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 910aad13-a03f-43df-999c-dfb436e01e3c · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13d7ffae-51a9-4fa2-8687-069a284d29d5 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE conference on computer vision and pattern recognition
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2f2b7b3-2aab-4796-87f4-1b828f9644d8 · outbound
How Far Are Video Models from True Multimodal Reasoning? FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7710b13-6fff-41bf-9363-2582b403ccf6 · outbound
How Far Are Video Models from True Multimodal Reasoning? Advances in neural information processing systems35, 22199–22213 (2022)
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d398dce-f81e-45df-b318-02285048d72c · outbound
How Far Are Video Models from True Multimodal Reasoning? HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e58883ef-2f62-4f55-b899-ea9d05938812 · outbound
How Far Are Video Models from True Multimodal Reasoning? AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27cd712f-594d-4bad-b1f0-99009a0320ef · outbound
How Far Are Video Models from True Multimodal Reasoning? LLaVA-OneVision: Easy Visual Task Transfer
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c24059df-e2fd-4adf-9f32-8f6d1bfc3bdb · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92b21136-59fa-4d3a-ba62-09cf07365568 · outbound
How Far Are Video Models from True Multimodal Reasoning? Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5ef4e07-2da1-43d8-94f8-0df1aa567665 · outbound
How Far Are Video Models from True Multimodal Reasoning? Zero-shot Voice Conversion with Diffusion Transformers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1afd451f-3ce8-435d-8514-c72ede652ee4 · outbound
How Far Are Video Models from True Multimodal Reasoning? EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 216da338-cbcc-4efc-955b-3425461459dd · outbound
How Far Are Video Models from True Multimodal Reasoning? FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6558052-1313-4c04-a51c-3db5ee0b87b5 · outbound
How Far Are Video Models from True Multimodal Reasoning? Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb649c18-4c9f-4a86-a8b6-fca162f97b67 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92c38536-02bc-4500-abd8-fac819c910b9 · outbound
How Far Are Video Models from True Multimodal Reasoning? pixverse.ai/(2023), accessed: March 3, 2026 4
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02e2aa0d-d6f7-4bd1-83f0-2c3ea9eb37ea · outbound
How Far Are Video Models from True Multimodal Reasoning? Movie Gen: A Cast of Media Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65de20bc-0713-4a33-8362-aa43b89c55bc · outbound
How Far Are Video Models from True Multimodal Reasoning? AoPS Wiki,https: //artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions, accessed: 2026-03-01
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b51b408-1063-40fc-9ae3-178ebd680f84 · outbound
How Far Are Video Models from True Multimodal Reasoning? GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97326679-9731-4b33-8030-ee6873eff826 · outbound
How Far Are Video Models from True Multimodal Reasoning? Videoworld 2: Learning transferable knowledge from real-world videos.arXiv preprint arXiv:2602.10102
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 298a8e46-9e1a-422b-948e-049929df1a85 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d55cae8b-62ff-4634-900c-ff5b9440ef24 · outbound
How Far Are Video Models from True Multimodal Reasoning? Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53934287-e432-4800-ae0d-82469cf8ea96 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceed- ings of the Computer Vision and Pattern Recognition Conference
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b824c7d1-6a8f-4896-8aa3-d9e972877e5b · outbound
How Far Are Video Models from True Multimodal Reasoning? In: The Twelfth In- ternational Conference on Learning Representations (2024)
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af536bcd-cec5-404f-969a-ee91dc10dedc · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the AAAI Conference on Artificial Intelligence
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02ed66ca-ac1b-418d-a9da-7cac3eb0bd4a · outbound
How Far Are Video Models from True Multimodal Reasoning? Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7443b29-6b35-43ed-98c2-d638fe66a287 · outbound
How Far Are Video Models from True Multimodal Reasoning? Kling-Omni Technical Report
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06ff4caa-f75b-473e-bc92-ddbc2fb7deae · outbound
How Far Are Video Models from True Multimodal Reasoning? HunyuanVideo 1.5 Technical Report
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a6c97c1-9ca6-42de-8a7b-98ab45fc1b44 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a28a771e-2934-455c-8ec8-31c4868f3364 · outbound
How Far Are Video Models from True Multimodal Reasoning? Wan: Open and Advanced Large-Scale Video Generative Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 813c54ac-e2a3-40fb-b85c-767a1d1d83a3 · outbound
How Far Are Video Models from True Multimodal Reasoning? ModelScope Text-to-Video Technical Report
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da7dfb3e-0f17-42a7-a9c7-69e5613ac49a · outbound
How Far Are Video Models from True Multimodal Reasoning? A very big video reasoning suite
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30a700d8-1665-4b73-854b-8111b55dfd0b · outbound
How Far Are Video Models from True Multimodal Reasoning? In: The Fourteenth International Conference on Learning Representations (2025) 4
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e32b48d-9b2e-4270-9b5e-c5e67edcf36d · outbound
How Far Are Video Models from True Multimodal Reasoning? Emu3: Next-Token Prediction is All You Need
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2fe887d-48f1-4ba8-971d-401d1f1b8c96 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: International Conference on Medical Image Computing and Computer-Assisted Intervention
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9122528-1e0e-48b7-b0d4-4eca101c9181 · outbound
How Far Are Video Models from True Multimodal Reasoning? UniVideo: Unified Understanding, Generation, and Editing for Videos
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eae2bd51-5c8f-4e39-9f79-38d4f658dd8d · outbound
How Far Are Video Models from True Multimodal Reasoning? Emergent Abilities of Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4daf5438-bf57-48ba-a6d8-7872d39fc58c · outbound
How Far Are Video Models from True Multimodal Reasoning? Video models are zero-shot learners and reasoners
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a224fdf9-b2c8-4cc0-8d01-244147e58636 · outbound
How Far Are Video Models from True Multimodal Reasoning? Univbench: Towards unified evaluation for video foundation models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b6c1e91-4471-490d-8255-80a75efc2616 · outbound
How Far Are Video Models from True Multimodal Reasoning? Visual generation unlocks human-like reasoning through multimodal world models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4670310-d722-47b3-bff7-4fbc42650372 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9330d049-5038-4673-ac99-0dafde9484b0 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: The Thirteenth International Conference on Learning Representations (2025)
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a12dcc8f-7077-49b0-ab8a-b4fb99cdd0b7 · outbound
How Far Are Video Models from True Multimodal Reasoning? VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d14a08f-2413-403e-b909-d341021b9fda · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE conference on computer vision and pattern recognition
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f594a219-3105-44df-9bdb-d0002bddcddb · outbound
How Far Are Video Models from True Multimodal Reasoning? VideoGen-Eval: Agent-based System for Video Generation Evaluation
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5b6c987-7af5-4a54-bfef-cda3f245644a · outbound
How Far Are Video Models from True Multimodal Reasoning? HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b37ae508-993f-4659-90a4-5529f0eefdfb · outbound
How Far Are Video Models from True Multimodal Reasoning? CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 389bbf94-25db-4377-be12-2d390b611fab · outbound
How Far Are Video Models from True Multimodal Reasoning? UNIC: Unified In-Context Video Editing
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfc04d27-65d8-4eab-9edc-1a5dca5b8bea · outbound
How Far Are Video Models from True Multimodal Reasoning? TextGrad: Automatic "Differentiation" via Text
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cd7821b-5897-4ccd-9f09-285ea8ed333c · outbound
How Far Are Video Models from True Multimodal Reasoning? VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation babd594d-de25-4a6a-9780-46494b97dec0 · outbound
How Far Are Video Models from True Multimodal Reasoning? International Journal of Computer Vision133(4), 1879–1893 (2025)
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01e1f432-0791-4a24-a933-309174507fc5 · outbound
How Far Are Video Models from True Multimodal Reasoning? Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f906ba5-9822-4877-8bf1-634876cf3c39 · outbound
How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE conference on computer vision and pattern recognition
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b226fd3-8636-40ff-8536-1c12f81a988a · outbound
How Far Are Video Models from True Multimodal Reasoning? Dynamic Diffusion Transformer
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb6f6736-e25d-42ce-8de7-fde8528464fb · outbound
How Far Are Video Models from True Multimodal Reasoning? In: The Thirteenth International Conference on Learning Representations (2025)
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b820b6b5-9595-4b4c-8f4b-43be5c6e4f99 · outbound
How Far Are Video Models from True Multimodal Reasoning? Large Language Models Are Human-Level Prompt Engineers
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27ff338b-f6df-4a15-bddf-c75c8d9dd94c · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 092f5470-8455-4fdb-81cd-7f5cabcb7460 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74ae64fb-f936-4b7e-8bde-9318a5bf4f7c · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97b84dee-3dba-4f33-a149-9423021955ea · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5aa8b00-35ba-4a68-99bc-e4eb20c77a43 · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c2bf64f-dca5-4700-9d3e-e826571b4d72 · outbound
How Far Are Video Models from True Multimodal Reasoning? 0[integer represents the data point]
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77462e89-3a5d-4f8f-9141-5900d412e6f8 · outbound
How Far Are Video Models from True Multimodal Reasoning? You should locate the most important and recurring weakness_correlation
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fa71fd8-6ec9-4c3b-b0a2-9f6bedb88f52 · outbound
How Far Are Video Models from True Multimodal Reasoning? You can reflect on the mistakes of the reasoning process in the model_evaluation grounded on the video
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 033c3046-f612-4d23-990c-6f8727207943 · outbound
How Far Are Video Models from True Multimodal Reasoning? Output Format:\
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efe822ac-6993-46af-abc6-028907ae918c · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22c4f212-68a8-48ef-b513-b693e3bbd36e · outbound
How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b32484b9-03d1-4b51-b22d-a33b8f6059b7 · outbound
How Far Are Video Models from True Multimodal Reasoning? reasoning
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b1a68f5-1ad5-4324-855c-56d12e797c63 · outbound
How Far Are Video Models from True Multimodal Reasoning? It details: what the video has presented now (the weakness) and what it should have presented
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cbeb9eb-2b60-4a1a-8552-02e89fa97245 · outbound
How Far Are Video Models from True Multimodal Reasoning? It details the weakness of the video
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7379099-2bc2-4004-882f-8a29378c8449 · inbound
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization How Far Are Video Models from True Multimodal Reasoning?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 856341ca-efa8-4bab-a6bb-7b1eeec1c6ec · inbound
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization How Far Are Video Models from True Multimodal Reasoning?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93c4d9c3-d5fb-4301-94ac-3f3d8c06a502 · inbound
VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation How Far Are Video Models from True Multimodal Reasoning?
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.