Pith. sign in

Paper Citation Record · LEDGER

How Far Are Video Models from True Multimodal Reasoning?

As of 7 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 3 inbound Pith citation observations for arXiv:2604.19193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19193 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T02:44:52.920816Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T10:38:22.619277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T20:57:23.153286Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact11
  • verified fuzzy33
  • unresolved14
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch41

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e4d32ba-24bc-45bd-934e-607289e37cb2 · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

How Far Are Video Models from True Multimodal Reasoning? GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:26:20.435474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:40afdc99ba563c4d05267c85faddbb447a4239143781ae7e3084ab192264f226

Observation deaab920-39a6-4d1a-baee-42e2dafd5374 · outbound

This paper cites Oxford University Press (1984) 6.

How Far Are Video Models from True Multimodal Reasoning? Oxford University Press (1984) 6

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.757394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:4c31e093f31f42b4cf2572c02b02ac6c621103d366558a1ee73616f23ac5fb00

Observation 0e09d2cf-76c8-41de-a044-490ba13a73aa · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

How Far Are Video Models from True Multimodal Reasoning? VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:34:37.994076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:97e50522c6c3ea1f11b698baf3fffd7753b1b687028c5c3993718c81dfbe0179

Observation 9864ebcd-f30e-4cc2-9c58-0e545b7f4650 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

How Far Are Video Models from True Multimodal Reasoning? VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:03.643358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:96309885cf982fe08837a83fdf29288e2235f773d68ad3076ebf0270326ad0e1

Observation ec496d56-a06b-4ec3-840b-117c60f06c02 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.755695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:1420fbfcb056afd276d0ac8511744b5ede14e927d2266931cc7f7f0c4f7b8143

Observation 3d42086c-960a-483a-9a37-1dccf399e2d2 · outbound

This paper cites Advances in neural information processing systems33, 1877–1901 (2020) 1.

How Far Are Video Models from True Multimodal Reasoning? Advances in neural information processing systems33, 1877–1901 (2020) 1

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.851393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:996f8826e6632263fbfc7300807816d0dd1db59ce7d4c15e4e872b857c2a79e5

Observation 4ce4eb6d-a50e-488d-8029-e814b9663fc4 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2015) 4.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2015) 4

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.848543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:205d6ce042bd12b0f14c80963edb441f3a2a694bab5ae64e60268aa2382e033a

Observation f8520187-7f24-407e-9afb-5b38b685c028 · outbound

This paper cites Training-free group relative policy optimization, October 2025.

How Far Are Video Models from True Multimodal Reasoning? Training-free group relative policy optimization, October 2025

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:04.145213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:5fac66cb9288932f928a50d50fca4c5292eaff465c1a4dfd8a795730827be35c

Observation 7b9ddc2a-5337-4393-a4c5-beb2b6385b24 · outbound

This paper cites Emerging Properties in Self-Supervised Vision Transformers.

How Far Are Video Models from True Multimodal Reasoning? Emerging Properties in Self-Supervised Vision Transformers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:04:51.784333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:aa44a30d6fcbe5655b0ed76614d7fb2423f7ef4d215b9aceffb197ac7b17fa4f

Observation b591f5e2-e6ce-48ce-ad74-ac81762f24d8 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

How Far Are Video Models from True Multimodal Reasoning? Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:04.015545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:68cf5927b45ede21218fea2bf26dea6d2dc242f352b0aae73a4ac02bafb2d992

Observation d1822d7c-af3c-46ce-9b39-9104241dc763 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

How Far Are Video Models from True Multimodal Reasoning? Emerging Properties in Unified Multimodal Pretraining

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:04.132440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:d5bc6082a4ee347504dc179343e826811e485fb4b9c2ff407587b042ededb6d8

Observation e99af556-8697-44f9-9c12-c6262988f3fb · outbound

This paper cites In: Proceedings of the IEEE international conference on computer vision.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE international conference on computer vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.834545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:ddfe5c1d11de86b0d1fe2107115947bef556eb40ad67a25db48524969ec1b7e7

Observation c2445045-e6d9-4182-b38a-59f9449287e6 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

How Far Are Video Models from True Multimodal Reasoning? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.216352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:3b366250b9e4f001f340572168c92b6523a74f31d25d4fd1ac0a9ecb25a55b68

Observation 973a294d-f044-40c4-a8dd-3fa76d11c722 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

How Far Are Video Models from True Multimodal Reasoning? VILA$^2$: VILA Augmented VILA

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.007239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:3f96916c60486e49ecba08d03dd624219c6aced795669bdd763f66a7528ba1f5

Observation fb159a9e-4f83-41e7-80b6-8c522d38ac45 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.832892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:d1fb1a15751c415e36e53bd21d4995f948389c71ba7b4ed6113b2b945f9d4eec

Observation b8fc3120-7dfe-4d30-9945-c2666a8df059 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.760815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:7f9cd610d3c06731f115557b7e98bc79f57b7775b1111e60b2141138c518c2fe

Observation 5cd44c22-9197-40fc-b21d-a9360fd3201a · outbound

This paper cites RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs.

How Far Are Video Models from True Multimodal Reasoning? RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:03.698359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:12d7cf19bd54fde095d506eb7fd69eb4119379dcf8d2d1af6069bc2d158bb376

Observation e4c9c31a-6721-4d31-9c71-30050e497e9c · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

How Far Are Video Models from True Multimodal Reasoning? LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.198821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:687c489483f311aa7904342ad2a8d00e235e20b21eb71e276b397457961e5492

Observation 4f54c3ac-48db-49f1-a83b-c091f267e172 · outbound

This paper cites Video-Bench: Human-Aligned Video Generation Benchmark.

How Far Are Video Models from True Multimodal Reasoning? Video-Bench: Human-Aligned Video Generation Benchmark

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.036366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:0f8c996b4b75d534a97271fd6cf88e4edd37f15f3299d741b13369654e62d8a4

Observation 09227181-e68c-485b-b41f-72da79208551 · outbound

This paper cites arXiv preprint arXiv:2512.07826 , year=.

How Far Are Video Models from True Multimodal Reasoning? arXiv preprint arXiv:2512.07826 , year=

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:03.765354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:50d6943351faec4b873f0d1ae547f6f141e2cfd81acb61922633b4419fe334bf

Observation 6e4307a2-df9b-4922-a401-24adf4841b6e · outbound

This paper cites Huynh-Thu, Q.

How Far Are Video Models from True Multimodal Reasoning? Huynh-Thu, Q

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.053309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:7444b0c9bf0cc34f26b7a1780448297e4f31e5cc69d08b59732f57e75d8b7d64

Observation c9665ed3-4b2b-4ad7-9ce5-ef91a69b4ee7 · outbound

This paper cites VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation.

How Far Are Video Models from True Multimodal Reasoning? VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.183684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:3f6d5db7912bc87fb829ee2a93f673d030bb6653ffdf6f5636c2741e0274f7b5

Observation f5af0c13-da3b-42c4-adea-d3bae9dc93c3 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How Far Are Video Models from True Multimodal Reasoning? Measuring Massive Multitask Language Understanding

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:04.125403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:6e5e521085eecc0dc80b4225c0f54db841ae7726858f40a5e0bf3fbcb230d6db

Observation 5d17476e-fcf0-4501-88e5-6652d083fae3 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.828914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:257520b200c621467a3b808c4c97144bd83764fe0f21bc20e48481b36ada0fca

Observation e8c6461c-8aed-4a14-a096-0e57e14776d0 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.831211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:49377cc9918c61342363734e52592f67a95fc57eaf7af7a0352724006ef90da2

Observation 6a2e8c76-5c9c-48f5-9488-ea7bd9579f38 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024) 2, 4.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024) 2, 4

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.836603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:36b1b742e52dd5794c19f13e3c24be608e41497e3ab154a793e1652113230f90

Observation bede773c-2c44-4bbc-934d-2cbadb9f6fa4 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.839918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:d6c089540d59534abaeac10fbb36722a939a8894559250b3236cd039c681636d

Observation 6ac7acfb-08be-49ff-b396-a1aeca5256d3 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.762583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:727b0d9de9cff1fdfb59fd5bd8cea27fee13adac7d28a239eb506c20091cda18

Observation 1aa385aa-bef8-4eb9-a418-1eeca0bfcf08 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.759089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:4e74c7738ebc65425e83063dc8adddcc1bdc861f9797deb69d180f826f84fd01

Observation 84eab3ba-4a50-4a12-b338-cbb29f745f3e · outbound

This paper cites Jiang, Y.

How Far Are Video Models from True Multimodal Reasoning? Jiang, Y

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.243053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:4b31402936429079566e14051eb42a07fff7becfff77f436544f805f16db0d31

Observation 98a7fe84-2002-485f-a39e-436583c22303 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.825065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:3d67415b8ff6891bf147eb398d9f457299bf3834c12f4178b371543c81b14fbe

Observation 910aad13-a03f-43df-999c-dfb436e01e3c · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.820101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:2b8e87713f2570efe832e589f3ef80d4ab44533a0fa480ba013499f88a9cf605

Observation 13d7ffae-51a9-4fa2-8687-069a284d29d5 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.766354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:999ef57c65ab8484104764e4c90e9a168f45f1b3e66c292bbc217ef1d61fe68b

Observation a2f2b7b3-2aab-4796-87f4-1b828f9644d8 · outbound

This paper cites FullDiT: Multi-Task Video Generative Foundation Model with Full Attention.

How Far Are Video Models from True Multimodal Reasoning? FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.097620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:e0ef2ee60332fc5108216acb24101edcedd6d11925ee36ede3c041dc1429601a

Observation c7710b13-6fff-41bf-9363-2582b403ccf6 · outbound

This paper cites Advances in neural information processing systems35, 22199–22213 (2022).

How Far Are Video Models from True Multimodal Reasoning? Advances in neural information processing systems35, 22199–22213 (2022)

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.818531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:2eba6c82550ec5936efe2ea64f3c11e8bb65e3bb608aab2dfa2949b9f8631fff

Observation 0d398dce-f81e-45df-b318-02285048d72c · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

How Far Are Video Models from True Multimodal Reasoning? HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:03.861371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:16593134d1152e718ab0918739bf978efe160c2ad4f4bbe7a02d5b9ee3bd99b0

Observation e58883ef-2f62-4f55-b899-ea9d05938812 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

How Far Are Video Models from True Multimodal Reasoning? AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:04.287346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:a02443294b4d02ae97e691f3533896f8a20b2aa0caa3593a6fedb44d40a042b4

Observation 27cd712f-594d-4bad-b1f0-99009a0320ef · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

How Far Are Video Models from True Multimodal Reasoning? LLaVA-OneVision: Easy Visual Task Transfer

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:04.303140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:c99a21b742147bd0a912912aca61d2f3577359f6f980a2c52a0a893f14e3e744

Observation c24059df-e2fd-4adf-9f32-8f6d1bfc3bdb · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.823460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:76b0c1a5c6821af86d9e0fa92ca0017a94e34ec56cf7c4fc7b3b884ed9b8c3ae

Observation 92b21136-59fa-4d3a-ba62-09cf07365568 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

How Far Are Video Models from True Multimodal Reasoning? Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:24:05.047503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:192641e57452b9fd155e275688d6bc3953708c193d87bad07e8721a1f6583944

Observation a5ef4e07-2da1-43d8-94f8-0df1aa567665 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

How Far Are Video Models from True Multimodal Reasoning? Zero-shot Voice Conversion with Diffusion Transformers

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.023382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:c9724f9cb859167a444a2f1d772835e09c9ba4f24a207cc9f63ed105f342acef

Observation 1afd451f-3ce8-435d-8514-c72ede652ee4 · outbound

This paper cites EvalCrafter: Benchmarking and Evaluating Large Video Generation Models.

How Far Are Video Models from True Multimodal Reasoning? EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.086174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:88fd2fc1de29998425a1942b3610cb384f8015396244fef7de08d252baf0ff5e

Observation 216da338-cbcc-4efc-955b-3425461459dd · outbound

This paper cites FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation.

How Far Are Video Models from True Multimodal Reasoning? FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:03.923156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:31b9e850586c12e4b2c5b8aff561846f49a2abd18c70122d40d6fead00864b3e

Observation a6558052-1313-4c04-a51c-3db5ee0b87b5 · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

How Far Are Video Models from True Multimodal Reasoning? Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:40:00.221207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:5ad1ab4291d06d03d34f8f8e1e97281fbee14870c7b1408f98720c8e6a7da337

Observation eb649c18-4c9f-4a86-a8b6-fca162f97b67 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.813929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:f98f19b843412383aa8177033925eac636b514758565653daebaea2f0cc33a76

Observation 92c38536-02bc-4500-abd8-fac819c910b9 · outbound

This paper cites pixverse.ai/(2023), accessed: March 3, 2026 4.

How Far Are Video Models from True Multimodal Reasoning? pixverse.ai/(2023), accessed: March 3, 2026 4

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.764311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:2979b61ab1fecef8b9df7ac8a83ea5ae7b0f1fd9c20fb61258ad66955de12901

Observation 02e2aa0d-d6f7-4bd1-83f0-2c3ea9eb37ea · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

How Far Are Video Models from True Multimodal Reasoning? Movie Gen: A Cast of Media Foundation Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:26.778393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:fd0279f2ded9313ebe6cbe1647f52b5de7e091fc922e9cc669f7daebd90d6374

Observation 65de20bc-0713-4a33-8362-aa43b89c55bc · outbound

This paper cites AoPS Wiki,https: //artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions, accessed: 2026-03-01.

How Far Are Video Models from True Multimodal Reasoning? AoPS Wiki,https: //artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions, accessed: 2026-03-01

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.810709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:7bc74786d8f762a7df8da9eb056e66243b6d5e5c55d9f37795fe5d394fb6c9a1

Observation 4b51b408-1063-40fc-9ae3-178ebd680f84 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

How Far Are Video Models from True Multimodal Reasoning? GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:51:04.295877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:09aaecc066e9989d0e585cf6fccf684c519ef517a152f1bb4740286251cf123f

Observation 97326679-9731-4b33-8030-ee6873eff826 · outbound

This paper cites Videoworld 2: Learning transferable knowledge from real-world videos.arXiv preprint arXiv:2602.10102.

How Far Are Video Models from True Multimodal Reasoning? Videoworld 2: Learning transferable knowledge from real-world videos.arXiv preprint arXiv:2602.10102

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:04.104580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:74b1432737f15a273a9ae40c05e51297058d891682629a88f9440a8f369ffac0

Observation 298a8e46-9e1a-422b-948e-049929df1a85 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.812378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:a68b489acad2bbbf6432ba8392aca9ae385fb6aac69dc8a0ae5dad537182f5bb

Observation d55cae8b-62ff-4634-900c-ff5b9440ef24 · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

How Far Are Video Models from True Multimodal Reasoning? Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T01:35:37.953991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:9050068b318351ca82451879b914b478d4281ac04da77c952b69f6337b437ef9

Observation 53934287-e432-4800-ae0d-82469cf8ea96 · outbound

This paper cites In: Proceed- ings of the Computer Vision and Pattern Recognition Conference.

How Far Are Video Models from True Multimodal Reasoning? In: Proceed- ings of the Computer Vision and Pattern Recognition Conference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.808678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:d0658a5a840109504974cbbd5a9e9b598a0b72d5e69358772aa31a61a983cc13

Observation b824c7d1-6a8f-4896-8aa3-d9e972877e5b · outbound

This paper cites In: The Twelfth In- ternational Conference on Learning Representations (2024).

How Far Are Video Models from True Multimodal Reasoning? In: The Twelfth In- ternational Conference on Learning Representations (2024)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.806535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:fab44dc70459e60883b043fca640e87538a734b20b93f22a4be00100f3c833ad

Observation af536bcd-cec5-404f-969a-ee91dc10dedc · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.821749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:f067eca4ec2dda7b7c5e0e53b71a5e4a7a7304b9cec09316162d62cf3442bade

Observation 02ed66ca-ac1b-418d-a9da-7cac3eb0bd4a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

How Far Are Video Models from True Multimodal Reasoning? Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:51:04.042262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:b3aee5392912d877674067f163d794855979f026b0fd2521114be473760cf194

Observation d7443b29-6b35-43ed-98c2-d638fe66a287 · outbound

This paper cites Kling-Omni Technical Report.

How Far Are Video Models from True Multimodal Reasoning? Kling-Omni Technical Report

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:00:58.771713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:8812e41022616bd97019e30f303f485314e37c5ea5c6aab420dc95f9fd4229eb

Observation 06ff4caa-f75b-473e-bc92-ddbc2fb7deae · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

How Far Are Video Models from True Multimodal Reasoning? HunyuanVideo 1.5 Technical Report

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:eb7952e20d1e8189e0d1783f3dbf0bbdf502c5f85bba10c909494c0c56429782

Observation 5a6c97c1-9ca6-42de-8a7b-98ab45fc1b44 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.838122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:671f48c15d2e91c7d9e17b0f74f5246092d59b4f7c54a8f8b1ef13f24016472d

Observation a28a771e-2934-455c-8ec8-31c4868f3364 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

How Far Are Video Models from True Multimodal Reasoning? Wan: Open and Advanced Large-Scale Video Generative Models

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:03.884926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:29de3e9bdbc691c519586bd5d707854739f8323512d2ae1adb622a62514cfc14

Observation 813c54ac-e2a3-40fb-b85c-767a1d1d83a3 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

How Far Are Video Models from True Multimodal Reasoning? ModelScope Text-to-Video Technical Report

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T19:47:29.701736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:7c654646ef32c3ef632616d43ca5b260ccb095e4116e8efe6c7d68d1bc8e7441

Observation da7dfb3e-0f17-42a7-a9c7-69e5613ac49a · outbound

This paper cites A very big video reasoning suite.

How Far Are Video Models from True Multimodal Reasoning? A very big video reasoning suite

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.313284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:b085990064327130e9c0a539056e7f1a5672ab069561cca4c8a820b92bbaf3e9

Observation 30a700d8-1665-4b73-854b-8111b55dfd0b · outbound

This paper cites In: The Fourteenth International Conference on Learning Representations (2025) 4.

How Far Are Video Models from True Multimodal Reasoning? In: The Fourteenth International Conference on Learning Representations (2025) 4

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.801830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:158c07a49fc255eeaac15c486e10a29400d33dbdd98dab74d22f1488e350ba18

Observation 8e32b48d-9b2e-4270-9b5e-c5e67edcf36d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

How Far Are Video Models from True Multimodal Reasoning? Emu3: Next-Token Prediction is All You Need

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:04.152990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:35db28a59cf401c2f3c59c2005e07551b3b9d2162809e422aa8ffd945bca9a51

Observation b2fe887d-48f1-4ba8-971d-401d1f1b8c96 · outbound

This paper cites In: International Conference on Medical Image Computing and Computer-Assisted Intervention.

How Far Are Video Models from True Multimodal Reasoning? In: International Conference on Medical Image Computing and Computer-Assisted Intervention

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.797728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:aa1d0ae8134eab803145f625d07ec1ba9fece467567fb804e902d9c81fa54ccc

Observation f9122528-1e0e-48b7-b0d4-4eca101c9181 · outbound

This paper cites UniVideo: Unified Understanding, Generation, and Editing for Videos.

How Far Are Video Models from True Multimodal Reasoning? UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:47e14ff761352382ed7ae3be170f4d272ef26e09ce3594ed60698a8669fe0871

Observation eae2bd51-5c8f-4e39-9f79-38d4f658dd8d · outbound

This paper cites Emergent Abilities of Large Language Models.

How Far Are Video Models from True Multimodal Reasoning? Emergent Abilities of Large Language Models

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:03.896102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:5994c457a656fcde955855173fb76cc79c9c818b8f26c14c7b182fd702f5994e

Observation 4daf5438-bf57-48ba-a6d8-7872d39fc58c · outbound

This paper cites Video models are zero-shot learners and reasoners.

How Far Are Video Models from True Multimodal Reasoning? Video models are zero-shot learners and reasoners

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:257064ec061e84e52f8ae1b20cace77a5475414e51528389b35557197d402ea4

Observation a224fdf9-b2c8-4cc0-8d01-244147e58636 · outbound

This paper cites Univbench: Towards unified evaluation for video foundation models.

How Far Are Video Models from True Multimodal Reasoning? Univbench: Towards unified evaluation for video foundation models

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.276881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:ec3e6a7840179e2719b4279fbe0d92f62ba93dff0580034a7ae174f9e987320f

Observation 1b6c1e91-4471-490d-8255-80a75efc2616 · outbound

This paper cites Visual generation unlocks human-like reasoning through multimodal world models.

How Far Are Video Models from True Multimodal Reasoning? Visual generation unlocks human-like reasoning through multimodal world models

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:03.999629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:e182e9908b027d732337e2cd3e3cb508d19def062d9b2d1ce7b79d8d63b9f982

Observation a4670310-d722-47b3-bff7-4fbc42650372 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.827623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:61b2fa0d776a5aaf525b4665c50ee148fe2fe2ed67d38a24427657c5d3f8ac6b

Observation 9330d049-5038-4673-ac99-0dafde9484b0 · outbound

This paper cites In: The Thirteenth International Conference on Learning Representations (2025).

How Far Are Video Models from True Multimodal Reasoning? In: The Thirteenth International Conference on Learning Representations (2025)

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.826854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:14a9e71b11e0e85ffc8b071f2d33e5d32bc70583c5e06847f6e39a00cc247d25

Observation a12dcc8f-7077-49b0-ab8a-b4fb99cdd0b7 · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

How Far Are Video Models from True Multimodal Reasoning? VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:49:14.981353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:2aa913cbab183191bb2617d3274735b4e28849e94583b7a44d73ce8a9426bf8b

Observation 8d14a08f-2413-403e-b909-d341021b9fda · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.799893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:946c8c268e9e6a209dbc3656f4305fe2204f3c7070b2d9bca2d5c9da23ba5c74

Observation f594a219-3105-44df-9bdb-d0002bddcddb · outbound

This paper cites VideoGen-Eval: Agent-based System for Video Generation Evaluation.

How Far Are Video Models from True Multimodal Reasoning? VideoGen-Eval: Agent-based System for Video Generation Evaluation

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:03.671354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:c175d1a937d04a0b191c073894590e2b2dcceb769483759249db38ca65523783

Observation a5b6c987-7af5-4a54-bfef-cda3f245644a · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

How Far Are Video Models from True Multimodal Reasoning? HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:05:12.673916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:7afbf5373b3b5e03e1f6e4c07a1c2f00798209f0a1e49427adcf8f1764156d36

Observation b37ae508-993f-4659-90a4-5529f0eefdfb · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

How Far Are Video Models from True Multimodal Reasoning? CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:03.934837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:406a30f052a7e1fc40f8c526c7415234739b00fc91695ad21f0968856f911452

Observation 389bbf94-25db-4377-be12-2d390b611fab · outbound

This paper cites UNIC: Unified In-Context Video Editing.

How Far Are Video Models from True Multimodal Reasoning? UNIC: Unified In-Context Video Editing

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:03.953580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:edb914f848963366d21930805b22e691cb976666c28099a0220d1e52e4485430

Observation dfc04d27-65d8-4eab-9edc-1a5dca5b8bea · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

How Far Are Video Models from True Multimodal Reasoning? TextGrad: Automatic "Differentiation" via Text

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T11:27:59.484812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:24d3d1d2c8e9058e887852e4250ecc2fe1cb47bc11b24e942b261dc24884588a

Observation 0cd7821b-5897-4ccd-9f09-285ea8ed333c · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

How Far Are Video Models from True Multimodal Reasoning? VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:51:03.979870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:78d6f1bd91912256a9e90b8f907b4e9b694faae275dea2fbd35a2a7f4239a570

Observation babd594d-de25-4a6a-9780-46494b97dec0 · outbound

This paper cites International Journal of Computer Vision133(4), 1879–1893 (2025).

How Far Are Video Models from True Multimodal Reasoning? International Journal of Computer Vision133(4), 1879–1893 (2025)

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.798095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:fb0294b73c0deda19322f852814e0608d75deab85b08815ad84b5152fbd77e82

Observation 01e1f432-0791-4a24-a933-309174507fc5 · outbound

This paper cites Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models.

How Far Are Video Models from True Multimodal Reasoning? Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:44:08.223014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:1f80553da7fe5b0d4990b901c1b28a6c8dbf18a97e78486ec42c0b619328f00f

Observation 2f906ba5-9822-4877-8bf1-634876cf3c39 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

How Far Are Video Models from True Multimodal Reasoning? In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.805628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:9df6d119c88123fc114cc7bdf91d74e6cc96705684dedee4f26f583884fb4118

Observation 6b226fd3-8636-40ff-8536-1c12f81a988a · outbound

This paper cites Dynamic Diffusion Transformer.

How Far Are Video Models from True Multimodal Reasoning? Dynamic Diffusion Transformer

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.119874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:d4ff4db4e9e7168a1e4cd9a7f55bbc8dd5d221723c6dce42e8a7647604ab16dd

Observation fb6f6736-e25d-42ce-8de7-fde8528464fb · outbound

This paper cites In: The Thirteenth International Conference on Learning Representations (2025).

How Far Are Video Models from True Multimodal Reasoning? In: The Thirteenth International Conference on Learning Representations (2025)

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.784661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:3ceb434eca5d9e2a54e1d59a62d0571121b8ad8ce615f3abd6c88450ea8029e5

Observation b820b6b5-9595-4b4c-8f4b-43be5c6e4f99 · outbound

This paper cites Large Language Models Are Human-Level Prompt Engineers.

How Far Are Video Models from True Multimodal Reasoning? Large Language Models Are Human-Level Prompt Engineers

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:43:26.534305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:11f74dfa097f0c25aacca859353174fa9910e7d6984577679f89232795e8d9b0

Observation 27ff338b-f6df-4a15-bddf-c75c8d9dd94c · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.793229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:26d0108579ab400e90912c4c07f00f93b394e808f4ac93bc6ba6cdc8461cbce3

Observation 092f5470-8455-4fdb-81cd-7f5cabcb7460 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.781175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:45531bfd5d1754ab3bcf7bcf7a5e0e48b2ed2d9ee4a5dd053fb6c6896cae8743

Observation 74ae64fb-f936-4b7e-8bde-9318a5bf4f7c · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 99

Resolution
parse uncertain
raw_fallback, observed 2026-05-22T19:45:04.777744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:e31e604d430a804ca7c4fe6b868cf6c9327e7b8fbc4dfcaf6c62a74002a89cf7

Observation 97b84dee-3dba-4f33-a149-9423021955ea · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.815900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:35fb437b752c89db621d8c837b62a32635e2ea8b00e0c1b5ebc649ffe0dfc4c5

Observation c5aa8b00-35ba-4a68-99bc-e4eb20c77a43 · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 101

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.783091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:279845ac584a0922f3f11313a0723e6a059577cf4841f4aa1d59c78323bb5c82

Observation 9c2bf64f-dca5-4700-9d3e-e826571b4d72 · outbound

This paper cites 0[integer represents the data point].

How Far Are Video Models from True Multimodal Reasoning? 0[integer represents the data point]

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.788472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:4691def7a4ccfc7477c537750016d1da6769c510c76d478adb570841e290d7a6

Observation 77462e89-3a5d-4f8f-9141-5900d412e6f8 · outbound

This paper cites You should locate the most important and recurring weakness_correlation.

How Far Are Video Models from True Multimodal Reasoning? You should locate the most important and recurring weakness_correlation

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.803777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:46c6bbe724d054520211fa5902994ccaaf67ff4c7ad9cce8c3599cca100aef0c

Observation 0fa71fd8-6ec9-4c3b-b0a2-9f6bedb88f52 · outbound

This paper cites You can reflect on the mistakes of the reasoning process in the model_evaluation grounded on the video.

How Far Are Video Models from True Multimodal Reasoning? You can reflect on the mistakes of the reasoning process in the model_evaluation grounded on the video

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.776267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:ccf9ffd3721c3b01778e36797952f8308e6d746a61e024caaed473073b57e73f

Observation 033c3046-f612-4d23-990c-6f8727207943 · outbound

This paper cites Output Format:\.

How Far Are Video Models from True Multimodal Reasoning? Output Format:\

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.843248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:bbf27488f3b14604c8ba5dd1ae0a7170878a01bd82983fbd49e5ebbf079cba61

Observation efe822ac-6993-46af-abc6-028907ae918c · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 106

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.841390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:ba97c7b816ed3ef507ffe35395e7dc72779dbc811daed220d9414e00efc4d45a

Observation 22c4f212-68a8-48ef-b513-b693e3bbd36e · outbound

This paper cites an unresolved cited work.

How Far Are Video Models from True Multimodal Reasoning? Unresolved cited work

Reference 107

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:45:04.772828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:5cac317d5d38a13ff39ce8fe81cd2eca96056afdc83de5f7be20af0433734497

Observation b32484b9-03d1-4b51-b22d-a33b8f6059b7 · outbound

This paper cites reasoning.

How Far Are Video Models from True Multimodal Reasoning? reasoning

Reference 108

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.773819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:fb2a54edcca1a0ab67816bb89b03b32d2c3996b4a89b87371c6999e9786baf60

Observation 1b1a68f5-1ad5-4324-855c-56d12e797c63 · outbound

This paper cites It details: what the video has presented now (the weakness) and what it should have presented.

How Far Are Video Models from True Multimodal Reasoning? It details: what the video has presented now (the weakness) and what it should have presented

Reference 109

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.846946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:c2d44037d18c5240566f01a878e0e8a512e04f7cb7c9aca9c9b0afb0eb7414f2

Observation 4cbeb9eb-2b60-4a1a-8552-02e89fa97245 · outbound

This paper cites It details the weakness of the video.

How Far Are Video Models from True Multimodal Reasoning? It details the weakness of the video

Reference 110

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:45:04.767877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:d63a0b9bba0ce1a3d785c1d0f8b818e960a9e35268ec29f2f39f07640e9f8a06

Pith citing papers

Observation f7379099-2bc2-4004-882f-8a29378c8449 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization How Far Are Video Models from True Multimodal Reasoning?

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.309495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:ab0b551109a377c31a9d6d1d5e7b2586dfd22b728abcdf136bf414d875abd83f

Observation 856341ca-efa8-4bab-a6bb-7b1eeec1c6ec · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization How Far Are Video Models from True Multimodal Reasoning?

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:44:36.838077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:92f58bf4c583d5b29bfc1d7de670c2aa4f9e7e4154c7b2bdd51d2b416707292b

Observation 93c4d9c3-d5fb-4301-94ac-3f3d8c06a502 · inbound

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation cites this paper.

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation How Far Are Video Models from True Multimodal Reasoning?

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T20:57:23.154343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T20:04:10.711238Z digest=sha256:dd1cc74ec96c6a64e26f883c79733e7d16192ea4eaf213c7f80b65fd57891c1b