Pith. sign in

Paper Citation Record · LEDGER

Understanding Complexity in VideoQA via Visual Program Generation

As of 16 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 2 inbound Pith citation observations for arXiv:2505.13429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13429 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:18:53.143452Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:32:22.239838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:02:34.105509Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy54
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89d10c2e-1c25-4f44-848f-5fe60e5a602c · outbound

This paper cites write newline.

Understanding Complexity in VideoQA via Visual Program Generation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:50.991012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:50.991012Z digest=sha256:a9c8b788a1c403307ff784f1b2c6e20fefe74ecf055c71faf736d1878db52f87

Observation 0bd956ba-abd7-406e-9e0a-4b1cdbb95458 · outbound

This paper cites A shared neural substrate for action verbs and observed actions in human posterior parietal cortex.

Understanding Complexity in VideoQA via Visual Program Generation A shared neural substrate for action verbs and observed actions in human posterior parietal cortex

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.059253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.059253Z digest=sha256:564d3d85e976d6ce6712daed131afe8215d2074dc5cc1e12aa092d0d1fe45a51

Observation 4e4bd6e0-9009-433b-9681-1475a3df77d4 · outbound

This paper cites Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications.

Understanding Complexity in VideoQA via Visual Program Generation Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.064334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.064334Z digest=sha256:bd0b8c387d15292d12027e1ab6dd9a317ba4e89cdd9d9d5b6a707acb4f4d78e8

Observation cc8c7136-bec3-4ceb-bc97-11ae44aaad58 · outbound

This paper cites Neural module networks.

Understanding Complexity in VideoQA via Visual Program Generation Neural module networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.068935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.068935Z digest=sha256:9a219ddeefc18c6179e0ca279b311871d05417b9a62b7b600b32784f002a20d4

Observation 8a7bc1fc-7307-4c53-8d3b-808d616c1fc6 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

Understanding Complexity in VideoQA via Visual Program Generation Is space-time attention all you need for video understanding? In ICML, 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.072802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.072802Z digest=sha256:4f615dc356d94f5bc699f574bc4517200550f4ab7b5c942515ad882247feec96

Observation a0cf77f5-ca21-45f3-b652-b62bcca0071f · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.076952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.076952Z digest=sha256:a45e17e0b91d55c3b5f7b7ea493eb93874681a59d77ada77761199a58d500104

Observation f7092f22-b78f-4178-9d96-c073b9fd1a44 · outbound

This paper cites H., Lu, P., Nocedal, J., and Zhu, C.

Understanding Complexity in VideoQA via Visual Program Generation H., Lu, P., Nocedal, J., and Zhu, C

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.081807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.081807Z digest=sha256:f79fbbcab8ed302b718e34b4314f28ac3369f11da4b696dd2f8a3f7976de1ffe

Observation 27f51f1a-24da-4516-addf-e93ff3112c37 · outbound

This paper cites ActivityNet : A large-scale video benchmark for human activity understanding.

Understanding Complexity in VideoQA via Visual Program Generation ActivityNet : A large-scale video benchmark for human activity understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.196727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.196727Z digest=sha256:e92559d42771a0f6574dc5bdd71f57ec657c14fe435be78444e3d989e7398f35

Observation 7d088cd8-c151-44b2-b710-7a69da522442 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Understanding Complexity in VideoQA via Visual Program Generation Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.223294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.223294Z digest=sha256:c7ac90be34e6061dd5a32bebbc38d93b16869678fb361a395569e9c08179c18e

Observation 60cc9b1c-45ba-4a8b-b171-89c597e44e29 · outbound

This paper cites E., et al.

Understanding Complexity in VideoQA via Visual Program Generation E., et al

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.227980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.227980Z digest=sha256:4d3fc2b94ef19fb4d72eb7fc0556f3529cc583c5d9e6ef4a27884770ad85e77a

Observation 40d5f1a4-85a6-4567-9246-d9a0ef99a0b7 · outbound

This paper cites and Gr \`e zes, J.

Understanding Complexity in VideoQA via Visual Program Generation and Gr \`e zes, J

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.231760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.231760Z digest=sha256:d7314c4e3eb6ab1a33282fb145f2269a4c11f901ae66bf36bf047275b54c302c

Observation 1b166968-310d-4ddb-a329-fb8689f7c3ea · outbound

This paper cites Brain activity during observation of actions.

Understanding Complexity in VideoQA via Visual Program Generation Brain activity during observation of actions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.566769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.235391Z digest=sha256:a381d02d2f34e7a807afd4f560c1a26ee90a6d29b5d44431930042f07bc4b7a2

Observation 5887e24f-ebc2-42d2-ac0c-16ec36401406 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:56.553800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.240031Z digest=sha256:132882870ae125c36e08b0b33eae6d448c56c9bc5cc4d4e52132991cac9b3f4d

Observation ca03a4fd-f291-4ce6-b059-8afaa07ecb69 · outbound

This paper cites K., Winn, J., and Zisserman, A.

Understanding Complexity in VideoQA via Visual Program Generation K., Winn, J., and Zisserman, A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.540693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.244319Z digest=sha256:3d018a3b1d2074a6972925a563178fa4f3d10133caaf900c4152f3bc10620f2f

Observation 7130db4c-09a2-4529-ae70-5129474e11da · outbound

This paper cites and Soto, A.

Understanding Complexity in VideoQA via Visual Program Generation and Soto, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.413268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.310511Z digest=sha256:dff149e9f5fab2afa6e439590b71f2bafdcc842cb531b16a4ee6a7026415f5b7

Observation d2e8376d-b64f-42d8-919d-f2117f9d7be6 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Understanding Complexity in VideoQA via Visual Program Generation Masked autoencoders as spatiotemporal learners

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.329880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.313786Z digest=sha256:faf74a81ec0930ad01d1c7b03b4a4a9875b310671294293d22f70597c8282007

Observation ecacacbb-e0c7-4e8b-921e-e8b30b1338f9 · outbound

This paper cites GPTScore: Evaluate as You Desire.

Understanding Complexity in VideoQA via Visual Program Generation GPTScore: Evaluate as You Desire

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.317153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.317153Z digest=sha256:20cf29c408ff7e85daf55d686ccd3d2415449f647cf0008222aac601fa323a4b

Observation 0fa37ad3-6d7a-43d0-bcc8-f59a618acda0 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Understanding Complexity in VideoQA via Visual Program Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.322345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.322345Z digest=sha256:08df463295a4e2e71e5a182b3bfbdda999b88594bdb7754c0085a515739799f1

Observation c108a262-3a8f-4caf-8cd2-0b48b7b7fc61 · outbound

This paper cites Recursive visual programming.

Understanding Complexity in VideoQA via Visual Program Generation Recursive visual programming

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.318898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.392721Z digest=sha256:2c490f4d307015738c8a909e210efd8e312931a68625556051c7078792ccbd5c

Observation b0e85dfe-50ca-49f7-af29-7c23e729e380 · outbound

This paper cites Adaptive computation time for recurrent neural networks.

Understanding Complexity in VideoQA via Visual Program Generation Adaptive computation time for recurrent neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.308405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.477662Z digest=sha256:9ec65ef6bd3512348ef4abf95ee299bbfa0fb46cce3c74859927a55e5ecc2869

Observation 1fd69439-8776-4337-9730-9ad50c48a07b · outbound

This paper cites AgQA : A benchmark for compositional spatio-temporal reasoning.

Understanding Complexity in VideoQA via Visual Program Generation AgQA : A benchmark for compositional spatio-temporal reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.183194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.544947Z digest=sha256:9468faffd15f8131f158c6293ddc793e3f319e2800d4010befd86331dd3fa730

Observation 0765632e-91f3-44ac-b645-9ed3a96339d4 · outbound

This paper cites and Kembhavi, A.

Understanding Complexity in VideoQA via Visual Program Generation and Kembhavi, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.174661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.549995Z digest=sha256:1a5e2e849e30ed336d8e378b89535eccb4eeecf06866971f495cba152a59921a

Observation 77b9ce6b-182f-48ed-83b7-62ea30d7dc28 · outbound

This paper cites V., Sethi, R., and Ullman, J.

Understanding Complexity in VideoQA via Visual Program Generation V., Sethi, R., and Ullman, J

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.140553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.553834Z digest=sha256:6dc1b84a26a29163cfc3d3fce1eff3c87af86e0fe1efc029f833a1eaa009ef41

Observation bb2e49d6-2f87-4124-8aa8-6bdd98dfe28a · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering.

Understanding Complexity in VideoQA via Visual Program Generation Learning to reason: End-to-end module networks for visual question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:56.022028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.557803Z digest=sha256:16cafb4e7c456c59757d1385480603ebf7e6fe2cc3a9cdc7464d8baacbc15a6e

Observation db603616-1975-4174-a68f-6b4da6f4e151 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:56.011441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.562304Z digest=sha256:a6188a26c183434da1c30d656e16aa72e93c0d2f4ec85e62814731ca3940b56e

Observation b11e4c7e-41b0-48a1-acb0-8711aed44e0e · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.999582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.566362Z digest=sha256:6b9ad4c1b27cd24cabcaac7e2223a6f892c1923c7fcfa60290ebcfd70fe916a2

Observation bad9a63c-f821-409f-9078-d7a0fd2d04a1 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.990558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.571755Z digest=sha256:35dd98954fc1060e8fd044adf28ad7e36567f2237669d7104d11c100fad85987

Observation 7be27510-0989-4a80-a577-fae7d25c3287 · outbound

This paper cites and Han, Y.

Understanding Complexity in VideoQA via Visual Program Generation and Han, Y

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.831349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.636026Z digest=sha256:5d5d73437f20b73a0cc1b2b5d8e14e1f65a9effd7cca23957a2b58537a5238a6

Observation 52d0685e-e369-426c-8407-dc37a16de43f · outbound

This paper cites CLEVR : A diagnostic dataset for compositional language and elementary visual reasoning.

Understanding Complexity in VideoQA via Visual Program Generation CLEVR : A diagnostic dataset for compositional language and elementary visual reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.787570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.682106Z digest=sha256:b8ed6204e32881dae222311578a40f6321077bc373e7027945ffadb10e9004bf

Observation a4a7cd14-7886-4e6f-bf97-dabeafb351df · outbound

This paper cites Inferring and executing programs for visual reasoning.

Understanding Complexity in VideoQA via Visual Program Generation Inferring and executing programs for visual reasoning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.777869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.780729Z digest=sha256:fc56c37bddbe3f69d004e3481240a4d80a152cf5510ee1b9a580fa007c573098

Observation b0c07739-aa8c-4d9d-8cd4-07464fbab446 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.766798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.784668Z digest=sha256:57c9b2c251f716110e54a2cde79ccedeb23d88c5709e2a2c4fe8fa1a6756dbd8

Observation cf7058e5-aedf-4cb1-9b87-09d8d8c7e458 · outbound

This paper cites W., Tapaswi, M., and Fidler, S.

Understanding Complexity in VideoQA via Visual Program Generation W., Tapaswi, M., and Fidler, S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.688625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.788366Z digest=sha256:f2be19b0551a979ee32c0d54d3fd3a42c678d322287b9f0a2526959f692f3e52

Observation c3beb552-4555-4324-a1cd-0a9299354cdb · outbound

This paper cites and Bojar, O.

Understanding Complexity in VideoQA via Visual Program Generation and Bojar, O

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.677424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.791691Z digest=sha256:0c19a4cffc6b873999772bd12fbcd5acb74a1e392b7df72b8d2f5010fdc33787

Observation 99390403-e02d-41ab-9bb4-aa4900e56ad2 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.664823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.796153Z digest=sha256:cbf8db74733b28e20e39d2ff3876b6db6bf5adaabc158338a42032edb3243bf9

Observation 59170970-a60e-427d-b6ea-254d76a67255 · outbound

This paper cites and Kramer, O.

Understanding Complexity in VideoQA via Visual Program Generation and Kramer, O

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.586622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.801724Z digest=sha256:a3b83a0f520fb77dd32f22e812012f039211ec24cbbf56a0ac2afbed6457c258

Observation eb2bc160-0bf3-4a8e-94f7-99ef0991233a · outbound

This paper cites Dense-captioning events in videos.

Understanding Complexity in VideoQA via Visual Program Generation Dense-captioning events in videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.505228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.806322Z digest=sha256:c15058897d1c99258cfdfb20e8046118091f2c1dfbc649fe2a1d524571c1cb38

Observation b345012c-50fb-41b0-bbc4-56890447f830 · outbound

This paper cites BLIP-2 : Bootstrapping language-image pre-training with frozen image encoders and large language models.

Understanding Complexity in VideoQA via Visual Program Generation BLIP-2 : Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.491312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.897852Z digest=sha256:88394e00d5e85a9bd4e682d9e707815ce8b8d019a93bebdb738b5987be1f3680

Observation 6004c26b-7436-4087-bf83-ba831b93b451 · outbound

This paper cites MVBench : A comprehensive multi-modal video understanding benchmark.

Understanding Complexity in VideoQA via Visual Program Generation MVBench : A comprehensive multi-modal video understanding benchmark

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.440557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.973337Z digest=sha256:4709910991bf2b393a35c1c9dd0706033e93dc61cb31a96899c6d6e0f75931f2

Observation 08bbdb7b-4e67-43fd-9767-878955cf28ac · outbound

This paper cites A technique for the measurement of attitudes.

Understanding Complexity in VideoQA via Visual Program Generation A technique for the measurement of attitudes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.976804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.976804Z digest=sha256:73a939d1b638893173e0b0b50fa99ce3d03e0d8a9813b1143f7ed1795727f61c

Observation d852b290-6d99-4c5b-aa41-734f0a6094bc · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.359443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.980311Z digest=sha256:54a147c485483d9d73ce58e1e9b064f958e6fe4de01587642f0cadf50b0e1057

Observation 7e17724d-cf08-4d74-af08-42ae8fdd86df · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.984196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.984196Z digest=sha256:61edd47f85a8d447a29e1af04bc75ee99c4f3bbe0f49a7966a5e4acb44444943

Observation 0ccd3f8a-0eff-47ba-b4fb-0de746ebedaf · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.341479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.988101Z digest=sha256:a4c6c20f28ae8e6ae9689fac2324dd993d387a909c7bdb50e77f6527435b6507

Observation ad5bd453-9680-40d6-b2c8-8c243b968d21 · outbound

This paper cites L., Nejadasl, F.

Understanding Complexity in VideoQA via Visual Program Generation L., Nejadasl, F

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.328942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.992410Z digest=sha256:7cfb0b520202eb15be3fe8cae0da91342add62dc06bc2856c2be582ac33588ee

Observation 1a2eef59-ab74-4526-aee2-f79c3c8ecddd · outbound

This paper cites C., Adeli, E., and Li, F.-F.

Understanding Complexity in VideoQA via Visual Program Generation C., Adeli, E., and Li, F.-F

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.294349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:51.996538Z digest=sha256:02bf415a7fabf292ba0a6076d8994e34058ad5f1c0a9032edc8f5c07b848c085

Observation 4e3f0a8a-6d99-4add-8081-45b11def5ea2 · outbound

This paper cites Y., Wu, J., Niebles, J.

Understanding Complexity in VideoQA via Visual Program Generation Y., Wu, J., Niebles, J

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.251931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.154088Z digest=sha256:362a52813601e4f09c37a62178b1d34ff6f50b30ac6278dc585b2b7c6f12c57e

Observation 2956f30d-8cff-437f-bb74-40007ba61fad · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

Understanding Complexity in VideoQA via Visual Program Generation VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.268106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.268106Z digest=sha256:ea1a7cd498d4699c9f056a3638b392a2574564600c7e473d4bb94e53c7ceee20

Observation 04dd940f-c62f-4b38-98a0-001d23921706 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Understanding Complexity in VideoQA via Visual Program Generation Self-refine: Iterative refinement with self-feedback

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.241956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.272442Z digest=sha256:1b3397f8b4ce3ddd4f4a7dd223cb8072da122128ba80436bcbe55cad5bf5c85f

Observation 3e63888b-157e-4a7b-8945-445c70b22ef1 · outbound

This paper cites EgoSchema : A diagnostic benchmark for very long-form video language understanding.

Understanding Complexity in VideoQA via Visual Program Generation EgoSchema : A diagnostic benchmark for very long-form video language understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.229170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.276215Z digest=sha256:c4ce947ce7348027e3b01183e55139ef92412d64451cc6b1f0d347b0dc9fbeed

Observation 157573ae-15cc-4564-a454-05c2a6117481 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.218861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.281674Z digest=sha256:fbd13fc2d34dc5d1d5d75c86d8fe19bc9f42cdd7f1e7313427665538bfe64f9f

Observation bd40ece6-92b4-4b2a-b863-56dd25c2bbb5 · outbound

This paper cites and Vondrick, C.

Understanding Complexity in VideoQA via Visual Program Generation and Vondrick, C

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.179229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.285183Z digest=sha256:fef6badcd0a215b373dba81ff2c62d511a03f793fa7bf2ca5b6bf8a90f0d29cc

Observation 0f745b9c-c4a7-4e3a-936d-d8e0b923dc59 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:55.084828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.290047Z digest=sha256:70cf743e94039afa5636e634fd4407dd2ea32e3df6cef2d1b39bcb99bbb29e1a

Observation 4abd0291-46ce-4fce-bd37-46f94e1b3f70 · outbound

This paper cites GPT-4 technical report, 2023 b.

Understanding Complexity in VideoQA via Visual Program Generation GPT-4 technical report, 2023 b

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.072753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.294290Z digest=sha256:8f6930d1eca24523a4a257d15fff9db5a522fb454bf2eaaab69708d4c0651c85

Observation c016c4a7-fb51-4f2e-83c6-d55dffdf232f · outbound

This paper cites and Schitter, C.

Understanding Complexity in VideoQA via Visual Program Generation and Schitter, C

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:55.035573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.355937Z digest=sha256:ad61f348d5d4c92a8c382e5d6933bf885435d8843fec517e71413b6167c7b3fd

Observation ca391c39-d642-4a39-be9c-1360b4519c2f · outbound

This paper cites A., Stretcu, O., Neubig, G., Pocz \'o s, B., and Mitchell, T.

Understanding Complexity in VideoQA via Visual Program Generation A., Stretcu, O., Neubig, G., Pocz \'o s, B., and Mitchell, T

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.935382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.439614Z digest=sha256:6b1aeb8b6c1a25d5fdf3d66f0878df4f023096b4bcd9f9d459b3e90fed821817

Observation 9bb79496-4021-4868-be98-b342926e1830 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Understanding Complexity in VideoQA via Visual Program Generation W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.505588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.505588Z digest=sha256:c969ba3ee3d0df1c5f47bbf0baba5d72e7fa67f3245b2171d8079c1868d97eac

Observation 66c93eda-5306-4770-b782-9afe71995b4d · outbound

This paper cites D., Ermon, S., and Finn, C.

Understanding Complexity in VideoQA via Visual Program Generation D., Ermon, S., and Finn, C

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.510472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.510472Z digest=sha256:90110cb493893c72caed74a9c810ee40dcc89341950ddf6a4f7b23daa39a1bb7

Observation 42d4c87c-6a6a-44b6-827e-3703364546a6 · outbound

This paper cites Annotating objects and relations in user-generated videos.

Understanding Complexity in VideoQA via Visual Program Generation Annotating objects and relations in user-generated videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.892020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.513979Z digest=sha256:56d69dc4bb13c91d3731480be21fbea62c9d29bca47fdf389a72c3c851b05453

Observation 2d485fa0-c56d-454b-b223-ea5a89d6485d · outbound

This paper cites HuggingGPT : Solving AI tasks with ChatGPT and its friends in hugging face.

Understanding Complexity in VideoQA via Visual Program Generation HuggingGPT : Solving AI tasks with ChatGPT and its friends in hugging face

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.835259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.517327Z digest=sha256:113710b81c11ffc2965cda17e279131e455e2f3fe05055bb487c3778b793bf96

Observation fb75f1b2-6544-412a-8f2e-20e9e56f825f · outbound

This paper cites A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A.

Understanding Complexity in VideoQA via Visual Program Generation A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.821822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.522238Z digest=sha256:282832406967abad415dc5296c1a1b4aa6d573571826dd599ce584c917a1a895

Observation bf749fed-8985-4223-b74e-56a9e968d7a8 · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:54.729547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.526964Z digest=sha256:c9eceb37aec323836893959ec381dbe29beae3e0fd580e93110efad59b52c131

Observation f403c41a-248b-4a1e-a39c-ce172b6e01e8 · outbound

This paper cites T., and Leordeanu, M.

Understanding Complexity in VideoQA via Visual Program Generation T., and Leordeanu, M

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.687545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.531242Z digest=sha256:eadcaf83f19176d2fb861e1de02367066250ca94255a031966d5ff415022140f

Observation 86f92031-0c26-4206-bfcd-adab7d323854 · outbound

This paper cites less is more.

Understanding Complexity in VideoQA via Visual Program Generation less is more

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.616302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.535344Z digest=sha256:95a0bacf21c9f638f1396ffb8807252c6999846d15241d0e6a04a7a13a04c4e1

Observation 6719b62a-8ec8-449d-ab37-3ca6793abaa9 · outbound

This paper cites Modular visual question answering via code generation.

Understanding Complexity in VideoQA via Visual Program Generation Modular visual question answering via code generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.548302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.616141Z digest=sha256:fa2adab119fa775ce8ae8e63125cecc542fda0ef70e7fad0d92e22aa39ee5985

Observation ca8dff22-2529-4b59-8ff5-755e67f4c2ac · outbound

This paper cites ViperGPT : Visual inference via python execution for reasoning.

Understanding Complexity in VideoQA via Visual Program Generation ViperGPT : Visual inference via python execution for reasoning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.536876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.677195Z digest=sha256:bb4e991864fa71556edbec00340cfe803e9e2678e5b692f4d25a43828ebd51c5

Observation fcb49ab8-56ad-40a6-a2ba-042cc377d8ee · outbound

This paper cites T., Fu, J., Phan, M.

Understanding Complexity in VideoQA via Visual Program Generation T., Fu, J., Phan, M

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.472736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.680634Z digest=sha256:432d0ad94005ac56734a303230bc7127280b9bcf9972808a9728738b8054e048

Observation 9b720306-7b7f-4986-94b3-e57a2bb6c320 · outbound

This paper cites A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J.

Understanding Complexity in VideoQA via Visual Program Generation A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.441325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.684219Z digest=sha256:bfdd2f986faa01174346b0dc3f941994dc976ec77dea45d7303b29c8737e09c3

Observation f9125a1f-35b7-4ff1-83ea-0b0b3f591949 · outbound

This paper cites Learning the curriculum with bayesian optimization for task-specific word representation learning.

Understanding Complexity in VideoQA via Visual Program Generation Learning the curriculum with bayesian optimization for task-specific word representation learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.430869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.689888Z digest=sha256:4c0b1242ce789caf538688411fcb5418fb05d0c85c06b2083d0feb2db092d101

Observation 16537687-d975-4425-be25-714d1e1780e8 · outbound

This paper cites Learning the curriculum with bayesian optimization for task-specific word representation learning.

Understanding Complexity in VideoQA via Visual Program Generation Learning the curriculum with bayesian optimization for task-specific word representation learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.419868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.693265Z digest=sha256:97e8bb17a6a38c0b047dd1264b9fbae45e248545f3e2a01cbf361d8d0a031dcf

Observation d4511928-59ec-49e5-b35c-d7b93c0d0c51 · outbound

This paper cites P., and Ferrari, V.

Understanding Complexity in VideoQA via Visual Program Generation P., and Ferrari, V

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.395810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.697978Z digest=sha256:61a9eaed982bc3ade33f96496b18cd9458cd71339a119121d3abf5b47f25efd8

Observation bb1fa786-6272-40f2-8ba2-23a4e27b8b15 · outbound

This paper cites Neural discrete representation learning.

Understanding Complexity in VideoQA via Visual Program Generation Neural discrete representation learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.365097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.737006Z digest=sha256:343bb8a969057e20d61f9aa0517f961002c4ff4a7d4538e015959f9d097b4572

Observation 2c1be96e-0b4a-4d08-ac31-3363eb4e34f0 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Understanding Complexity in VideoQA via Visual Program Generation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.797005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.797005Z digest=sha256:e8af01dcf404a1d09fabb5caefc5fa2c9b03df4ad93cd81f3db35fdee1b23c83

Observation 9d6faa57-5ac2-4c06-9732-2142efba93d2 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Understanding Complexity in VideoQA via Visual Program Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.824535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.824535Z digest=sha256:d4a2082d6cb1aed0dc66d5aa11e61a059b85b4936d3ade4534a36feaf8143b98

Observation d3c51b07-1823-4f27-bf79-d91550cafb26 · outbound

This paper cites Language models with image descriptors are strong few-shot video-language learners.

Understanding Complexity in VideoQA via Visual Program Generation Language models with image descriptors are strong few-shot video-language learners

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.251351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.829436Z digest=sha256:f9a8966417743a963f77e3f006b050de64e6283ac0b25d806a963203ccbae17d

Observation 024758c1-057c-4296-af9d-9cd960e4c43e · outbound

This paper cites STC : A simple to complex framework for weakly-supervised semantic segmentation.

Understanding Complexity in VideoQA via Visual Program Generation STC : A simple to complex framework for weakly-supervised semantic segmentation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.240302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.833774Z digest=sha256:13ec2708d1cca8c15598c14dce5de9f986690e4c9ea5178ea6d044c7d909f994

Observation 5a5aa68f-7ab6-46df-96c9-f8cad68acceb · outbound

This paper cites B., and Gan, C.

Understanding Complexity in VideoQA via Visual Program Generation B., and Gan, C

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.170867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.838175Z digest=sha256:03c075a0454af1784ad6de861f64d2fe800ecf869c86efa4c772c70b8d5cd1ac

Observation 6c4c1dfd-2fd1-44bd-bff1-c6f74c7d7b0d · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:54.091843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.842464Z digest=sha256:84dd3524e101812be5d012182b80cde3ec15f15dfb31c6f886ecd3a4d992942f

Observation e52093e4-2b9e-4fb0-ad5c-b5dd892ec27d · outbound

This paper cites Next-QA : Next phase of question-answering to explaining temporal actions.

Understanding Complexity in VideoQA via Visual Program Generation Next-QA : Next phase of question-answering to explaining temporal actions

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.081918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.875308Z digest=sha256:32d320823cfe8e84ad1d27653415cd35b0fd62741f5f027114262d7b62a84c15

Observation cdf6bca2-8547-46d6-9c05-8ee636428641 · outbound

This paper cites mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video.

Understanding Complexity in VideoQA via Visual Program Generation mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.964977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.964977Z digest=sha256:a826cbceb5e0f307fd19a79755e833cdbe3a943eca4abdc2e7f0c3beaad7f1a2

Observation 0303ce46-2eb6-4197-9eed-bb9a81059ffd · outbound

This paper cites W., Salakhutdinov, R., and Manning, C.

Understanding Complexity in VideoQA via Visual Program Generation W., Salakhutdinov, R., and Manning, C

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:54.002089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.969649Z digest=sha256:629fa515533d8f134a18add99c9ccf5e2873249231db58f8f567028ed2fd9fc2

Observation 7a2ab9c0-d6f5-4a9e-83a0-62150c5bf623 · outbound

This paper cites Neural-symbolic VQA : Disentangling reasoning from vision and language understanding.

Understanding Complexity in VideoQA via Visual Program Generation Neural-symbolic VQA : Disentangling reasoning from vision and language understanding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.931880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.972655Z digest=sha256:cdf2efb6a292c271175159ed6d94882d27b3fbb08a82ffabe2ed2b5fe8d5229b

Observation 40c08bc2-3a66-42bc-b321-6d9ddfb3a8a1 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Understanding Complexity in VideoQA via Visual Program Generation Self-chained image-language model for video localization and question answering

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.920673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.977650Z digest=sha256:809e1462e1a938812f383be95cbe8c80028706d1dcf96315b9669e8b18d975c6

Observation 1b879339-726b-4b7f-8ee6-5397d7ae8b1b · outbound

This paper cites ANetQA : A large-scale benchmark for fine-grained compositional reasoning over untrimmed videos.

Understanding Complexity in VideoQA via Visual Program Generation ANetQA : A large-scale benchmark for fine-grained compositional reasoning over untrimmed videos

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.740461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.982250Z digest=sha256:144786afedc639b104f5bba6a3e5673f44a48255d10c606d3efb098c052e1580

Observation 662ab958-dd81-444f-abc3-bf481bfcf122 · outbound

This paper cites S., Cao, J., Farhadi, A., and Choi, Y.

Understanding Complexity in VideoQA via Visual Program Generation S., Cao, J., Farhadi, A., and Choi, Y

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.728692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:52.987034Z digest=sha256:45103943869b2db3ae58e96fb66bd8c29e6f6f045d0bb1a75ba5734bb69d99ff

Observation 4902b6f6-d0f6-4c1f-8293-e79c098a773d · outbound

This paper cites Socratic models: Composing zero-shot multimodal reasoning with language.

Understanding Complexity in VideoQA via Visual Program Generation Socratic models: Composing zero-shot multimodal reasoning with language

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.662002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.039709Z digest=sha256:4f454927c3e7c2902fd452f937e80a961cb432222b39016865c896c959bf3678

Observation a80565bb-cabc-40ef-b32b-3474db711920 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Understanding Complexity in VideoQA via Visual Program Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:53.128703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:53.128703Z digest=sha256:3eba874c5443f6c07363bad957fd46018d7577a6bcd7dd7ac7c3bd1ba97b047c

Observation 3f09dded-6a1b-4309-98b5-1a3fe53b8741 · outbound

This paper cites Where does it exist: Spatio-temporal video grounding for multi-form sentences.

Understanding Complexity in VideoQA via Visual Program Generation Where does it exist: Spatio-temporal video grounding for multi-form sentences

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.493343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.132142Z digest=sha256:af64f5e0e10f05bebd121d07ac7ca7b9ca874b347386b2b3547a01545c68e514

Observation cd2bf090-c967-4181-93b4-2b902cba8d16 · outbound

This paper cites Video question answering: Datasets, algorithms and challenges.

Understanding Complexity in VideoQA via Visual Program Generation Video question answering: Datasets, algorithms and challenges

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.481138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.135905Z digest=sha256:31f5b2707aeb16fda540183b3940dbc9232e1171007f01476214c92571dcb817

Observation ed530913-ea58-4888-b9e1-a821b2d8c0ef · outbound

This paper cites J., and Rohrbach, M.

Understanding Complexity in VideoQA via Visual Program Generation J., and Rohrbach, M

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:18:53.467634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.139787Z digest=sha256:0c2ca2fbadd37f30bd1571aa5a62a191697dd20026579fbc91dddd9264985810

Observation 5c21cffe-8e23-465d-901e-3412ac9b505c · outbound

This paper cites an unresolved cited work.

Understanding Complexity in VideoQA via Visual Program Generation Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:18:53.276370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:18:53.143452Z digest=sha256:b0f4560bc428f10ad0c5aaaf712f57c5f5696d680fad29e40094625bf5c9b290

Pith citing papers

Observation 8a371f5e-a9cb-4c5e-89fe-76a533985358 · inbound

An Attribute-Based Measure of Video Complexity cites this paper.

An Attribute-Based Measure of Video Complexity Understanding Complexity in VideoQA via Visual Program Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.107311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T19:00:54.718177Z digest=sha256:ab1770db532053862c567b88de8530a999a607e1253b18702d9b45ee7cfbf806

Observation e4a04b15-1e58-4a35-83c1-cd3e39cf05e4 · inbound

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding cites this paper.

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding Understanding Complexity in VideoQA via Visual Program Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:22.239838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:32:22.239838Z digest=sha256:6d06718ea42d9f1e0e1dc06a74f1f1931c934d7e438dfaca969c1561c3ce7c9d