Pith. sign in

Paper Citation Record · LEDGER

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks

As of 22 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2412.01132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01132 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:08.780881Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:16.204012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:15:52.849428Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 129e97fc-118e-48b7-a6ed-537ce680ce53 · outbound

This paper cites Video question answering: Datasets, algorithms and challenges, 2022.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video question answering: Datasets, algorithms and challenges, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.628106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.628106Z digest=sha256:35e3f4628f9d7b636b2928d4e6bc768fb8598df0071e009757418a28e2b99790

Observation dfad524f-48b6-47a8-843c-d10e471f9518 · outbound

This paper cites Videoqa in the era of llms: An empirical study, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Videoqa in the era of llms: An empirical study, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.158989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.631391Z digest=sha256:c22263c525eefa99d92bbad6eeb76f1d1dd1d164c5a5ef32ebe13368ab09802f

Observation 44aedcde-13e9-4ce0-96ad-770869e13167 · outbound

This paper cites Francis, and Alessandro Oltramari.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Francis, and Alessandro Oltramari

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.152971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.634468Z digest=sha256:df6dd4b69aa1227386214f0c211960e21aa749729f84d7615cd2d144852e2e12

Observation 7f603df1-b510-4b89-9090-6979884feb52 · outbound

This paper cites Carla: An open urban driving simulator, 2017.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Carla: An open urban driving simulator, 2017

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.638216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.638216Z digest=sha256:ab35e73835e8ae37aad10fe01b11f382adfd235a84fc841b4c128bed33c1f3b2

Observation 4761f4e7-63f7-48f4-8915-06f870a3c798 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Lingoqa: Visual question answering for autonomous driving, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.142418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.640880Z digest=sha256:e31da6383477c6116f5a5d43abfe0e99af20960742a52ea162a92eb0fe1af8d6

Observation 20c7a9e5-ca6e-4f0d-967c-a60cf68fbd61 · outbound

This paper cites Align and aggregate: Compositional reasoning with video alignment and answer aggregation for video question-answering, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Align and aggregate: Compositional reasoning with video alignment and answer aggregation for video question-answering, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.135003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.643534Z digest=sha256:5743fe7d23eaf3c829b941117dad6cceed6c83cc3382c206044cfd24f66f1ad1

Observation c9995d22-7b77-4980-8a48-530f2f7d47e3 · outbound

This paper cites GPT-4o: System Card.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks GPT-4o: System Card

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.126898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.646199Z digest=sha256:f2404ed38b0618498eb058ddd8b246e68444eed9c1a57358ff333d4438bb9a56

Observation 872b3005-0b4b-4fe1-a58a-d34a9b1d47f3 · outbound

This paper cites Dynamic spatio-temporal graph reasoning for videoqa with self-supervised event recognition.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Dynamic spatio-temporal graph reasoning for videoqa with self-supervised event recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.119502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.649449Z digest=sha256:632d3b41f238fd9458dc3e2655752ba8230670256d9597e5c7ab99c03f26fccf

Observation 5248eca9-14de-4926-9522-8ba932fb945a · outbound

This paper cites Recognizing an action using its name: A knowledge-based approach.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Recognizing an action using its name: A knowledge-based approach

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.112135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.651945Z digest=sha256:a51f7f8feae6b71f26b60a1bb2ef19bc5775934466484d351dd445fc6f80e1ee

Observation 1f6253e7-a2ae-458e-8b2b-88901f668cc4 · outbound

This paper cites Video representation learning with deep neural networks.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video representation learning with deep neural networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.104141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.654520Z digest=sha256:17d412bc8fc23ced1f1dda5f622bc1661a17530d9c540449ecb0dcd6d21ce7aa

Observation f79878ee-26e8-49f2-a65b-bff1bcd36640 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video question answering via gradually refined attention over appearance and motion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.657336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.657336Z digest=sha256:5cb29b2cd3e29445b82f9dd3e026fada8fa1e561844ca3b5a4a3bbb630b31712

Observation 288832a7-3821-4f8b-a103-91085ed142b5 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representa- tions for vision-and-language tasks.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Vilbert: Pretraining task-agnostic visiolinguistic representa- tions for vision-and-language tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.092676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.659970Z digest=sha256:620a9aba32bd78c0fd129b424280d358d1ba34d587dbc26ccc42113bc431391d

Observation b4658756-8c62-4aa0-967b-6cc2596ce35a · outbound

This paper cites Compositional attention networks with two-stream fusion for video question answering.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Compositional attention networks with two-stream fusion for video question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.083934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.663000Z digest=sha256:b56f27159ef324991a0b8ae7c8a0c71fc6abfe158a46f5cfe172d7251656663b

Observation eff6c892-b8f5-4b73-bede-f3b5f70725ca · outbound

This paper cites Visualbert: A simple and performant baseline for vision and language, 2019.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Visualbert: A simple and performant baseline for vision and language, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.075921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.665457Z digest=sha256:ec6e76ced6b587959442351724698e042b679d590d969d3453905285510bc3dc

Observation e93753ee-4fb1-42ea-af3b-7a21806dd383 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Activitynet: A large-scale video benchmark for human activity understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.068111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.668187Z digest=sha256:7545d18b145643ed9a519ccd2d3cb775c954eb9e409d1bc2cee4bd2fcf24c195

Observation c73926d5-4343-4d0b-af80-0c22f0799a30 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description, 2016.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Tgif: A new dataset and benchmark on animated gif description, 2016

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.061016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.671405Z digest=sha256:375c74130a93067ad2073a5a317441066370f871098e2eef925492e7745f0acf

Observation 5d00888a-1c89-44a0-af8f-a916f2ee03a5 · outbound

This paper cites Next-qa:next phase of question-answering to explaining temporal actions, 2021.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Next-qa:next phase of question-answering to explaining temporal actions, 2021

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.052959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.673233Z digest=sha256:70d4b2ed710ed7363ce756b5aa6e0163a4e5d85ed5244f8a126b997a7bac6c28

Observation 79cac09c-c190-41f0-b3dd-f544ee9bf5d8 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Gemini: A family of highly capable multimodal models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.045004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.675015Z digest=sha256:2a3633e3a2b2edade0387ec883c974cc00a133d57dfa7930423a55e01a045013

Observation 08f1446b-302c-487e-971f-60fdc007da69 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video instruction tuning with synthetic data, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.677145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.677145Z digest=sha256:fc959034cbfef136542a43709121b1a9c272a494aba7e03a520437da2faa968c

Observation c726ee96-410e-46ec-9cbb-1a3f447843dd · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video, 2022.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Ego4d: Around the world in 3,000 hours of egocentric video, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.035218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.679543Z digest=sha256:a2be0facccf6a59d0dff2f2b022d85c698979f49409cb8f184e4254580824801

Observation e9d99f4f-780b-4100-9884-961ae8c0217f · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.028613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.681380Z digest=sha256:00a6d5235c2f2d60efb97c6bfc71e5a3d40f237af01b14d5a46d1b248c13586a

Observation 7573b07a-f401-4f63-8698-9066dda52739 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.683715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.683715Z digest=sha256:3023e54495177fc85b37a03dc0f51dd4a7d9437394147ceb356f5ed7eed7ad6f

Observation 8d4d09cf-3d89-4d80-96f0-5dc03a22fba5 · outbound

This paper cites Mist: Medical image segmenta- tion transformer with convolutional attention mixing (cam) decoder.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Mist: Medical image segmenta- tion transformer with convolutional attention mixing (cam) decoder

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.020688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.687529Z digest=sha256:33f7f6a2488efe3247cbb50e420094144c8e7360d8da4e5d524d4f7e753370e6

Observation 757f607f-27dc-47d1-b85b-0de084d1c596 · outbound

This paper cites A multi-world approach to question answering about real-world scenes based on uncertain input.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks A multi-world approach to question answering about real-world scenes based on uncertain input

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.011374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.689912Z digest=sha256:93e15c704056364cdcdd0a488fdf1cf4b7b1cced4dc7dd12685d9fabed25bb6b

Observation 5da1ab1b-8c51-41a9-b746-be81af226e7f · outbound

This paper cites an unresolved cited work.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:44:09.002771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.692251Z digest=sha256:5a4ebb5468cc336a39e5831cbecd8121bb96c7fa4eb9c8754fa388e313d62428

Observation 3b0ec725-2eeb-40b2-bb12-70ed4eb13ed5 · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.994916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.694585Z digest=sha256:0e9acf181ea5ef9b91a97a40f21c91cb14348ff57349bbda1fad45313a5e87df

Observation 5d18766e-e981-475a-bb57-8f40c2e22232 · outbound

This paper cites Argos vision: Advanced computer vision solutions, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Argos vision: Advanced computer vision solutions, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.988205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.697182Z digest=sha256:f2be4a19898a1a30f33adfda26f2f949e28e3fe15a286250251814e99cde5384

Observation 702591f2-84be-46f0-8228-b9617433be06 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Can i trust your answer? visually grounded video question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.981283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.699688Z digest=sha256:cac50a987f650426a4d85d125fcbed39e9c7935aa241c20f445ef94cb1c76e35

Observation 61059adc-ffe0-4a6c-9a91-a63ae3a9cb7b · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.703083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.703083Z digest=sha256:d3fadcff1b42d73dbabd4cda08530c102f25afff8640102b69e54c59b91eefef

Observation 1dc40cf6-3779-4f84-adfb-f31fc1f4023c · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.706359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.706359Z digest=sha256:a154123ac81769962bdd1a69287557eb7451860d142366e4f00859b3ce8013d8

Observation 9366bec2-442a-4b46-9588-cd422ecda1d4 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.709264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.709264Z digest=sha256:6d4c4ef419fdd9c50af1bd0af8e26e299d2f9e7c17d7c142863b1bb226d37609

Observation 2562f93c-e1f9-4275-88af-c45b25505be9 · outbound

This paper cites Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.712087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.712087Z digest=sha256:bc07c020bcba6653d14247b38c5bf595251639178a68638f8c08e27d52a64b0b

Observation 534606d4-8008-4126-b81e-c515c2655036 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.714649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.714649Z digest=sha256:5e92b822587373228e7f802e318774428fa1ac0136de2cc22c69e6b71d785c05

Observation 5cc90ef3-4729-45c5-a76b-58a6d51ae024 · outbound

This paper cites Improved baselines with visual instruction tuning.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Improved baselines with visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.969301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.718434Z digest=sha256:b10b56add214a22aedda37cba5c995b2f9e70d2b3015813019a8d5c02fc68fa1

Observation bcfcc618-7819-43ac-b05d-3f39601e250d · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.720642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.720642Z digest=sha256:e4f0b38f54012938700348136c9826f85ab1b4bcf6fcd3a90394e652fb175e8e

Observation 09fad8df-814b-4bc4-a41d-6e3584624b82 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.722805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.722805Z digest=sha256:6683611a1114717f3d80ab34f555aea3e70dace71162f36d9e945c8371ff5a00

Observation 609fe016-13b9-49e0-91f5-1451828ffdfa · outbound

This paper cites InternLM2 Technical Report.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternLM2 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.725326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.725326Z digest=sha256:b6268ceb70337148a6207bf36562648cde3c1c9024b9bd26f31a6506d651ee53

Observation d87acb9a-73b4-425c-936a-2e63e780a536 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.728515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.728515Z digest=sha256:42d18a8b16f8d55adce6a525b6b9822b2aba54e4a3e6330d95040c4080d47c46

Observation 185f6007-f6c3-4765-aa97-855bd88e3305 · outbound

This paper cites Minesh Mathew.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Minesh Mathew

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.961214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.731416Z digest=sha256:05de73d15e4ec1f6ad51608335b6056cff8e8aab05370b362fd4f6820c81b754

Observation d2ac603b-7010-4854-ad84-de5d025caa2d · outbound

This paper cites an unresolved cited work.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:44:08.954331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.733741Z digest=sha256:ee0dd55d1ec4ff882a5b89cf5f78c5a27cccf6b23b886217dca641f95de4b368

Observation 2987ac11-a28d-4c06-8098-a0ae3a7b7454 · outbound

This paper cites Infographicvqa.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Infographicvqa

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.947975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.737210Z digest=sha256:933e1bd2104b14ae641706d747964eed666f5970754ccbc195cc0d30e08e2454

Observation 59ea8c61-fff6-4480-89f1-6123798168da · outbound

This paper cites Towards vqa models that can read.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Towards vqa models that can read

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.941345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.739917Z digest=sha256:c3dbe92f7c9959e8667f5cb1dfeef9acb16cda6d1eb2e484b17992765eb62564

Observation 3a21812c-f164-45cd-97ac-fd6c8ce900a6 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.742181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.742181Z digest=sha256:aef476cccfb16ef30d3b284a604e566ce53b8df1155b60de04a8e2d56f24a307

Observation e0bb42d2-7d75-4433-9fe4-2636430cfaee · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.746244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.746244Z digest=sha256:7bcfcc5287b294adf3239c60fc6cbd926d3f3c64091bbcf7ad8173c7c4fdf567

Observation 2f90c4b2-758f-4fed-ba1c-122d09534858 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Learning transferable visual models from natural language supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.932031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.749583Z digest=sha256:12f63597e3d27347b7e367d223886595c48dccc0943732db67523e9bbb5f434c

Observation 824b2fb0-b4b9-43eb-ae50-573ec8937a95 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Beats: Audio pre-training with acoustic tokenizers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.923785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.752199Z digest=sha256:11a3b11ffbaa168609f86cf1e6181bce5f6862932229561f6ae7f1a6f5e11e52

Observation 40d87ff2-f9ca-4180-a1f5-dfbe488898d5 · outbound

This paper cites an unresolved cited work.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.754826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.754826Z digest=sha256:579991617203682ef64f906f2c9f4488974175ab7b72a611626570cbd0250c87

Observation 3f8fb43b-7709-4f52-b0a4-a0b2a3109792 · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.757039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.757039Z digest=sha256:1a14792952e97b939c5b805ff6d434ca1073f9a6388b6b677b64c14a99aa727d

Observation fd311754-fd02-46e3-b8ed-c8f7b2ddd62a · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.759431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.759431Z digest=sha256:6266c437557894e176626fd690ef88279c19df670eb2246e2589bba16bcc9d4e

Observation f0de5971-23ab-4a76-a6ac-dc29b659efeb · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.911658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.762072Z digest=sha256:0a3c2e918efffe9381f0e0b512065f50a0a408166cbeaace4291aa1e54daf289

Observation bdeeb4e0-ac6c-41f9-a072-332ba319b5f3 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.765337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.765337Z digest=sha256:0e34dda71f209fd149f7532bff499f95a6d571b0892d3d7a9b298a14c9fc356c

Observation 45b27448-3032-4290-8177-014191844ef8 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.903229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.768099Z digest=sha256:9e0f3e31d0c29f97eb7eb13c48b9a528175fa1b97c74b322e563e3d326af2501

Observation f6484ece-c312-4ca4-8b39-8a7d9917b913 · outbound

This paper cites A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.770534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.770534Z digest=sha256:7c06806b158d1a547c358814b7f574eaa95e730544d3354f3fc727290dce4cff

Observation ac9974fc-3fc7-408f-802f-d469af7619dd · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.772811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.772811Z digest=sha256:629d20725028d2620d86cbbf6a706764bae9b182416fb63c4c07edbdd2c5ecbd

Observation 683195b3-4202-4a26-8627-2b7242cd4580 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Gonzalez, Ion Stoica, and Eric P

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.775565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.775565Z digest=sha256:0c0bfbe00a07015b55b5c2ecb8b450bee01d8f52fde4949f6a7e742f0a97730c

Observation 59587831-6f21-4d66-b0dd-7251b5939791 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.778585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.778585Z digest=sha256:6e05bf19d015271d372521251776dd7946cf833ac8794f91084761b8244f3534

Observation ce52ce43-89c2-4ffe-a81b-12acc45d3999 · outbound

This paper cites How many cars can you see?.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks How many cars can you see?

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.883904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:44:08.780881Z digest=sha256:9d027072401edd20b62bb0c2d49932739c1fa16b53d610391cc26589b480a514

Pith citing papers

Observation f316b1d1-7c9a-494e-96fa-35b68ef33342 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:15:52.858304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:fa5849419f15e960233c901c3ad293156671d8fde5d9ab7584a882ec92524ed3