Pith. sign in

Paper Citation Record · LEDGER

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks

As of 14 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2412.01132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01132 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:08.780881Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:16.204012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:15:52.849428Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 129e97fc-118e-48b7-a6ed-537ce680ce53 · outbound

This paper cites Video question answering: Datasets, algorithms and challenges, 2022.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video question answering: Datasets, algorithms and challenges, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.628106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.628106Z digest=sha256:a16b860d440b3c73f273f25e8db8d50839b94a2d6556922471e7a47b01d42491

Observation dfad524f-48b6-47a8-843c-d10e471f9518 · outbound

This paper cites Videoqa in the era of llms: An empirical study, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Videoqa in the era of llms: An empirical study, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.158989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.631391Z digest=sha256:f4fee1da8d757e6b5b050921cde8f4bd82a44ec06069cee66a86d7e4b4299410

Observation 44aedcde-13e9-4ce0-96ad-770869e13167 · outbound

This paper cites Francis, and Alessandro Oltramari.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Francis, and Alessandro Oltramari

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.152971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.634468Z digest=sha256:cc13d557ffecbdf183612b6e17018e92552dff07ef153731b17fbf8d33f6bc3e

Observation 7f603df1-b510-4b89-9090-6979884feb52 · outbound

This paper cites Carla: An open urban driving simulator, 2017.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Carla: An open urban driving simulator, 2017

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.638216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.638216Z digest=sha256:eabb670f8f73ea9c1282aab71f91df486128e56fd27325fde0172f54f6df9dec

Observation 4761f4e7-63f7-48f4-8915-06f870a3c798 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Lingoqa: Visual question answering for autonomous driving, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.142418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.640880Z digest=sha256:87d2144290b5e28de40faa23d8e53b1862395c5699f37f191b98b3ead42bf942

Observation 20c7a9e5-ca6e-4f0d-967c-a60cf68fbd61 · outbound

This paper cites Align and aggregate: Compositional reasoning with video alignment and answer aggregation for video question-answering, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Align and aggregate: Compositional reasoning with video alignment and answer aggregation for video question-answering, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.135003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.643534Z digest=sha256:ba953f1642ba7226312dd8586286bde69d3226f12d77312a0df2f37b15f2d7b3

Observation c9995d22-7b77-4980-8a48-530f2f7d47e3 · outbound

This paper cites GPT-4o: System Card.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks GPT-4o: System Card

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.126898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.646199Z digest=sha256:3a6d55633a991349046d35693ee69e2429bcb1bd0b2f9c3d49009c67b2dc6e18

Observation 872b3005-0b4b-4fe1-a58a-d34a9b1d47f3 · outbound

This paper cites Dynamic spatio-temporal graph reasoning for videoqa with self-supervised event recognition.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Dynamic spatio-temporal graph reasoning for videoqa with self-supervised event recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.119502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.649449Z digest=sha256:b342b5458454e5e1a4118626936ae97a5436c360dd00ce01873247e1ec48ca63

Observation 5248eca9-14de-4926-9522-8ba932fb945a · outbound

This paper cites Recognizing an action using its name: A knowledge-based approach.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Recognizing an action using its name: A knowledge-based approach

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.112135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.651945Z digest=sha256:8947f52ddc414b44db34c2ee105065cee7a86b41da57c6eb329aaf1c33e0ebd2

Observation 1f6253e7-a2ae-458e-8b2b-88901f668cc4 · outbound

This paper cites Video representation learning with deep neural networks.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video representation learning with deep neural networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.104141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.654520Z digest=sha256:484c319d8c6d35df8107e338dd16870b7ddd79b01d7dde9e82c6366d5ded458e

Observation f79878ee-26e8-49f2-a65b-bff1bcd36640 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video question answering via gradually refined attention over appearance and motion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.657336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.657336Z digest=sha256:66733218ed984368497c239bd1d80344344667f2de5f0872ded9fe94f35dcfde

Observation 288832a7-3821-4f8b-a103-91085ed142b5 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representa- tions for vision-and-language tasks.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Vilbert: Pretraining task-agnostic visiolinguistic representa- tions for vision-and-language tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.092676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.659970Z digest=sha256:8ef217ab8a6a9ebc52c23286f78f7ebcbc095f5d65c113eed1840c2b9b1c8529

Observation b4658756-8c62-4aa0-967b-6cc2596ce35a · outbound

This paper cites Compositional attention networks with two-stream fusion for video question answering.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Compositional attention networks with two-stream fusion for video question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.083934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.663000Z digest=sha256:71c2d603c49a7317033d802c01eca3bf37bb6f3c418882f829ee4eebba4ca223

Observation eff6c892-b8f5-4b73-bede-f3b5f70725ca · outbound

This paper cites Visualbert: A simple and performant baseline for vision and language, 2019.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Visualbert: A simple and performant baseline for vision and language, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.075921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.665457Z digest=sha256:28b33eb64d49b399fe856b5cbd71135e5b8c4acac43ff28323269bb7a637b298

Observation e93753ee-4fb1-42ea-af3b-7a21806dd383 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Activitynet: A large-scale video benchmark for human activity understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.068111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.668187Z digest=sha256:3c91725eec28393c8d388b26816aab01a60d9f4ad2bbc716b8ec9a1e965a8453

Observation c73926d5-4343-4d0b-af80-0c22f0799a30 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description, 2016.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Tgif: A new dataset and benchmark on animated gif description, 2016

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.061016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.671405Z digest=sha256:21b5aba01165b9e5e97fb04e55c73c1f0e1757b09b9497f7424191c108896d3e

Observation 5d00888a-1c89-44a0-af8f-a916f2ee03a5 · outbound

This paper cites Next-qa:next phase of question-answering to explaining temporal actions, 2021.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Next-qa:next phase of question-answering to explaining temporal actions, 2021

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.052959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.673233Z digest=sha256:18b6a4f4a8f451ce121e272b9328ec111b4fb25e423ba0a7c4f1ea1895e58d94

Observation 79cac09c-c190-41f0-b3dd-f544ee9bf5d8 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Gemini: A family of highly capable multimodal models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.045004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.675015Z digest=sha256:b91b1967a4611aaa635d4bff6bff286daceb8fb24b7cd3d4e2e3582cc27ac2a5

Observation 08f1446b-302c-487e-971f-60fdc007da69 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video instruction tuning with synthetic data, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.677145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.677145Z digest=sha256:fb69fa88a6edf3308212ac32d3c1d229ed90dc5909063d08bd8286db45c3d236

Observation c726ee96-410e-46ec-9cbb-1a3f447843dd · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video, 2022.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Ego4d: Around the world in 3,000 hours of egocentric video, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.035218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.679543Z digest=sha256:6e7e431f0f88de06fe75b306bd57748e2162d14aa03ef6941420ddb3fe2a98f6

Observation e9d99f4f-780b-4100-9884-961ae8c0217f · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.028613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.681380Z digest=sha256:ccfbabad262972bba717e9bc3e70adbd7a53edc26f591bec6ceeecff2840a2db

Observation 7573b07a-f401-4f63-8698-9066dda52739 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.683715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.683715Z digest=sha256:83b89599cbe9756c23ea954601089611cb3e61efb4235b80f4ba7c02bb3397b5

Observation 8d4d09cf-3d89-4d80-96f0-5dc03a22fba5 · outbound

This paper cites Mist: Medical image segmenta- tion transformer with convolutional attention mixing (cam) decoder.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Mist: Medical image segmenta- tion transformer with convolutional attention mixing (cam) decoder

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.020688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.687529Z digest=sha256:8e434b7ed72d9676ab39dcba16c259f28ac22c6e3d36afa0179dbf203abb920d

Observation 757f607f-27dc-47d1-b85b-0de084d1c596 · outbound

This paper cites A multi-world approach to question answering about real-world scenes based on uncertain input.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks A multi-world approach to question answering about real-world scenes based on uncertain input

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:09.011374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.689912Z digest=sha256:33f5fea77f05651e6ba9097aeff59a2eae072b07a91c5518f4131aa921f66da7

Observation 5da1ab1b-8c51-41a9-b746-be81af226e7f · outbound

This paper cites an unresolved cited work.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:44:09.002771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.692251Z digest=sha256:af2f0b57cb291a5c01faa4878c2ea1923e4ab8647de311ce6e798b135f94cbba

Observation 3b0ec725-2eeb-40b2-bb12-70ed4eb13ed5 · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.994916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.694585Z digest=sha256:cc65ecd2cae47921ec71da486f127f7179845171b2ab2dde07e9038a0f15f3e0

Observation 5d18766e-e981-475a-bb57-8f40c2e22232 · outbound

This paper cites Argos vision: Advanced computer vision solutions, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Argos vision: Advanced computer vision solutions, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.988205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.697182Z digest=sha256:055312bdbb43dd6899902cd2a571d5489fc487e0d2b0667811cb808484f8c8be

Observation 702591f2-84be-46f0-8228-b9617433be06 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Can i trust your answer? visually grounded video question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.981283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.699688Z digest=sha256:3ea9792e041ac583f8026c9e2e720b13b161050a32d9f387f321864186e7fd08

Observation 61059adc-ffe0-4a6c-9a91-a63ae3a9cb7b · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.703083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.703083Z digest=sha256:7d7e1271ad5ac83bfbcd226a5b243177088b294834cd95604dd3abe03c48b559

Observation 1dc40cf6-3779-4f84-adfb-f31fc1f4023c · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.706359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.706359Z digest=sha256:c866dfa4e68e25d74acff7d98930a6352aa5c8d96296c434969b8d1bf45402c8

Observation 9366bec2-442a-4b46-9588-cd422ecda1d4 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.709264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.709264Z digest=sha256:f4cce70d9122b2ec139b50d51e7df12a0b73f99e7beeb97e1472cd5e69662986

Observation 2562f93c-e1f9-4275-88af-c45b25505be9 · outbound

This paper cites Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.712087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.712087Z digest=sha256:1ff49149ce3a2983249e1d1d2639b9a1f3fc2966317f4b392c8b624ec4fed385

Observation 534606d4-8008-4126-b81e-c515c2655036 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.714649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.714649Z digest=sha256:c17bcbd98016b0e0562caa40bf437820944bcff959ea86f31fddb37c644e696b

Observation 5cc90ef3-4729-45c5-a76b-58a6d51ae024 · outbound

This paper cites Improved baselines with visual instruction tuning.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Improved baselines with visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.969301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.718434Z digest=sha256:ec36a3716629e4f5254aaaa1fe27c5ecd9badf0f7df63651c985424a4a62adab

Observation bcfcc618-7819-43ac-b05d-3f39601e250d · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.720642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.720642Z digest=sha256:2b8a39c52a69ab01905a33cdd0360ef66d71620f4c4edd818dfdb678f1e67260

Observation 09fad8df-814b-4bc4-a41d-6e3584624b82 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.722805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.722805Z digest=sha256:af8239e7ad0aa43b0c14e4d0cbdc2142f3494e7f9474285ee427965d7552fc7c

Observation 609fe016-13b9-49e0-91f5-1451828ffdfa · outbound

This paper cites InternLM2 Technical Report.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternLM2 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.725326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.725326Z digest=sha256:1cd7ace9041d00b4063bcd4399ff3f43f592444463d46ab55654e3b8c14070f6

Observation d87acb9a-73b4-425c-936a-2e63e780a536 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.728515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.728515Z digest=sha256:a80aa50ca628b76d511d3280c645bb8549e81b18ea552de124a8bb4b6c32f945

Observation 185f6007-f6c3-4765-aa97-855bd88e3305 · outbound

This paper cites Minesh Mathew.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Minesh Mathew

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.961214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.731416Z digest=sha256:e95e5a43a8f332d75d283567d3ff21b1559c3f324a5e7bbe586a44488be03ae4

Observation d2ac603b-7010-4854-ad84-de5d025caa2d · outbound

This paper cites an unresolved cited work.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:44:08.954331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.733741Z digest=sha256:b14413623d118c65451d20df7f3348a1543d61375b9110070ce217947d7c05ba

Observation 2987ac11-a28d-4c06-8098-a0ae3a7b7454 · outbound

This paper cites Infographicvqa.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Infographicvqa

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.947975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.737210Z digest=sha256:fad5cc3f34ef649f42650d0d427e2dcf930aefa921fc008ab1efacb83b72e779

Observation 59ea8c61-fff6-4480-89f1-6123798168da · outbound

This paper cites Towards vqa models that can read.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Towards vqa models that can read

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.941345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.739917Z digest=sha256:a2675b329076ec76ebdd364fd6c5ef972e536bf9cae73781467b542fa868e70a

Observation 3a21812c-f164-45cd-97ac-fd6c8ce900a6 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.742181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.742181Z digest=sha256:3d27d976dd0f12384a8698ecdb5fc0ccda837986ba77a66bdfc52f7aac508378

Observation e0bb42d2-7d75-4433-9fe4-2636430cfaee · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.746244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.746244Z digest=sha256:9be4254ace780e127d0c85b7c1e0f52982e9e019bf764bd6db26b407410f7fa4

Observation 2f90c4b2-758f-4fed-ba1c-122d09534858 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Learning transferable visual models from natural language supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.932031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.749583Z digest=sha256:9756e10ca13b14ddd2b713ab1fd5764a0fd3262fdb8683fb7cd6ef380a0c03df

Observation 824b2fb0-b4b9-43eb-ae50-573ec8937a95 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Beats: Audio pre-training with acoustic tokenizers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.923785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.752199Z digest=sha256:7df46635ddc87fdf07ceae21c795d651bc803197f7bf3b5e2bbfb6f9e4d5eedf

Observation 40d87ff2-f9ca-4180-a1f5-dfbe488898d5 · outbound

This paper cites an unresolved cited work.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.754826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.754826Z digest=sha256:6d041cdeed587a4b2dcdbebdb8f7cd4a42ccca7930534c680a4be19b25563a85

Observation 3f8fb43b-7709-4f52-b0a4-a0b2a3109792 · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.757039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.757039Z digest=sha256:8772edf53c0e5b20b4faaa9e50f989bd4fd6d510eada29ef6d90ac364fc1c70a

Observation fd311754-fd02-46e3-b8ed-c8f7b2ddd62a · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.759431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.759431Z digest=sha256:c52534a1ab7e98f1ae809f1136c069629dc6a9eda2d44b9f56008eaa6a78832e

Observation f0de5971-23ab-4a76-a6ac-dc29b659efeb · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.911658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.762072Z digest=sha256:96ae7be70a0b450cb7ee8eb6c87d643dabda4820e2f5731a8a0dcac8c6990494

Observation bdeeb4e0-ac6c-41f9-a072-332ba319b5f3 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.765337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.765337Z digest=sha256:32794929484f663c3f81959f9c1515bc23db922bf7d8d9c26d4813423903ccd2

Observation 45b27448-3032-4290-8177-014191844ef8 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.903229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.768099Z digest=sha256:376ad13ffa88409f995086bbcb517669bbc3848a4f17d8aae29142a7a0d83110

Observation f6484ece-c312-4ca4-8b39-8a7d9917b913 · outbound

This paper cites A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.770534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.770534Z digest=sha256:4216401aa4c7481e1bbd1524bd526d1d5d67e7b44923a6f5071b57d31976953a

Observation ac9974fc-3fc7-408f-802f-d469af7619dd · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.772811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.772811Z digest=sha256:6048a4712cbd530a85b88817b2fd7df93a26585e8f6a1a6016a3a5d9eeb2e5a0

Observation 683195b3-4202-4a26-8627-2b7242cd4580 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Gonzalez, Ion Stoica, and Eric P

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.775565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.775565Z digest=sha256:4ab69ea7c022999b591e6b2560d4006803b4a2f0b4a58ea5ac50ca15dec91b9e

Observation 59587831-6f21-4d66-b0dd-7251b5939791 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.778585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.778585Z digest=sha256:9774fbb29f2c8be3d25ff797c5f1a59692df446e9d2b1bbd9ac3f02ba0487286

Observation ce52ce43-89c2-4ffe-a81b-12acc45d3999 · outbound

This paper cites How many cars can you see?.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks How many cars can you see?

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:08.883904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:44:08.780881Z digest=sha256:5a2e3dffeb96bbae4e2714d6d924a936d2fadaa74701150a22f4449dffee9a43

Pith citing papers

Observation f316b1d1-7c9a-494e-96fa-35b68ef33342 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:15:52.858304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:aac0a6ad9c1dc93d3f9b48c1e27b1ad27400429b59a2287d7bc61263003a8d51