Pith. sign in

Paper Citation Record · LEDGER

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models

As of 18 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2505.08455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08455 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:33.853704Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved52
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b89677d-3b55-4f0b-b0fa-46f60380ab2e · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Robovqa: Multimodal long-horizon reasoning for robotics

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.239931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.444017Z digest=sha256:80b467c7b9cbfc0ec791a7f9bf668f7572af1237a50c3ae2d7c40ea01c658dd9

Observation c878ab8e-0a4b-4286-9bb6-0c705479fbd3 · outbound

This paper cites MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.450018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.450018Z digest=sha256:760f99df727ebb89d207b0208db7e1bc5237baa4d16493b1cdb2a0641e98958c

Observation 236615c5-5e79-4ada-ba59-d2b445ddf7cc · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.455775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.455775Z digest=sha256:4da69f15f1e63168d3c8991a9deda659d2aa5f240025015d431f035fcca8a62a

Observation 144edce9-8fb6-4610-bc43-0d037e06a919 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Openeqa: Embodied question answering in the era of foundation models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.224949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.461137Z digest=sha256:f6b14ed8e0abf62812fd459a753a480fabbcfc8139e8668322cdd11de42ee0ab

Observation d80e3b03-69f1-440e-8910-a0d0c85dc28a · outbound

This paper cites Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.466110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.466110Z digest=sha256:c9ac5425e712da8e8a12dffdadb5ed635d4760e4d9d7576a1765bc5d7da821c5

Observation 6f59abb2-0476-4a6d-a816-b4485f291313 · outbound

This paper cites A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.471828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.471828Z digest=sha256:38bcee95cabf220f14a808908bacd492abdce1d9542ddd5feafff67a7c3f8656

Observation 45022173-519c-455a-bc87-57177bbf8318 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.208399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.477225Z digest=sha256:a02e2a2e136a9fec377d549c69d494a07f1a0a22ad77b9edb772bc3fec0e5564

Observation 6cb698aa-a7bd-4323-8b10-e43c707f888d · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.483140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.483140Z digest=sha256:0d5362aac968f793ccc5a1d1529787b92c549e5d59ecf6853b2079bca8f69405

Observation 46c22944-33a7-4cec-936e-4a58d7efeb44 · outbound

This paper cites A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.489753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.489753Z digest=sha256:860a31e8663a21097ed231c50c4bed1839c0f66a797ff2595f970dfa780c6fb4

Observation 81b998e8-84cd-4ebd-82ad-3ab02563bc02 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.495030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.495030Z digest=sha256:644fb3acc17a524cd448c43afba24b89a2800701b5c028661967313f3833374a

Observation e50418e7-bf9e-4d51-86fd-1264daa20d9b · outbound

This paper cites Learning transferable visual models from natural language supervision.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Learning transferable visual models from natural language supervision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.193579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.500010Z digest=sha256:05dc72d9f141af69f29ebf349a8358a710c7349fe5aa2529a52b2c42a06b3db9

Observation cdff0e5f-ee66-4d28-a96b-3f5716613f57 · outbound

This paper cites Sigmoid loss for language image pre-training.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Sigmoid loss for language image pre-training

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.178154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.504583Z digest=sha256:7dfbbb134aabc4e76f2882228a20fe22a44cb4a390cdb20624f4fd0234f115d5

Observation 82f532e7-71c2-45e0-b07e-e60c41ae0a9c · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.509378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.509378Z digest=sha256:5602eacb991844cd0a3ff310dc534e4cd42512fc6db0d6d0d40558e9fd7e5c12

Observation 96598b34-9e8d-4787-afd8-98e1cfe4e53e · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.514987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.514987Z digest=sha256:cafa7f3c14c06a52797f1b81ac09c0a40a0a71bbb2f6820307179630ca95d441

Observation 7270f04b-5650-4b02-bbef-ae2a4a921ff8 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.520336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.520336Z digest=sha256:a17a849e22c8b53bf04249006650a454258d937546e2391ff27ace29dc33a56c

Observation 605c03e1-405b-4afc-b7ee-1ac3815bde4b · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.162399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.525739Z digest=sha256:9361ba327292630e60ed036b766e6bb44cddfe81c18ada832201d131e86670d5

Observation a418e650-accc-4d06-8ae2-41cfbd230895 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.532560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.532560Z digest=sha256:178ecf48dbe7ce1e786617921464d3d87f5a385bb336c6e6d2cdac9c83470618

Observation ae912172-be3e-4f6d-ac41-6e33bfe252ab · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.537472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.537472Z digest=sha256:0242eacfbb063da6af294a664f051fcf99c375f76fff54b7ab3ef129335fd261

Observation 50caf98d-008e-4c68-b7f7-a27bc88dd2a9 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.543100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.543100Z digest=sha256:f70a780464e587c9af5d668e7d1735f2b4608dbb72c4810b602490891d07472f

Observation 9cf4d981-c735-4225-bb99-f5f4764d190b · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.147759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.547817Z digest=sha256:746ef2a364f81be12cb032bb211bb82e5e1698eb9d4c002f0d1d7717e551b8f3

Observation bb79dd26-bed0-4924-b083-bc95c8cfb7b3 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.552216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.552216Z digest=sha256:fc5ecf770c1040ae108384f41b3a456d8b0d3b56ffc945ea133ac79bc4402e57

Observation bfdc15e9-adde-4fe9-a13d-624b11eb2769 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.557150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.557150Z digest=sha256:ae3537c915186f236c647e6d9e818c0c6e4a0b83bd84675d3063f221d665cc07

Observation e5439363-9c59-44d3-a269-b8b7db28832c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.562745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.562745Z digest=sha256:91ee971f44009f9ef2d06ead6c7b8e91b49413dc6000dfe5a40b94df1c865626

Observation 578520e5-b3bb-4afe-9cc2-8bb64ce22f25 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.567661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.567661Z digest=sha256:cf50c609a087d447491f1a1b299e5303da3e7d032d2efc39c6676480b296516d

Observation 43c1aa7a-e644-415c-9837-9a358b3c61b3 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.132123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.572655Z digest=sha256:35043350189371da3699b5d5cb6bc070144e04159df353b4a31f0a06c7982cd4

Observation 1ebae743-b61f-48fa-b400-fbdbd49df6e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.577918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.577918Z digest=sha256:083394b512fd68ed1f55e5fc1634cd1ce7416a3b9ae0aff2b4dc9b322f190a1e

Observation abc00dbc-3b25-4895-8253-99514193ed0c · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen-vl: A versatile vision-language model for understanding, localization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.116420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.582957Z digest=sha256:c67d8ee1da494f9a5104ab7768a57057a45f14284a449bc42adc7f5973a911ba

Observation ed7d0b5a-a7ca-4ac5-b33c-f162c980f986 · outbound

This paper cites Qwen2.5-VL Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen2.5-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.587609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.587609Z digest=sha256:47d6dc3469ac76046d1c3eda6f60250a6650878c9bf6ad8f362afbcef062a76f

Observation 8d2f1091-ee92-4f41-b884-c575b4d27450 · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Intentqa: Context-aware video intent reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.084753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.597957Z digest=sha256:1f3a7d3e03781fa13506f12fbf21a7e881c9f8ad1098d3436a42609ff33ca435

Observation 852a0e58-1b22-4e8e-903b-0ef34fe7db90 · outbound

This paper cites Rex- time: A benchmark suite for reasoning-across-time in videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Rex- time: A benchmark suite for reasoning-across-time in videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.069675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.602270Z digest=sha256:2122bb7c8c823b922797174ca88215a8087434c205810b8648a5859f2f7e0f80

Observation 95eb2e73-826e-45ec-ba3f-6e5aa90e6e69 · outbound

This paper cites From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.052199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.607218Z digest=sha256:a0f178f203a1be4ba03f9de9247470ab1f35390d77175d8708b8cd8efe572b89

Observation a9ef2c91-79b1-4a9a-ab86-7279b66c08b5 · outbound

This paper cites Long Context Transfer from Language to Vision.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Long Context Transfer from Language to Vision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.611835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.611835Z digest=sha256:cf10c6c8e845436305e01c6fb244612b0e0580ba7e65058cd15f4632eaa11dbe

Observation 906e0ef9-b0e3-40b6-a7d0-d3b554e4d1d6 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.036159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.617442Z digest=sha256:9e8da506197321da7855ce7f9296b47518eefb7dedb197b18547dff0d818ec6e

Observation c725de8b-4d3c-42b8-bbb1-e77eda694019 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.622058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.622058Z digest=sha256:3c4b68ce86a243a5f3e64f6e0dc24d214c63725761d127d15bbee833ddd7350e

Observation ec3dd096-487a-4ce7-8741-b02ab8275c7c · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.626864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.626864Z digest=sha256:0470d9660590483a2b2df18d790a18238eccdd0e4c5e6501057485cd53b6249e

Observation 01a1d4b7-fae9-465e-a4fd-a4c44af26b47 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.631527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.631527Z digest=sha256:2c852fc4758fe3c257bf456ded3050f7134c3239691d51f0e181018652c6e053

Observation 4e1f5f85-d232-4aab-8ea3-c116692f5b5e · outbound

This paper cites Video understanding with large language models: A survey.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video understanding with large language models: A survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.637682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.637682Z digest=sha256:808ea2845764343df02e032f4a46197ee3797399d0ab3b0779c0172cd455356a

Observation 920f45ec-7d6e-4994-a205-3e17288a260c · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Foundation Models for Video Understanding: A Survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.642396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.642396Z digest=sha256:7ee534355d490c2fe63173b8f03d1b1592c8542d118592a5bda7b6227a71efee

Observation 6ae3780d-0c83-40ce-a81c-fa06d730fe3f · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.647351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.647351Z digest=sha256:e0b4bbf98248275c2c73fe77711bca93cfac1421379c739c3a6ca387ca9d01e2

Observation eca70cb1-cdc6-4f87-93d4-b83935edb694 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video question answering via gradually refined attention over appearance and motion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.008796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.653452Z digest=sha256:3b66320697b776452b0dd127d790619cdf3dd78c81bc06c930983c7ce2a257b7

Observation 2a023add-8aef-4e6e-8b2f-07d2cf103131 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Next-qa: Next phase of question- answering to explaining temporal actions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.990702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.658131Z digest=sha256:210a87e96ead10ef47d5f359a203cffe5d0f1d3172b3ab273e2ee6141edcefd4

Observation acea9151-c1c2-4783-9c9f-1d659178ccc0 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.975540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.663554Z digest=sha256:7d8e68ed44ea58000a162dd3bb56dabd56d2cefc139e9177f0c75a497893446a

Observation e75fe3f7-da28-4ef7-8baf-7a88741e82dc · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.668469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.668469Z digest=sha256:a9bd13512cb954edc238a3b633b2c6263ecb0693b9fe6a5bfcc67fe0e50b6ac0

Observation e39db794-266f-4f71-b8c6-afea1340d505 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.674311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.674311Z digest=sha256:a90f55c49da965e82d21a506a3a00e170bec5c712cc3a82053d42fc9b12f69a3

Observation 0efe9d6d-0a96-4d3d-b9df-09d0c1139a53 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models TempCompass: Do Video LLMs Really Understand Videos?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.679506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.679506Z digest=sha256:cf03f24f587652e0df501b4b687975383e270980857b270b9f8f6994ef15025c

Observation dd700fdb-2a23-43c3-9a61-04b000dd0e50 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.684509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.684509Z digest=sha256:e90e963dbf0ab7d2ac88ccb28110d85169f04ccd6532da332c5eec6ff9c68d54

Observation 861cd7f6-85bf-478e-bea8-bf3dc51f9482 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.689913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.689913Z digest=sha256:5a529ce2055268ebe711ff0417d35a22a861733c8c26ea3bad79b984bf604c93

Observation bb3b54af-c5be-4614-b2b4-575ef17ec51b · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Longvideobench: A benchmark for long- context interleaved video-language understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.959836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.696191Z digest=sha256:c30358f07fc4b1313551ee4a9ccced67978b97da778e635a620b4054294402a5

Observation 4af261fa-79dd-48c9-a700-bc67770e9b37 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.944426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.701111Z digest=sha256:e0a707a9153653612068371311e9abb6957b72554a0f98baf395773172b6bf77

Observation d0eb0d0b-aa1a-4a6d-83de-645e6b7d1a3a · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.705892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.705892Z digest=sha256:66034edc0fd0f5fbc6ebba6d868e07c22472ce4b4cf2c3c5e18a6f0b14b4cadb

Observation 3d5e342c-f4bb-4d8b-a56c-51bb49b5abcd · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:59:33.711050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.711050Z digest=sha256:e6644b9a536c7a2bf4916c323ba910b27811b1e39ecc973ccca870b61ecbf59e

Observation b934793b-b950-48ce-93e9-5c77b1a3a9c0 · outbound

This paper cites Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.929308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.716606Z digest=sha256:4d677b3af4c28b3695faa199471f7a00725633497459375b3e0a5aabbbe7034a

Observation fa4837b5-47b0-42d8-8d65-e4e6cd901d7a · outbound

This paper cites MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.721333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.721333Z digest=sha256:507f4ebf66c6b9d1ab03b8d7132d1865873a8048483a2fd1ff7ac9a6c935e74b

Observation d78f15c0-b7e3-4db3-8dc0-380def64992e · outbound

This paper cites ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.726301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.726301Z digest=sha256:6f936b31a4d7c9b2f511226e1ea9a02d4d2a8c7c4dfd701749cc14654bc2d33a

Observation ac8e6e87-15f4-430c-9521-2b29d9083cbf · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.731448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.731448Z digest=sha256:88f549547e86548c9d3f28fe0f5589822fc546fc6601f9662641470be9b36d28

Observation 8580e34b-d30c-4b28-b1e6-c0848ed85041 · outbound

This paper cites Deep reinforcement learning from human preferences.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Deep reinforcement learning from human preferences

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.904330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.736464Z digest=sha256:3aa7ce0979cdb253c1b39d8e6d91998878aaee434857f689916a3f6df273af50

Observation 89df4697-23fa-4319-9127-b731d1cc5aac · outbound

This paper cites Proximal Policy Optimization Algorithms.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Proximal Policy Optimization Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.741587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.741587Z digest=sha256:c44851ce94d5a8ee4fc11023941df2a72712b043ebd8d93bfab10be8e3296c65

Observation 97814645-f452-4faa-8251-db811fa094db · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.748580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.748580Z digest=sha256:914b1719d21522418018da40c7b60f09e61c647bf9b4205f6ec0b1022bcfc25f

Observation 24680908-ecc4-4b80-bf19-e826a679e737 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.753498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.753498Z digest=sha256:dc19f27212377e45a5075d3d33fee2baf896e6140edceb84e944821a35a92a20

Observation 9e1edf25-7ef5-4916-a861-67fdbd79e8fd · outbound

This paper cites Learning to summarize with human feedback.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Learning to summarize with human feedback

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.889209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.758432Z digest=sha256:a2d712c65b0731a1dd741769e557eea5a118cf1b1480523bab91cec1f41ef6ac

Observation a8c36977-6c37-45b7-9b12-17e04a61dd27 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models WebGPT: Browser-assisted question-answering with human feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.763413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.763413Z digest=sha256:0ecd28a7581a7f59c960f09ce798e16da08453b9cae169556ef9aaad5d2d9fe6

Observation 95e00432-1208-4893-87e9-0d48cb358aa9 · outbound

This paper cites Decomposed Prompting: A Modular Approach for Solving Complex Tasks.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Decomposed Prompting: A Modular Approach for Solving Complex Tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.768463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.768463Z digest=sha256:2a0766b84b74753d47469b688b7179dfb1d1f17bb2494e3f60c4a0d69f2cfe81

Observation 7501a49a-7ab1-4c2a-b32f-8c0ba6a17baf · outbound

This paper cites Cross-task weakly supervised learning from instructional videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Cross-task weakly supervised learning from instructional videos

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.774187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.774187Z digest=sha256:3d54e55e38972fa9e43d1863db41c298f102dce01f6d335a56aeedfa790d948e

Observation bfcfd53b-1f87-4dc8-b886-63934c8076a9 · outbound

This paper cites GPT-4 Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models GPT-4 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.779074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.779074Z digest=sha256:a5abcd24ba5eecdc2ae837a817af931ab6094aba6d4344a78b71648c3b71ab91

Observation 1d7f9d09-8164-4e31-b0ff-cf7cb4912456 · outbound

This paper cites Procedure planning in instructional videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Procedure planning in instructional videos

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.862986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.783601Z digest=sha256:9df67bc698aad2da44191828e74e7c038ab314d801372d0d1f2171a93bdc352e

Observation 45816025-77a8-4996-8343-f5f185a6bce3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.788339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.788339Z digest=sha256:d940bc2c0bb2dedefb85a0eb505f378ce7a69f5d01373cc1b3ba53a482f8957d

Observation 6b696177-b007-40e3-a160-eb62669a7c28 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.793129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.793129Z digest=sha256:840cb0bf5e8e57f51483b5b91bb63af9c16b4d9b5e859428f15e069e709b2bdd

Observation ca8f6bae-b1e5-4f52-8979-691d0a85fb06 · outbound

This paper cites The Llama 3 Herd of Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models The Llama 3 Herd of Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.797702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.797702Z digest=sha256:bc93c6b666331f085c116468a06ad33416578158d66d8846a139456de721b079

Observation 577db0fc-0d97-4733-8a58-afaccc14980c · outbound

This paper cites Identifying and mitigating vulnerabilities in llm-integrated applications.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Identifying and mitigating vulnerabilities in llm-integrated applications

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.847612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.803407Z digest=sha256:83cc24e7ecbe2ba1464cb22839acca3fce6e5164e489d589e2ec0f5e50bc7fd7

Observation cf3fe687-b0a0-444a-87ea-e87f80ecc16e · outbound

This paper cites Qwen Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.808139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.808139Z digest=sha256:de0d2667781a5fae44841a31d8c7e7cfba987126fea218900da6d3ebb34a1175

Observation 047be0b7-5194-461e-9856-6be2066e33e2 · outbound

This paper cites Qwen2.5 Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen2.5 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.812991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.812991Z digest=sha256:23952d6f85769857e9e40afb944d19f8e40fb29a726d1222e450507df0053e32

Observation de0015e8-0aeb-4867-97b3-f8b93ab7d2bc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.832119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.817963Z digest=sha256:4a33c6e0b53490b78fa0e8a0d341f75c587725ab7e7a0b2d7106c7c15371272f

Observation 5ef27645-4dab-43fc-85a8-edf929a3dbc1 · outbound

This paper cites Improved baselines with visual instruction tuning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Improved baselines with visual instruction tuning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.815963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.822529Z digest=sha256:b34e7a6bbda4b3972365137a6089bd7df55dd4b50cb30b7f626d7bed736331c0

Observation faf3c1a1-325b-4625-ae12-0e58912ab2c3 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.827196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.827196Z digest=sha256:8e84af968af035501efb145f2ccc8b84b4dfd72d1a86b9239d353bf57a98944e

Observation 53afbc61-2f7a-425b-a164-06134bde3f79 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Unmasked teacher: Towards training-efficient video foundation models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.799739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.833255Z digest=sha256:842f29cc0ef0c505dcfd783e0d36e5fe46f452fe353f4fa5e3c2d88f290a6825

Observation 443f75ea-042d-40ab-bbbf-82c5d5f71f07 · outbound

This paper cites Self-alignment of large video language models with refined regularized preference optimization.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Self-alignment of large video language models with refined regularized preference optimization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.838263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.838263Z digest=sha256:6e6a7b0553a9c7fba7d93651ac7c53a8464250fd6725d319a12280737a588ff1

Observation e95008de-bb64-4ce2-8fda-886c9f83364d · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models NVILA: Efficient Frontier Visual Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.842742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.842742Z digest=sha256:31f261a09de9110c4e3a95f9f9baf79fc5eede9e6096d814dbd034609bd4f25e

Observation 7df600f0-39b1-42a5-8f74-b0b2ab60c761 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.848016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.848016Z digest=sha256:e21d3c0aa3fbe30bf97fc15103ede97f61cb4b87c5d5263491249a813a9e3359

Observation c2d2c112-4f3f-477f-9dd0-beaee1a5d3d4 · outbound

This paper cites GPT-4o System Card.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models GPT-4o System Card

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.853704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.853704Z digest=sha256:071770b2e8be019220a29eefadb185034502923296ff307682e3b7b54d5a583e

Observation f3caf5fc-588c-41b9-9866-da8cedf6f0d6 · outbound

This paper cites an unresolved cited work.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T21:59:35.100239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:59:33.592973Z digest=sha256:c9d5d035151e76472ad0f35eda13edcf860f7293241fb125e73847bce6772aaa

Pith citing papers

No inbound Pith citation observations are available.