Pith. sign in

Paper Citation Record · LEDGER

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2508.13470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13470 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:06:25.327098Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.379714Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1ccbd97-d541-4f94-8c02-1e42ddfd6af8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:22.867967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:22.867967Z digest=sha256:854de3123237feb5119d97b86c73548b13cfe1eb9d6e9478b17a3e6d45b3059a

Observation 273c710f-7f77-4bba-87e7-804e49fe16e2 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Qwen2.5-vl technical report, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.641597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:22.954759Z digest=sha256:2acb4c1ccc4e1a0ac158717257b176e308a60e59a0c4585f3147475eb7e16642

Observation 2ee33d05-da6e-4be0-89fc-4f0f0d3a715c · outbound

This paper cites Maplm: A real-world large-scale vision-language benchmark for map and traffic scene un- derstanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Maplm: A real-world large-scale vision-language benchmark for map and traffic scene un- derstanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.503846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.004322Z digest=sha256:67833e43ad597a64d387d11144a6a2f4fb5507578769baf2c2095bafdb622a8e

Observation afc6916c-e97f-43e6-b66b-f4a376c76f49 · outbound

This paper cites Cityllava: Efficient fine- tuning for vlms in city scenario.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cityllava: Efficient fine- tuning for vlms in city scenario

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.338854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.047424Z digest=sha256:7a504f6c99b1cf80466f637078889593827f9d70be268ad7952f786b2bef7e58

Observation 97a332a3-6fda-4503-ab1c-039cbffc7f9f · outbound

This paper cites Cityllava: Efficient fine-tuning for vlms in city scenario.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cityllava: Efficient fine-tuning for vlms in city scenario

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.215166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.121343Z digest=sha256:aeedf60e7afeada2dbdb3e9ebb61757c80c0ce58856b7280057a7c2d65f6a776

Observation 683612bf-d7a1-485c-b5c1-8417c7c23426 · outbound

This paper cites Optimal gradient checkpoint search for arbitrary computation graphs.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Optimal gradient checkpoint search for arbitrary computation graphs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:29.085557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.188507Z digest=sha256:30bb8ee8a721a39d57acfcc70d6164176c324edde2d3086992905e0cfa87aa45

Observation 73c536d8-c260-472d-896d-4d0875328523 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Lora: Low-rank adaptation of large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.871406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.240284Z digest=sha256:a6f44274c3e66f10ee65434b49654f8aee9913c5cceddea03ad331d2d28220a5

Observation cc28c621-9006-4e38-b4a2-a4e01cc313bb · outbound

This paper cites Better zero-shot reasoning with role-play prompting, 2024.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Better zero-shot reasoning with role-play prompting, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.728846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.340402Z digest=sha256:c0ed5412b241aa03ee9776026ae77fade8aa3eb82d01085bb6d87019e7d46b17

Observation 18c180c3-c576-499b-b6df-f33e92a77418 · outbound

This paper cites Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understand- ing.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understand- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.580566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.441211Z digest=sha256:1624a506f16a63ffa2c052d1d6db34e3cf9ea718147c4995e1285368d32d6d0c

Observation 323ef9e9-47cc-41dc-8691-b32fc052179e · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.616871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.616871Z digest=sha256:cda5e77def86984dbcc75ce0e62b921bafd9f853c9d8c0f33f78fc110c7ee89e

Observation 4b115999-15e5-4575-8c98-d5ceb4a12853 · outbound

This paper cites SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.701939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.701939Z digest=sha256:1ca400fd280c89b197d639a2a1fd13d60e4302a6c9352df6ca06a780b9115a62

Observation 56de8027-29a1-481e-8680-aa6ead3cb860 · outbound

This paper cites Improved baselines with visual instruction tuning.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.256587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.781156Z digest=sha256:c761d0b5dcb62273fb963bc352abfd24ebb4590ef383ea2d058d9d7d5ab7bdc2

Observation 7f6b4a87-5206-469f-9fef-79ed1d70dd9d · outbound

This paper cites Improving generalization in visual reasoning via self-ensemble, 2024.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Improving generalization in visual reasoning via self-ensemble, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.857493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.857493Z digest=sha256:ff7aa39d4d8ebad3c351881fa633c84343e5f72b848a6c5810fe82bc8d5d270d

Observation 6fd23e2d-cd9c-40d0-9a4d-56cb10567450 · outbound

This paper cites Hybrid, unified and itera- tive: A novel framework for text-based person anomaly re- trieval.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Hybrid, unified and itera- tive: A novel framework for text-based person anomaly re- trieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.124916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.920763Z digest=sha256:5a21310204b3ed84bc7435f95eda9a6f6b54704adfafec77f1536adf582e2e3d

Observation a6764805-6aa6-4800-b0fd-0c01c7d9b09c · outbound

This paper cites Le, and Quang-Vinh Dinh.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Le, and Quang-Vinh Dinh

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.921115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.979582Z digest=sha256:7b92f8c74b2def58c57da67b84c1a1db2682ddd0a9470f4ba55417da9fb51d0d

Observation 87a592ff-d671-481f-8425-e07c624a32d4 · outbound

This paper cites Gpt-4v(ision).

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Gpt-4v(ision)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.797475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.016832Z digest=sha256:af52662a953576c3df18b9df65848963a3ff8662b61b752157bbe0b9e38021ae

Observation b69b4d08-62d1-41bc-95b4-b66e6049efc0 · outbound

This paper cites Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from so- cial video narratives.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from so- cial video narratives

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.678940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.075072Z digest=sha256:71cf54e6a026198ebe04120471809b2852ad58fae39c74182a7dee2096074126

Observation 5b6bc513-0163-4956-85fc-894ebf82f210 · outbound

This paper cites 1, 2, 3, 5.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models 1, 2, 3, 5

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:28.401999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:23.527136Z digest=sha256:078590dbc89f78d471d92a6fa82402f00a03e35c4f13ff5aa86ccc15cbe81169

Observation 6f9b1c7d-cf42-4e37-a67e-0d548291fcac · outbound

This paper cites Safeplug: Empow- ering multimodal llms with pixel-level insight and temporal grounding for traffic accident understanding, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Safeplug: Empow- ering multimodal llms with pixel-level insight and temporal grounding for traffic accident understanding, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.509429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.137473Z digest=sha256:09381b4ca88a078664565f1ad8e85bdea2495e62ac42ad9a71319ea36f062e90

Observation 6184360f-bbac-4af5-ab0c-a8d93fdd0657 · outbound

This paper cites Scvlm: Enhancing vision-language model for safety-critical event understanding, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Scvlm: Enhancing vision-language model for safety-critical event understanding, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.365440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.221004Z digest=sha256:b655b578f9c565fc55d929072fd19190ff3e1a925caefe83ad94a3bf4c089c63

Observation 84a8b82a-efdc-4e8e-b6b4-d69547a4cf23 · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:27.181553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.286445Z digest=sha256:7014ab084935f2e9b0da887930df31c970b9129bcd27dccea3b2aea7e54bed75

Observation 7fed3d03-dbc0-40ed-b51f-eec4ec332711 · outbound

This paper cites Le, and Quang- Vinh Dinh.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Le, and Quang- Vinh Dinh

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.981603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.346433Z digest=sha256:6e74144309b619be96ebbf62087ba2bbca09d5a057af3a7d0be6577b462fffd7

Observation 8977a4f4-d9bc-4267-9041-e6ab6ab98974 · outbound

This paper cites Accidentgpt: A v2x environmental perception multi-modal large model for acci- dent analysis and prevention.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Accidentgpt: A v2x environmental perception multi-modal large model for acci- dent analysis and prevention

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.832408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.381382Z digest=sha256:edc8ef861e4203eba148c9e52f0b0b67396ef90ec6b613b425fc355808814686

Observation 0cf37018-ff54-4d37-8b33-bd3c090bbe75 · outbound

This paper cites Anastasiu, Zheng Tang, Ming- Ching Chang, Yue Yao, Liang Zheng, Mohammed Shaiqur Rahman, Meenakshi S.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Anastasiu, Zheng Tang, Ming- Ching Chang, Yue Yao, Liang Zheng, Mohammed Shaiqur Rahman, Meenakshi S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.668464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.470008Z digest=sha256:81f0e6e0aeb48f571773c7a29cc46a2aaf62b866b579a1aacf9564f058b15176

Observation 23da6946-b704-42ff-8433-1fadd6fe12e4 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Cogvlm: Visual expert for pretrained language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.496382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.572649Z digest=sha256:ac7440d86b8b6b917ecb91e0af30f41522f7605a35d40ad8562f3d92b6be9b13

Observation ebbe07ed-b501-4724-9d36-7ff24729e2b8 · outbound

This paper cites SUTD-TrafficQA: A Ques- tion Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SUTD-TrafficQA: A Ques- tion Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.379777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.632142Z digest=sha256:43f8820b788fcfcbe1d9c4b8152383732f12d3d2ad6987cb3472836808e1d456

Observation 117f22c4-c441-4ad8-9258-a6695f8841c0 · outbound

This paper cites Di- vide and conquer boosting for enhanced traffic safety de- scription and analysis with large vision language model.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Di- vide and conquer boosting for enhanced traffic safety de- scription and analysis with large vision language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.263756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.730131Z digest=sha256:77ad9e4b4facbb8c4eacc46b78184b403d9a2a623e5f1a98a8b3a0ee7feb4f21

Observation 688e393a-b896-4e19-b5df-d7386dda7b8a · outbound

This paper cites CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:06:25.771980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.791229Z digest=sha256:425db9f8e7de38731b59f47952ccdaac418143d5ee4db5500c20c6feebb4efee

Observation ea8fc92d-7dbb-49f8-9275-9a6d9b86247a · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:26.105600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:24.856071Z digest=sha256:db70a4cb01bc8bbb6b0ff53dd6ccb1bed11c4b966518af3aee9cc9d8f8734c20

Observation 84d2afda-c73a-4913-a5ae-9c33b2006d0b · outbound

This paper cites A Study of Situational Reasoning for Traffic Understanding.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models A Study of Situational Reasoning for Traffic Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:24.973033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:24.973033Z digest=sha256:b5ba0d5fef16edc1b30047edf47bea93f2942e48aadda4cfee5f4cfa8dc7d91d

Observation b0cbf26b-8b6a-47ac-92ab-18eea1b1bb09 · outbound

This paper cites When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:25.069634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:25.069634Z digest=sha256:a3816085024125c16083c961b1bebae6796c0f3e6f4d995978fef3816b7cc779

Observation 41fe9ed6-f67d-4514-a3ac-efed3ca5e00f · outbound

This paper cites CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:06:25.539818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:25.134699Z digest=sha256:1278c971fba195c15fb643cd15bbc8b2dbfe2983a299a93ea1aca5d6c697fbde

Observation e66332e3-e7c8-49e4-b33e-672f422eade7 · outbound

This paper cites TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:25.219559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:25.219559Z digest=sha256:2a5e8f73a933023553af3b47852d31fdd8ac3bfcba14eda765f6054a91cf49b9

Observation 3610eae5-e259-4029-b6d7-b2658d72b531 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:06:25.927402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:06:25.327098Z digest=sha256:c9cc87531e1b5abafc571cbaac1eb19c30f6a421cb94d9b6b866dcc78cfddee1

Pith citing papers

Observation ca2de7d7-d1d7-4c89-9448-115a09479bea · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:d09c67d3524fa6f2c81bb32434296c14764f9c863d8a953e611c42866b8addc3