Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:04:14.254349Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.03916.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:04:14.254349Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c395a067-53ef-4ccb-b30a-dc08fa59f74d · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Docvqa: A dataset for vqa on document images,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b28438b-a3fe-4128-8c23-9f9353e9a0cb · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Infographicvqa,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2df9f805-09b6-4fb0-a073-b514c5a137e0 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Slidevqa: A dataset for document visual question answering on multiple images,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4563e43-4b88-4be2-8990-a3fceb74b5aa · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Msr-vtt: A large video description dataset for bridging video and language,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08d47dbe-8fcf-45b2-995e-772c54db6ee1 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Dense-captioning events in videos,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e442f33-eac8-4dd1-8aa5-3fc047c20570 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards vqa models that can read,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f1a2db1-2a3a-4add-82ed-3acfec3fd233 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qwen2.5-VL Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb972a4-28be-4dd1-8bba-bf8a6557453c · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Lora: Low-rank adaptation of large language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5a2dbb4-ea70-4edf-a9e3-95b4227a3f6e · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models python-pptx: Create open xml powerpoint documents in python,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33776ba4-c58e-4dfe-86d9-32512bd56db7 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Bleu: a method for automatic evaluation of machine translation,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e77378b-d7c8-4df0-bdf9-81e1c5430b33 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Rouge: A package for automatic evaluation of summaries,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc245fc-f4d6-41aa-bc62-5fa3879789f2 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Spice: Semantic propositional image caption evaluation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8eb1cfb-6104-433e-9f23-ed5f6634ba94 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3d8a2c0-3142-4172-9299-1b0e655d37b1 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 25579411-a4fa-4a53-8687-69c83adccef3 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea5f9202-a5dc-4ec1-b97a-393df857d521 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39d791f-4869-43b5-95af-3875843b3abc · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Flamingo: a visual language model for few-shot learning,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d36aad-f476-499a-bdc2-7e0f6b44d1e6 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlm: Pre-training of text and layout for document image understanding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ceb7149-7d64-49fe-b295-13df1a98afe5 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9ecb96-d720-4380-80fc-6464a625a2d8 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93069f4d-ff76-4b0a-a1f6-fc35d3299d62 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlmv3: Pre-training for document ai with unified text and image masking,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d641d5f-7ee8-4f4f-80cb-971c6b5ceb37 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Unifying vision, text, and layout for universal document processing,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab129cd5-8c04-4f1b-875a-cd566038aa19 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards automatic learning of procedures from web instructional videos,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce4ed28f-4ecf-4329-80cf-bd6c990a50dc · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e60d0b-1127-4e64-9a6c-cc9b0317ddf2 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d6141b0-04e1-40ac-91a3-b340e99843a4 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f9c5f80-2681-4dc3-aed7-67647c00322f · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b633c2c2-2d2e-4609-b6f5-f9cef1479bbb · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2191883-4d52-4f65-9ef9-e8cb41d415a1 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a41c1694-56cf-4873-a39a-98715f7f79ea · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qlora: Efficient finetuning of quantized llms,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241ac941-1d25-4a6f-9925-8e59ec464c91 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Prefix-Tuning: Optimizing Continuous Prompts for Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 516f1b29-e1f9-4c7b-89f5-87d015388132 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffec0415-ea13-4779-ae3b-7ebd73c1e1f9 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Evaluation metrics for video captioning: A survey,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 18ef3acb-5841-415a-b5a1-f8a70a19b3d7 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f4cd553-01dc-4ddb-a5b0-af92db97c562 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mvbench: A comprehensive multi-modal video understanding benchmark,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a774904-43d6-4660-b4f7-18abba67c15c · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mmbench-video: A long-form multi-shot benchmark for holistic video understanding,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 618594b0-6d38-4c7e-a528-0ed42817d561 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models TempCompass: Do Video LLMs Really Understand Videos?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b657244-8cae-4f8c-bbd8-ae81bfa0b2b9 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbae45d3-4c04-40d1-9178-959b0eee45f9 · outbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.