Pith. sign in

Paper Citation Record · LEDGER

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.03916.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03916 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:04:14.254349Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c395a067-53ef-4ccb-b30a-dc08fa59f74d · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Docvqa: A dataset for vqa on document images,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:19.199956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:09.955672Z digest=sha256:7d5e0b8a2f39187c10b0eb1ddfa8ee24f8a405072b974f72b738f903f3d7640e

Observation 1b28438b-a3fe-4128-8c23-9f9353e9a0cb · outbound

This paper cites Infographicvqa,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Infographicvqa,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.920954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:10.072039Z digest=sha256:6d708daca12ff8d56c03b6edfb85e80dac362e0a19471ce5b6490ee63e9e9d92

Observation 2df9f805-09b6-4fb0-a073-b514c5a137e0 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Slidevqa: A dataset for document visual question answering on multiple images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.704656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:10.189362Z digest=sha256:fb97777ad8944190d5fd8f3cfb22b54d712f293c9498d0492c0ae956b4d98e58

Observation f4563e43-4b88-4be2-8990-a3fceb74b5aa · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Msr-vtt: A large video description dataset for bridging video and language,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.495421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:10.342478Z digest=sha256:1d319311976669dc535965c1f90956eab25ac33f7ca7700eace1fbf8137292db

Observation 08d47dbe-8fcf-45b2-995e-772c54db6ee1 · outbound

This paper cites Dense-captioning events in videos,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Dense-captioning events in videos,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.280657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:10.475082Z digest=sha256:cdbf49a51d2f1d39177cbc34be811d29cf1cb9adfa73f1ff47431b63aa401fbc

Observation 5e442f33-eac8-4dd1-8aa5-3fc047c20570 · outbound

This paper cites Towards vqa models that can read,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards vqa models that can read,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.037729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:10.636457Z digest=sha256:76327c5fba06938abdf2010a86ea7e3d8181563c8d92c8e6e0b89c3cb8156f74

Observation 8f1a2db1-2a3a-4add-82ed-3acfec3fd233 · outbound

This paper cites Qwen2.5-VL Technical Report.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:10.751674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:10.751674Z digest=sha256:db54a40c8ce2c11bc408566877b803780bc1466a79964980a7ca8bbcf8b83858

Observation afb972a4-28be-4dd1-8bba-bf8a6557453c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Lora: Low-rank adaptation of large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.727409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:10.850014Z digest=sha256:ac63289f4a4f1802e5d7d553004e6a3f3dc441c1279a1ad76a278048092cfb32

Observation b5a2dbb4-ea70-4edf-a9e3-95b4227a3f6e · outbound

This paper cites python-pptx: Create open xml powerpoint documents in python,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models python-pptx: Create open xml powerpoint documents in python,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.400977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:10.980238Z digest=sha256:f233f9d48cce3b93f3d8552489b4a3e572cb741df44e5651147d6c44d8ea879a

Observation 33776ba4-c58e-4dfe-86d9-32512bd56db7 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Bleu: a method for automatic evaluation of machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.119793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.119793Z digest=sha256:9788c6693de7fb2634364500043c8ec8836f7cfc5e3b8dfbf6913631cc193896

Observation 7e77378b-d7c8-4df0-bdf9-81e1c5430b33 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Rouge: A package for automatic evaluation of summaries,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.237427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.237427Z digest=sha256:8a9a215ce892f5059ab63c4ebfe496975f87c0076c7f6df5f9671da193b1f272

Observation 4fc245fc-f4d6-41aa-bc62-5fa3879789f2 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Spice: Semantic propositional image caption evaluation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.185503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:11.361447Z digest=sha256:d7c5b5957cf2eee73e50a11911f1935169f7681fa633eaea8c0cebccd0428f06

Observation c8eb1cfb-6104-433e-9f23-ed5f6634ba94 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.505717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.505717Z digest=sha256:a8b7ba27e078245bd88afd9f620c619bc154328fabde8c4689dcc88e23a2f364

Observation f3d8a2c0-3142-4172-9299-1b0e655d37b1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.918071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:11.654830Z digest=sha256:c16c4aa31303e3ce63858e248e551a8bb6faf3d676e0f0e0db8abe12d4825efb

Observation 25579411-a4fa-4a53-8687-69c83adccef3 · outbound

This paper cites GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.787732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.787732Z digest=sha256:e3e20b3c83f8abc4f7e2949994a186cee202618cfb4cbc34a11437461af165ee

Observation ea5f9202-a5dc-4ec1-b97a-393df857d521 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.888949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.888949Z digest=sha256:a1132a7b5980d194b74ec5f29b52360b9393cf0e4d5343bb92941cc355c93c4f

Observation b39d791f-4869-43b5-95af-3875843b3abc · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Flamingo: a visual language model for few-shot learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.035982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.035982Z digest=sha256:85ae2da92964ffb89e8e08e49467e182c609354b228d8f1cb65301b1af77d903

Observation 32d36aad-f476-499a-bdc2-7e0f6b44d1e6 · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlm: Pre-training of text and layout for document image understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.638138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:12.156363Z digest=sha256:efbdf9b68493e6205a61748335ce494946812eddc7b24dbd5fea4c572a2153c5

Observation 3ceb7149-7d64-49fe-b295-13df1a98afe5 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.315928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.315928Z digest=sha256:b7a5494c5d2bbdf3cf1535842a6d0d1bff779349650910689e0c551f90be1138

Observation 3a9ecb96-d720-4380-80fc-6464a625a2d8 · outbound

This paper cites LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.416233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.416233Z digest=sha256:493e95b990598975810c2847d68e7403ee13f8718cfe63968519c0e33d3b74fa

Observation 93069f4d-ff76-4b0a-a1f6-fc35d3299d62 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlmv3: Pre-training for document ai with unified text and image masking,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.381185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:12.517861Z digest=sha256:17ce5ab9d828275d1dd0fee0527044b26e06b1275f378d4514c588c3ba449c75

Observation 5d641d5f-7ee8-4f4f-80cb-971c6b5ceb37 · outbound

This paper cites Unifying vision, text, and layout for universal document processing,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Unifying vision, text, and layout for universal document processing,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.189850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:12.637123Z digest=sha256:4e8116c762c591795cd8e65a39338ebc91d57816c6dcf5f061b134a4bf2c6cd5

Observation ab129cd5-8c04-4f1b-875a-cd566038aa19 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards automatic learning of procedures from web instructional videos,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.953547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:12.751843Z digest=sha256:fe06cf41fbaebfad8dcb1cd6d393315f1f438d439fd8c4e2b0ad1f6a301b1c7d

Observation ce4ed28f-4ecf-4329-80cf-bd6c990a50dc · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.865362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.865362Z digest=sha256:7063a428706286ea8d71f3db9c2298f26a35fe3cad4734dd626b931d85ba96fc

Observation 28e60d0b-1127-4e64-9a6c-cc9b0317ddf2 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.002034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.002034Z digest=sha256:92a565a08dfe5f799e56991741bd703ced62ee6d86cb08463ec258ab9b436a19

Observation 9d6141b0-04e1-40ac-91a3-b340e99843a4 · outbound

This paper cites X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.646879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:13.153091Z digest=sha256:e175c76495421358ae02ef23c8a493a1f635c57b0f8bc66e2010fe7d765456b7

Observation 8f9c5f80-2681-4dc3-aed7-67647c00322f · outbound

This paper cites VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.314128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.314128Z digest=sha256:1548e310fb391fa789027267e360bd2c54e1ab4de02771e55ab5c8335971377c

Observation b633c2c2-2d2e-4609-b6f5-f9cef1479bbb · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.425644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.425644Z digest=sha256:acf4066c6ce4783c5a4d08729f1dcfd09f64eab0edd1f04ba0a9e8ea930dcbac

Observation e2191883-4d52-4f65-9ef9-e8cb41d415a1 · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.503234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.503234Z digest=sha256:4ead2ef525f0ceaea1b7b2ff88532346a88bdfb0e732649fadc62d9d1b98a037

Observation a41c1694-56cf-4873-a39a-98715f7f79ea · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qlora: Efficient finetuning of quantized llms,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.548641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.548641Z digest=sha256:404f827ed9a318743e5e7caad28c94f9f7f114398bdfee9b6d46ec271cf24b34

Observation 241ac941-1d25-4a6f-9925-8e59ec464c91 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.653416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.653416Z digest=sha256:05a7a2ad733b9280d6394cd0d51360374129e2b6b7b88fea6400a7e76414decc

Observation 516f1b29-e1f9-4c7b-89f5-87d015388132 · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.722971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.722971Z digest=sha256:d01342d8becda5fdad5c49a1a2ee6e4609606a4ffdb03ccac321148864aefe93

Observation ffec0415-ea13-4779-ae3b-7ebd73c1e1f9 · outbound

This paper cites Evaluation metrics for video captioning: A survey,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Evaluation metrics for video captioning: A survey,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.358915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:13.815372Z digest=sha256:0e95fd79cb8b2ff316ac46abddb55bbe36cbfe3653f4daa69485fbf21e1a729f

Observation 18ef3acb-5841-415a-b5a1-f8a70a19b3d7 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.061021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:13.893403Z digest=sha256:5a6bc061893d87824f9f1d48fc7aa0d270ed1ca1a18ae89c25b189e3ecaf35c8

Observation 2f4cd553-01dc-4ddb-a5b0-af92db97c562 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:14.812717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:13.974323Z digest=sha256:9c209f613d8ca57cf212baccb6e579b256cdb5d769df80acd7f319f7655697e5

Observation 0a774904-43d6-4660-b4f7-18abba67c15c · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mmbench-video: A long-form multi-shot benchmark for holistic video understanding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:14.672111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:04:14.050908Z digest=sha256:4aeb9739a889c8609cddac31017b8c4dbfe741452d98aa38c15aaac64b970f4b

Observation 618594b0-6d38-4c7e-a528-0ed42817d561 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models TempCompass: Do Video LLMs Really Understand Videos?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.128716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.128716Z digest=sha256:8e08ac8b270cdddb3eda80e68c1e2c3c3da6d0e02003e0fd1d0971f8ec2245d0

Observation 0b657244-8cae-4f8c-bbd8-ae81bfa0b2b9 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.186124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.186124Z digest=sha256:0e50f4bc46d456e4d20bbb3719e89f7974ec36c0c51f8784f6ee27f7d17a8f51

Observation cbae45d3-4c04-40d1-9178-959b0eee45f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.254349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.254349Z digest=sha256:b0ca1e52ae37a2c5ad7ff092b8ef8cb1359635312f1175a2712a5ef2e77975c2

Pith citing papers

No inbound Pith citation observations are available.