Pith. sign in

Paper Citation Record · LEDGER

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation

As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:2411.19951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19951 v5

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:44:06.086591Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.816809Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:32:36.496230Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e07829f0-d914-4553-b6cf-d96e6ab80b11 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.763113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.763113Z digest=sha256:1d5fcd345c257afff110cb7130fe7e8fa9a6a4578121dab32946e8a8f169b9c0

Observation 8ed1952a-5f1c-4997-b5b4-82435fc4dc65 · outbound

This paper cites A survey on multimodal large language models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation A survey on multimodal large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.769019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.769019Z digest=sha256:61b19ae9411bba0a8bdb950f41d696dc6d36d797d617b898ca3dd621efb558a4

Observation c1c2058b-d59c-4dff-8b15-f469b4dfc314 · outbound

This paper cites 3ur-llm: An end-to-end multimodal large lan- guage model for 3d scene understanding.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation 3ur-llm: An end-to-end multimodal large lan- guage model for 3d scene understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.196342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.774526Z digest=sha256:68025f1671c211392c9ca4429605abf4da6ddd083b32f7100fe90026f2158e09

Observation 6dae7588-575f-4172-84bb-c9fa8a4bd768 · outbound

This paper cites Mul- timodal large models are effective action anticipators.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Mul- timodal large models are effective action anticipators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.183459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.779438Z digest=sha256:2fbb3b08c367f0f5feb7afddba54b0abd717d17e1ab20bdc7dc6f8add086c8e9

Observation b3220637-3663-4ba3-b8b6-e2eb4cc4010b · outbound

This paper cites Lmeye: An interactive perception network for large language models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Lmeye: An interactive perception network for large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.169785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.784089Z digest=sha256:0754b4f0875b1adce0dd47fc163ed5f10dd802488e41c25dff2e0ae9bf3a0395

Observation 74d29327-4e37-4e25-9037-4f91baf31f74 · outbound

This paper cites Effi- cient transfer from image-based large multimodal models to video tasks.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Effi- cient transfer from image-based large multimodal models to video tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.156125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.788034Z digest=sha256:dc00117d53ff4106814ab71574afbea377bbcefa5ee38b8b8105bc0079ac0d83

Observation 85d6ac83-0355-4520-b337-bc3a72a4c147 · outbound

This paper cites Shapegpt: 3d shape generation with a unified multi-modal language model.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Shapegpt: 3d shape generation with a unified multi-modal language model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.142865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.792491Z digest=sha256:595c98e7d6b08ae3d1c1c441a67ae130f1ca6d1dee7b46778442716be8c33e3f

Observation 40d58c35-37a6-4c80-83d7-ef63baea8315 · outbound

This paper cites Context-enhanced video moment retrieval with large language models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Context-enhanced video moment retrieval with large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.127812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.796519Z digest=sha256:4c23fb4adb42387c1c8a5230551413063c305a3c81f6e50153503e65c606f3eb

Observation fab5becb-3c87-418d-9197-4cbac2800371 · outbound

This paper cites Etc: Temporal boundary expand then clarify for weakly supervised video grounding with multimodal large 11 language model.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Etc: Temporal boundary expand then clarify for weakly supervised video grounding with multimodal large 11 language model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.115176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.800695Z digest=sha256:8eaf28e1701dcf6bda9ad44ce5e51e38a23e5d74fa03303743206fbbc541b0e7

Observation 63769707-bbaa-4f1c-b43d-8ef6ac04a625 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gener- ation image-text models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Laion-5b: An open large-scale dataset for training next gener- ation image-text models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.101237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.805414Z digest=sha256:957638668d9292099f566b5705bf95304c52bfd2d641eff2d6d616f264903799

Observation 3486ccea-cc5f-4a5c-907d-61a761d5a2ba · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.086575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.809861Z digest=sha256:6f941aa5abefeeb369a7c0c2718193c65cd0aea32fea9ecf6d7249e12a3ebeed

Observation c3d17f2f-f83f-49d9-815b-b750d910f892 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.072635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.813766Z digest=sha256:8143965879d6d061bb8eb4703e61e4c98ce44d165ea3f57414abb379345de6ea

Observation a5226424-df39-49fa-ba0a-4ac0c7dd85da · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.817913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.817913Z digest=sha256:4fe4e75fb9868eaca2ef941ef04ddae5f955a3151cf563acb2f5e0c46cd50b3a

Observation cec04faf-e648-422e-a96f-ac3e5ffe86b1 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.822868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.822868Z digest=sha256:75f0c5100fbd8a345af6616f9349933ac8482021f819148459d3504c702d6bfb

Observation a00808dd-34c9-434f-997e-540ce43734b7 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.057806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.827667Z digest=sha256:5445c4c4a6f23298d70845cd2af70c22c717983e3d77a1ecf5adc5b0ab0c55ca

Observation 1fe94f63-6b90-45e1-a2c6-f9d55e4dc02d · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation An image grid can be worth a video: Zero-shot video question answering using a vlm

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.042781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.832184Z digest=sha256:1459016c6e399ef4d994bffb069b9b95feca26b9c40ba2fc45416fa23f4d9b1f

Observation d96dadec-8a78-48f9-950f-320a2629dd55 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.836761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.836761Z digest=sha256:e2f00e34954cb4850bb3e2e88a376a06cdac0858ea13322e96ab97bc2d8d6c19

Observation cd6bc861-253c-4f62-a883-eb7f56059d91 · outbound

This paper cites Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.841307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.841307Z digest=sha256:c4c344531198e55bb0b14ef07cf854bb1e2d4fd9c31fa3ad7871ca0d045737d4

Observation 03595745-8325-4e69-b452-6fde4368b674 · outbound

This paper cites Video-llava: Learning united visual representa- tion by alignment before projection.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-llava: Learning united visual representa- tion by alignment before projection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.028203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.845836Z digest=sha256:0d9fb0be36157fedd2277e34c133094781dc62d8b02e8713da908fbaf111f5c7

Observation 4e393e50-18e0-4408-996b-98e7147a69a2 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.014510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.850295Z digest=sha256:ba34f1775631d320e7e5f895aaa0968d4d6cdf54a5e2b1e3966ee804fd8c62f5

Observation f6a9524b-136b-4c10-9151-789b82f9fb4d · outbound

This paper cites Visual instruction tuning.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:07.001085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.855207Z digest=sha256:b10c264afd3099e592067526b0562c5f70846d2eceec674439302664f318917b

Observation 20a019b2-03ba-47e7-9b7e-255680ca90e6 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.986777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.859429Z digest=sha256:5d83490f6bb90040c2d6e1531c9e7f628678b054185fe520b20f3d2a2a2c3a99

Observation df7d79a3-755a-4fae-a71e-6aba50a051df · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Vtimellm: Empower llm to grasp video moments

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.972992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.863667Z digest=sha256:918c80d47b3db1292ade70d570ec33d7d38c9b381be57ef3f5b7b913ad45bfbe

Observation d07bae86-25a3-4507-8e8d-368a00307c3e · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.867573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.867573Z digest=sha256:a028b5a06cc82810147689f53f90fc6384d1a9428ca09bee4272cb2c1936ddad

Observation cf50a6a8-49b9-423d-8fdc-8038fe203b48 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Llava-next: A strong zero-shot video understanding model

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.959840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.871792Z digest=sha256:e19b0ae883acb27c52557469523061acdb02bba157315c54370b330d84d9af9a

Observation 5acad12b-786e-4b2f-8a64-fdb290a0eb4f · outbound

This paper cites Long Context Transfer from Language to Vision.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Long Context Transfer from Language to Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.875906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.875906Z digest=sha256:71fa84e777086e719fa62fbe99927a858c2bf86fb74269cbc53fe5502764049d

Observation a8e24842-51ba-493b-a40a-7b8fb6c75356 · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Longvila: Scaling long-context visual language models for long videos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.945846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.879602Z digest=sha256:fb4e3239374b1a2726f331090ef317d39a8b0d96b89d4af63dd6df7c70f043a9

Observation b304a8e9-c627-465c-9a19-19f0c90d267f · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.883956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.883956Z digest=sha256:50352e6479b5a731d88c76f4699716561a845f953cc668c4bce1c88cffaaeca0

Observation db3b625c-bd5b-480b-a776-206f61a13d6c · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Llama-vid: An image is worth 2 tokens in large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.932460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.889614Z digest=sha256:6c7267acf20173122d3a461d7e684621a84776cde3fcc6e5e4868e2d7a0e9e3a

Observation 726c8d2c-69de-496f-aee9-e2bd67324906 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Moviechat: From dense token to sparse memory for long video understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.916921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.894925Z digest=sha256:b506e6340e9970d083e3216ff53ea11d9e4999b545c5dd7d0a30ac2ce3937f3d

Observation c40a6ee4-98fa-4b30-ab10-f5b05db0c5ce · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.902365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.899483Z digest=sha256:ea8954e7e8d6d0be695ea2a299e56c9db7f02a8d65064aded3c6f4d559372182

Observation 2a6ebf20-8279-4bbd-8d68-0b9d28fd76a4 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.903874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.903874Z digest=sha256:83fd964c3b86e9c2d2a285cff97a4102532807335ed27be4dcb9a9bcb613ebe9

Observation 2e63a0c8-c49d-426e-baef-ceea104c4611 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-rag: Visually-aligned retrieval-augmented long video comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.908171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.908171Z digest=sha256:5f5bcc0cd5168434ce694249d7631fce6be98c6a868a7e315a32df5bbc64307a

Observation 67bcc6be-f8be-4b10-abf2-696bed1ec14f · outbound

This paper cites Rethinking temporal context in video-qa: A comprehensive study of single-frame static bias.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Rethinking temporal context in video-qa: A comprehensive study of single-frame static bias

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.887200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.912256Z digest=sha256:0bb9658025a2f2f2f5af515d5ec5c57f871f44099847b3253fd7a4a1eb42c70c

Observation c6146fb5-1615-40c9-80fd-8e44bd6af1e8 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video question answering via gradually refined attention over appearance and motion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.871709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.917082Z digest=sha256:3417f6c5d264094111108a6ad8720daf3ececf3c0e560f845c8cf1ab468f98fa

Observation 8bfbbff2-ca46-4202-a7ab-f0823f8bd84f · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.854971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.922044Z digest=sha256:e0b009bfa8de50718893b21d3e725a621ccee0b9157b9ed582d68da7da894a7a

Observation 115511c5-ee09-431a-b8d2-6530b699b622 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.837851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.926451Z digest=sha256:3c3a78e679f06ecbce394a4563b32d5a397c924e503044818428aa5eff3d65e2

Observation 79f51786-fda2-4be9-a433-cb49271c43a3 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.821753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.931161Z digest=sha256:db015731c2f6ee703ae5eb863162915857d692cec7e651d652908a47c2969665

Observation 8f3a0d8f-19d9-449f-aa85-32036b29bd5c · outbound

This paper cites Temp- compass: Do video llms really understand videos? In ACL (Findings), 2024.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Temp- compass: Do video llms really understand videos? In ACL (Findings), 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.806433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.935575Z digest=sha256:979012c130501e2dde50decbda5b7beefe922724221fc9c1c3ff1a8e0d064212

Observation 82eb4c28-8f4a-45c8-bbde-e88a1f6e0611 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.939793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.939793Z digest=sha256:c69fdb150ef828a8913f6a72d4a29ad656bab86370b76e41abc997f1467d2543

Observation 9c74765c-bcd1-4bc9-9998-0b3e9b2878ee · outbound

This paper cites Topa: Extending large language models for video understanding via text-only pre-alignment.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Topa: Extending large language models for video understanding via text-only pre-alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.792640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.945173Z digest=sha256:db73863dae0850da689f86859a6a4a504f05b3ad8c6a5205e3c322032c560ec6

Observation c598d106-8ba8-4aab-8c07-e3574204b94b · outbound

This paper cites Temporal reasoning transfer from text to video.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Temporal reasoning transfer from text to video

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.779026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.949057Z digest=sha256:ba03832dee2bc4689e3bec22bbacfae28d4d8ae20abc84943161dd29ae90e619

Observation 89237e06-3f77-4a33-a5ee-f309802bd30e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Learning transferable visual models from natural language supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.954188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.954188Z digest=sha256:cdbb2d816dae4feaa04bda7f5fbe16b1c450394a9295aea15d26fd6c43310205

Observation 58c98233-b42c-4319-a34f-3e942d50f9c1 · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.756735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.959234Z digest=sha256:6755474c1515b3c464b1da8ffc6240afa036f4237422946212ba0a14a0e1160c

Observation 7215ecba-6f80-41ee-9cdd-6ff1eb9251a7 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Monkey: Image resolution and text label are important things for large multi-modal models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.740582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.963888Z digest=sha256:79b6792e8513d2118600ad13fd3b1f5f51cf1c5b95fd609cf8217797896480e6

Observation 2dca21c6-f390-4882-9d9d-4c56fd005234 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.967424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.967424Z digest=sha256:f69686cb83bc403c9d9d5c85736d422dbe3d9995329e175ae1985402f4be8b0a

Observation 977cb2b0-9c0d-45f2-b100-444cd6a4cfc0 · outbound

This paper cites Sharegemini: Scaling up video caption data for multimodal large language models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Sharegemini: Scaling up video caption data for multimodal large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.724755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.971675Z digest=sha256:8d22c6e43dc41ace1bd04e7228359af3db1eaa7e7056b542ebe6e1ec6b64590c

Observation 76a7b6ef-69ac-45e3-98e5-64a13013db83 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.708995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.975461Z digest=sha256:f466bcaf6eea131fdd14af06c360bfaf7df6133ee1095487b9d894c80a93dff8

Observation f5cf478f-5fd6-44a6-b3a9-11f40d84a00e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.979369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.979369Z digest=sha256:a525b2564497aec101ba560e2e1a072eb3d4a406d88fbb2024023721aacb871b

Observation a07fd844-0e75-440f-8f3c-bf8e6e3757ac · outbound

This paper cites Token merging: Your vit but faster.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Token merging: Your vit but faster

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.695834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.984373Z digest=sha256:43e8fad7e7dba5332c6e98c4b60eb8dc899ba6c133ba9a8b98d00cc52a3a0c42

Observation a5451fef-f59f-46e4-9aa2-93c320acd294 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Activitynet: A large-scale video benchmark for human activity understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.681841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.988523Z digest=sha256:f5fc7daeee110cc22e6df967ac306b4a44154c6ac7525a964098d5c0152e7a0e

Observation 9b46a8eb-1676-40ce-bbd5-78aa5a2384e2 · outbound

This paper cites Lima: Less is more for alignment.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Lima: Less is more for alignment

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.667267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.992391Z digest=sha256:eb54dab2382ce202afac7d969c37b05c908a1b1dc6c659649c62aee7068968ba

Observation 22002286-d833-4b76-8495-e5c0b9a76887 · outbound

This paper cites What matters in training a gpt4-style language model with multimodal inputs? In NAACL, 2024.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation What matters in training a gpt4-style language model with multimodal inputs? In NAACL, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.653645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:05.996593Z digest=sha256:db800fa138554dab11c5081b042c165ea654360c51251ba9554aadb0f01f1c47

Observation 75c06f77-0aa4-459f-8d02-6d88fa4a8072 · outbound

This paper cites Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.638231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.001029Z digest=sha256:3d77debd91b90d5d65d51e28880a597a5bc9f70a1aeba6a91215c056de0f82ee

Observation cebed98c-4814-4ccb-8672-11b932082178 · outbound

This paper cites Wildchat: 1m chatgpt interaction logs in the wild.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Wildchat: 1m chatgpt interaction logs in the wild

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.622905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.005568Z digest=sha256:44335fccf9c600c36dc029a7231ccd42ec748409153222eab2b138195c6e09db

Observation a6d26911-c6a0-4260-9bb1-e8b3e4f789cb · outbound

This paper cites Gpt-4v(ision) system card.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Gpt-4v(ision) system card

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.607972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.009491Z digest=sha256:0a5df11894dd6dc98e599cb41534c47e8cb3f10840f9d8c5e6e7d3cee1fb8bad

Observation 37e6f01e-a358-49de-a7f8-e126d4ef68b6 · outbound

This paper cites Introducing claude 3.5 sonnet.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Introducing claude 3.5 sonnet

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.592924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.014829Z digest=sha256:7a564d6dbd707820aff5aa549be9d7f3ace6b829631963836f06bf6ace1f3b21

Observation f5ba2e3b-bcea-446b-bc0d-487f2723449b · outbound

This paper cites Hello gpt-4o.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Hello gpt-4o

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.576087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.019106Z digest=sha256:6e531f2a20b193d2fe8438c5eafb8fc03cb24d19cbb671b536d39efec313d25d

Observation 1425fc8a-67e5-41a1-8b22-a7054dc8f63f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Gemini: A Family of Highly Capable Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:06.023714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:06.023714Z digest=sha256:2014741320968f4589d53c5909cde6ee65f214e95957890ee7bb14db132bd65f

Observation a524375b-aa76-453c-8836-06d677f80dee · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:06.029230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:06.029230Z digest=sha256:3fa3892f6730852dc06884aeb6c7682aac16ac07b9ecc7b25efaa11f4b40dcef

Observation 205c701b-4378-4652-827e-40a68f618c15 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:06.034630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:06.034630Z digest=sha256:ef4d9486ea90ebf797595f39043232c3bb3b8db898938bc1f8db0ac82fbe6ad6

Observation aabfc9e4-6f09-4dee-86bc-69cea001412c · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.560247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.039294Z digest=sha256:ffab2efb5b36ad5a241381aede620704f0476fe9b489736571d437ade92690e0

Observation 9f87d359-c568-489b-97c6-a61df3107b10 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation MLVU: Benchmarking Multi-task Long Video Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:06.043632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:06.043632Z digest=sha256:bb55b4b053ca92a1913f7ce50233ba4cff1692ef5a22ab2119fdec84a78aa40f

Observation 2ea2620b-cc6e-43ca-a5e9-38ab47faf8d9 · outbound

This paper cites Introducing llama 3.1: Our most capable models to date.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Introducing llama 3.1: Our most capable models to date

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.544862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.047279Z digest=sha256:eaaadd9a542a625e7a524f0743d22fba611e7aaef64a0662eff9a8650b73f4bf

Observation 43247aff-68cf-44ff-8279-4ded7179045e · outbound

This paper cites Answer with the option’s letter from the given choices directly.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Answer with the option’s letter from the given choices directly

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.526568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.051839Z digest=sha256:1def677fe7d5f2f178c8a5674da11f263248e689053add562e838339e9b9769c

Observation 7c1e11d8-d880-4f2b-84ee-9656d872e886 · outbound

This paper cites an unresolved cited work.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:44:06.458666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.069253Z digest=sha256:cc9458ead288645a481cd16973bbc092bc39ffb60fa25a744b1db1d34bd8869a

Observation 576604f9-e7f2-4a53-b0cb-53139af6f489 · outbound

This paper cites an unresolved cited work.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:44:06.506201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.073288Z digest=sha256:22814cce9f152b732bc6bd94af8d84c654e136201a4b0c5f65d0b73f32aea2fc

Observation bcf4b791-8b11-4e26-92a7-84fb535fe895 · outbound

This paper cites an unresolved cited work.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:44:06.488331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.077548Z digest=sha256:9716b29cac327a7f312489e35efa10be958acf1f7114694b6c5882a3ccbeab9b

Observation c0aca438-5fb8-41f2-928e-fcd5d0cc2185 · outbound

This paper cites an unresolved cited work.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:44:06.473859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.082055Z digest=sha256:7614b01bb8a8cbb01dad93a0fa480277f115f992a42f14e94637d4966cc2613a

Observation f68fd608-d6c1-4088-bc66-1b7dd0149384 · outbound

This paper cites Therefore, the answer is A.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Therefore, the answer is A

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:44:06.443945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T05:44:06.086591Z digest=sha256:219f10709a1f8131a9a1713497d91ff1cb0682e3f4caf9be4ba9ca1d4878cbf8

Pith citing papers

Observation 1803023a-38b0-4f32-97fc-5a1faeda3cd7 · inbound

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey cites this paper.

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-10T04:36:37.577135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:36:37.577135Z digest=sha256:8081205f5cfdd63b689a1f3a0f949ec82d2b47218e41e4b57156c3edc20bb9b3

Observation cf517a71-3707-41cd-8b24-fd8c2f92ba05 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.816809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.816809Z digest=sha256:e47a09470039fc113d84050a9862b0b1f9c28c8798990ba581769cecdb2abe58

Observation fd186912-810c-4e4b-a4bb-b44001cb2a61 · inbound

LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment cites this paper.

LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:36.551624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:35.144267Z digest=sha256:6c75d995671dc6c638d2d893d49d7576ebab2a63dccd0ff0ddbaa21cc43dbe82