Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:44:06.086591Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:2411.19951.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:44:06.086591Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.816809Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T11:32:36.496230Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e07829f0-d914-4553-b6cf-d96e6ab80b11 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed1952a-5f1c-4997-b5b4-82435fc4dc65 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation A survey on multimodal large language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1c2058b-d59c-4dff-8b15-f469b4dfc314 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation 3ur-llm: An end-to-end multimodal large lan- guage model for 3d scene understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6dae7588-575f-4172-84bb-c9fa8a4bd768 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Mul- timodal large models are effective action anticipators
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b3220637-3663-4ba3-b8b6-e2eb4cc4010b · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Lmeye: An interactive perception network for large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 74d29327-4e37-4e25-9037-4f91baf31f74 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Effi- cient transfer from image-based large multimodal models to video tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 85d6ac83-0355-4520-b337-bc3a72a4c147 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Shapegpt: 3d shape generation with a unified multi-modal language model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40d58c35-37a6-4c80-83d7-ef63baea8315 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Context-enhanced video moment retrieval with large language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fab5becb-3c87-418d-9197-4cbac2800371 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Etc: Temporal boundary expand then clarify for weakly supervised video grounding with multimodal large 11 language model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 63769707-bbaa-4f1c-b43d-8ef6ac04a625 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Laion-5b: An open large-scale dataset for training next gener- ation image-text models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3486ccea-cc5f-4a5c-907d-61a761d5a2ba · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c3d17f2f-f83f-49d9-815b-b750d910f892 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5226424-df39-49fa-ba0a-4ac0c7dd85da · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec04faf-e648-422e-a96f-ac3e5ffe86b1 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00808dd-34c9-434f-997e-540ce43734b7 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1fe94f63-6b90-45e1-a2c6-f9d55e4dc02d · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation An image grid can be worth a video: Zero-shot video question answering using a vlm
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d96dadec-8a78-48f9-950f-320a2629dd55 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd6bc861-253c-4f62-a883-eb7f56059d91 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03595745-8325-4e69-b452-6fde4368b674 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-llava: Learning united visual representa- tion by alignment before projection
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4e393e50-18e0-4408-996b-98e7147a69a2 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f6a9524b-136b-4c10-9151-789b82f9fb4d · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Visual instruction tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 20a019b2-03ba-47e7-9b7e-255680ca90e6 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df7d79a3-755a-4fae-a71e-6aba50a051df · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Vtimellm: Empower llm to grasp video moments
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d07bae86-25a3-4507-8e8d-368a00307c3e · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf50a6a8-49b9-423d-8fdc-8038fe203b48 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Llava-next: A strong zero-shot video understanding model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5acad12b-786e-4b2f-8a64-fdb290a0eb4f · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Long Context Transfer from Language to Vision
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e24842-51ba-493b-a40a-7b8fb6c75356 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Longvila: Scaling long-context visual language models for long videos
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b304a8e9-c627-465c-9a19-19f0c90d267f · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3b625c-bd5b-480b-a776-206f61a13d6c · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Llama-vid: An image is worth 2 tokens in large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 726c8d2c-69de-496f-aee9-e2bd67324906 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Moviechat: From dense token to sparse memory for long video understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c40a6ee4-98fa-4b30-ab10-f5b05db0c5ce · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2a6ebf20-8279-4bbd-8d68-0b9d28fd76a4 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation KeyVideoLLM: Towards Large-scale Video Keyframe Selection
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e63a0c8-c49d-426e-baef-ceea104c4611 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-rag: Visually-aligned retrieval-augmented long video comprehension
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67bcc6be-f8be-4b10-abf2-696bed1ec14f · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Rethinking temporal context in video-qa: A comprehensive study of single-frame static bias
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6146fb5-1615-40c9-80fd-8e44bd6af1e8 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video question answering via gradually refined attention over appearance and motion
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8bfbbff2-ca46-4202-a7ab-f0823f8bd84f · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 115511c5-ee09-431a-b8d2-6530b699b622 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79f51786-fda2-4be9-a433-cb49271c43a3 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f3a0d8f-19d9-449f-aa85-32036b29bd5c · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Temp- compass: Do video llms really understand videos? In ACL (Findings), 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 82eb4c28-8f4a-45c8-bbde-e88a1f6e0611 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c74765c-bcd1-4bc9-9998-0b3e9b2878ee · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Topa: Extending large language models for video understanding via text-only pre-alignment
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c598d106-8ba8-4aab-8c07-e3574204b94b · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Temporal reasoning transfer from text to video
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 89237e06-3f77-4a33-a5ee-f309802bd30e · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Learning transferable visual models from natural language supervision
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58c98233-b42c-4319-a34f-3e942d50f9c1 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7215ecba-6f80-41ee-9cdd-6ff1eb9251a7 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Monkey: Image resolution and text label are important things for large multi-modal models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2dca21c6-f390-4882-9d9d-4c56fd005234 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977cb2b0-9c0d-45f2-b100-444cd6a4cfc0 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Sharegemini: Scaling up video caption data for multimodal large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 76a7b6ef-69ac-45e3-98e5-64a13013db83 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Frozen in time: A joint video and image encoder for end-to- end retrieval
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f5cf478f-5fd6-44a6-b3a9-11f40d84a00e · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a07fd844-0e75-440f-8f3c-bf8e6e3757ac · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Token merging: Your vit but faster
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5451fef-f59f-46e4-9aa2-93c320acd294 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Activitynet: A large-scale video benchmark for human activity understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b46a8eb-1676-40ce-bbd5-78aa5a2384e2 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Lima: Less is more for alignment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 22002286-d833-4b76-8495-e5c0b9a76887 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation What matters in training a gpt4-style language model with multimodal inputs? In NAACL, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75c06f77-0aa4-459f-8d02-6d88fa4a8072 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cebed98c-4814-4ccb-8672-11b932082178 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Wildchat: 1m chatgpt interaction logs in the wild
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6d26911-c6a0-4260-9bb1-e8b3e4f789cb · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Gpt-4v(ision) system card
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 37e6f01e-a358-49de-a7f8-e126d4ef68b6 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Introducing claude 3.5 sonnet
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f5ba2e3b-bcea-446b-bc0d-487f2723449b · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Hello gpt-4o
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1425fc8a-67e5-41a1-8b22-a7054dc8f63f · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Gemini: A Family of Highly Capable Multimodal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a524375b-aa76-453c-8836-06d677f80dee · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 205c701b-4378-4652-827e-40a68f618c15 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aabfc9e4-6f09-4dee-86bc-69cea001412c · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f87d359-c568-489b-97c6-a61df3107b10 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation MLVU: Benchmarking Multi-task Long Video Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea2620b-cc6e-43ca-a5e9-38ab47faf8d9 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Introducing llama 3.1: Our most capable models to date
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 43247aff-68cf-44ff-8279-4ded7179045e · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Answer with the option’s letter from the given choices directly
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c1e11d8-d880-4f2b-84ee-9656d872e886 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 576604f9-e7f2-4a53-b0cb-53139af6f489 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bcf4b791-8b11-4e26-92a7-84fb535fe895 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c0aca438-5fb8-41f2-928e-fcd5d0cc2185 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f68fd608-d6c1-4088-bc66-1b7dd0149384 · outbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Therefore, the answer is A
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1803023a-38b0-4f32-97fc-5a1faeda3cd7 · inbound
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
Reference 134
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf517a71-3707-41cd-8b24-fd8c2f92ba05 · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
Reference 151
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd186912-810c-4e4b-a4bb-b44001cb2a61 · inbound
LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.