Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T23:20:32.330351Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 190 outbound references and 100 inbound Pith citation observations for arXiv:2410.02713.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T23:20:32.330351Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:29:51.267504Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T21:36:34.348434Z
100 of 190 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4301052c-2903-45d7-9e0d-cba4e2288d59 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66b0dd14-3493-4daf-8f35-012140a2a899 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Localizing moments in video with natural language
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 655bad6a-4350-4479-b131-88e785e2e9f5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Localizing moments in video with natural language
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 23827f82-95b0-4124-9dea-f53285e95fac · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 76d02b7a-a755-42fe-b659-ce2066e3eec5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Activitynet: A large-scale video benchmark for human activity understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbae0af6-2579-407e-a892-124f161217a3 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Collecting highly parallel data for paraphrase evaluation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 502ca65c-a6f0-457e-9304-9af098457771 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 00a4b2d9-fc92-42a8-a72f-f26fb623fc9a · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Slowfast networks for video recognition
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31902d00-b60d-4d16-83c7-29f32c0fd786 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data something something
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73279b0b-7784-475a-835f-2be857f8aed9 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Ego4d: Around the world in 3,000 hours of egocentric video
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02bbcd81-0380-4091-8d7b-d1525f9018fd · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Agqa: A benchmark for compositional spatio-temporal reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b0bd1c8-7485-45f5-980a-f4846de20a99 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 230286e9-5613-4870-a364-f26a277a7dd8 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Less is more: Clipbert for video-and-language learning via sparse sampling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06030416-0d62-4587-91a3-5c3df9d660e8 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Llava-next: What else influences visual instruction tuning beyond data?, May 2024 a
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c17b7a02-819e-4e02-a2c8-aaa5535a4ee5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024 b
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3d63041-360f-4226-9e92-71874183a63f · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Multimodal foundation models: From specialists to general-purpose assistants
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f9f5c995-439f-48ae-af2b-edd84930ef81 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 945090e3-98eb-48ff-a09f-1ec94a9f1394 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data VideoChat: Chat-Centric Video Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1a439b78-4e3e-4342-99f1-27e05dc28378 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Vila: On pre-training for visual language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d78ea22e-5e74-4614-be36-1dd0f7fc7107 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Visual instruction tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17e17dbd-6315-4b2c-9073-1492eeaf970b · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Video detail caption
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b3105ae8-a139-46f9-b45c-ebed021ea7ff · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 759fdf4c-38fb-41b6-a66e-290f5ae01270 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 471d5b74-6463-4ae3-9a81-01fb4e080289 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data How T o100 M : L earning a T ext- V ideo E mbedding by W atching H undred M illion N arrated V ideo C lips
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3bd7134-00d2-47b8-af6c-0d0d1a4c37fd · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85fd65e4-af8a-4e1c-b389-d5de12345a19 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Hello gpt-4o
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 659eb492-fffb-4a12-9c03-e000901435ef · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Perception test: A diagnostic benchmark for multimodal video models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c185b1b6-8ce5-490a-a6ad-bcf64fd3ec7d · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Learning transferable visual models from natural language supervision
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0510af63-23de-47e2-a2ab-9d33ba745e87 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6828b804-2371-4c0a-92bc-c2a141f862b8 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data A dataset for movie description
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1044fd9b-52f3-431a-aa5c-7a565aed18b4 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Annotating objects and relations in user-generated videos
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 881e83fa-9b05-4b89-ae6f-f6e0a7292f14 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Hollywood in homes: Crowdsourcing data collection for activity understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0be139c6-460b-4b87-ae5c-c34e807f64df · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95d1cba9-dfb0-440c-9a95-63c8b7808591 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Internvid: A large-scale video-text dataset for multimodal understanding and generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13c7293b-8856-4dbe-af2f-33d45d80e6b2 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e48d6f90-c0d8-4c76-9fae-0e74d949dfa9 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Next-qa: Next phase of question-answering to explaining temporal actions
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b10e2c7-aa1d-43e2-a878-7e361e573fd0 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Video question answering via gradually refined attention over appearance and motion
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7650c80a-2a37-4b39-94bb-361ea1bdb894 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Msr-vtt: A large video description dataset for bridging video and language
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 247feb16-86fd-47c0-a678-b8d9db22fd23 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Advancing high-resolution video-language representation with large-scale video transcriptions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8076cbd7-5638-48d4-b610-64ead95d16c3 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d876a09d-f8a6-44ce-9a46-70e84343ad8f · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Social-iq: A question answering benchmark for artificial social intelligence
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d2c33ba-49f1-4ca1-86c9-9915819579b5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Merlot: Multimodal neural script knowledge models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 978b6f3c-4301-4ef1-8d03-d9067de561a5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Sigmoid loss for language image pre-training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc265d68-56a3-4e44-be49-5edbc41c42b3 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Direct preference optimization of video large multimodal models from language model reward, 2024 d
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 989f9f34-5446-4f2a-bed0-7bcefbbe4e4f · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Llava-next: A strong zero-shot video understanding model, April 2024 e
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 038e60f8-da92-414b-9ef3-356f2041dc95 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ac8ac4ed-4906-4c74-9462-d53e945408da · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment, 2023 a
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 815f9ce0-8cdd-4806-a4f9-961be6afd4e3 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Visual Prompt Tuning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 851caa27-a1fa-421d-9265-2d9f1a246e0b · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12e5c4b7-7708-480b-a5db-e7190070fa95 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data LoRA: Low-Rank Adaptation of Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a7721538-11c9-406b-af25-5b10904daabe · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Towards a Unified View of Parameter-Efficient Transfer Learning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bf8445e4-ab2f-4cda-869f-ba74c271fd6a · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Factual Probing Is [MASK]: Learning vs. Learning to Recall
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8f9d25e-f5a3-4096-b93e-d1a79eb07e64 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13ac991c-6321-4026-bd09-d1972bb5af00 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Advances in neural information processing systems (NeuIPS) , volume=
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 65e560d1-d32c-4388-9613-7d8015d3bbf0 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 548183ac-c428-4bac-9a42-8f1c91d45264 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Advances in neural information processing systems (NeuIPS) , volume=
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99b7857a-520c-418c-a08c-269cfde97a3f · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Prefix-Tuning: Optimizing Continuous Prompts for Generation
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c57b2c5c-6d40-46db-8a70-2d7dd357c908 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0dabb73c-8813-4f71-a429-74a23878f851 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f1cce10b-33e4-4261-92e6-125f6063e2b1 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data International Journal of Computer Vision (IJCV) , year=
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed9d954f-8fea-4fa6-a205-fd88e978b7b9 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b1a8a24-e2bb-41d5-82a3-04139c63d618 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b65256a4-561d-4e0a-8e13-4fbf232632f5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation af0df508-c6f1-432f-a240-b6d7eadb743b · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Neural Architecture Search with Reinforcement Learning
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44187c4e-6bee-4547-b0c5-c33f95d3041a · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b8cfa88-2e5f-4dc9-9118-c979bc5ed0a1 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , year=
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cff8ec30-f885-4feb-8412-25c83ffc151e · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da59bf03-ebb7-48cc-96a0-65743ee8bead · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data International conference on machine learning (ICML) , pages=
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d6537fa-6fc8-4b2e-b057-07ce0c2cb512 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data DARTS: Differentiable Architecture Search
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f4feed9b-5d73-4ac1-a7cb-507cc782c787 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2488925-6126-4f6d-8399-b4319aca5bcc · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7cbb515c-9b65-4e04-9fc8-a7d8d0718e9e · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb77a37b-9a32-43d9-b740-b43eeaf694e8 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data DeepMind Lab
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 079d7a4b-b31b-487a-a4f2-96cb863087f5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation becd0eca-edd6-48e1-a4cb-0ac702bedc2c · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 386789d2-454e-4bc1-864d-34a9202ebb3e · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data A Generalist Agent
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e151b766-d10c-4e28-8602-7d59752067f5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , volume=
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6151fcf9-2484-4cc6-99a4-9ea3ac4333f4 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data European conference on computer vision (ECCV) , pages=
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 15138065-4b00-4e76-bf61-966ddadb9054 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Domain Generalization: A Survey
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3f589d9b-43fd-44f4-b4d8-c27517ce5713 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Florence: A New Foundation Model for Computer Vision
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fa01dc7-494e-42cb-8ae2-601bb5a9a1f1 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine Synergy
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2031154a-0a67-4953-9281-e3ff212fcf98 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data The International Journal of Robotics Research , volume=
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6468b77b-5232-4c11-b814-83c8a6f35115 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data European conference on computer vision (ECCV) , pages=
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6699bcf1-ff8e-4feb-b031-aefe71b4b598 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , volume=
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d4769f5-2af0-4189-a26e-0905173bd6c6 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE international conference on computer vision workshops , pages=
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 177e2e96-0336-49b6-b186-406a605a0d42 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1bc02244-54b3-42ed-bc46-4ebb6ea780d4 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Fine-Grained Visual Classification of Aircraft
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 11037149-2c26-4f9c-ac01-070d7d6b68e5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fe8658e-da36-4b0b-b222-812b1063b19f · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Advances in Neural Information Processing Systems (NeuIPS) , volume=
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c12945ab-129a-4116-8b49-11ef163ed4ad · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 492ef995-fcaf-449d-847e-1f697b9d95e2 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c4c79ab6-9f45-4398-a24e-defaa9e5fae1 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ecb94d70-e088-48f9-85ff-aca8c816f96a · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da8c7e21-26a4-44d9-8fcd-85d1c68b5489 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ba6f137f-115a-4c0f-9c18-fd65ee659570 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Masked Autoencoders Are Scalable Vision Learners
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T15:38:38.318698+00:00.
Observation 8eab20aa-1df1-4fff-a0f6-48c1e6598ab5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data 2009 , publisher=
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 823fbbd0-89ce-4027-9d77-4e4df0d39b4b · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data conference on computer vision and pattern recognition workshop , pages=
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 868691b8-0914-4978-bde2-a6ffbf9a5ed5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Cimpoi and S
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation db0b8a0e-8d0b-4442-a52f-fc3ec12764d5 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1cc1e7cb-037b-4b85-b3c4-d45c773c5660 · outbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data 2010 IEEE computer society conference on computer vision and pattern recognition , pages=
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a3218de-1b04-4090-986d-e0b103dd39fa · inbound
NVILA: Efficient Frontier Visual Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1e36f211-c965-4223-a32c-3496cfddc48c · inbound
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 976db303-747f-4c80-bb8d-17d02eaadee7 · inbound
VidCtx: Context-aware Video Question Answering with Image Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84022ade-c828-472f-a4f0-b9028aa05dfa · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10a626aa-374f-43e8-b0d1-20146cd3af0c · inbound
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c991fea-0792-4d41-a4f4-08a52fb229d2 · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2daef667-331a-46dd-8bfb-59fbce9c1109 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b5267eb6-df57-4449-8bc5-3b6215c980dc · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86efaa9-45dc-49a9-bece-0435d3fd93e1 · inbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf57c0c-21c5-4ba8-9e87-41ac7bf780c8 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e5a4491-57c8-48bc-8d32-d175db17ac2d · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10ba334a-bb4c-4f00-b439-b1f38cf09b18 · inbound
Temporal Preference Optimization for Long-Form Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce12156-d8cc-4ea6-ba88-ee60aab2a8c3 · inbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0aa3591-a22b-459d-80b9-f05cf52d2255 · inbound
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1d6fdb8-f008-4e9b-bf41-016f8e7b30e9 · inbound
HD-EPIC: A Highly-Detailed Egocentric Video Dataset LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a483b44-5e09-4120-b675-e950839ff857 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73da6a09-7f01-4c25-b393-89a6b3c6cfe2 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ed3bcc-cb45-4f90-b7e0-6f2b455b14b6 · inbound
Unified Reward Model for Multimodal Understanding and Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 28953e7a-4738-41e5-8848-730399223951 · inbound
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b57506b3-c9fb-473e-94f1-db33721ee035 · inbound
SmolVLM: Redefining small and efficient multimodal models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3caa9af6-8bcb-4af9-a24a-5ae85ccc00ba · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aece33d9-7349-46da-b737-66bdd2dbf49b · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f6e468-d31e-4e85-a9f6-ac9bc1835124 · inbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75ba5317-b5ea-4f21-b12a-517e7916add4 · inbound
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eacf7cc-13ab-4b30-9181-67af9725b7e1 · inbound
Clapper: Compact Learning and Video Representation in VLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331ed5d7-2084-415c-bfeb-bc831b564724 · inbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08d2a24f-71fa-421f-afa5-f205f726ef01 · inbound
Inference Compute-Optimal Video Vision Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1acf10f4-d309-4cd2-9127-00201219f46b · inbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a25c67-145d-45dc-be12-5c0ffaeba027 · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3a496cd-ad53-4c95-84c7-41cd139ded15 · inbound
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a20b4b-0c71-4321-a18f-87146e3c31f0 · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3976c4-30c8-426f-a89b-b2cd0d54fcb2 · inbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 123279b9-48f2-448e-98d8-9f2e0c7aed8f · inbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ee6a8e-3681-43ed-8861-7810bfeeed6b · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5065b583-6453-4089-8ddf-b0e1f36b6f9d · inbound
Fostering Video Reasoning via Next-Event Prediction LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 120b31ac-4054-4861-882b-b452c3ddd294 · inbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604cc1a4-c2bd-4554-9788-ffd5f7ee5bf6 · inbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c641c0-38e4-4aeb-a236-fcdb37168342 · inbound
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7b318f2c-c5ba-4ee7-b673-9e291a49eed9 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0051c6d8-b97b-495d-8d5f-7257e36b7dad · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b30ebedc-b77a-4463-b3a6-ef8c7a773f6d · inbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5762f209-5a99-4b11-9c1c-55cd14564bbe · inbound
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6d8b62-bb23-49b7-b3fc-4b9f9fd99e74 · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9410b52a-a52a-4521-9fcd-8ae683fdb1c7 · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759fd701-e7b1-4f1a-86f2-762e8626f470 · inbound
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ab9372-e3a4-457d-8ffd-1ef64bb7114f · inbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3078cda2-ac9b-406d-9bb7-d6b3f3c12bea · inbound
Is Extending Modality The Right Path Towards Omni-Modality? LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 500a1ece-fc59-4481-a16d-988122648280 · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e09aaaa5-b745-416b-afaa-963f8e4c80b2 · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d568b3-97ea-4d33-bb01-cfcee5cc0e46 · inbound
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c26d1035-6bcc-4792-800a-cc0fad81dce9 · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc43a348-2f3f-474c-a8be-147f491a82d5 · inbound
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 94c33dc7-c577-4677-a974-ecd3400ac393 · inbound
Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b014608-1e4e-4c3d-bae1-f2fe67a62cb5 · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6f15c9d-18d7-4e73-9bcc-4a2501bf2bc6 · inbound
How Important are Videos for Training Video LLMs? LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac43119-00d2-4316-80ae-1ac7c5cda771 · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a25c88f-a587-46d3-8cd0-b7f0467bacc4 · inbound
Audio-Sync Video Generation with Multi-Stream Temporal Control LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d0b0a7-8220-479c-b170-bdf75b947399 · inbound
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 28c45976-fa18-49ff-ad2e-50198112175c · inbound
GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9abb6d21-375d-41a4-a060-60b64bdc1499 · inbound
Show-o2: Improved Native Unified Multimodal Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d883cd6-d552-4a7f-a565-86d5b4dc3436 · inbound
MMSearch-R1: Incentivizing LMMs to Search LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c307f2e-7cdd-4c08-b992-c4fc22cea03a · inbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6e43932-032a-497b-a74f-48c42b77f00f · inbound
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113072f7-b149-48b4-98a8-8e9f4e1254e5 · inbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f8ab5d-e1e0-4dc3-a967-4662dd2d825b · inbound
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1e3cba-0d00-4e50-a1e7-d1699e39c8d3 · inbound
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80aecfca-840e-43ab-83a4-dc31abeba65e · inbound
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f86124a1-b5ac-4052-a9e4-c56c69400990 · inbound
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd7c2631-a2ca-43c3-be08-e90da971380a · inbound
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c59fd0ac-0659-49c7-88c2-f56a3efe1277 · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904869a4-664e-4648-b434-788a107d910f · inbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da2a2cfb-ce36-4457-8938-c5ab46742040 · inbound
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 393861fd-ebd2-4293-882b-fc6cb6c5ca62 · inbound
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f43ed20-4d5f-404e-a453-983c4777fad8 · inbound
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cb2719-8b81-48f7-adeb-ca73373abd9c · inbound
Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 542037b8-ea75-435c-8d9c-f7c41a471fe3 · inbound
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc61fe64-0088-4c2f-b95f-349400158007 · inbound
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8afe1ff4-893c-4b1b-a864-e129ce931309 · inbound
IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed91865-da65-4d6e-a1db-a804263ec9bb · inbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a81c394-c805-4be3-bdd9-a4156c2916be · inbound
Training-Free Multimodal Large Language Model Orchestration LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5df480c-2700-4203-8757-002aa20bbbf0 · inbound
Training-Free Multimodal Large Language Model Orchestration LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 90743981-e64d-4adc-b237-3e7eea28453e · inbound
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22aecef3-2f8e-4ee9-b934-8a0d42a3adde · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03a14a9-e0c2-4e9f-8a77-f03e31f886ec · inbound
DiscussLLM: Teaching Large Language Models When to Speak LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 495fcdb7-43d6-4704-b7bb-7734549ef176 · inbound
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96eb1452-6673-46ff-b315-def636bb0db6 · inbound
Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2394c300-a708-4382-b347-33c74c3f59ff · inbound
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61073f3f-8553-4916-b6cb-ec7067afd79e · inbound
CAViAR: Critic-Augmented Video Agentic Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc89babc-b975-4d58-9f7b-b9b38196ef01 · inbound
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e48f6b-027d-4227-996f-575df6d5d810 · inbound
AdsQA: Towards Advertisement Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6b7a38-761c-4dc1-80c4-79e67f6e1854 · inbound
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eaca711-a83a-4a5a-a29c-ad30b18f7463 · inbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8bfc4c-c46f-4564-aac7-aeb8fe4723b5 · inbound
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3f1a1a3e-e5ca-41fa-9fe4-c25a10aaeab6 · inbound
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a320d18-c23d-4e6d-8f39-76695c00133f · inbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 117
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d30f68c-5231-4f13-94d8-2390bf3c4d29 · inbound
Vision-Language Memory for Spatial Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfb4cde-dffd-4ca0-800e-78e26826345d · inbound
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67934772-bf1f-4c7c-b7cf-dd600c3842b5 · inbound
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 129f71a8-9fbb-4def-97d4-53c46bd94599 · inbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd94a8c5-6131-4c9a-85df-3dc92179f7b5 · inbound
Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.