Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:31:10.814698Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2411.16201.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:31:10.814698Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eca11ac4-b118-4b84-b791-5d36912e82cd · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22bd658a-6679-4d56-b378-404ad7ada325 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ace32fdb-527b-4fd2-946b-bb00af8afd99 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd396c2f-d606-4388-97d0-7a2753996647 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcb0f39c-00e6-47a1-8c81-6bcd4cf07bb0 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2222748-57a1-49c7-8b5f-2c75b0e93df1 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853a2a25-2629-437f-8fe0-cda44967dfa9 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08710526-0657-4472-8222-192f6e320dff · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72d6367-31f0-4582-9c47-499e9a7b9109 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b977978-c600-4e54-9f9e-2b06636631bb · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models PaLM-E: An Embodied Multimodal Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ad9708-b22c-4188-9ed3-a8cfdd831e80 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0eebc09-dac3-4705-9906-85b88d1675ab · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Evaluating Object Hallucination in Large Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c70d2ac-9b1d-4da8-b970-aff52c43e887 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e687df-5993-40e2-9eaa-6a61212db21c · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models GPT-4o System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0118922d-762f-4fdc-bcff-11827726ed82 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Self-Rewarding Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6897a2d7-2a33-4c9b-8678-c23035d08cca · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32855b2e-1994-4440-8ade-2291b1ab1ee9 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Statistical Rejection Sampling Improves Preference Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c071df3-41bb-493c-a256-f81fcce82abc · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf776f6-ceca-4d22-9551-1afc8ad641ff · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9332e79f-d7ce-46ff-82ee-da5bfda82e81 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Silkie: Preference Distillation for Large Visual Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e5fbeb-dd43-4e40-a0fe-5cad899f7f54 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88b3a2d7-0385-44ec-8608-466cfb54084c · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Rank analysis of incomplete block designs: I
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a7dfcbb-7853-45c8-8df7-051d97f7878c · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Model Extrapolation Expedites Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3a7c618-105f-479a-b8e0-f5dd33172327 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Activitynet: A large-scale video benchmark for human activity understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191dc859-7ff7-4a0e-998f-720dd03601ae · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b80de64-9711-4c68-befb-f88f3430b5dc · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation da689573-6cbe-4ae8-b9ce-9f1e8505ddf7 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Llava-next: A strong zero-shot video understanding model, April 2024
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea46d369-dad9-4098-a159-5b1315716da8 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0c5fab-4a56-46fa-ad5d-29576a414594 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Learning transferable visual models from natural language supervision
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae951733-419c-4de0-8bee-6c3d47550f85 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6da3eb-9320-488a-b35a-d615b36ee3fd · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Understanding Reference Policies in Direct Preference Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c1f849-6227-4133-b55c-5fdca36ea965 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Averaging Weights Leads to Wider Optima and Better Generalization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b0144e-f424-4ec9-89cc-34b2917dd750 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Mitigating the alignment tax of rlhf
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d9f3d4ec-cfa5-40b9-8788-4841e330076c · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Spurious Feature Diversification Improves Out-of-distribution Generalization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fed1565-deae-4028-ac33-0bf978d6e7f1 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08502e6c-3882-4202-b606-44fdd5c2b591 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Msr-vtt: A large video description dataset for bridging video and language
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4448ab51-afb0-46ab-a2be-246f99138b23 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Collecting highly parallel data for paraphrase evaluation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b7491e61-2905-4ed1-9a7f-8e0f7626bc67 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Tgif: A new dataset and benchmark on animated gif description
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9000a92-d14b-4285-8f02-7219c1b1f778 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models something something
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b5a80dc-c49c-460d-96ba-9fa08f53d769 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Eva: Exploring the limits of masked visual representation learning at scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4499c109-e71c-4b1a-a225-533eb47d1388 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Vision transformer with quadrangle attention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dc08f4c8-2f55-45f4-801d-1e6d1db9056e · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea81f325-c03c-41fb-8884-478455918ba4 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Decoupled Weight Decay Regularization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c2e256-00f5-4b48-bbb5-c1a4675cc820 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb050356-f8a8-4c6f-96bb-d281a68530e1 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 43217553-0d2e-4953-9722-a874f9a72be0 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Consider the following criteria for evaluation: -**Relevance**:Evaluate how relevant the model's predicted answer is to the question asked
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ddb17ddb-16c1-414f-9989-3f1d82ec9922 · outbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Temporal Consistency:Does the answer appropriately reflect the temporal progression and events of the video?3
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.