Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T05:08:20.197862Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.28509.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T05:08:20.197862Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9e5ef122-d5b2-4e1c-ac12-b90a3088b2d5 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc44a2e2-fbe3-4938-89cb-ed8987f1fb75 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001d8a6d-16d8-4517-a577-0e49064afd8d · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb30ee5-66a8-43c3-8085-14267234ec98 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Multi-subject open-set personalization in video generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96553f7-95ff-41d1-97b1-a146976fc099 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning VidCapBench: A compre- hensive benchmark of video captioning for controllable text- to-video generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7439c6ac-9db8-49f3-b763-50921b4a1af2 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c3ba9da-2809-43e7-ab9e-ef0315868f10 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning MA- GREF: Masked guidance for any-reference video generation with subject disentanglement
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d25759c-5f23-49ae-9276-95bc2c367031 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e930ff39-68a4-4b18-a102-22993e62df74 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video Re- Cap: Recursive captioning of hour-long videos
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 530a17a1-e33c-4a7a-99f5-0ece07dd6da9 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Movie Weaver: Tuning-free multi-concept video personalization with anchored prompts
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44944426-2480-4b27-9826-53eead274374 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video-LLaV A: Learning united visual repre- sentation by alignment before projection
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 649bffd7-b00c-4762-8276-9ebe87696f6c · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning SwinBERT: End-to-end transformers with sparse attention for video cap- tioning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381b38d9-d150-4018-8fed-6a2b508ea59d · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936c2fd8-3e98-48ca-8f67-3c4ceecf8b7c · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f69209-b7f3-4682-af80-3f0d9950324d · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68f4953-7a95-42ca-948f-95555a8fc5cf · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615d3bc5-6a95-4e2d-b61d-a1eebcacc36c · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e5cb05-37fb-4071-bc4d-d9d543497470 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Mavors: Multi-granularity video representation for multimodal large language model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe5d3fb-b499-4e26-bc43-6dd95317afaf · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Mme-videoocr: Evaluating ocr-based capabilities of multimodal llms in video scenarios.Advances in Neural Information Processing Systems, 38, 2026
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e622f811-a402-40e9-a67e-af992d2c6438 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning MV-S2V: Multi-View Subject-Consistent Video Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6eb2ebc-b5d9-4832-800d-c3cb37952d86 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning VideoBERT: A joint model for video and language representation learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e950d10-d95c-4277-b5ac-2d157881787b · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1320625-5011-40a3-83c6-97bb1683d67f · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e088145-2503-4500-af3d-d9708d3888b2 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Qwen3.6-35B-A3B: Agentic coding power, now open to all, 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c040a0-a24c-46dd-bfd9-f94d9c9b30a2 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Qwen3.6-27B: Flagship-level coding in a 27B dense model, 2026
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9643093e-bfed-4251-9083-b15b50c20ec5 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504f917d-82c7-4a01-a168-bbe21f75d59e · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Monet: Reasoning in latent visual space beyond images and language.arXiv preprint arXiv:2511.21395, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb0bcbb-9b50-40f2-9e0a-9b548a1b3769 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8e0b6b-1218-4c10-b85e-65931206b4b5 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning VaTeX: A large-scale, high- quality multilingual dataset for video-and-language research
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832921ac-748d-4087-8db4-f2144980eff4 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Intern- Video2: Scaling foundation models for multimodal video understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5dae559-0b17-4e16-a728-7434a8665c0f · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a12be9-d250-4a05-8de4-ccfca55f846b · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning MiMo-VL Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75d6403e-93ba-456f-aa11-c6fd1c55050f · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning LumosX: Relate any identities with their attributes for personalized video generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe59067-2732-403c-a53a-3905d0203d1b · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning MSR-VTT: A large video description dataset for bridging video and lan- guage
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff35ee6-1a97-4754-850f-c57dd0aea30f · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Progress-aware video frame captioning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d57d85-9cea-4001-bb59-f7aac00a37b4 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Kwai Keye-VL 1.5 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdba910-7052-43bc-a0ce-292612471ffe · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58e3c8b7-1217-441c-92a7-d2575c5889ea · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26dda30b-25a9-4516-8989-979dba28bf11 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1d58b1-77ce-4a71-b977-9a7e07eb5853 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning VCapsBench: A large-scale fine-grained benchmark for video caption quality evaluation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ef230b-d319-41b9-8b12-9628295b96e0 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9799e92a-b957-4f5b-84e3-496b094fb6fb · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Debiasing multimodal large language models via penal- ization of language priors
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3b8fe5-73e7-4103-9538-2a8d605865c7 · outbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning from <Image_N>
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.