Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T15:22:31.310003Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2606.07639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T15:22:31.310003Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 534d687e-f1f2-4106-8aff-bed773e4c5be · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Visual Instruction Tuning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85416010-3373-4ff7-89de-6aa7fdbc2272 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b4edbaed-a809-41bb-8800-3283032b198b · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60fafe70-7195-453f-91c4-e1289c0ed508 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention VideoLLM-online: Online Video Large Language Model for Streaming Video
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb9648f3-4db2-45db-9ff8-f7030cfec6b2 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 16ce67ef-091d-4468-a7a9-0547ad079b39 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23e42558-8c68-4240-b219-8f04add94e01 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Qwen3-VL Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a584e931-e40a-462b-bbde-022ef2d05f03 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5aa55b11-e67e-4d3f-a6c2-e7d06f5661d2 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Flamingo: a Visual Language Model for Few-Shot Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e5e0311c-cb62-4fb8-a2ce-2eda2d3962da · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 447d702c-84d8-44c7-8173-805dae8e90c5 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f9ca443-8094-4aa5-aabb-c23d93a03b40 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d6e6827-dc1e-47e4-b382-c73b69027ce5 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Qwen2.5-VL Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d505dc38-4621-4f4f-9ce5-87459b303448 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6453e69f-deaf-40f8-811f-7358abe13d2d · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1ac21c0-ec49-4733-85b9-06aeb3e62c5e · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0701350a-dcd4-404e-ad4c-7204e4177d82 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention UnifiedVisual: A framework for constructing unified vision- language datasets.arXiv preprint arXiv:2509.14738, 2025.https://arxiv.org/abs/2509.14738
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c4d1470-6226-4ca6-b937-0bda961159d8 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9cb0997-5ccb-4b5f-89c3-df9657029bc5 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f77726b6-5cd0-4e28-82dd-7801f22a5c80 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention DecoupledProxyAlignment: Mitigatinglanguagepriorconflictfor multimodal alignment in MLLM.arXiv preprint arXiv:2509.14735, 2025.https://arxiv.org/abs/2509.14735
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4cb2791f-4d7a-4090-bdd8-ee8f967339d4 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67bf9386-ea11-4b0f-99dd-e994e8c4f025 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 371c0666-998b-4f2c-a4cc-9596ea918a4a · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c87cbb39-249b-4dbb-8d9e-978cbcacd063 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d7ff161-b207-4f19-91ff-0fc4fcd01313 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MMBench: Is Your Multi-modal Model an All-around Player?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3ff5ea9-bbf9-4409-89c5-16109cac203c · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6bdc7b07-3731-431c-8f1e-af0354dab525 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention RealWorldQA: A new benchmark for real-world multimodal understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 097fb2a1-fe27-49a9-b5c3-89cd04b0f600 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff3a96f7-ff3e-48c4-86c5-e012a0edf410 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73fe40b3-7ea2-4124-a1b0-8a33a705c4f6 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cedf9d78-4bc6-4fbc-b4b2-628871e7c466 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Evaluating Object Hallucination in Large Vision-Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84b003d3-ed6e-47e1-a363-22ad5f129c1f · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79c455c0-6e7c-4123-848c-947c636d6fd2 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac6e24e2-42c4-40c5-b845-0546edbf288e · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention A Diagram Is Worth A Dozen Images
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd534004-9d1b-4be6-bba6-129580f50e80 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8068ad65-ea6a-480c-9149-ba38d04dddc3 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Vision language models are blind: Failing to translate detailed visual features into words
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0a912cf-793d-4413-8cd4-8a37758a026f · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 105ca07b-9b98-47b6-b929-cba92bd9d553 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4d7ccb-aa1d-498d-b469-0e2de9b20030 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ceec0df-85e7-4ce4-8fab-9895482f6318 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a6cfca76-03ff-4d9c-8ff4-4126eae2738e · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MLVU: Benchmarking Multi-task Long Video Understanding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 63c5fe9e-b0b1-4fc7-ae31-c4533b20762f · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LVBench: An Extreme Long Video Understanding Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90702f37-b6a1-43ac-b4c9-8a60aa5978b5 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention TempCompass: Do Video LLMs Really Understand Videos?
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2a3e8e24-21ee-441d-843d-3676576839c0 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eb17668b-b0cc-4e76-afc5-6aa23abbb5d1 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 345cf5bf-79cf-4c81-a392-20a444a110de · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed4aedf6-16e5-45f9-8473-46938f380488 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Proximal Policy Optimization Algorithms
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7a305ea6-81a4-47e7-8f6a-0496b716f718 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1ab2ef0-4a6a-430f-bc6b-f17efd817e57 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 702f9e44-561d-495f-a5ef-ae30dc9bc281 · outbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.