Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:45:10.619719Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 25 inbound Pith citation observations for arXiv:2507.01949.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:45:10.619719Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:43:52.914738Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T06:34:41.887069Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa035023-0b60-4337-812d-97d4bfb14777 · outbound
Kwai Keye-VL Technical Report Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00049c26-0138-46dd-b2b8-e54e677f408a · outbound
Kwai Keye-VL Technical Report Silent Data Corruptions at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 716d0469-a98c-4a3e-bcc1-54772dad37b5 · outbound
Kwai Keye-VL Technical Report An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d02b21c7-e89c-4c32-8c30-562d73b9d075 · outbound
Kwai Keye-VL Technical Report Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e894e23-6a02-4cc3-b213-d15f08c5d49d · outbound
Kwai Keye-VL Technical Report VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8194bf4b-eb1b-4054-9f14-338be4474a75 · outbound
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa497178-df82-495f-98b7-c896c5ffffa8 · outbound
Kwai Keye-VL Technical Report Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac710382-fece-46a8-8d34-9745e74abd7e · outbound
Kwai Keye-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0949158-1432-48a1-8008-fccfa2df428c · outbound
Kwai Keye-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fb6e9b-cda6-458a-b94e-7628f175970b · outbound
Kwai Keye-VL Technical Report Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46baefde-b270-4f83-bc5e-a32daa32ac28 · outbound
Kwai Keye-VL Technical Report ReferItGame: Referring to objects in photographs of natural scenes
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e21e9a31-6b36-4a6f-90ac-6cc483401ae9 · outbound
Kwai Keye-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6acbc47d-b25b-403a-95df-24f4be8505bb · outbound
Kwai Keye-VL Technical Report DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04dde608-7a4e-4db6-9a65-e5d8367a23ee · outbound
Kwai Keye-VL Technical Report Model Merging in Pre-training of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65153693-4adc-4351-a78e-412e08ebf6d2 · outbound
Kwai Keye-VL Technical Report Microsoft coco: Common objects in context
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1b6e616f-8bed-4505-b16b-a1568f75d97a · outbound
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34afaad0-0b3f-4547-a0ae-8aa6e5659f28 · outbound
Kwai Keye-VL Technical Report VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff3c3d0-80d0-40bb-9d61-d79f4b14c792 · outbound
Kwai Keye-VL Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2bbe7d-61f3-4cca-97ff-014e3b93bda0 · outbound
Kwai Keye-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9098439-5a61-4d74-aec9-75f4fab69c10 · outbound
Kwai Keye-VL Technical Report Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb6d1d28-9ee8-4bd0-a9c1-13e33cb6de4b · outbound
Kwai Keye-VL Technical Report We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f374187d-769d-4084-8ca9-cceac74b5a30 · outbound
Kwai Keye-VL Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3969843-e511-4422-85a7-14803d4a5744 · outbound
Kwai Keye-VL Technical Report Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74984ca9-6330-4310-8652-e903a9c7d26f · outbound
Kwai Keye-VL Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1907430b-df54-483a-a628-d402608b8833 · outbound
Kwai Keye-VL Technical Report Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7851bc2-5050-40d9-b16a-82833ddc17b7 · outbound
Kwai Keye-VL Technical Report OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 100de4f0-9f26-4ac2-aa83-2cf668567b9b · outbound
Kwai Keye-VL Technical Report Gemini: A Family of Highly Capable Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bafcd25-29f8-4585-bf20-6b413f5803a5 · outbound
Kwai Keye-VL Technical Report Gemini Robotics: Bringing AI into the Physical World
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b30d42-d9b3-4dd6-9872-dddba4b0effe · outbound
Kwai Keye-VL Technical Report Toloka Visual Question Answering Benchmark
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c53f9e-2649-4dfc-86f2-095298371dd6 · outbound
Kwai Keye-VL Technical Report Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dacee50-f4b9-4936-9f49-36f001a0265f · outbound
Kwai Keye-VL Technical Report LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c8541a-7675-4cd5-9fc8-2cb2f4bb202a · outbound
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c4812d-0371-429d-bf7f-9cc53426813c · outbound
Kwai Keye-VL Technical Report MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a136efc8-662b-450c-90c5-22835e454eb3 · outbound
Kwai Keye-VL Technical Report Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39613d2d-a4c9-482a-b3a3-f5002c3b70f0 · outbound
Kwai Keye-VL Technical Report Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6ba017-1eab-41ca-a6ef-395dada80f49 · outbound
Kwai Keye-VL Technical Report DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f229ab41-a038-40bc-8d33-c8e51bbdda16 · outbound
Kwai Keye-VL Technical Report Onerec technical report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0999d1ab-1d5a-4155-9634-3b779e585c12 · outbound
Kwai Keye-VL Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da45d34-702a-41cc-9fa9-a0241d60580d · outbound
Kwai Keye-VL Technical Report DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1206d2-9301-4662-ba78-b61db910c575 · outbound
Kwai Keye-VL Technical Report likes" a video receives within a specific timeframe after being uploaded. Using a predetermined threshold, we classify videos into two categories:
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f494b009-3de5-49a9-838b-72a05811c8fe · outbound
Kwai Keye-VL Technical Report doi: 10.3115/v1/D14-1086
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60df3888-994a-4b31-af8a-eeda25178d8f · outbound
Kwai Keye-VL Technical Report URL https://doi.org/10.1007/ s11263-016-0981-7
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba8ed5f-3bb5-4fc0-a0b3-d68040702d1a · outbound
Kwai Keye-VL Technical Report Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b4da9b-480f-4490-9f66-1bc60373c293 · outbound
Kwai Keye-VL Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee409d9-1a98-41f7-9cf8-738df6aa7406 · outbound
Kwai Keye-VL Technical Report MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a2f78a4-44eb-401e-b382-cd342e824503 · outbound
Kwai Keye-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7762f938-92ce-4402-8d56-14b7a3f31c8a · outbound
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a8ab2c4-d19b-4943-8a45-02d4599946b5 · outbound
Kwai Keye-VL Technical Report Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 574deb97-d5bb-408c-a6d8-0403b5d74e9d · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Kwai Keye-VL Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e2cfcd-788a-4d8c-85b5-7d2114fb4bea · inbound
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Kwai Keye-VL Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1053a9d7-4770-40b1-bef5-dad6792d701f · inbound
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Kwai Keye-VL Technical Report
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e0d2528-1958-4e3f-b77f-1efc4f76eaf6 · inbound
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models Kwai Keye-VL Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd1f391-b638-4c21-87cd-1a87b6ef29b1 · inbound
Grounding Multilingual Multimodal LLMs With Cultural Knowledge Kwai Keye-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63dcac3-964e-4c56-b3bb-68c9f5b1a2a2 · inbound
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning Kwai Keye-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b15dedef-a39d-4731-a99c-989ebb43ac39 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Kwai Keye-VL Technical Report
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c6c3555-0a03-4eba-adc3-da888b32384c · inbound
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Kwai Keye-VL Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92fdacf7-f999-4184-8d6f-32b1d1fb72b0 · inbound
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach Kwai Keye-VL Technical Report
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6eda1611-25b6-45d4-8231-261136b9ec68 · inbound
CodePercept: Code-Grounded Visual STEM Perception for MLLMs Kwai Keye-VL Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d38d343-5adc-45cf-91d4-2f39702ba63f · inbound
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing Kwai Keye-VL Technical Report
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14d5b53b-7dab-4662-a12b-6e333a1b4f5e · inbound
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Kwai Keye-VL Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a97d56bc-32f5-4899-830a-8043487d9530 · inbound
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Kwai Keye-VL Technical Report
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45a011ee-1423-40a9-b718-111bf788ccd6 · inbound
Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts Kwai Keye-VL Technical Report
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6368d297-4abc-4852-a53d-fb2764703461 · inbound
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization Kwai Keye-VL Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d2d63b6b-68f5-43d4-b9c8-a755bfa8d64a · inbound
Swift Sampling: Selecting Temporal Surprises via Taylor Series Kwai Keye-VL Technical Report
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7ae43e5-0197-46aa-8c87-c70a43fe5a4b · inbound
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Kwai Keye-VL Technical Report
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b595bd2-1395-4842-ae56-7f9724efadc4 · inbound
Kwai Keye-VL-2.0 Technical Report Kwai Keye-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 52563904-d40b-4292-b9d6-3d1ef6285ac4 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Kwai Keye-VL Technical Report
Reference 278
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b83f04a-2d77-4d8d-af95-6ebbfa54e923 · inbound
HoloCount: A Holistic Visual Counting Benchmark for MLLMs Kwai Keye-VL Technical Report
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bf371730-be48-44f0-93b4-d31f8282461e · inbound
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Kwai Keye-VL Technical Report
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75f4b92-3907-478c-bcfe-b74fa1f47dc4 · inbound
Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Kwai Keye-VL Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02027fe2-c7a0-44ef-895c-b3b3600d5672 · inbound
LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos Kwai Keye-VL Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673afc8b-8945-4379-ab7e-8b4a898275e2 · inbound
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification Kwai Keye-VL Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853230a5-6d66-4020-abd8-1d0f9c3354fd · inbound
InSight-doc: Agentic Visual Perception for Long-Document Understanding Kwai Keye-VL Technical Report
Reference 128
Source-reported events for the cited work
Unavailable: canonical work link unavailable.