Pith. sign in

Paper Citation Record · LEDGER

Kwai Keye-VL Technical Report

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 25 inbound Pith citation observations for arXiv:2507.01949.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01949 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:45:10.619719Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:43:52.914738Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T06:34:41.887069Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa035023-0b60-4337-812d-97d4bfb14777 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Kwai Keye-VL Technical Report Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.006947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.006947Z digest=sha256:8ab4078a48be09d70577f97455eeacbeac884ba97acb6fa1515294cd028c8c03

Observation 00049c26-0138-46dd-b2b8-e54e677f408a · outbound

This paper cites Silent Data Corruptions at Scale.

Kwai Keye-VL Technical Report Silent Data Corruptions at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.525781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.525781Z digest=sha256:829f1a0cd1fb704423c40ebdfe1a07626f8250e39c5fdcf766dec4db4cfd5167

Observation 716d0469-a98c-4a3e-bcc1-54772dad37b5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Kwai Keye-VL Technical Report An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.643045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.643045Z digest=sha256:137db9170bb5e1cb3ffbdb043a2f099e875435d435f655da5f0328450bb903a9

Observation d02b21c7-e89c-4c32-8c30-562d73b9d075 · outbound

This paper cites Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement.

Kwai Keye-VL Technical Report Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.817314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.817314Z digest=sha256:08f97a6960f1ee66c85ed61d0f0261704effb0c597fd840624111d865e829a72

Observation 3e894e23-6a02-4cc3-b213-d15f08c5d49d · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Kwai Keye-VL Technical Report VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.935971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.935971Z digest=sha256:91e7e0c9bf2e9b4ff9ad450af65c75083ca23ce94c77711b6ec4b61ba5cf8831

Observation 8194bf4b-eb1b-4054-9f14-338be4474a75 · outbound

This paper cites The Llama 3 Herd of Models.

Kwai Keye-VL Technical Report The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.043474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.043474Z digest=sha256:736a4e8bf7e2b5f52ff16cd4e43fcf6432993204b9f90a1cd1f0eefc188840b9

Observation aa497178-df82-495f-98b7-c896c5ffffa8 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

Kwai Keye-VL Technical Report Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.165354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.165354Z digest=sha256:7c98d0a63ca7dbbba7c7ba1c8fa433bac1076660a67143bde6befbc5b626b841

Observation ac710382-fece-46a8-8d34-9745e74abd7e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Kwai Keye-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.272357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.272357Z digest=sha256:bc38d39a19e221b628383a4ee72d8ae8f4649e97c158c5df3358de407941ca4c

Observation f0949158-1432-48a1-8008-fccfa2df428c · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Kwai Keye-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.401157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.401157Z digest=sha256:d66a421322724357587edcd532b9dbe41ef9b1a1d59495c8b587985b8267b6c7

Observation 68fb6e9b-cda6-458a-b94e-7628f175970b · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Kwai Keye-VL Technical Report Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.512122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.512122Z digest=sha256:b4d4c22b8c81ba2550ed65ca60b12350a62882cf13fc908371ca6f53f3bc4a60

Observation 46baefde-b270-4f83-bc5e-a32daa32ac28 · outbound

This paper cites ReferItGame: Referring to objects in photographs of natural scenes.

Kwai Keye-VL Technical Report ReferItGame: Referring to objects in photographs of natural scenes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.601283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.601283Z digest=sha256:ff0dba52afd8b2163f391082d46cfb222f47ee4fa0b2abaede5fdd56330de225

Observation e21e9a31-6b36-4a6f-90ac-6cc483401ae9 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Kwai Keye-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.955283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.955283Z digest=sha256:cf77bc9ef3d80dc4a5fab09be66beb09a64f1f784ccb7e847bd4582886ea4c12

Observation 6acbc47d-b25b-403a-95df-24f4be8505bb · outbound

This paper cites DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World.

Kwai Keye-VL Technical Report DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.054292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.054292Z digest=sha256:5c609807954e5d5724d8336655715ed44865102e500cad8871bd874a0028d1db

Observation 04dde608-7a4e-4db6-9a65-e5d8367a23ee · outbound

This paper cites Model Merging in Pre-training of Large Language Models.

Kwai Keye-VL Technical Report Model Merging in Pre-training of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.181516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.181516Z digest=sha256:f6177f266303dbc05dcb5ecce1812712a0a3398bfbc3b4ba9d2ebadf1d6c3e96

Observation 65153693-4adc-4351-a78e-412e08ebf6d2 · outbound

This paper cites Microsoft coco: Common objects in context.

Kwai Keye-VL Technical Report Microsoft coco: Common objects in context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:45:11.727828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T20:45:08.311935Z digest=sha256:1fe0ea35ad4e24b441384e9221302f92077c1662102190cec7c6dd83df0b14ef

Observation 1b6e616f-8bed-4505-b16b-a1568f75d97a · outbound

This paper cites DeepSeek-V3 Technical Report.

Kwai Keye-VL Technical Report DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.421413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.421413Z digest=sha256:8a7f4595bfc5358572d6b8dc21906d116201357228b0981d32df6c8f31fd5953

Observation 34afaad0-0b3f-4547-a0ae-8aa6e5659f28 · outbound

This paper cites VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform.

Kwai Keye-VL Technical Report VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.488038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.488038Z digest=sha256:6e55190cfb42fc9088ad2eb7fbb88ced75a9dab0e72d61e19d6cff3b10404716

Observation 3ff3c3d0-80d0-40bb-9d61-d79f4b14c792 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Kwai Keye-VL Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.557208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.557208Z digest=sha256:2e88acc08cd231bd2495b8f962c39da68cf5ebcb0bd1c799e9f3a2d7c7a233e8

Observation 4a2bbe7d-61f3-4cca-97ff-014e3b93bda0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Kwai Keye-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.608342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.608342Z digest=sha256:5e4a74d666ca349516e33a075300d5049056a533e70a0712e968fe3e772ba235

Observation f9098439-5a61-4d74-aec9-75f4fab69c10 · outbound

This paper cites Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms.

Kwai Keye-VL Technical Report Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.709183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.709183Z digest=sha256:940f5c7967fd33dfafa0c882f76b22472ee32292ac085ccd9513e0e15a9f7edf

Observation cb6d1d28-9ee8-4bd0-a9c1-13e33cb6de4b · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Kwai Keye-VL Technical Report We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.799775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.799775Z digest=sha256:1f3615b07d4cd0cb78d5a2fe9ecf7b633a0af96fa331f591879f161c16e34949

Observation f374187d-769d-4084-8ca9-cceac74b5a30 · outbound

This paper cites ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models.

Kwai Keye-VL Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.961173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.961173Z digest=sha256:4525021ab4e70559f15479478347e566e7ca92794ad4a57cdf3c1d94632951c7

Observation f3969843-e511-4422-85a7-14803d4a5744 · outbound

This paper cites an unresolved cited work.

Kwai Keye-VL Technical Report Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.052688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.052688Z digest=sha256:e33b9fa3ea865966cf482950d2f5811a9fb9551e2dde577c5d6cde04ab4456d7

Observation 74984ca9-6330-4310-8652-e903a9c7d26f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Kwai Keye-VL Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.105293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.105293Z digest=sha256:44eb65641c9c61148b057c69ef5ce484eae4f68769afa1c66d1aab552a19bad9

Observation 1907430b-df54-483a-a628-d402608b8833 · outbound

This paper cites Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.

Kwai Keye-VL Technical Report Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.227497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.227497Z digest=sha256:5264c939a359082495d3381098295263231a5bd3e6e7c99cc9ebcf03b5cd9ffc

Observation a7851bc2-5050-40d9-b16a-82833ddc17b7 · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Kwai Keye-VL Technical Report OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.310019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.310019Z digest=sha256:954a82b41fdfb9221c314431b312a004f3474176e5a02f445e33949002c58a07

Observation 100de4f0-9f26-4ac2-aa83-2cf668567b9b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Kwai Keye-VL Technical Report Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.382354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.382354Z digest=sha256:142a091c6b012c929b4b584d0117af41907ed7d3049b78980297a8ae17cfd859

Observation 1bafcd25-29f8-4585-bf20-6b413f5803a5 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Kwai Keye-VL Technical Report Gemini Robotics: Bringing AI into the Physical World

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.477913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.477913Z digest=sha256:208e27773e81c170134c62d545d6804a30982df6649653d00493ea429be57c71

Observation d3b30d42-d9b3-4dd6-9872-dddba4b0effe · outbound

This paper cites Toloka Visual Question Answering Benchmark.

Kwai Keye-VL Technical Report Toloka Visual Question Answering Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.544304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.544304Z digest=sha256:ac2f0388f042210a750742c6d3d9d9a69ac799e2ca6cb2ffd459d0079a5dd172

Observation e1c53f9e-2649-4dfc-86f2-095298371dd6 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Kwai Keye-VL Technical Report Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.636302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.636302Z digest=sha256:e6a5bf64be00909af39b03d97e96ba98324ab32a4feb92a30657c59982d8ac31

Observation 9dacee50-f4b9-4936-9f49-36f001a0265f · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Kwai Keye-VL Technical Report LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.744728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.744728Z digest=sha256:6088b07a4cde36703a0b8df4b3807d158e065e036721cf7b16d5c1b98f3a4634

Observation b2c8541a-7675-4cd5-9fc8-2cb2f4bb202a · outbound

This paper cites MiMo-VL Technical Report.

Kwai Keye-VL Technical Report MiMo-VL Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.848642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.848642Z digest=sha256:5ee28d455b2e95ed9193af1a73e3b31100ac304c175885a78e067deda3635bd4

Observation a2c4812d-0371-429d-bf7f-9cc53426813c · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Kwai Keye-VL Technical Report MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.040837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.040837Z digest=sha256:2fabbf22527eb879f78efa2f5c38e49c1d051befc21a1b90f83d3a480bf4895d

Observation a136efc8-662b-450c-90c5-22835e454eb3 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

Kwai Keye-VL Technical Report Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.093677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.093677Z digest=sha256:0734ec875fdacab821d08a74364b24e492ea74664f013dcad52553a19bad54d1

Observation 39613d2d-a4c9-482a-b3a3-f5002c3b70f0 · outbound

This paper cites Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs.

Kwai Keye-VL Technical Report Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.173386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.173386Z digest=sha256:adf9607878447d70bd73ecc4d478183bd1cf4d93b166d7c7753e4d3226d3c986

Observation 5f6ba017-1eab-41ca-a6ef-395dada80f49 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Kwai Keye-VL Technical Report DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.261769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.261769Z digest=sha256:b21d7402ed092630bf91c327b3bf8a46bc8028d0c4316c504e174f8eefaf65de

Observation f229ab41-a038-40bc-8d33-c8e51bbdda16 · outbound

This paper cites Onerec technical report.

Kwai Keye-VL Technical Report Onerec technical report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.352461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.352461Z digest=sha256:8aa4e08fbb422f0997b511d3ea771297702095eb65d8c77c915bfa543e8e9519

Observation 0999d1ab-1d5a-4155-9634-3b779e585c12 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Kwai Keye-VL Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.445998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.445998Z digest=sha256:16d9257449e90c15072e66726ada94b64d4539caa7db1f21381c5e615ddb495b

Observation 9da45d34-702a-41cc-9fa9-a0241d60580d · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Kwai Keye-VL Technical Report DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.546613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.546613Z digest=sha256:4d0a72d05c88555e3ed2b068e13b870b29681dee4aea740db5c5a723986808c9

Observation 8a1206d2-9301-4662-ba78-b61db910c575 · outbound

This paper cites likes" a video receives within a specific timeframe after being uploaded. Using a predetermined threshold, we classify videos into two categories:.

Kwai Keye-VL Technical Report likes" a video receives within a specific timeframe after being uploaded. Using a predetermined threshold, we classify videos into two categories:

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:45:11.561478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T20:45:10.619719Z digest=sha256:47466388461099fe4a9415761ce5c9d222438d0b1d9443a684c76f430a426da3

Observation f494b009-3de5-49a9-838b-72a05811c8fe · outbound

This paper cites doi: 10.3115/v1/D14-1086.

Kwai Keye-VL Technical Report doi: 10.3115/v1/D14-1086

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.721860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.721860Z digest=sha256:0babbf0a5b4b2cf7c8ab14ed29ae5a059e2a2a739cf40066fd3ca64c156d6a69

Observation 60df3888-994a-4b31-af8a-eeda25178d8f · outbound

This paper cites URL https://doi.org/10.1007/ s11263-016-0981-7.

Kwai Keye-VL Technical Report URL https://doi.org/10.1007/ s11263-016-0981-7

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.833490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.833490Z digest=sha256:17f4a8947252b2aedb920d8b9c032a3b519a9f4b1ce4241b97a7dd4fa13419a8

Observation 2ba8ed5f-3bb5-4fc0-a0b3-d68040702d1a · outbound

This paper cites Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts.

Kwai Keye-VL Technical Report Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.886712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.886712Z digest=sha256:bae8686c7ca08a9f35357d9142cb0543b054e599d36b1471a1274f1200dcccdf

Observation 87b4da9b-480f-4490-9f66-1bc60373c293 · outbound

This paper cites TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types.

Kwai Keye-VL Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.207778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.207778Z digest=sha256:d1b396b39db32b573a68338db392f1d0d3edcbc3a43d3418d39b40368519f4fd

Observation 6ee409d9-1a98-41f7-9cf8-738df6aa7406 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Kwai Keye-VL Technical Report MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.654026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.654026Z digest=sha256:f8d145955b0c1a243356c19cfaf001ebf297c4a62626f441724f19363b73dc17

Observation 8a2f78a4-44eb-401e-b382-cd342e824503 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Kwai Keye-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.380141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.380141Z digest=sha256:aa8ab00683d3becefe5ecc7e5754556c00942809f715090301ec8baec75666d0

Observation 7762f938-92ce-4402-8d56-14b7a3f31c8a · outbound

This paper cites Qwen2.5-VL Technical Report.

Kwai Keye-VL Technical Report Qwen2.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.065307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.065307Z digest=sha256:25da45c62f828ad1ff5abd25094a1630a850f9f4b30aff4a924fc8cf4d101222

Observation 9a8ab2c4-d19b-4943-8a45-02d4599946b5 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Kwai Keye-VL Technical Report Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.295016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.295016Z digest=sha256:33b5e9876cfba4a2629e8bde90ca4e8628176bcd7d3690841aa73226ee61b9b7

Pith citing papers

Observation 574deb97-d5bb-408c-a6d8-0403b5d74e9d · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Kwai Keye-VL Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:56.004930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:56.004930Z digest=sha256:b656a0ccdc67d8aec554b1ce963e00ef0a848ff9b25b1ea2cff620b7bc42ec33

Observation 36e2cfcd-788a-4d8c-85b5-7d2114fb4bea · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Kwai Keye-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.675805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.675805Z digest=sha256:e9cd546b100e89795b34de66f3166ecd31814f686949e57b522f5b7a44e46df6

Observation 1053a9d7-4770-40b1-bef5-dad6792d701f · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Kwai Keye-VL Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:16.696524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:16.696524Z digest=sha256:56ad0ae8c5ebe362764685985e5d91015fd8b931d26a8481199687daa48f530d

Observation 1e0d2528-1958-4e3f-b77f-1efc4f76eaf6 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models Kwai Keye-VL Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.118939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.118939Z digest=sha256:ee6932defe1a2a1892c34e717f5567d828b0373a494dbce55e7a2a07ca29d358

Observation 0bd1f391-b638-4c21-87cd-1a87b6ef29b1 · inbound

Grounding Multilingual Multimodal LLMs With Cultural Knowledge cites this paper.

Grounding Multilingual Multimodal LLMs With Cultural Knowledge Kwai Keye-VL Technical Report

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-05T22:10:51.210312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:51.210312Z digest=sha256:6da63eac88b3a5c2e009ac1a9102eb96b3e5ab64460276f7e8d7e4ddb2418234

Observation b63dcac3-964e-4c56-b3bb-68c9f5b1a2a2 · inbound

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning cites this paper.

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning Kwai Keye-VL Technical Report

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:16:53.778625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T23:16:13.165715Z digest=sha256:eb32bb2c9bb76e2147ab767c9a445af699a726be23462c95810054792882d3fb

Observation b15dedef-a39d-4731-a99c-989ebb43ac39 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Kwai Keye-VL Technical Report

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.805001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:7f38dc59cd2bb6450c620a7fe09fa84540a96812232a177f81a50f546962978a

Observation 6c6c3555-0a03-4eba-adc3-da888b32384c · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Kwai Keye-VL Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.099214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.099214Z digest=sha256:e84ec70d35d69b67ab07b946af627830166cc0311eff58bd2420953e8a2d3666

Observation 92fdacf7-f999-4184-8d6f-32b1d1fb72b0 · inbound

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach cites this paper.

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach Kwai Keye-VL Technical Report

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.649654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T12:39:56.396231Z digest=sha256:d6518352a81bc66b333ce099e8c958a7d31fe9ccb1bb49114a3019428635262d

Observation 6eda1611-25b6-45d4-8231-261136b9ec68 · inbound

CodePercept: Code-Grounded Visual STEM Perception for MLLMs cites this paper.

CodePercept: Code-Grounded Visual STEM Perception for MLLMs Kwai Keye-VL Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T23:22:13.847876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:22:13.847876Z digest=sha256:dc8bea36504cb45ae37b4d1816cec311681a8167b868498e7f2cba0d0a4d4205

Observation 6d38d343-5adc-45cf-91d4-2f39702ba63f · inbound

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing cites this paper.

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing Kwai Keye-VL Technical Report

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:28:23.543654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T00:25:19.782732Z digest=sha256:57bb1ea2b4528fd1d40bf4b2800a72f6e9592c52bace79211b1bf4dc22986935

Observation 14d5b53b-7dab-4662-a12b-6e333a1b4f5e · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Kwai Keye-VL Technical Report

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.095758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:868aee27a2ef92387017dcd45c9c60180b581e62da9db3ca62647164682da884

Observation a97d56bc-32f5-4899-830a-8043487d9530 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Kwai Keye-VL Technical Report

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.311855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:a85c296d0434d638d4f27866b9385da0ffc961472ed8d82a0f6fc93e14ecb1c8

Observation 45a011ee-1423-40a9-b718-111bf788ccd6 · inbound

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts cites this paper.

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts Kwai Keye-VL Technical Report

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:02.841158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:08:15.393476Z digest=sha256:120fdbcd76e8559d08b924e0505e650340d3188b4fd5502d3118143560eecf06

Observation 6368d297-4abc-4852-a53d-fb2764703461 · inbound

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization cites this paper.

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization Kwai Keye-VL Technical Report

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:21:12.756346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T07:19:57.063802Z digest=sha256:4fc0f83013ed832c7a83b6f17674f4e375dc4c08b9990b98bcb13803a87f4437

Observation d2d63b6b-68f5-43d4-b9c8-a755bfa8d64a · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series Kwai Keye-VL Technical Report

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.868280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:d6caa9802d4812d9bc7c862fef47ff0a9100015275c21b55eb449a3a17fb450d

Observation b7ae43e5-0197-46aa-8c87-c70a43fe5a4b · inbound

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction cites this paper.

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Kwai Keye-VL Technical Report

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.928968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T13:25:04.053283Z digest=sha256:fd07603fc023489565c65df54bc0454ed7435f801ec4dd8d5ee1027126f41ea9

Observation 9b595bd2-1395-4842-ae56-7f9724efadc4 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Kwai Keye-VL Technical Report

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:37.036364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:b819c967789c540142f91b3ebf4e3df1f5a73acf42cc456f837a4a2a3f897ddc

Observation 52563904-d40b-4292-b9d6-3d1ef6285ac4 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Kwai Keye-VL Technical Report

Reference 278

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:b3653860f4d294c34eb634551e1c2b5103b332917981d37b40b82157bf255ded

Observation 8b83f04a-2d77-4d8d-af95-6ebbfa54e923 · inbound

HoloCount: A Holistic Visual Counting Benchmark for MLLMs cites this paper.

HoloCount: A Holistic Visual Counting Benchmark for MLLMs Kwai Keye-VL Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-08T06:34:41.888370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T06:26:22.635106Z digest=sha256:73f1f20fcfecf829588bb72ea9149764f8cebddaad8f7adf79be3a734e8d4d72

Observation bf371730-be48-44f0-93b4-d31f8282461e · inbound

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI cites this paper.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Kwai Keye-VL Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:13c333d2563150c7565a95c1d01a494e603f44fee8668e86114ac4bbee61b0d1

Observation b75f4b92-3907-478c-bcfe-b74fa1f47dc4 · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Kwai Keye-VL Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.344173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.344173Z digest=sha256:84c5c4e724c03433878c6d5bec019fefcf2b3af45431ae9171b6e12ae5326e14

Observation 02027fe2-c7a0-44ef-895c-b3b3600d5672 · inbound

LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos cites this paper.

LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos Kwai Keye-VL Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T00:41:43.585577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:41:43.585577Z digest=sha256:af18363fcc4c8e395b4d95d9761831d541378f21f63673fe205f27379e2487a8

Observation 673afc8b-8945-4379-ab7e-8b4a898275e2 · inbound

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification cites this paper.

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification Kwai Keye-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T14:08:11.220405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:08:11.220405Z digest=sha256:767a97d20da4d757eb9dc5ae4182051202d08a78fb466928a9ad8256f79caf63

Observation 853230a5-6d66-4020-abd8-1d0f9c3354fd · inbound

InSight-doc: Agentic Visual Perception for Long-Document Understanding cites this paper.

InSight-doc: Agentic Visual Perception for Long-Document Understanding Kwai Keye-VL Technical Report

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-12T20:43:52.914738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:43:52.914738Z digest=sha256:b091b395d8e00e3b07f973004be137c25ebdebde9ad45f96183d1806333c0b14