Pith. sign in

Paper Citation Record · LEDGER

Kimi-VL Technical Report

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2504.07491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07491 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 184 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:17:17.568652Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T06:34:41.884360Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 55779ed8-40cf-4c77-9c59-708f0a491d66 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models Kimi-VL Technical Report

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:70bea1a70dd8ac595fbe5d9ed6f452298a6f077459ba2d935378ef81c995111a

Observation 4d03bda5-c45e-47a9-8b8c-6b989d066cfd · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Kimi-VL Technical Report

Reference 156

Resolution
malformed identifier
local_arxiv, observed 2026-05-17T20:33:26.936643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:17b2b2dccfec13603d04e15f331a9592032bb0d170295ce2d5aa03dc80105270

Observation f93d5bae-3de3-40f4-9fa5-a20bb918df81 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Kimi-VL Technical Report

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.266803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:6e060266678a668b4729b0c8996beba90e2e3006fe2fac0ec3695832c85adaaa

Observation 2339702a-3672-4d47-83e6-3c371489d463 · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning Kimi-VL Technical Report

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:18:43.820420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:11acf05826434402df1418ab1db035f32f79688fdf1a917821583efd18988ebb

Observation 65203e92-b1e1-48c1-af8a-f5a14bacb9ba · inbound

InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners cites this paper.

InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners Kimi-VL Technical Report

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:54:44.166120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T13:54:44.011048Z digest=sha256:9ba8b13663b759e36e0f2cfdaed27c46fdd9d49ace6e676557b9b845717cf634

Observation 95c2cca8-0b97-4cd3-96cd-a6608d91a571 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Kimi-VL Technical Report

Reference 131

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.528482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:acb4ce2e8af20ab90867cf55172b74247bb8c4f8c72772e2182ecbfb58cb6d97

Observation 43fb7d13-7bce-4a2f-836d-933acc57486c · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining Kimi-VL Technical Report

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:67578bb1069e72e5fc7edc4f1cee0054c52ddf28ef2607913ef81aa993fe5d22

Observation d5a84dc6-6683-4aa7-a25a-2c590443bfc5 · inbound

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence cites this paper.

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence Kimi-VL Technical Report

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:11:35.700108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T13:07:11.548885Z digest=sha256:79cf246f7328ac61cc1d9857f7e35cc645c9e44297354d50a3ad6825e9638f9b

Observation e00a243f-4920-475c-b18c-07329bf16a9f · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Kimi-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.025955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:32a7034cec669bbb8118bd48ff891c5c2406d25690e35ace4588536c196bfdb9

Observation 297854d3-26f4-4e7e-b3b7-7d2a5aebcd13 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Kimi-VL Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:51.891067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:c05aff19ce7d3fbeeb0ffd6aa57b73966489871162a72b234fc39c30c8925d4a

Observation 9afad529-973b-4530-9c32-f0f0ef0dd51d · inbound

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts cites this paper.

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts Kimi-VL Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:47:15.155684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:43:02.601014Z digest=sha256:b4a76f59d2af7d7b703aebca356af5b99812fd66a50cc9b465c3578f44bb5ddb

Observation 9556cc53-a69b-4904-92f2-650f750ec813 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Kimi-VL Technical Report

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:27:04.310839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:9e283ce8616bb1e5dbfe0c8aa65248c654a8c4bd9972ca0ffffbe1ea02f43e7e

Observation 2e5864fe-54cc-4b85-b577-67c86d72d40b · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Kimi-VL Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.007047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:92a0a6825ea4564481262640c93a4892705dabd3bcb3c1620b9f3fb7ee9f7496

Observation cf55cf4f-b851-4170-b693-177f4ad91a9c · inbound

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning cites this paper.

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning Kimi-VL Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:21:53.252865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T22:17:41.758059Z digest=sha256:0bcbca4c75b997428bcd93619d7acf1f0b236a82b92116b1a2272e46d9c45196

Observation e6701644-17ca-449c-afab-0e4c9da3fa28 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Kimi-VL Technical Report

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:0d2218fe17d29182cc884cd40019570bd1e75bf14f351fa76f5988e08793a0d1

Observation 0709c628-21b7-4af1-8f8e-f4cacba41191 · inbound

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration cites this paper.

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration Kimi-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:17:17.568652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:17:17.568652Z digest=sha256:e7a49b67a6c9abccd43db56632e8888b555ab180a10cce3b7b15df7febe4c967

Observation 3d3e06c0-d74d-46b7-9b6b-e3fcd231fe75 · inbound

From Charts to Code: A Hierarchical Benchmark for Multimodal Models cites this paper.

From Charts to Code: A Hierarchical Benchmark for Multimodal Models Kimi-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T06:10:57.783014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c304ae2caa1ef59121acb44543f50feffc73152d608ead5324eeef78e7e8f606

Observation ab9d455c-f058-420f-9c91-3be7a5cae8a0 · inbound

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning cites this paper.

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning Kimi-VL Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:02:43.587509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:02:43.587509Z digest=sha256:ce0e98470380eb72ce91b7aff3926695682a4a7a12f9ebe0a02ba7d557443b25

Observation ebf26015-b100-4edf-b917-5bb22bcb914e · inbound

Grounding Computer Use Agents on Human Demonstrations cites this paper.

Grounding Computer Use Agents on Human Demonstrations Kimi-VL Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T23:06:04.886225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:06:04.886225Z digest=sha256:19074f2557b11c6ac8654a034029f0e786e6e827afd5f72ef0e4fcfbb5988f34

Observation f4106c95-603e-4526-b91a-3109f8fdc443 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards Kimi-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:49.425505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:49.425505Z digest=sha256:2302d9580203c4132fef1db541949e57f81ac88da996db48b547f794e9c55b88

Observation a5cad98f-d99c-48b3-8804-b6d80881b188 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Kimi-VL Technical Report

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:54:20.195365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T19:51:04.983299Z digest=sha256:7cccd46beacd84a8efe5b8d7e7dd1869b9d1e41e35dc91e5cceadbb80ce1ccb7

Observation 3259e867-3b33-4db2-ace3-fe0b8d1b766f · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Kimi-VL Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T21:44:02.104425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:44:02.104425Z digest=sha256:69460c0d0b0a7b6ab3dc2f0eaaa0cce6af4dbf3ab0df7ade973ee01d3816f271

Observation 2a7b7d8d-67ee-4f80-8e10-e28e67b5abfe · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models Kimi-VL Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:50:15.080914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:7ae3d5298f9e646683a3aba6902f29f28441e3d2cacfa24723ccee7383e1922c

Observation 2e0875c6-fdbf-4465-a492-8569bf22967a · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model Kimi-VL Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:04:09.178208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:435fa4f55300fa78b09be739c34029d2f2f6f462a4cd56b796707867fcbf916c

Observation c68224d0-4ad7-4081-a6b7-44ce98252ea2 · inbound

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models cites this paper.

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models Kimi-VL Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:41:51.622305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:41:51.622305Z digest=sha256:200178272caf06e6b6adf7a5086ed26d57e630d3596fc8230dbc45ce7dd8f9ac

Observation acdfc20d-93b7-4397-a17f-edbafef311f4 · inbound

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos cites this paper.

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos Kimi-VL Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:08:56.583164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:05:42.773193Z digest=sha256:06702d0a579a61c50f3877192bd382facc4a434471c2b66e213d0ca2c7bcd75c

Observation 84a80519-a6b0-4f5d-a9a4-f08b08442d62 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Kimi-VL Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.839790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.839790Z digest=sha256:d433ac9c398beea9f6d0481681cfa94ba2e696a95f646f39ff69a25386ac80fb

Observation db78e159-c9e5-4d5a-b7be-ba64ded36bda · inbound

Weather-R1: Logically Consistent Reinforcement Fine-Tuning for Multimodal Reasoning in Meteorology cites this paper.

Weather-R1: Logically Consistent Reinforcement Fine-Tuning for Multimodal Reasoning in Meteorology Kimi-VL Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:37:53.067665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:36:23.661583Z digest=sha256:a35b611795b7eca845f74d3a4752025ccdaf13b4df4b549c8a08ca19e4452889

Observation 60f25c16-e848-4c3e-ba7f-83f8a5fe2ba9 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Kimi-VL Technical Report

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:961740f05f18a3083cf8b3aae010e970aeb9d005e81726bbc61454d8fd7893f8

Observation 0105c6cb-8537-4c49-ae5d-384c59aa0c56 · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? Kimi-VL Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:40:12.554830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:36:28.844714Z digest=sha256:9a7dd2cd656492841c3f8c2185fdf3cc836977811f6a4b7545b8b37701b36f7b

Observation 8d9669c1-4a0b-4094-acb6-4f00d770d1d9 · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? Kimi-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T04:31:21.357882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:31:21.357882Z digest=sha256:11e51b992fc3643065a980f39d96f938b31e2859672ba6a566177f09e92e9ebd

Observation 5c82d91f-46ef-414b-b25b-f9b160fed38d · inbound

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models cites this paper.

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models Kimi-VL Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:45.980403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:57:45.980403Z digest=sha256:05d37e950ec032c89c20e38975380848ffc891b21943803d2310f54c3335a451

Observation 76438f37-5a28-4ff7-b9d0-5233ad183ec3 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Kimi-VL Technical Report

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:16:34.519578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:193686c0c0f5a9e7f8b527b854d9b162c10530be4d2119dfa9fa4faa23b04ed5

Observation 055bd07d-91a0-4474-a9d2-e17016b4c4ce · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Kimi-VL Technical Report

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:31:25.304135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:74b0ad8a6dc8c05d459f3e02487a99838501c56f4de2fc8b234bcbc8ac8a3c32

Observation 3d042823-38ec-4a82-96de-2a6eeda89d5d · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Kimi-VL Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:09.601701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:09.601701Z digest=sha256:0e6033d7a3f97ebd6c4a70109c1f82db243c158e14c510ae1c6bb10434135059

Observation ff69048e-71cd-4460-a2e1-18f8d04d9a29 · inbound

MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training cites this paper.

MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training Kimi-VL Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:26:31.388717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T19:24:22.699712Z digest=sha256:69ef0b3f56700ba26507d7f9dc53341078a5cfdf050475d6de68e522c6f57211

Observation 9009df79-bb90-43c7-b0eb-7db5a41098fb · inbound

CodePercept: Code-Grounded Visual STEM Perception for MLLMs cites this paper.

CodePercept: Code-Grounded Visual STEM Perception for MLLMs Kimi-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T23:22:13.847876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:22:13.847876Z digest=sha256:e6a5ee262894972d15d16fbc133433a302b9788e4b2f09f0b2943fbe9cbe4039

Observation 6cb78262-1f02-45ed-b11d-712a227bc1d5 · inbound

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next cites this paper.

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next Kimi-VL Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T05:50:27.244335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:50:27.244335Z digest=sha256:d0b5b45ae827c73ea7fcf2e356efbca34b2f60830a9031375392ec844c2001e1

Observation ffd491ef-fa29-43c1-a66a-58ee5588d094 · inbound

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding cites this paper.

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding Kimi-VL Technical Report

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:39:32.946024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T22:39:06.113655Z digest=sha256:6a56aa64c938958623e982ec81a7f291e3cd80fd5169686e170d9b35cca7de51

Observation 852a61a8-ef38-42ca-8b8c-1ff165f2f52c · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Kimi-VL Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:08:04.276718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:fc982bb1f00e7abb0380ecaba32a601974549b131b4618592d8e7071592c9df1

Observation 0f9777ea-4649-4051-8c43-e588ba82ee54 · inbound

Optimal Projection-Free Adaptive SGD for Matrix Optimization cites this paper.

Optimal Projection-Free Adaptive SGD for Matrix Optimization Kimi-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:58:15.749456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:56:41.466669Z digest=sha256:75c32557a2310850a726d4c387f38de32bab16e1900e017bf29c9682fede2869

Observation d2d870cd-bd50-40b1-be9a-5953943306ca · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Kimi-VL Technical Report

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:08:17.289912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:371dd874cf08d097beacf2e9ae2b096d045a154e12c8aa15018c38a7af9c7315

Observation be5414fa-f4ba-49be-a3d3-a87196db679e · inbound

Discrete Prototypical Memories for Federated Time Series Foundation Models cites this paper.

Discrete Prototypical Memories for Federated Time Series Foundation Models Kimi-VL Technical Report

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:59:18.819953Z digest=sha256:d49fc711f48bc73df96c1f48ebd9f30c898eb2a02a86fc97c9926f418b5d65f1

Observation 0f89481f-bca4-46bf-a9c0-d6c1bf413243 · inbound

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding cites this paper.

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding Kimi-VL Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:47:32.778695Z digest=sha256:7b4e64e1e4e5cba88e75641debe9218e5a18a31d5ac0568b32a7d3e0517aebb5

Observation ff2e63ab-bbc6-438b-be97-c7efc1c984e5 · inbound

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence cites this paper.

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence Kimi-VL Technical Report

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:20:56.623575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:41:31.544716Z digest=sha256:73412a896219c45fc97b39fbe20e4d12383e2a01ac158d68b6693b291bf57b82

Observation c3fd5b0d-0524-49f8-8bf7-2ffe0449a1e9 · inbound

Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding cites this paper.

Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding Kimi-VL Technical Report

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:24:17.872120Z digest=sha256:2fb7934b1d759584dfb4c5c1b3f8c2c91764772491e52e22d280582eed0be2ce

Observation 272b9161-9183-4ac1-86ec-d7f8351168c3 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Kimi-VL Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:b69db4b42a16b29ade9f21a48769b83bb6679a9158acd4f04a2b7e9802fcbda8

Observation 7800a3c8-c310-42dd-8eaa-cc6b5a8455bf · inbound

Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts cites this paper.

Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts Kimi-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:41:00.571757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:03:13.911088Z digest=sha256:afbede691b71d1f91ed29bd43f364cab7caaf2111cfe8c9884e46a62efee1e0b

Observation 9d575264-c386-48eb-9efb-7bf23cef0413 · inbound

HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing cites this paper.

HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing Kimi-VL Technical Report

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:30:57.733616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:06:49.114269Z digest=sha256:a6a50b8a6a0d015e2bd77e0ea705d0cd57f0964d981bf2ebd9b381b3b99acf9d

Observation 92ce7279-8e24-4644-9123-9ea32d23c8ca · inbound

Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning cites this paper.

Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning Kimi-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:41:01.624156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:00:53.682217Z digest=sha256:551753f3212bb39a8257cec8cd607761a792129ec48def9b1f122db5d1ad4c8a

Observation 4f5cf572-9618-4155-bd6f-47e414bb41f1 · inbound

Omnimodal Dataset Distillation via High-order Proxy Alignment cites this paper.

Omnimodal Dataset Distillation via High-order Proxy Alignment Kimi-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:59.633360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:15:46.835576Z digest=sha256:48820a422bc709dfef95280270dbe4dc23264452d2945e3bc89e27515608a72b

Observation 6875ff6c-b4e0-43b9-8754-ddd2f43fe59d · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Kimi-VL Technical Report

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:41:03.867263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:0ce138dce9159b164b1d041bd63be170d0a09cb11d7125c984f1dc107eeb3378

Observation ce46f2f9-95b5-4755-b1ac-f921006d1639 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding Kimi-VL Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:31:01.314372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:7cb642757f8564ec899614a8eae84f3497fc06dd6e5f1a262f94fa150f2cfce0

Observation ded5dacf-80c2-4e36-8c2c-4d918bbd7462 · inbound

Towards Scalable Lightweight GUI Agents via Multi-role Orchestration cites this paper.

Towards Scalable Lightweight GUI Agents via Multi-role Orchestration Kimi-VL Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:45:17.686098Z digest=sha256:aa66863630cc7a407a292b79c94afaf2f8888e4c448a99449d3f0df8af3503e6

Observation 69bac594-d3a6-4c8e-8be4-259563a5fd3f · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Kimi-VL Technical Report

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:79f8a63b5b3409aa69e9613f941b902537dc43d709febc960ed0a3e65743651e

Observation 8acf0d45-7092-4679-8faf-46fd139958f2 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Kimi-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:24.048781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:24.048781Z digest=sha256:4a3538e1422285af88440203f8d1fd6c09f3429830d5a2d6546594759e486b60

Observation e85b1f34-676a-4e5c-a11e-9b42ba5eaf4e · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers Kimi-VL Technical Report

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:b110ad4b8c291f39c21a1a69f4da5d760f07fe174a7fb6d4bdea7109ae7683cf

Observation aff05f45-09d0-4abe-baaa-de789768a30f · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding Kimi-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:39377ef521163963b919ac51a83aed1b8235d907ed03b4f3a5419bc5102cccc4

Observation 4e557216-bc91-4672-b8f6-963f6d8c9e27 · inbound

UniMesh: Unifying 3D Mesh Understanding and Generation cites this paper.

UniMesh: Unifying 3D Mesh Understanding and Generation Kimi-VL Technical Report

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T06:55:42.679323Z digest=sha256:15bc0dab10e24a1d1ee18053f959acd9714ea3aed8dfee44bfb494b94bfc31cf

Observation e20c7f5c-ca36-4a5a-9b29-196b4022684f · inbound

Class-specific diffusion models improve military object detection in a low-data domain cites this paper.

Class-specific diffusion models improve military object detection in a low-data domain Kimi-VL Technical Report

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:34:29.790101Z digest=sha256:bcc5f97ad17855e6676145de4417d011400cd2d5f0014f5a6e4ca50b5676902f

Observation d9e07943-8640-41be-8e92-0c9d796e0a08 · inbound

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs cites this paper.

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs Kimi-VL Technical Report

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:31:25.826205Z digest=sha256:cbc01d351f365746fee7ba87a8fba3c1a2edcf8b016ae71c8af1b0b56d240631

Observation 1cc4c9cd-31a2-4e2f-9f28-1cd4dd94b57f · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Kimi-VL Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:03.196041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:940a6f49176c219ffd343fa141ef2f68485235d5617203e7c667ed343e0d5362

Observation 90311577-66e0-4f69-a77c-51dd9ff15b44 · inbound

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models cites this paper.

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models Kimi-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:00.217783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:10:32.154619Z digest=sha256:0fedb33bfd510f5f3374c1bb2d26fcbd9c323b48a1cb7ef468ff12fd22ab228e

Observation c8f3be58-bc35-47d5-bb88-b618a19a1f0e · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Kimi-VL Technical Report

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:06.002603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:8ac2e5328b182f780cdafdb64c14c3a77a0f20e0594296edf9898dbdb9e21015

Observation e4506c0b-26e3-4342-9f6d-67390857b338 · inbound

Can Multimodal Large Language Models Truly Understand Small Objects? cites this paper.

Can Multimodal Large Language Models Truly Understand Small Objects? Kimi-VL Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:01:19.432151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T12:49:53.645987Z digest=sha256:ab4f0d0093abba7d062397935beb012b1f7d67b8882230f01b5ad5f56d235870

Observation 50affc38-1529-484a-a1c5-8fb34d30ef20 · inbound

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs cites this paper.

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs Kimi-VL Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:13.660981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:41:52.098355Z digest=sha256:35207c3473097da51f721dd2578906239b7b807e6bbd55dc233e5c79662be8eb

Observation ed11a268-f0c8-4ac5-905e-ef92496eb14d · inbound

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ cites this paper.

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ Kimi-VL Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:25:40.494402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T09:08:06.592577Z digest=sha256:3ca88abe882bde003b813eabb3983976c1d4be6b37f2444da343b157b4571a33

Observation ed0422f9-8741-40da-9316-9ee88988cc23 · inbound

FCMBench-Video: Benchmarking Document Video Intelligence cites this paper.

FCMBench-Video: Benchmarking Document Video Intelligence Kimi-VL Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:21:41.296487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T17:14:19.186123Z digest=sha256:3be5ac317f80fdc62086f55e0584ae3edb3c9106389f2ddcc71f960ded731030

Observation 6dd2b337-08df-4c83-a4bd-9e7b2726aa17 · inbound

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding cites this paper.

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding Kimi-VL Technical Report

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:46:26.715698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:22:12.446406Z digest=sha256:87ba640a76c1476c10d8f831afe62d08e23001cab600fc503ee83a5f5c9783a1

Observation 58639a19-deb9-40d5-a6bb-c7c8caa849ca · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training Kimi-VL Technical Report

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:01:11.335758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:30edc1f5b45c7e59f91d7cdd20b1d990a2185a68a27f70b0959832df188ecb96

Observation 944ba6da-b60f-4260-a002-7c7f43d851b4 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training Kimi-VL Technical Report

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:28.739389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:5453d33c950da602f484857ea65c89eefa04e9ef94bbdfd271d1880540bfa99a

Observation 44864580-5ae5-4ff7-88d0-26ecb75c141b · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Kimi-VL Technical Report

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:01:22.615220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:be0caf80417f66bf4c0ec8d916e91fbf36177537a6f7b1a1bc0c440b5a146a0c

Observation dc9dfde6-5282-4a23-938b-f37f6b6758d2 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Kimi-VL Technical Report

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:50:51.647475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:548846cfc1ccc9b1ebdd63b6d403ecbc9e9a5e348be93a304d4987ec95f0889e

Observation 143290ce-9240-42f3-8932-36e849ee8ca6 · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning Kimi-VL Technical Report

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:3ad6f010d8b198a3f5af5b7e45a07daee1120ac8feae2cfec6b82485176fea34

Observation 519b8b23-d68c-4d19-98a7-c5f6c4829e94 · inbound

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs cites this paper.

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs Kimi-VL Technical Report

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:46:17.330355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T19:22:00.217729Z digest=sha256:b3852a2b5b0ef5f8e66ec882c31adc26183e40c9b82d8a151f933086a0893b41

Observation f4079c6e-4bc2-4c78-99f2-a13b59b81248 · inbound

Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts cites this paper.

Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts Kimi-VL Technical Report

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T17:58:10.323354Z digest=sha256:fcb2840525039f0504c6e61c521aeff5baafdf1fffab943ae5bb17285b02ef86

Observation 81e106a4-0252-4c02-b40c-6dde3db2cc46 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Kimi-VL Technical Report

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T06:47:25.155437Z digest=sha256:8138133ac2161cf6892f026e98cb483690dffe349d296ff78caffa355e5a91d5

Observation ebfafc57-1960-43aa-a73b-f0267bca95a0 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Kimi-VL Technical Report

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T01:45:52.113756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:21be172d6da3c6f6d9535bffe621f9e3a9ce456917a6d3bf5ee0d1beb47bb4dc

Observation e585d113-8417-4a7b-a537-96be0a90b489 · inbound

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning cites this paper.

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning Kimi-VL Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:46:09.414376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:02:29.480442Z digest=sha256:b2346de73dc78026c38ee1d9eb0496218a6a9c1cbaab184319d79d0ff0291c01

Observation 06c72520-bd69-4a5f-a3c0-d13b3844838a · inbound

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning cites this paper.

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning Kimi-VL Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:45:51.805531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:29:47.384341Z digest=sha256:c33be2c1668c7cb433cc53568bbc63c8c9b76ee53602d17805b9f01c594beb7f

Observation 1d024309-8fee-40c4-b3b7-cd40d590e0f9 · inbound

LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning cites this paper.

LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning Kimi-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:45:57.935347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T02:18:57.917353Z digest=sha256:53bee8944fe634381c17c80c7e6080d4dc5168c9ed2adf3cbe62bf04799ff3f5

Observation c8b7d24c-e20a-45b2-a407-a24adb0b967f · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Kimi-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:15:55.952698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:a6d9f63aa5ffb31c0d3ebc4ea3956685feafb466bdec3bf9e504ec1c349a887d

Observation 89a52820-41d9-43c5-b413-02b7bfa3c68b · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Kimi-VL Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:29.289741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:23:37.488121Z digest=sha256:edfa5dad6968c390c756bb97356238d4c701755522b093447af51826cae8a2ec

Observation 9bd34d24-beb3-4142-801a-5db5382fec4d · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Kimi-VL Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:07:00.414633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:03:30.995808Z digest=sha256:343c0ad488e96011b6a1a9edf1a31c832972f01e593f55b5dd25bc85e1c4203f

Observation f63da5b3-c95e-4e3c-9ef0-06db2e7d397f · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Kimi-VL Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:29:49.799179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T06:25:15.510083Z digest=sha256:1aaed4356af124c0c4b1e6570e5a5954ba7ce97e0a19d14a4ce2bbc56866fb52

Observation 75c478d7-5255-4e2e-89e0-e06041dbd440 · inbound

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? cites this paper.

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? Kimi-VL Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:42:07.786299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:06:45.858231Z digest=sha256:f0e8397fbcae712672d3974fb298fb8dfe0db37f8c00826ff25708537afedf04

Observation 7fa43a1f-9e9b-4fb4-b746-cd98dd54ced6 · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search Kimi-VL Technical Report

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:18.316815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:141f6780d885c6fe6018e20bc43b00d11d79b891e5ff3a0a4b7993b54719b902

Observation 44c20eae-566d-41b6-a340-09f19488e962 · inbound

Dimension-Free Saddle-Point Escape in Muon cites this paper.

Dimension-Free Saddle-Point Escape in Muon Kimi-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:26.439415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:38:25.480020Z digest=sha256:08312b2ac529b2690546c231011fdde5b2daeed71ebdd59589548c8e781802af

Observation 5fc7e60e-ac09-43a2-b9d3-790bff6e4f49 · inbound

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs cites this paper.

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs Kimi-VL Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:56:44.903344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:43:51.900721Z digest=sha256:ffe1452eabe28eff4fcafcf4400fd615eb175e5596c9182f9c35f1a802609fd2

Observation ef612a3b-b402-47e2-81c5-d3f09adf6cac · inbound

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning cites this paper.

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning Kimi-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:41:25.948392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:24:12.349405Z digest=sha256:04122f2228c91aa88cfde3ec86dddca498146dca3fea112d5123e8ae6665b76d

Observation 44535bfc-7a27-43d3-9996-ed6a4c7d08ad · inbound

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models cites this paper.

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models Kimi-VL Technical Report

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:46:35.081227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:00:50.940530Z digest=sha256:d04b338e854a73de8ce7e8569504d8fc0efde5e95e018cdc3bb4bff327f07fa7

Observation 02956506-94bd-4833-8c62-b86720e4cc37 · inbound

NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation cites this paper.

NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation Kimi-VL Technical Report

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:31:27.353807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:12:49.742272Z digest=sha256:2292dd1103dca866eac6f84e14330aee07bf96f74b88bd0fd500cc19deac5837

Observation 75586e0a-b514-4f9e-aedb-8ce5ac422ef5 · inbound

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale cites this paper.

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale Kimi-VL Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-12T17:04:17.588929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:04:17.588929Z digest=sha256:b869cf8ea2215d25e1cfcdc23a2b941d7a6986e0fa4134178adc2e0d8291b8d2

Observation 04bdf586-7b0c-472e-b631-475e7fc8be21 · inbound

Count Anything at Any Granularity cites this paper.

Count Anything at Any Granularity Kimi-VL Technical Report

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T06:21:27.742404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:19:44.335405Z digest=sha256:46885c88891a5d236a0401d7691bd5b376b3e2d51391a6286d70838ea94beff6

Observation d597b91d-b9da-4b38-a13a-89bcda16b691 · inbound

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction cites this paper.

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction Kimi-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:17:06.857745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T02:12:42.449883Z digest=sha256:a8ed7baa96829f1a91583d63f40fd6e6a0336447557aff1e640caf998b2cd940

Observation e077f81d-82a4-431a-a405-3fc7b9a93245 · inbound

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction cites this paper.

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction Kimi-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:57:59.213371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T20:56:56.216287Z digest=sha256:c2e0736d785dd3636793e77a7c367b81ee76ea12c5afbfb78b08c478895653ed

Observation e479df73-081e-4129-a889-de4397b03b0a · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Kimi-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:22:54.848190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:21:26.640493Z digest=sha256:5d37f17764fbecde83f1e31eb6566b53f0025c13e7468560d8f2d257c26543fe

Observation c96b2197-4b03-408d-8ed1-ed372401f371 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Kimi-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:43:45.403517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T21:42:49.452347Z digest=sha256:ce190e59e8c2f77544e17688a7028e14536835a23deef42dba486a73a384d371

Observation 829c43ba-081b-4050-a06b-900a72d59be9 · inbound

Large Language Models Lack Temporal Awareness of Medical Knowledge cites this paper.

Large Language Models Lack Temporal Awareness of Medical Knowledge Kimi-VL Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:19:27.211816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:18:08.160768Z digest=sha256:8e4967acd5f9b7a31305b41f1149ded2411884e33ce69a19d5e853ce452841ca

Observation bbe3dc4e-f5f0-474f-99b9-b7395d0b208b · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models Kimi-VL Technical Report

Reference 144

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:22:56.332311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:570c49b9c1bc773711e2e968594e0a17f0c756059a2ff0fd4eb3958e5fcbf1bc