{"id":"3bbfc7c1-1d28-4071-b641-b577f73f40fc","arxiv_id":"2507.01949","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Kwai Keye-VL shows that a five-mode chain-of-thought cold-start plus mix-mode reinforcement learning can push an 8B multimodal model to strong short-video and general vision-language performance.","lead":"Kwai Keye-VL is an 8-billion-parameter multimodal model built for short-video understanding, trained on more than 600 billion tokens and a five-mode reasoning mixture. The report claims state-of-the-art results on several public video benchmarks and introduces a new short-video benchmark, KC-MMBench, where the model leads by a wide margin.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"KC-MMBench is not an independent probe of short-video understanding: Step II.2 RL is trained on the same Kuaishou task family, so the 68.03 vs 57.62 gap may measure task-specific optimization.","rationale":"The reader's weakest assumption was that KC-MMBench is a valid, contamination-free proxy. My concern is the same benchmark but with a more specific mechanism: Section 4.2.2 describes RL training on short-video tasks with ground-truth labels from the same platform, and KC-MMBench is built from that same platform and task family. This is an evidentiary-independence problem, not an accusation of fraud. Public benchmarks and human evaluation provide partial support, but they are mixed: on Video-MME, GPT-4o (71.9) exceeds Keye-VL (67.7); on LongVideoBench, InternVL3 (63.9) exceeds Keye-VL (62.8); and in the internal human evaluation, Keye-VL's video correctness (3.34) trails Qwen2.5-VL (3.41) while its temporal understanding (2.92) is not clearly superior. Thus the distinctive short-video lead rests heavily on KC-MMBench, which is exactly where task-overlap is most plausible. A conditional acceptance with a concrete release-and-independent-replication condition is appropriate; if the audit reveals substantial overlap or a non-significant confidence interval, the headline claim should be rejected. I therefore retain the reader's CONDITIONAL verdict while sharpening the condition. Agreement is marked partial because the reader emphasized contamination thresholds, whereas the load-bearing mechanism is same-task RL training.","tokens_in":36227,"tokens_out":8382,"duration_ms":92880,"concrete_test":"Release the complete KC-MMBench items, per-task labels, and evaluation code. Have an independent group run Keye-VL-8B and MiMo-VL-7B-RL with identical 64-frame sampling, compute per-task accuracy and a bootstrap 95% CI for the average gap, and audit task overlap by embedding each KC-MMBench prompt against the Step II.2 short-video RL training corpus, reporting nearest-neighbor cosine at thresholds 0.3-0.9. If the CI includes zero, or more than 5% of KC-MMBench items have nearest-neighbor cosine above 0.7 to any RL prompt, the claimed advantage is not established. The definitive version of this test is a newly sampled, later-time-window set of the same six tasks with no RL training exposure.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The centerpiece evidence for short-video superiority is the self-built KC-MMBench (Table 4, Appendix B), which contains six business tasks drawn from Kuaishou platform data: CPV, SPU, Collection Order, Hot Videos Aggregation, High Like, and Pornographic Comment. Section 4.2.2 ('RL for short video understanding') states that the model is trained with GRPO on ground-truth/annotated labels from 'various short video content understanding tasks' and explicitly aligns outputs with 'desired value orientations.' If those RL examples come from the same annotation pipelines, task formats, and value-orientation rubrics as KC-MMBench, then the benchmark measures how well the model was optimized for this exact task family, not a general short-video capability. Appendix A.2's decontamination only removes near-duplicate image-question pairs at author-chosen thresholds (image cosine 0.98, text cosine 0.50) and gives no sensitivity analysis; it does not address task-family overlap or prompt-template mimicry, and the paper does not state that KC-MMBench was included in the 29-benchmark decontamination set. The 10.41-point average gap (68.03 vs 57.62) in Table 3 therefore lacks the independent evidentiary weight needed for the headline claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Kwai Keye-VL, an 8-billion-parameter multimodal large language model built on a SigLIP-400M vision encoder and Qwen3-8B language decoder, with a native-resolution vision transformer, a 600B-token four-stage pre-training pipeline, and a two-phase post-training recipe. The post-training combines no-reasoning SFT and preference optimization with a five-mode CoT cold-start mixture and mix-mode GRPO reinforcement learning, followed by iterative alignment. The authors report state-of-the-art results on several public video benchmarks, competitive performance on general image and reasoning benchmarks, a large advantage on their newly released KC-MMBench for short-video tasks, and a small-scale human evaluation of user experience.","tokens_in":36432,"tokens_out":4789,"duration_ms":52176,"significance":"If the claims hold, the paper shows that an 8B open-source model can lead short-video understanding while remaining competitive on general vision-language tasks, which is practically significant for deployment on video-centric platforms. The paper's strengths include a detailed and transparent description of the data pipeline and training recipe, explicit reporting of decontamination procedures, the release of the KC-MMBench benchmark, and an honest limitations section. The central short-video claim, however, rests on a self-constructed benchmark whose task family overlaps with the RL training data, and the quantitative evidence base is weakened by single-run benchmark numbers and a small human evaluation.","major_comments":[{"comment":"The KC-MMBench advantage is not independent evidence for general short-video capability. The RL stage in §4.2.2 is trained on ground-truth or annotated labels from 'various short video content understanding tasks' and explicitly aligns outputs with 'desired value orientations,' while the six KC-MMBench tasks in Appendix B and Table 4 (CPV, SPU, Collection Order, Hot Videos Aggregation, High Like, Pornographic Comment) are drawn from the same Kuaishou platform data and annotation pipelines. The decontamination in Appendix A.2 covers 29 benchmarks but does not state that KC-MMBench is among them, and the image/text similarity thresholds (0.98 and 0.50) are author-chosen with no sensitivity analysis. Under these conditions, the 68.03 vs. 57.62 gap in Table 3 may measure task-specific optimization rather than a general short-video understanding capability. I recommend either releasing the benchmark and reporting an ablation with the short-video RL data held out, or demonstrating transfer to independently constructed short-video tasks.","section":"§4.2.2, §6.3.3, Appendix B"},{"comment":"The claim of 'significantly outperforms' on public video benchmarks is stronger than the table supports. On LongVideoBench in thinking mode Keye-VL scores 62.8, below InternVL3's 63.9; on Video-MMME the margin over MiMo-VL is 0.3 percentage points (67.7 vs. 67.4), and GPT-4o scores 71.9. All numbers are single runs with no error bars or significance tests, so margins this small cannot support the word 'significantly.' The abstract should also carry the qualification 'among models of similar scale' that appears in Figure 1.","section":"§6.2, Table 3"},{"comment":"The human evaluation is too small and weakly reported to substantiate 'superior user experience.' It uses only 150 question-answer pairs per modality, gives no inter-annotator agreement statistics, no confidence intervals, and no significance tests; the overall video composite difference is 0.02 (3.33 vs. 3.31). This claim should be softened or the evaluation should be expanded with appropriate statistical analysis.","section":"§6.3.2, Tables 5-6"}],"minor_comments":[{"comment":"'existing a large gap' should be 'there exists a large gap.'","section":"§2.1"},{"comment":"The benchmark name 'MathVistaMINI' should be written as 'MathVista-Mini' for consistency with the original benchmark naming.","section":"§6.2"},{"comment":"Figure 6 contains emojis that do not render in the text, making the example difficult to follow.","section":"Appendix C.1.1"},{"comment":"The task is called 'High-Like Video Classification' in Appendix B but 'High Like' in Table 4; the naming should be unified.","section":"Appendix B, Table 4"},{"comment":"The dataset 'MMPR' is cited several times but is not clearly defined in the reference list; the MPO paper by Wang et al. (2024b) does not appear to be the MMPR dataset reference.","section":"References"},{"comment":"The text says 'MathVista' but Table 7 and the public benchmark table use 'MathVistaMINI' or 'MathVista-Mini'; please use consistent terminology.","section":"§7.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a technical report from a major platform, and the overlap between the RL training tasks and the self-built KC-MMBench is the most serious evidentiary concern for the headline short-video claim. The public benchmark results are mostly independent and support a strong but not unqualified claim; the revision should focus on decontamination transparency and an ablation that isolates the effect of the short-video RL data. The small human evaluation is a secondary issue but should be reported with appropriate uncertainty measures."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know up front. First, this is a competent, unusually detailed technical report about building an 8B video-centric MLLM; the training recipe is described in enough detail to be a useful reference for anyone doing similar work. Second, the headline number that supposedly proves short-video superiority—68.03% on their own KC-MMBench vs 57.62% for MiMo-VL—is much weaker evidence than it looks, because the model was RL-trained on the same Kuaishou task family that the benchmark samples. That gap likely measures task-specific optimization, not general short-video ability.\n\nWhat is genuinely new: the five-mode cold-start mixture (thinking, non-thinking, auto-think, think-with-image, and high-quality video) is a thoughtful combination, and the mix-mode RL with outcome and consistency rewards is a reasonable recipe. The paper is transparent about its own limitations, reports decontamination efforts, and releases a new benchmark. Those are real contributions.\n\nThe soft spots are mostly about evidence quality. The central public-benchmark claim is actually grounded in external benchmarks like Video-MMME, TempCompass, and LongVideoBench, where Keye-VL does lead comparable open models; I find that plausible. But the KC-MMBench result is not an independent probe. The decontamination in Appendix A.2 removes near-duplicate image-question pairs at author-chosen thresholds; it does nothing about task-family overlap or prompt-template mimicry. The RL short-video data in Section 4.2.2 is drawn from the same annotation pipelines and 'value orientations' as the benchmark tasks, so the 10-point gap is exactly what you'd expect from overfitting to a task distribution. Also, all benchmark tables are single runs with no error bars, and the human eval covers only 150 items per modality. These are not fatal flaws for a technical report, but they should be stated more carefully.\n\nWho gets value? Practitioners building video MLLMs, especially at platforms with similar commercial tasks, will find the data and training details useful. For a general ML audience, the paper is a solid but incremental data-and-recipe contribution.\n\nBottom line: I'd send it to peer review as a technical report, but I'd ask for evaluation code, uncertainty estimates, and an explicit statement about whether KC-MMBench tasks overlap with the RL training data. The authors' own claim of 'significant advantage' on KC-MMBench should be toned down until that overlap is addressed.","headline":"A detailed, useful video-MLLM technical report whose public-benchmark results are plausible but whose self-built short-video benchmark is likely measuring RL overfitting to the same task family it samples.","tokens_in":37251,"tokens_out":2573,"would_cite":false,"duration_ms":27686,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Kwai Keye-VL, an 8B multimodal model, claims the lead in short-video understanding with 68.03% on KC-MMBench.","keywords":["Kwai Keye-VL","multimodal large language model","short-video understanding","chain-of-thought reasoning","auto-think mode","reinforcement learning","preference optimization","KC-MMBench"],"falsifier":"Take KC-MMBench and re-annotate it on a freshly collected set of short videos with reworded questions by annotators who did not build the benchmark; if the 10-point gap over MiMo-VL narrows materially, the advantage is partly benchmark-specific. A cheaper check is to vary the text cosine decontamination threshold around 0.50 and count how many excluded training samples would then overlap with KC-MMBench.","tokens_in":35983,"feed_emoji":"🎬","tokens_out":9736,"duration_ms":88777,"temperature":0.7,"pith_summary":"This paper introduces Kwai Keye-VL, an 8-billion-parameter multimodal model built for short-form video understanding, and argues that a video-heavy training corpus paired with a staged reasoning recipe is what lets a model of this size lead on video benchmarks. The authors report state-of-the-art results on public video benchmarks and a 68.03% average accuracy on their released KC-MMBench, against 57.62% for the next-best model at the same scale. The central claim is that the model learns when to reason and when to answer directly, taught through a five-mode cold-start data mixture followed by reinforcement learning and iterative preference alignment. If correct, this means an open 8B model can handle content moderation, product attribute prediction, event ordering, and similar short-video tasks without sacrificing general image and reasoning performance.","feed_headline":"8B video model leads short-video benchmark by 10 points","feed_subtitle":"Kwai Keye-VL scores 68.03% on KC-MMBench, beating MiMo-VL's 57.62%, while staying competitive on general tasks.","key_machinery":"The load-bearing mechanism is the five-mode cold-start data mixture used in the second post-training phase: 330,000 non-reasoning samples, 230,000 thinking samples, 20,000 auto-think samples, 100,000 agentic think-with-image samples, and curated high-quality video data. This mixture teaches the model to decide when to emit a long chain of thought and when to answer directly, which the paper credits for simultaneous gains in Thinking and Non-Thinking modes. Reinforcement learning then applies GRPO with outcome and consistency rewards, including a dedicated short-video RL stage using ground-truth labels, and the final iterative alignment step uses rule-based and model-based scores to build preference pairs that correct repetitive or illogical outputs.","core_discovery":"On the paper's own terms, the discovery is that short-video understanding is primarily a data and reasoning-scheduling problem rather than a vision-encoder problem. Keye-VL starts from the Qwen3-8B language decoder and a native-resolution SigLIP vision encoder, pre-trains on over 600 billion tokens with a strong video component, and then runs a two-phase post-training recipe: supervised fine-tuning plus mixed preference optimization for foundational skills, followed by a five-mode cold-start mixture (thinking, non-thinking, auto-think, think-with-image, and high-quality video), GRPO reinforcement learning, and iterative alignment. Reported results put Keye-VL ahead of similar-scale open models on Video-MMMU, TempCompass, LongVideoBench, and MMVU, and give it a 10-point edge on KC-MMBench, a new short-video benchmark the authors built from in-house platform data and released.","pith_inferences":["If the Auto-Think mechanism is as general as reported, the same five-mode scheduling could be applied to other mixed-difficulty modalities, such as long-form video or live audio, where the model routes simple inputs to direct answers and complex ones to chain-of-thought.","The 10-point KC-MMBench gap comes from a benchmark the authors built from their own platform; it would be more convincing if it survived independent re-annotation or a third-party short-video suite, since question wording and data decontamination choices can favor the model that trained on similar data.","The reported mutual enhancement between reasoning and non-reasoning data suggests long chain-of-thought data may act as a general regularizer, a hypothesis worth testing on smaller models where per-task ablations are cheaper."],"forward_implications":["An 8B model with video-tailored data can outperform similar-scale open models on Video-MMMU, TempCompass, LongVideoBench, and MMVU while staying competitive on general image benchmarks.","The five-mode cold-start plus RL recipe provides a template for letting a model choose reasoning depth by task difficulty, as shown by Auto-Think thinking ratios of 0.35 on MathVista, 0.34 on MMStar, 0.08 on HallusionBench, and 0.00 on OCRBench.","KC-MMBench opens a way to evaluate skills that public English benchmarks miss, including pornographic comment detection, collection order, hot-video aggregation, high-like prediction, and e-commerce product matching and attribute prediction.","The decontamination analysis flags MM-Eureka and MMPR as datasets with substantial overlap against common benchmarks, warning practitioners that training on them can inflate reported scores."],"supporting_citations":[{"why":"Supplies Qwen3-8B, the language decoder whose reasoning and world knowledge Keye-VL builds on.","marker":"Yang et al. (2025)"},{"why":"Provides the SigLIP-400M-384-14 vision encoder weights that Keye-VL adapts to native resolution with 2D RoPE.","marker":"Zhai et al. (2023)"},{"why":"Used to generate image and video re-captions, sample chain-of-thought paths, and create agentic code data during post-training.","marker":"Bai et al. (2025)"},{"why":"Provides the GRPO algorithm used in Mix-Mode RL to sharpen reasoning and short-video understanding.","marker":"Shao et al. (2024)"},{"why":"Provides the MPO algorithm used in both the no-reasoning and iterative-alignment stages; its MMPR data also feeds the reasoning mix.","marker":"Wang et al. (2024b)"},{"why":"Supplies the 47,000 DeepEyes samples used for agentic reasoning in the RL stage.","marker":"Zheng et al. (2025)"},{"why":"Supplies MM-Eureka data used for auto-think and RL training, and is flagged in the decontamination analysis for benchmark overlap.","marker":"Meng et al. (2025)"}],"fun_headline_variants":["Keye-VL: 8B model tops short-video benchmarks with 600B tokens","Five-mode cold-start reasoning unlocks short-video AI","Keye-VL: 8B video model beats MiMo-VL by 10 points","Short-video understanding needs data and reasoning, not bigger encoders","Keye-VL: two-phase training, five reasoning modes, top video scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline short-video advantage rests on the assumption that KC-MMBench, built and decontaminated by the authors with image similarity above 0.98 and text cosine similarity above 0.50 used as exclusion thresholds, measures real short-video skill rather than matching the model's own training data and prompt wording.","fun_headline_variants_meta":{"raw":{"variants":["Keye-VL: 8B model tops short-video benchmarks with 600B tokens","Five-mode cold-start reasoning unlocks short-video AI","Keye-VL: 8B video model beats MiMo-VL by 10 points","Short-video understanding needs data and reasoning, not bigger encoders","Keye-VL: two-phase training, five reasoning modes, top video scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000925,"raw_usage":{"total_tokens":4011,"prompt_tokens":1042,"completion_tokens":2969,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":2876}},"tokens_in":658,"tokens_out":2969,"duration_ms":23325,"temperature":1.0,"reasoning_tokens":2876,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:39:09.502192+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take KC-MMBench and re-annotate it on a freshly collected set of short videos with reworded questions by annotators who did not build the benchmark; if the 10-point gap over MiMo-VL narrows materially, the advantage is partly benchmark-specific. A cheaper check is to vary the text cosine decontamination threshold around 0.50 and count how many excluded training samples would then overlap with KC-MMBench.","supporting_citations":[],"review_version":1}