Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2501.12599.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:51:52.132024Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T22:47:36.964713Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6d708888-b4c7-438a-8898-fea272246142 · inbound
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 217
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8653bbd6-566a-47be-be2e-bcd8840e7dd3 · inbound
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1aecb095-1905-4f80-bfd2-616e467a813c · inbound
Process Reinforcement through Implicit Rewards Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c9711623-d425-4701-856b-75f9a58a784d · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a248f80a-a5e1-4951-93b9-13e1d36b7984 · inbound
Learning to Reason at the Frontier of Learnability Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f941a00-a54c-4ee6-979b-bb26acf1bf0d · inbound
MoBA: Mixture of Block Attention for Long-Context LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1de12cd-fede-4117-9244-e85470f4b94a · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 48d3ba51-6ecb-458c-842b-9ce19e4cb341 · inbound
Visual-RFT: Visual Reinforcement Fine-Tuning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 61f2d32a-2f4c-4f4a-a371-a3c811f9c7ee · inbound
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd497d56-4312-4bd4-8d76-2b089ab49a81 · inbound
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 89102e30-a3ca-44a5-9881-58c3533f5078 · inbound
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71a9e7e7-0dce-4ccb-a69d-5eaa52a1f685 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 269
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1796957-ff78-4880-9a77-b36772e2866b · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ed0712cc-0800-4619-b28c-bdfa039ab602 · inbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb2847bd-bbcb-4fdc-aa5a-2a26c3025500 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4bb04e3b-a2db-468c-a636-90096ca97130 · inbound
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e319ed2a-463a-442f-8c6e-cbee2b25a1e7 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f747e53-c4ee-49ba-b0f1-0f2fe085ca78 · inbound
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 095a3046-4b55-4e32-8369-ad3a6e39dbe8 · inbound
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 40925960-7ef7-471c-996d-d81765b6e95d · inbound
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d2beea4-e50d-438e-acf1-f56f250e6438 · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e7fddc5-25e0-44ad-a184-5248ab255269 · inbound
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 84ff15ce-ebef-48c2-9808-cad69e8cdd71 · inbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c409a1c8-e932-48b3-aa22-b95800169138 · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54e9dd96-825c-4069-a733-e74f78df2ff5 · inbound
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0edcb6b-d345-46ca-8a7e-518e59e634c0 · inbound
Reinforcement Learning from Human Feedback Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1b512694-8ec9-43ea-be6c-cd971f4682f6 · inbound
ToolRL: Reward is All Tool Learning Needs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e976d878-afad-4ede-b44f-1a7396f5fc7e · inbound
Learning to Reason under Off-Policy Guidance Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b7736300-0613-4043-9c1e-bd38fc37a277 · inbound
Kimi-Audio Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5411608b-f66e-48e2-bd8e-c4f42fd216c6 · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d51f35e-89ac-4b36-ade0-9ed494a3c128 · inbound
Phi-4-reasoning Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e865d535-2f92-40a2-8330-4064c2c4a11a · inbound
Group-in-Group Policy Optimization for LLM Agent Training Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 666e2039-f190-465b-a673-347e31f827fc · inbound
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a42a63b1-4bfd-4501-bbc5-69b026968ef4 · inbound
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a9dca9a3-4128-4376-acc7-12a18e9f0db2 · inbound
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 934b28bb-4b84-43ab-90b0-434af68644ec · inbound
Skywork Open Reasoner 1 Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cb35132-a5a3-451f-82d5-7527f8c1d3e8 · inbound
Grounded Reinforcement Learning for Visual Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a4fc95c6-5257-48ff-b5b5-465603ebac42 · inbound
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c56f9962-2490-41a2-8bd5-18ab8b5d60d3 · inbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a166b24c-de5e-4551-95c1-ce9de0be2e20 · inbound
Steering Your Diffusion Policy with Latent Space Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6c0fd85-71fb-496c-8b74-bfa8053f38ed · inbound
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72d823a9-7154-4c18-9dbd-b136a94a75e3 · inbound
MMSearch-R1: Incentivizing LMMs to Search Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0e63ede9-b656-4c97-8415-86db9d5cda0f · inbound
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f72f700c-fb04-4146-92b8-ef88c2b4f01a · inbound
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45a08cdb-c93e-46ad-a4d7-fa26f212ad43 · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 64a46139-cb99-4a58-be2c-d2cc0558fd1c · inbound
Kimi K2: Open Agentic Intelligence Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cdc8a545-7cdc-4b37-8d5b-ca1b96c3f2dc · inbound
ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e437c7e8-593e-4319-abf4-4ed102200f61 · inbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b5fe955-a88c-408b-91b0-14dc6f1ddaee · inbound
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 79a5917e-3321-462c-9d62-c4859bbe21ea · inbound
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 495ab55e-7b60-4b90-917a-3100a15e9bc2 · inbound
EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5343a5ab-74e2-4848-9367-89c845b1b52b · inbound
Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16b82acc-34d1-4641-88e5-503afeb4c90f · inbound
UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa45072b-0a9e-46aa-959d-8a7e4d6ac8f6 · inbound
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d9c33360-1dc0-4ee4-a6b2-ae2f6bebbbed · inbound
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 634e0119-9e7c-4483-b5ff-2b2d75fbe67f · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 225
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9537a6b7-3ac1-4c0c-bcc5-e1fa48a06faf · inbound
Self-Aligned Reward: Towards Effective and Efficient Reasoners Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff25f9ea-9d50-4bec-a63e-616b03eb82c8 · inbound
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8297df1d-caa8-46aa-9253-1e0b93e05d42 · inbound
Positional Encoding via Token-Aware Phase Attention Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 74123168-75ff-4d0c-ad17-f26f70ac480d · inbound
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e46a6aa9-4bfd-4f67-82dd-d0f155e4dc36 · inbound
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 454907b5-d3db-4382-ac25-dcc3279a8f0a · inbound
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 652fea36-7d7b-4267-b529-c2c7a6595417 · inbound
Structured In-context Environment Scaling for Large Language Model Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1edd9de9-3563-4cfd-909c-e5cc69260978 · inbound
AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74cf059-1eb6-41ef-94ac-0172c31491b6 · inbound
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8462615-fa77-4b81-901a-61b7829f41f3 · inbound
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4547ea09-eeec-48c5-af86-c79ed85c1dd6 · inbound
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ddc515c0-7ad1-46d2-b913-dab5683e33b3 · inbound
Which Heads Matter for Reasoning? RL-Guided KV Cache Compression Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24792794-07c6-4499-91f7-1e8a82d5f00f · inbound
On the optimization dynamics of RLVR: Gradient gap and step size thresholds Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c63b186b-05fb-47b6-a7b8-365e099d1d7e · inbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ecfa51-61e4-48fe-a525-8bfa6dc7a1a4 · inbound
Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8adb4402-be2c-45d4-8862-6d61638836ac · inbound
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1b29aa27-7954-4bf4-8343-a987911f4dc4 · inbound
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38770a1-69dd-46fe-9f6b-2e2af7817e33 · inbound
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6250a448-94fb-4daf-90d1-7da7064947fe · inbound
Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d269007c-82a1-4721-bcf9-8589d34c8e8a · inbound
Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c65c24ce-5113-413c-8eab-149900b53faf · inbound
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bfe9ebae-13de-45fb-888f-7bc8b4c40af2 · inbound
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee9956e-6d00-4ef4-aa60-664224a7a13f · inbound
Kimi Linear: An Expressive, Efficient Attention Architecture Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 890e057e-89b8-4774-8bee-3f1ad332755e · inbound
Sharpness-Guided Group Relative Policy Optimization via Probability Shaping Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b0dc3d5f-5cf6-4b9b-ba7f-6fd593e0749e · inbound
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32b2ea2-1a58-457f-979f-3b3d0eeae31a · inbound
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 09a89a1e-676c-4fb0-8bf3-f3927607c82b · inbound
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4fdfebe1-4ffc-406b-ab43-1fd957c17927 · inbound
Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14a228b8-0e3d-44e5-aee9-78dc66e2842c · inbound
Asking like Socrates: Socrates helps VLMs understand remote sensing images Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6dcaa061-8b6f-4c7c-ae79-ed424a7cf4af · inbound
GENIUS: An Agentic AI Framework for Autonomous Design and Execution of Simulation Protocols Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8f481df-406a-44b2-8ce7-2945308d9dbc · inbound
MOA: Multi-Objective Alignment for Role-Playing Agents Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a0019ae-4485-4311-a5ff-804f88b5e82c · inbound
AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8909b1bc-07cd-4c06-b838-24e173dc2860 · inbound
AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb83637-2f4a-4ae4-9a38-9c349983cf06 · inbound
Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7420bf05-d9d8-4cdc-b062-a7fee0a143f1 · inbound
AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91db51b0-4d4e-4b45-9ca0-0c52ebdd393b · inbound
NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2015259b-08df-4c51-9892-23bf62ba4781 · inbound
NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98645355-1915-40ff-b637-5116521dd9aa · inbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4b2c9cd-ba6e-4513-bdae-886c1419e46a · inbound
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d11159-1771-498d-947f-e2badaf31425 · inbound
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd878a42-f3c8-4de1-8349-2ca61243d6c6 · inbound
Agentic Reasoning for Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 241
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5bdd606-2987-4e21-b619-42fbc18bb306 · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919d3b82-5613-427c-b3c8-cae35a7f5003 · inbound
ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780500d3-4c3a-4b05-9736-390e4d6d965d · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.