Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T09:28:16.189617Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 100 inbound Pith citation observations for arXiv:2506.13585.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T09:28:16.189617Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T21:20:34.201020Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T22:26:37.204571Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e68edf5f-de20-434f-bd6c-dfeda10c4553 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Simple linear attention language models balance the recall-throughput tradeoff
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5b3e7e35-31d4-4657-b486-35d5ac92efa7 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9c05ef52-b361-4e23-b9fb-94875aad42da · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Titans: Learning to Memorize at Test Time
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ec78ae99-6974-4acb-bfb4-de531ad346aa · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Longformer: The Long-Document Transformer
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b29f544-0aea-41a5-8357-dcd458e463e4 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 56f03810-65e0-453b-8ed5-1da4fd01f4a0 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 593fad5d-b88f-46bc-837a-d9b2a412b8a8 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 21d32be9-0bca-4a8b-a714-84827c5d8f7e · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 60964a14-e308-481f-ae78-743bba18e333 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Jacob Dunefsky, Philippe Chlenski, and Neel Nanda
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9654d861-6642-4b76-88b4-5557070304b0 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Zamba: A Compact 7B SSM Hybrid Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e3659577-ad0c-4230-a534-68a211f70d99 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Efficiently modeling long sequences with structured state spaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation eafeb127-343a-4d17-87f4-02c7f3e9b5d8 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 37b134cb-284e-42a6-9895-08da0793bb76 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Measuring Mathematical Problem Solving With the MATH Dataset
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0f551b5e-2242-45cc-98a4-232f5899fd46 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f80e8e31-7220-4a05-bfc8-859117216744 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Jamba-1.5: Hybrid Transformer-Mamba Models at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c56f9962-2490-41a2-8bd5-18ab8b5d60d3 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 16742d04-cdeb-40b6-a297-36ef8ac30758 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 67260a79-9d62-4e4f-986a-a8294be5bfab · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 331728be-3b51-4110-b04a-41cbf0c5ed92 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Understanding R1-Zero-Like Training: A Critical Perspective
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 91649ab1-c636-4a27-89a8-ad28dba2fafb · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0dc7b9d1-f66e-4e67-a50d-3c2d18760fbe · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Parallelizing linear recurrent neural nets over sequence length
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d4ef5545-0d5b-4647-9c1f-e04512a18721 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 65ff9138-aa07-45f5-819a-2f8d2b8f82d2 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention A Theory on Adam Instability in Large-Scale Machine Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 16f2f937-4798-4a27-8407-540cfc123fe9 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Openai mrcr dataset
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation da63c11c-e263-4b29-bf34-bb5d6546aba4 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1c9fa2bd-f427-4dbb-a00c-f6156be8024b · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Humanity's Last Exam
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T22:08:32.844211+00:00.
Observation 0ca1a511-b229-49fe-8dcc-d6e0fd5a6318 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention The devil in linear transformer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 858af31e-a12c-47ee-8cee-705f48de85b2 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2abd9990-72fc-4936-8a72-533e2a802a8a · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 86f5c81e-a8b6-4b0a-97e8-c95199e0703b · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d3aa5dac-3974-4304-85ac-82d203f22792 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 57125760-c6af-406e-96dd-e4250e091bd0 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0ccd4f7e-c438-4461-ab6b-800a4b5c448e · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Scaling laws for linear complexity language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6b8eed45-b87b-48cf-bab6-8411df851386 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention HybridFlow: A Flexible and Efficient RLHF Framework
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5a94e21b-abd7-44fa-b043-1f857267facf · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation dc341404-69b9-44f8-917a-5b76f4040c2a · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Deltaproduct: Im- proving state-tracking in linear rnns via householder products
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0e6f16b8-a732-4ca0-90e5-adf7914b8582 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3a54b0c2-3903-4112-8156-c39f53c51fa1 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 798487c2-6b20-4806-aba4-8317a974ad20 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Learning to (Learn at Test Time): RNNs with Expressive Hidden States
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 624bdbb1-782b-4c71-b378-26dc17b3111c · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Retentive Network: A Successor to Transformer for Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ba43a89e-9bb5-4d29-85e7-7f2927d72a1f · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cabe8746-15d0-422b-ab9f-301c54448768 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e33cebfb-ea64-479b-89db-b92606590f77 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation aa04c82a-eca2-449a-a320-c53ea51779a6 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Measuring short-form factuality in large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 99e26778-80cd-4cc1-a9fa-d38686e4cf07 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Agentless: Demystifying LLM-based Software Engineering Agents
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f0db653e-4769-4fec-9817-2f1f7d4535b5 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 634b95bd-8e8e-4906-bf3a-5e771fea1811 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fac225b2-26b8-4575-bb79-0f177dd04143 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5fd4e1fa-3aa3-4669-be5d-55a1ad0c4583 · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 71b77b6d-887f-484d-b7d6-defffc02baff · outbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3e09aa0f-8239-440b-992d-2301121851ed · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6b011508-66c4-4c06-809c-a7c39b4c800c · inbound
Reinforcement Learning from Human Feedback MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c73d9f34-1448-4c09-92c8-74a51a692335 · inbound
Group Sequence Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 970fa807-a60c-49d6-a484-bfb9ec99dfb1 · inbound
The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 32620d74-df48-4bcd-a248-defddb64ab0e · inbound
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b4ae12b1-fc5f-4937-9669-dc814ac88bb9 · inbound
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7d1b4d7a-7c71-4105-8913-838fc92e4dd8 · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2183f214-97a3-4e5a-adfd-749caa639e52 · inbound
OneRec-V2 Technical Report MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6efab08b-9e24-4260-9bcf-bd7ac153841d · inbound
A Survey of Reinforcement Learning for Large Reasoning Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c6916d2d-64c6-4933-839e-720cf355372d · inbound
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 16e52cfc-216f-4d74-b354-52098de5ed16 · inbound
The Art of Scaling Reinforcement Learning Compute for LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 25b384bf-a634-4caf-a3b9-6bdb3acea3b1 · inbound
SSPO: Subsentence-level Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ffcfdc02-7916-4590-8457-9ce5e23b79b5 · inbound
From Ranking to Reasoning: Explainable Web API Recommendation via Semantic Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 49ab84e1-f9a9-48b3-bb11-c090793019c2 · inbound
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea722621-c861-4a8c-a7a6-6e6c108e1baf · inbound
Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fc2a974b-b169-4d1a-9190-b1f39d27718a · inbound
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 20298989-d242-4b22-9495-ce66dcc40aef · inbound
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bdb29954-907c-48f1-a257-01f58658a29b · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 58dbb434-80ab-4321-8210-2d25144aa89d · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ab3d3b-027d-4553-895e-a50d9a704d3c · inbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4443504-4134-4d24-bede-585fccf716e6 · inbound
Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2a83887d-1d26-4f19-b7d5-de748c4df8af · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de312d3a-41bd-4ffc-95b6-86e7050505b2 · inbound
When RL Meets Adaptive Speculative Training: A Unified Training-Serving System MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2e920905-a857-441f-801a-235ad2dd578a · inbound
SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2b21eaa7-957c-4b05-a7db-64ea4f9633dc · inbound
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d4238400-8aa8-40c3-aa10-5ab80151159a · inbound
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 02ec5989-ebb3-4bc6-927f-9c4a4c8737eb · inbound
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d229a8a5-acd8-4255-a874-362312e2bd32 · inbound
Soft Sequence Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd8f254-9ddd-4193-9ac6-e7ee3efeba62 · inbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404904fe-7036-4f2d-a49d-842460e67567 · inbound
Stabilizing Policy Optimization via Logits Convexity MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a1b6f1-47b9-48a6-912b-1279eefc91f6 · inbound
STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826de674-94c9-4999-ba5f-478f8030ec94 · inbound
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5ee11886-eb84-478a-9375-506b66d0ab64 · inbound
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab4f3d66-e961-4bf2-80bb-bcca4c1a1447 · inbound
Policy Improvement Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a8e293ed-5ea8-4fa0-a606-56021edcbb8e · inbound
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 47233751-3fc5-4d2f-86e9-1e806c38c06d · inbound
SAGE: A Service Agent Graph-guided Evaluation Benchmark MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4673c71a-edd0-4790-863a-1de397d6c26b · inbound
MEMENTO: Teaching LLMs to Manage Their Own Context MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1325ab61-421d-41af-b60d-d1e13849a183 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4761a986-7f68-48d8-9da8-d9427ce7edb6 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e4225109-fd0a-49b0-932a-51456e662cbc · inbound
Beyond Distribution Sharpening: The Importance of Task Rewards MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 592c9e31-cc86-43bb-879a-b8e8f9e117a8 · inbound
Scaling Self-Play with Self-Guidance MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d81cc469-82bb-4da0-ba32-304f0a53c70f · inbound
Building a Precise Video Language with Human-AI Oversight MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c3306625-b3bb-4f70-aa84-832274d4fed5 · inbound
Cost-Aware Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd34b05a-6c76-4482-87b8-fc79236e14c5 · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 13159349-17ac-41ca-ac49-0bd6baf553bb · inbound
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8f0e9aa9-b5d1-471d-aceb-7d0cedfb200d · inbound
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b3e036a8-8ebb-43b7-bced-178e70046601 · inbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9d543917-dfea-41ce-bc7e-b3bc3c1adb73 · inbound
ZAYA1-8B Technical Report MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 164
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 59b9e726-dfb3-404f-bc6a-cad1101ea784 · inbound
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e18527fe-41b6-422f-9a24-faf07d29ef46 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 43fff8e8-7ccc-47bc-9c6c-7260f7cd1fa7 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 016ba4b9-6d44-450f-bac6-581e01c96a98 · inbound
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation df9dd6c4-b80e-44f1-94e4-e6a3681bcc5f · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d87f61ed-4dd0-4d57-a869-e9d4312d91d4 · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ad615347-e005-4078-8bd6-bdcabfffad1e · inbound
KL for a KL: On-Policy Distillation with Control Variate Baseline MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 73390f12-bb59-4489-a813-743d78cc460c · inbound
Priming: Hybrid State Space Models From Pre-trained Transformers MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 528b2895-8b71-464c-92a7-c2eac2d97257 · inbound
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e7494e4d-7876-4197-ba33-222885f30225 · inbound
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b9beb139-041a-48c5-8ca1-e0ea276e2ab6 · inbound
ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f7d85ecf-0f21-4be1-bb20-4448c4dd40e9 · inbound
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 81dff8b9-4bfc-4a25-aa92-56ef2365ab0b · inbound
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a3e68dec-3125-4f92-92e4-0e351725efc6 · inbound
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 243d28b7-4cab-4dcb-aa7b-30b7bef80327 · inbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f4e1da17-d336-4264-ab27-989813359986 · inbound
Learning, Fast and Slow: Towards LLMs That Adapt Continually MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 535f1e6b-0f6e-4ca1-be81-d4523d591ec8 · inbound
Learning, Fast and Slow: Towards LLMs That Adapt Continually MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 973ec6af-f331-4fc5-8496-d04104a314fa · inbound
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 43feeb7a-5639-4b33-af6b-461beb8e7121 · inbound
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4ffc632f-bad1-4fe8-b374-29c830ddcb33 · inbound
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 219c7ddb-d5e2-42b4-976c-0e857e6f353b · inbound
MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1073cbd1-ab0b-49d0-ba04-7735432bf065 · inbound
D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6262dd51-39d7-44ee-85ed-93f7170c6f54 · inbound
D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e540ea9f-088f-4a8f-85db-d5f7c9dc722c · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d46f416a-957b-4119-96f0-016ef6d78a3b · inbound
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation dab4ccb6-f27a-449f-87aa-0db2c25d18af · inbound
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b8f3cae1-2e6f-429a-bcd8-e626ad505f2e · inbound
Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 28a354d0-1b5e-4951-b3ac-d8ea05325266 · inbound
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 972dc449-351d-45e5-b653-90d1a3d809b5 · inbound
BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a461d5d7-ff41-4cf7-9c3a-72fc634dacd6 · inbound
BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 56b928d6-8271-49f6-a332-83d15bb20553 · inbound
KVBuffer: IO-aware Serving for Linear Attention MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 219827f1-287f-4231-8e00-68d4379b366b · inbound
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b9511ce5-506c-4d82-ae68-815637aac990 · inbound
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 44906f4b-b4bc-4132-b2a4-7acccd1d3ac0 · inbound
One-Way Policy Optimization for Self-Evolving LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5e0859d6-e3d1-4d7d-a264-7c97c2c886c0 · inbound
Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ec7be72c-99cc-44f9-9651-00d30ef08ec7 · inbound
Extreme Region Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bde85b5b-d6b2-4f13-9184-fd7844d8d025 · inbound
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c0ee9ffe-4d9e-4825-88e0-d45f00133734 · inbound
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 29c8779f-fdda-409e-80f1-508e656c1a1f · inbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ef2d275e-88f3-46ad-b6ba-6e2b3a692a95 · inbound
Trust Region On-Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2dbb7668-e0ee-4cf6-b1c9-42c231482f3b · inbound
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3102dbf0-819d-480e-bf9a-ed70f4e3c6e5 · inbound
Building Better Activation Oracles MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1a887651-8d7f-4103-9808-8dffc4b75f73 · inbound
CodegenBench: Can LLMs Write Efficient Code Across Architectures? MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fe087c66-1ca6-4dc4-825d-69e0e5a4391e · inbound
Rollout-Level Advantage-Prioritized Experience Replay for GRPO MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1d23178c-503c-493e-9650-a536e97ea135 · inbound
AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8ae9c60e-72cb-40cd-b58a-313cf0ae28c2 · inbound
Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 79b6cc93-3e73-4351-b30f-205024855454 · inbound
Rethinking the Divergence Regularization in LLM RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 78afcae2-efc6-422d-97f7-29e17030d033 · inbound
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f3be0ae9-5ee3-49b2-a6ba-95ec85a59737 · inbound
Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 01ac48bd-6ae6-43fa-a5de-407bc10e5e52 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b6b38a68-9110-4cd3-bba9-22a29141e4f2 · inbound
Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 288
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fffc33e7-5c0c-41f3-9f00-c91e0f7b3f80 · inbound
APPO: Agentic Procedural Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.