{"total":18,"items":[{"citing_arxiv_id":"2606.29928","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Latent-CURE for Breast Cancer Diagnosis","primary_cat":"cs.CV","submitted_at":"2026-06-29T08:05:16+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Latent-CURE introduces latent-space chain-of-thought reasoning and dual-asymmetric optimization to produce transparent, robust breast cancer diagnoses in imbalanced cohorts.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.21943","ref_index":183,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning","primary_cat":"cs.LG","submitted_at":"2026-06-20T08:20:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"©2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning 3 Existing surveys and how this work differs.Several surveys address specific aspects of RL in LLMs. Sun and van der Schaar [183] focuses on inverse reinforcement learning (reward modeling), while Liu et al. [122] investigate practical implementation choices in policy optimization, extensively evaluating design decisions such as baseline estimation, advantage normalization, and clipping strategies at scale. Others treat RL as one component within broader topics: Plaat et al."},{"citing_arxiv_id":"2606.17924","ref_index":22,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space","primary_cat":"cs.RO","submitted_at":"2026-06-16T13:38:03+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"PearlVLA achieves SOTA on LIBERO by separating VLM representations into visual grounding and an iterative latent plan branch refined via world model queries and RefineNet with process-reward RL.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.07108","ref_index":32,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling","primary_cat":"cs.AI","submitted_at":"2026-06-05T10:02:19+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DyCon dynamically controls reasoning depth in LRMs by modeling evolving difficulty from step-level embeddings, reducing redundant steps across multiple benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.05859","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization","primary_cat":"cs.CL","submitted_at":"2026-06-04T08:30:53+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"TARPO is a pure RL framework using a token-wise action router to switch between discrete token generation and latent reasoning in LLMs, with joint optimization showing outperformance on benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02871","ref_index":26,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Adaptive Latent Agentic Reasoning","primary_cat":"cs.CL","submitted_at":"2026-06-01T20:36:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ALAR trains LLM agents to perform most reasoning in a latent space supervised by actions and escalates to explicit CoT only when needed, cutting tokens by up to 84.6% while preserving accuracy on search and tool-use benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02842","ref_index":21,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning","primary_cat":"cs.LG","submitted_at":"2026-06-01T20:06:50+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SpecFlow represents intermediate visual thoughts in fixed-size DCT space and uses classifier-free guidance to steer updates from textual thoughts, achieving up to 2.1x lower computation and KV cache costs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02248","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Geometric Latent Reasoning Induces Shorter Generations in LLMs","primary_cat":"cs.CL","submitted_at":"2026-06-01T13:40:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"GLR formulates latent reasoning as geometric path approximation in pretrained embedding space and reports shorter LLM generations on math tasks without an explicit length penalty.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.01532","ref_index":91,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete","primary_cat":"cs.LG","submitted_at":"2026-06-01T01:28:42+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"Sliding-window transformers without positional encodings are Turing complete because the sliding window breaks permutation symmetry and suffices to simulate Post machines via a constant-size histogram state.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.28600","ref_index":41,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Transformers Provably Learn to Internalize Chain-of-Thought","primary_cat":"cs.LG","submitted_at":"2026-05-27T15:17:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"L-layer transformers under Log-ICoT curriculum provably learn k-parity with poly(n) samples and log k stages, matching explicit CoT efficiency without inference overhead.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.14323","ref_index":42,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Dynamic Latent Routing","primary_cat":"cs.LG","submitted_at":"2026-05-14T03:35:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Dynamic Latent Routing jointly learns discrete latent codes, routing policies, and model parameters via dynamic search to match or exceed supervised fine-tuning by 6.6 points on average in low-data settings across four datasets and six models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.06165","ref_index":89,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost","primary_cat":"cs.AI","submitted_at":"2026-05-07T12:51:49+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Post-Reasoning boosts LLM accuracy by reversing the usual answer-after-reasoning order, delivering mean relative gains of 17.37% across 117 model-benchmark pairs with zero extra cost.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.26355","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens","primary_cat":"cs.CL","submitted_at":"2026-04-29T07:06:43+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.21027","ref_index":86,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering","primary_cat":"cs.AI","submitted_at":"2026-04-22T19:18:36+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A 22M-parameter hyperbolic model answers structured EHR questions with accuracy close to LLM-based systems (EHRXQA 89.5%, MIMIC-Instr 76.0%).","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.18486","ref_index":90,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation","primary_cat":"cs.CV","submitted_at":"2026-04-20T16:37:22+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"OneVL achieves superior accuracy to explicit chain-of-thought reasoning at answer-only latency by supervising latent tokens with a visual world model decoder that predicts future frames.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"of dense summary vectors prepended to the input, achieving substantial length reduction at only modest accuracy cost. CODI [86] adopts sequence-level self-distillation, training a student model to align its anchor latent hidden state, typically the final hidden representation before the answer, with the teacher model's full chain-of-thought sequence, narrowing the performance gap while preserving efficiency. Token Assorted [90] offers a flexible middle ground by interleaving discrete text tokens and continuous latent tokens within the same sequence, interpolating between fully explicit and fully implicit reasoning. SIM-CoT [105] identifies a latent instability problem, where representations collapse as the number of latent tokens grows without per-step supervision, and addresses it with a plug-and-play auxiliary decoder that aligns each implicit token"},{"citing_arxiv_id":"2604.17892","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LEPO: Latent Reasoning Policy Optimization for Large Language Models","primary_cat":"cs.LG","submitted_at":"2026-04-20T07:05:12+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"LEPO applies RL to continuous latent representations in LLMs by injecting Gumbel-Softmax stochasticity for diverse trajectory sampling and unified gradient estimation, outperforming existing discrete and latent RL methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.08299","ref_index":37,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SeLaR: Selective Latent Reasoning in Large Language Models","primary_cat":"cs.CL","submitted_at":"2026-04-09T14:32:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SeLaR selectively applies latent soft reasoning in LLMs via entropy gating and contrastive regularization, outperforming standard CoT on five benchmarks without training.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2503.16419","ref_index":163,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models","primary_cat":"cs.CL","submitted_at":"2025-03-20T17:59:38+00:00","verdict":"ACCEPT","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A survey organizing techniques to achieve efficient reasoning in LLMs by shortening chain-of-thought outputs.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Distilling 2-1 [219]; C3oT [78]; TokenSkip [194]; CoT-Valve [130]; Self-Training [133]; Learnto Skip [115]; Token-Budget [58]; Verbosity [72]; Stepwise [31]; Z1 [223]; Prune-on-Logic [243];LS-Mixture SFT [218]; DRP [75]; AutoL2S [125]; Assembly of Experts [79]; Ada-R1 [126];ConCISE [145]; VeriThinker [19]; R1-Compress [187]; CTS [226]; A∗-Thought [205]; TLDR [96];OThink-R1 [235]; PNS [220]; ReCUT [77]; StepEntropy [94]; ASAP [229]; ReasoningOutput-basedEfficient Reasoning LatentRepresentationCompression e.g. Coconut [60]; CODI [155]; CCoT [20]; Heima [153]; Token Assorted [163]; Loop [149];SoftCoT [206]; Back Attention [222]; CoLaR [167]; SEAL [14]; Overclocking [41];Controlling [106]; DynamicReasoningParadigm e.g. Speculative Rejection [165]; Sampling-Efficient TTS [188]; DPTS [38]; Certaindex [50];Dynasor-CoT [51]; Fast MCTS [89]; ST-BoN [188]; More is Less [193]; RSD [100]; SpeculativeThinking [213]; SCoT [181]; SpecSearch [189]; ValueFree [148]; GG [54]; VGS [183]; DO [61];FFS [1]; DORA [185]; SPECS [10]; Best-Route [37]; LightThinker [231]; INFTYTHINK [209];SCoT [195]; RASC [178]; Adaptive Reasoning [224]; AdaptiveStep [119]; Self-Calib [70];CISC [170]; ESC [92]; DSC [184]; PathC [247]; RPC [246]; Sleep-time Compute [104];SpecReason [138]; TOPS [214]; Retro-Se"}],"limit":50,"offset":0}