Entropy-guided supertokens from BPE on reasoning traces compress LLM outputs by 8.1% on average across models and math benchmarks with no accuracy loss while exposing strategy differences between correct and incorrect traces.
CtrlCoT: Dual-granularity chain-of-thought compression for controllable reasoning.arXiv preprint arXiv:2601.20467
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
E-MRL trains VLMs via RL on a diagnosis-localization-verification MDP with a novel cross-view consistency reward to ground 3D tumor reports in verifiable CT slices.
HMPO is a single-stage RL framework for CoT compression that reports 19-46% token reduction with negligible accuracy loss on models from 9B to 122B parameters across math, code, science, and instruction tasks.
citing papers explorer
-
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
Entropy-guided supertokens from BPE on reasoning traces compress LLM outputs by 8.1% on average across models and math benchmarks with no accuracy loss while exposing strategy differences between correct and incorrect traces.
-
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
E-MRL trains VLMs via RL on a diagnosis-localization-verification MDP with a novel cross-view consistency reward to ground 3D tumor reports in verifiable CT slices.
-
HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression
HMPO is a single-stage RL framework for CoT compression that reports 19-46% token reduction with negligible accuracy loss on models from 9B to 122B parameters across math, code, science, and instruction tasks.