IS-CoT framework interleaves planning, writing, and reflection in LLMs to prevent length collapse, yielding IS-Writer-8B that outperforms larger models on long-form benchmarks with better length compliance.
arXiv preprint arXiv:2412.10079 , year=
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 5roles
background 1polarities
background 1representative citing papers
LLMs exhibit a Weakest Link Effect in multi-hop QA where performance collapses to the least visible evidence position; MFAI resolves recognition bottlenecks with up to 11.49% gains in low-visibility spots.
Passages made from high-convergence sentences improve LLM performance on inferential questions compared to cosine similarity selection.
UserGPT introduces a generative LLM framework with a behavior simulation engine, semantization module, and DF-GRPO post-training that scores 0.7325 on tag prediction and 0.7528 on summary generation on HPR-Bench while compressing records by up to 97.9%.
citing papers explorer
-
IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking
IS-CoT framework interleaves planning, writing, and reflection in LLMs to prevent length collapse, yielding IS-Writer-8B that outperforms larger models on long-form benchmarks with better length compliance.
-
Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
LLMs exhibit a Weakest Link Effect in multi-hop QA where performance collapses to the least visible evidence position; MFAI resolves recognition bottlenecks with up to 11.49% gains in low-visibility spots.
-
Context Convergence Improves Answering Inferential Questions
Passages made from high-convergence sentences improve LLM performance on inferential questions compared to cosine similarity selection.
-
UserGPT Technical Report
UserGPT introduces a generative LLM framework with a behavior simulation engine, semantization module, and DF-GRPO post-training that scores 0.7325 on tag prediction and 0.7528 on summary generation on HPR-Bench while compressing records by up to 97.9%.
- The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes