{"total":15,"items":[{"citing_arxiv_id":"2607.01218","ref_index":35,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"The State-Prediction Separation Hypothesis","primary_cat":"cs.CL","submitted_at":"2026-07-01T17:55:09+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A two-stream Transformer variant that separates state storage from next-token prediction improves validation loss and downstream task performance by 2-3 points over standard Transformers.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01065","ref_index":29,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache","primary_cat":"cs.LG","submitted_at":"2026-07-01T15:25:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"GSRQ applies a gain-shape variant of K-means inside residual quantization to improve directional fidelity, raising LongBench accuracy from 11.34 to 33.54 at 1-bit on LLaMA-3-8B.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.31397","ref_index":30,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models","primary_cat":"cs.LG","submitted_at":"2026-06-30T09:25:37+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Mixture-of-Control adaptively combines local and global control states in transformer fine-tuning by treating per-block states as experts in a sparse MoE setup to improve cross-block communication while keeping memory and compute costs comparable to prior state-based methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.16600","ref_index":36,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space","primary_cat":"cs.LG","submitted_at":"2026-05-15T20:00:59+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Pretraining and alignment induce asymmetric geometric traces in transformer weights because alignment updates concentrate in read pathways due to activation covariance while write pathways inherit less structure from alignment losses.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.10468","ref_index":20,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Can Muon Fine-tune Adam-Pretrained Models?","primary_cat":"cs.LG","submitted_at":"2026-05-11T12:34:20+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Constraining fine-tuning updates with LoRA mitigates performance degradation when switching from Adam to Muon on pretrained models.","context_count":1,"top_context_role":"background","top_context_polarity":"unclear","context_text":"tical significance of LoRA's mismatch mitigation by com- puting the reduction in the Adam-Muon performance gap when switching from full fine-tuning to LoRA for each task, and aggregating across all tasks in Tables 2-4 using random- effects meta-analysis. The pooled gap reduction is 0.72% (95% CI: [0.41, 1.04], p <0.001 ) for Muon and 0.83% (95% CI: [0.45, 1.20], p <0.001 ) for Muon-PE, confirming that LoRA significantly mitigates the optimizer mismatch across tasks. 4.4. Effect of LoRA Rank Our analysis in Section 3 suggests that LoRA mitigates optimizer mismatch by limiting updates to the pretrained weights. A natural prediction is that this benefit may dimin- ish at higher ranks, as LoRA increasingly resembles full"},{"citing_arxiv_id":"2605.11007","ref_index":209,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"The Transformer as a Polar State Estimator","primary_cat":"cs.LG","submitted_at":"2026-05-10T08:14:14+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"The paper casts the standard Transformer block with RoPE as a first-order approximation of a radial–tangential state estimator and introduces a Polar Transformer variant that retains the discarded geometric corrections.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.07924","ref_index":35,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation","primary_cat":"cs.LG","submitted_at":"2026-05-08T15:58:22+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Energy-navigated trajectory shaping during training produces 8-step discrete flow matching students that achieve 32% lower perplexity than 1024-step teachers on 170M language models with unchanged inference cost.","context_count":1,"top_context_role":"background","top_context_polarity":"unclear","context_text":"out(x 0, x1)pairs and construct flow states at 20 randomly sampled timesteps and the clean target (t=1.0), producing4 200total evaluations per source distribution. We computeEϕ(xt)and report binned mean energy, monotonicity, rank correlation, and aggregate metrics in tables 9 and 10. Table 9Time-energy monotonicity.Mean energyE ϕ(xt)per time bin; both source distributions. Time bin Src[0,.1) [.1,.2) [.2,.3) [.3,.4) [.4,.5) [.5,.6) [.6,.7) [.7,.8) [.8,.9) [.9,1)t=1 Uni. ¯Eϕ −0.12−0.57−0.75−0.90−1.02−1.10−1.17−1.22−1.29−1.34−1.45 std0.50 0.07 0.06 0.06 0.08 0.07 0.12 0.21 0.21 0.38 0.38 Mask ¯Eϕ −1.94−2.24−2.35−2.43−2.47−2.55−2.59−2.66−2.68−2.75−3.20 std0.34 0.13 0.19 0.21 0.34 0.31 0.42 0.38 0.52 0.55 0.61 The binned means areperfectlymonotone for both source distributions: all 10 consecutive bin pairs show strictly decreasing energy. At the sample level, the uniform compass achieves 97.2% smoothness with only 44 of 200 samples exhibiting any violation, while the mask compass achieves 94.0% with 74 violating samples. The mask compass assigns overall lower (more negative) energies because[Mask]tokens provide a uniform, low-entropy background at unrevealed positions, making partially-revealed sequences more structured. This compresses the total energy dynamic range (1.26 vs. 1.33 for uniform) and the gap between adjacent timesteps, which explains both the higher per-sample violation rate and the weaker Pearson correlation (−0.54vs. −0.80)-the energy-time relationship is less linear for the mask source because the energy curve flattens"},{"citing_arxiv_id":"2605.04901","ref_index":1,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference","primary_cat":"cs.CR","submitted_at":"2026-05-06T13:31:15+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"An attack aligns differently shuffled intermediate activations from secure Transformer inference queries to recover model weights with low error using roughly one dollar of queries.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.04344","ref_index":36,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Perturbation is All You Need for Extrapolating Language Models","primary_cat":"stat.ML","submitted_at":"2026-05-05T23:03:33+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Perturbing the prefix before next-token prediction, during both training and inference, improves out-of-distribution language-model generation and yields a conditional extrapolation guarantee.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.04217","ref_index":17,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Jordan-RoPE: Non-Semisimple Relative Positional Encoding via Complex Jordan Blocks","primary_cat":"cs.LG","submitted_at":"2026-05-05T18:59:31+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Jordan-RoPE realizes a distance-modulated phase basis via non-semisimple Jordan blocks, generating features such as d e^{iωd} for relative positional encoding.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.03229","ref_index":7,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning","primary_cat":"cs.CL","submitted_at":"2026-05-04T23:46:40+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"SMF adds KV memory layers and sparsely updates only heavily-read rows, yielding +2.5pp on MedMCQA with near-zero drift on WikiText and TriviaQA probes versus larger gains but clear forgetting from LoRA and full finetuning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.21265","ref_index":10,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Listen and Chant Before You Read: The Ladder of Beauty in LM Pre-Training","primary_cat":"cs.CL","submitted_at":"2026-04-23T04:20:12+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A music-to-poetry-to-prose pre-training ladder improves small language model perplexity by 17.5% with faster convergence and lower plateau loss.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.19398","ref_index":34,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models","primary_cat":"cs.AI","submitted_at":"2026-04-21T12:26:16+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"GRASPrune removes 50% of parameters from LLaMA-2-7B via global gating and projected straight-through estimation, reaching 12.18 WikiText-2 perplexity and competitive zero-shot accuracy after four epochs on 512 calibration sequences.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2406.03736","ref_index":55,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data","primary_cat":"cs.LG","submitted_at":"2024-06-06T04:22:11+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Absorbing discrete diffusion models the conditional distributions of clean data; reparameterizing yields a time-independent RADD that unifies with AO-ARMs and reaches SOTA perplexity among diffusion models on zero-shot language benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2312.03732","ref_index":43,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA","primary_cat":"cs.CL","submitted_at":"2023-11-28T03:23:20+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"LoRA adapters should be scaled by 1/sqrt(rank) rather than 1/rank to stabilize learning and enable effective use of higher ranks during fine-tuning of large language models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}