REVIEW 2 major objections 3 minor 55 references
Around a trained multitask LoRA checkpoint, a single task perturbation is first-order predictable through scale 1e-2, but two sequential task updates stop commuting at a pair-specific scale set by a curvature commutator with onset ηκ≈0.1.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:50 UTC pith:YKU5H7TK
load-bearing objection A careful, honestly-scoped empirical atlas showing that one-step local adaptation is first-order predictable to 1e-2 while two-step order effects can break inside that window; the Lie-bracket law is solid in fixed LoRA coordinates, but gauge-dependence and the failed cross-fit limit its reach. the 2 major comments →
First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that the local geometry around a trained checkpoint separates cleanly by symmetry into three objects with different scales: self-curvature governs single-direction antisymmetry, a symmetric mixed derivative governs simultaneous activation additivity, and the antisymmetric Lie bracket H_B g_A − H_A g_B governs the order of two sequential memoryless gradient updates. In the fixed LoRA factor coordinates used by the optimizer, the normalized endpoint defect is c(η)=ηκ+O(η²) with κ = ‖H_B g_A − H_A g_B‖/‖g_A+g_B‖; when both the two-step paths and the Hessian-vector products use the same frozen minibatches, finite-difference and HVP routes agree at median ratio 1.002, and
What carries the argument
The Lie bracket of the two task-gradient vector fields, H_B g_A − H_A g_B, normalized by ‖g_A+g_B‖. It is the leading order-dependent difference between the two update orders: whichever task goes second takes its step from ground the first task has already moved, so the expansion produces the bracket at order η². Two Hessian-vector products at the operating point compute κ, and the onset prediction η†≈0.10/κ is parameter-free. The measurements all live in the fixed Euclidean coordinates of the LoRA factors, which the paper states explicitly is gauge-dependent.
Load-bearing premise
Every measured radius, window, and onset scale is quoted in the fixed Euclidean coordinates of the LoRA factors, and those coordinates are not unique: the same network can be represented with different factor scalings, so the headline numbers might be an artifact of that coordinate choice.
What would settle it
Re-run the fine-grid commutator protocol after applying a function-preserving gauge transformation to the LoRA factors (e.g., multiply A by c and B by 1/c) and recompute the onset for the same task pairs; if the products η†κ scatter outside ~0.10 and the onset moves materially, the law is coordinate-specific. A second falsifier is to compute the bracket in a function-space or Fisher metric and check whether the median ratio 1.002 survives.
If this is right
- For a single intended direction, probe-loss changes are first-order predictable through the tested grid on every model median, so random proposals can be ranked by their gradient projection before running any of them.
- For two sequential memoryless gradient steps, update-order sensitivity is set by κ: two Hessian-vector products give the warning onset η†≈0.10/κ, and pair-to-pair differences in κ reach up to ~15× within one model.
- No universal radius exists: on over a third of measured model–task-pair combinations, order sensitivity starts inside the single-update window, so composition claims must be checked per pair and per scale.
- Activation additivity at full task-vector scale fails on several models, including both held-out 7B models; additivity holds on all models only for fractions α≲0.3 under this probe.
- The bracket is a same-probe diagnostic, not a transferable predictor: cross-sample prediction of endpoints or task effects from the bracket was not validated.
Where Pith is reading between the lines
- Beyond the paper: because all scales are quoted in the fixed LoRA factor coordinates, a function-preserving rescaling of the factors could move the measured windows and onsets; the 1e-2 and 0.1/κ numbers should be expected to change under a gauge-invariant metric.
- Beyond the paper: the matched-route success and cross-sample failure together suggest that minibatch stochasticity, not just curvature, is what breaks transfer; a stochastic-bracket estimator with covariance correction is a natural next test.
- Beyond the paper: the symmetry split (self-curvature vs symmetric mixed derivative vs antisymmetric bracket) predicts that each adaptation tool needs its own diagnostic; a single 'linear regime' radius is unlikely to exist for any model.
- Beyond the paper: the η†≈0.10/κ law is cheap to check on new architectures and task pairs; if it reproduces, update-order sensitivity can be screened without running the two-step optimization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper measures eight local properties of the loss landscape around a fixed multitask LoRA operating point on nine transformers (82M–7B), using a prospectively registered property list, thresholds, development/held-out split, and a uniform harness. The main positive finding is single-update first-order predictability: along a task's gradient, probe-loss changes are dominated by the linear term through the tested scale 10^-2 on every model median, and first-order scores rank random perturbations. The main negative finding is pairwise fragility: order sensitivity of two sequential task-gradient steps sets in at pair-specific scales, sometimes well inside the single-update window. For two memoryless gradient steps, the leading order-dependent endpoint difference is the Lie bracket H_B g_A - H_A g_B, and with matched minibatches the normalized defect satisfies c(eta)=eta*kappa+O(eta^2) at median ratio 1.002; the onset eta_dag defined by c(eta)=0.10 gives eta_dag*kappa close to 0.10. The paper also reports several registered failures: the mixture-gradient covariance low-rank test is non-falsifiable at m=32 and shows no plateau up to m=256, activation additivity fails at full task-vector scale on several models including both held-out 7B models, the global mean-vector weight-to-steering correspondence fails on all model medians, and the cross-fitted functional forecast from the bracket does not validate.
Significance. The paper is methodologically strong in several respects: the pre-registration, explicit reporting of failed controls, per-seed tables, threshold-sensitivity analysis, and the clean shared-minibatch HVP/finite-difference agreement at median ratio 1.002 are all commendable. The atlas across nine models, with a single frozen harness, is a useful empirical resource, and the honest reporting of the failed cross-fitted forecast is scientifically valuable. If the results hold, the practical message—check the diagnostic that matches your operation at your scale and on your pair—is sensible. However, the central quantitative claim about the onset scale eta_dag ~ 0.10/kappa and its three-order-of-magnitude span across models and pairs is coordinate-specific and partly definitional, so its interpretation as a statement about task geometry requires additional support or careful reframing.
major comments (2)
- [§6.2, Table 3/Fig. 9 and §3 'Coordinate convention'] The central quantitative claim—that kappa is a pair-specific warning signal whose scale and ordering span three orders of magnitude across models—is measured entirely in fixed Euclidean coordinates of the LoRA factors. The paper acknowledges the gauge freedom (B,A) -> (BQ, Q^{-1}A) and states in Limitation 1 that no gauge-rescaling, balanced-gauge, Fisher, or output-KL control is reported. This is not merely a philosophical caveat: gradient, Hessian, and the HVP commutator are all coordinate-dependent, so both kappa and the measured onset eta_dag can change under a function-preserving reparametrization. As written, the abstract's claim that eta_dag ~ 0.10/kappa 'spans three orders of magnitude across models and task pairs' and the Section 8 advice to use kappa as a pair-specific order-sensitivity diagnostic are not yet supported as statements about task geometry; they are statements abou
- [§6.1, Eq. (10) and §6.2] The 'law' eta_dag*kappa ~ 0.10 is a consequence of the definition of eta_dag as the first point with c(eta) >= 0.10 together with the leading-order expansion c(eta) = eta*kappa + O(eta^2); it is not an independent empirical discovery. The paper acknowledges this ('Because eta_dag is read from the same defect curve... the substantive content is route-to-route consistency'), but the abstract and Section 6.2 headline the onset span as a main result. The actual empirical content is the pointwise agreement between finite-difference defects and HVP commutators at median ratio 1.002 (Fig. 8b), which is strong. Please restructure the presentation so that the route-to-route consistency is the claim, and the product eta_dag*kappa ~ 0.10 is presented as a consistency check of the threshold definition rather than as a separate falsifiable law.
minor comments (3)
- [Abstract and §5.4] The one-direction window is right-censored at 10^-2 on every model median, so the phrase 'through the tested scale 10^-2' is accurate but should be read as a lower bound. The common-ruler analysis in §5.4 is a descriptive reanalysis and is clearly labeled as such; consider stating this in the abstract as well.
- [Table 2 and §6] DistilGPT-2 and OPT-1.3B are excluded from the fine-grid commutator measurements because of failed HVP reconstruction checks, but Table 2 lists P8 onsets for them on the coarse grid. This should be stated more prominently, as the 'three orders of magnitude' span is over the seven passing models, not all nine.
- [§6.5] The lambda_max association is explicitly exploratory and based on seven model medians, with a failed out-of-sample forecast. The text already qualifies this appropriately; the one seed-level reproducibility mismatch (Pythia-160M seed 1) should also be mentioned in the main text rather than only in the scoring-script flag.
Circularity Check
Central Lie-bracket expansion is genuine and self-contained; the onset law eta^dag~0.10/kappa is a disclosed arithmetic corollary of the definition of eta^dag, not an independent prediction.
specific steps
-
self definitional
[Section 3 (P8: 'we report the onset η† = min{η : c(η) ≥ 0.10}'); Section 6.2; Table 3]
"we report the onset η† = min{η : c(η) ≥ 0.10} ... Because η† is read from the same defect curve, η†κ≈ 0.10 is arithmetic once the leading expansion holds pointwise; the substantive content is route-to-route consistency."
η† is defined as the crossing of the measured defect curve c(η) at 0.10, and κ is the measured slope of that same curve (c(η)=ηκ, verified with matched finite-difference/HVP estimators). At the crossing, 0.10=c(η†)≈η†·κ, so the pre-registered 'prediction' η†·κ≈0.10 (Table 3: every product in [0.096,0.109]) is an arithmetic corollary of the definition plus the verified slope law, not an independent test. The paper discloses this explicitly. Independent content survives in the route-to-route slope agreement (κ from double-backward HVPs vs finite-difference two-step endpoints, median ratio 1.002, 94% of mid-range points within 10%), and the three-orders-of-magnitude span is a real measured property of κ; but the onset relation itself adds no measurement beyond rescaling κ by the fixed thresho
full rationale
The core derivation is not circular: Section 6.1 re-derives the Lie-bracket leading term by elementary Taylor expansion, and the median ratio 1.002 between the HVP-computed κ and the finite-difference defect is a genuine two-route consistency check (no fitted parameters; the routes compute different objects). The single partially definitional element is the onset law η†≈0.10/κ, which reduces to the definition of η† plus the verified pointwise law; the paper itself states this ('arithmetic once the leading expansion holds pointwise'), so it is disclosed, partial, and a corollary rather than the central claim. No load-bearing self-citation: the companion best-of-N formula [31] is re-derived and corrected here with standard normal order statistics, and the bracket attribution to [34,39] does not carry the derivation since Section 6.1 re-derives it. The paper's own reported failures and scoping statements reduce circularity concerns: the cross-fitted forecast fails (R2_zero=-560.4), the λmax forecast is undecidable, P2's registered bar is exposed as algebraically non-falsifiable, P6 is disclosed as in-sample, and Limitation 1 concedes the fixed-coordinate, gauge-dependent scope. Gauge dependence is a validity/robustness concern about what the numbers mean, not a reduction of the derivation to its inputs, and per the reviewing rules it does not count as circularity. Score 3 reflects the one disclosed, definitional corollary while recognizing that the substantive two-route check stands.
Axiom & Free-Parameter Ledger
free parameters (4)
- commutator threshold c>=0.10 =
0.10
- P1 antisymmetry threshold A<0.10 =
0.10
- P4 additivity threshold epsilon_add<0.15 =
0.15
- P5 correspondence cosine bar =
0.30
axioms (5)
- ad hoc to paper Euclidean geometry in the fixed LoRA factor coordinates is the right metric for the measured radii and curvatures.
- domain assumption The loss and activation maps are smooth enough for second-order Taylor expansions.
- domain assumption The fixed 32-example multitask probe is representative of the task loss, gradient, and Hessian.
- domain assumption Memoryless gradient steps without optimizer state are the relevant update model for P8.
- domain assumption One 300-step multitask LoRA operating point represents adaptation behavior across model families and scales.
read the original abstract
Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space. We measure 8 such properties with the same harness around a multitask LoRA operating point, on 9 transformers (82M-7B), with a prospectively registered property list, thresholds, and test split. We find a shared one-direction validity window up to the tested scale $10^{-2}$, but no universal radius for pairwise composition or update ordering. Along individual directions, changes of the probe loss remain first-order predictable throughout the grid: a perturbation's effect on the loss is essentially its projection onto the gradient, which is also what makes local random search work. Pairwise structure, however, proves to be far more fragile: on over a third of the measured (model, task pair) combinations, two-update order sensitivity sets in strictly inside that window; task-gradient subspaces rotate within tens of steps; additivity under our fixed activation probe fails at full task-vector scale on several models, including both held-out 7B models; and no model median passes the registered global mean-vector weight-to-steering correspondence bar. For two sequential task-gradient steps, the leading order-dependent term is the Lie bracket $H_B\textbf{g}_A-H_A\textbf{g}_B$; its normalized prediction $c(\eta)=\eta\kappa+O(\eta^2)$ tracks the measured defect at median ratio 1.002, while the onset scale $\eta^\dagger\approx0.10/\kappa$ spans three orders of magnitude across models and task pairs.
Figures
Reference graph
Works this paper leans on
-
[1]
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. InAnnual Meeting of the Association for Computational Linguistics (ACL), 2021
2021
-
[2]
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.Annals of Probability, 33:1643–1697, 2005
Jinho Baik, Gerard Ben Arous, and Sandrine Peche. Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.Annals of Probability, 33:1643–1697, 2005
2005
-
[3]
Florent Benaych-Georges and Raj Rao Nadakuditi. The singular values and vectors of low rank perturbations of large rectangular random matrices.Advances in Mathematics (arXiv:0910.2120), 2011
Pith/arXiv arXiv 2011
-
[4]
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. Pythia: A suite for analyzing large language models across training and scaling. InInternational Conference on Machine Learning (...
Pith/arXiv arXiv 2023
-
[5]
GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow. Zenodo, 2021. URL https: //doi.org/10.5281/zenodo.5297715
-
[6]
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In NAACL-HLT, 2019
2019
-
[7]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
Pith/arXiv arXiv 2018
-
[8]
da Silva, Mohammed Adnan, Felix Dangel, and Sageev Oore
Marvin F. da Silva, Mohammed Adnan, Felix Dangel, and Sageev Oore. Generalizing the geometry of model merging through Fréchet averages.arXiv preprint arXiv:2604.27155, 2026. URLhttps://arxiv.org/abs/2604.27155
Pith/arXiv arXiv 2026
-
[9]
Sharp minima can generalize for deep nets.arXiv:1703.04933 (ICML), 2017
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets.arXiv:1703.04933 (ICML), 2017
Pith/arXiv arXiv 2017
-
[10]
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht. Essentially no barriers in neural network energy landscape. InInternational Conference on Machine Learning (ICML), 2018. URLhttps://arxiv.org/abs/1803.00885
Pith/arXiv arXiv 2018
-
[11]
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware mini- mization for efficiently improving generalization.arXiv:2010.01412 (ICLR), 2021
Pith/arXiv arXiv 2010
-
[12]
Deep ensembles: A loss landscape perspective.arXiv preprint arXiv:1912.02757, 2019
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. Deep ensembles: A loss landscape perspective.arXiv preprint arXiv:1912.02757, 2019
Pith/arXiv arXiv 1912
-
[13]
Yulu Gan and Phillip Isola. Neural thickets: Diverse task experts are dense around pretrained weights.arXiv preprint arXiv:2603.12228, 2026. URLhttps://arxiv.org/abs/2603.12228
arXiv 2026
-
[14]
Loss surfaces, mode connectivity, and fast ensembling of DNNs
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry Vetrov, and Andrew Gordon Wilson. Loss surfaces, mode connectivity, and fast ensembling of DNNs. InAdvances in Neural Information Processing Systems (NeurIPS), 2018. URLhttps://arxiv.org/abs/1802.10026
Pith/arXiv arXiv 2018
-
[15]
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao. An investigation into neural net optimization via hessian eigenvalue density. InICML, 2019
2019
-
[16]
OLMo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. OLMo: Accelerating the science of language models. InAnnual Meeting of the Association for Computational Linguistics (ACL), 2024. URLhttps://arxiv.org/abs/2402.00838
Pith/arXiv arXiv 2024
-
[17]
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer. Gradient descent happens in a tiny subspace. arXiv:1812.04754, 2018
Pith/arXiv arXiv 2018
-
[18]
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. InInternational Conference on Learning Representations (ICLR), 2022. URLhttps://arxiv.org/abs/2106. 09685. 27
2022
-
[19]
DistilGPT2 model card
Hugging Face. DistilGPT2 model card. Hugging Face model repository, 2019. URLhttps: //huggingface.co/distilbert/distilgpt2. Accessed 2026-07-10
2019
-
[20]
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. InInternational Conference on Learning Representations (ICLR), 2023. URLhttps://arxiv.org/abs/2212.04089
Pith/arXiv arXiv 2023
-
[21]
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. InInternational Conference on Learning Representations (ICLR), 2017. URLhttps: //arxiv.org/abs/1609.04836
Pith/arXiv arXiv 2017
-
[22]
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. Measuring the intrinsic dimension of objective landscapes. InInternational Conference on Learning Representations (ICLR), 2018. URLhttps://arxiv.org/abs/1804.08838
Pith/arXiv arXiv 2018
-
[23]
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. Visualizing the loss landscape of neural nets. InAdvances in Neural Information Processing Systems (NeurIPS), 2018
2018
-
[24]
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. InAdvances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[25]
Locating and editing factual associations in gpt.NeurIPS (arXiv:2202.05262), 2022
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in gpt.NeurIPS (arXiv:2202.05262), 2022
Pith/arXiv arXiv 2022
-
[26]
Phase transitions for feature learning in neural networks
Andrea Montanari and Zihao Wang. Phase transitions for feature learning in neural networks. arXiv:2602.01434, 2026. URLhttps://arxiv.org/abs/2602.01434
arXiv 2026
-
[27]
Julien Nicolas, Mohamed Maouche, Sonia Ben Mokhtar, and Mark Coates. Dome: Im- proving signal-to-noise in stochastic gradient descent via sharp-direction subspace filtering. arXiv:2507.03545, 2025. URLhttps://arxiv.org/abs/2507.03545
arXiv 2025
-
[28]
Vardan Papyan. Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians.arXiv:1901.08244 (ICML), 2019
Pith/arXiv arXiv 1901
-
[29]
Vardan Papyan. The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size.arXiv:1811.07062, 2019
Pith/arXiv arXiv 2019
-
[30]
Asymptotics of sample eigenstructure for a large dimensional spiked covariance model.Statistica Sinica, 17:1617–1642, 2007
Debashis Paul. Asymptotics of sample eigenstructure for a large dimensional spiked covariance model.Statistica Sinica, 17:1617–1642, 2007
2007
-
[31]
Irina Piontkovskaia and Sergey Nikolenko. Recoverable but not stationary: Local linear structures in weights and activations.arXiv preprint arXiv:2606.10929, 2026. URL https: //arxiv.org/abs/2606.10929
Pith/arXiv arXiv 2026
-
[32]
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. Steering llama 2 via contrastive activation addition. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15504–15522,
-
[33]
Discretization drift in two-player games
Mihaela Rosca, Yan Wu, Benoit Dherin, and David G T Barrett. Discretization drift in two-player games. InICML, 2021. arXiv:2105.13922
Pith/arXiv arXiv 2021
-
[35]
Ugur Guney, Yann Dauphin, and Leon Bottou
Levent Sagun, Utku Evci, V. Ugur Guney, Yann Dauphin, and Leon Bottou. Empirical analysis of the hessian of over-parametrized neural networks.arXiv:1706.04454, 2017
Pith/arXiv arXiv 2017
-
[36]
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning.arXiv preprint arXiv:1703.03864, 2017
Pith/arXiv arXiv 2017
-
[37]
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David G T Barrett, and Soham De. On the origin of implicit regularization in stochastic gradient descent. InICLR, 2021. arXiv:2101.12176
Pith/arXiv arXiv 2021
-
[38]
Does sgd really happen in tiny subspaces? In International Conference on Learning Representations (ICLR),2025
Minhak Song, Kwangjun Ahn, and Chulhee Yun. Does sgd really happen in tiny subspaces? In International Conference on Learning Representations (ICLR),2025. URLhttps://openreview. net/forum?id=v6iLQBoIJw
2025
-
[39]
The geometry of sequential learning: Lie-bracket prediction of transfer order
John Sweeney. The geometry of sequential learning: Lie-bracket prediction of transfer order. Proceedings of the 43rd International Conference on Machine Learning, 2026. URLhttps: //arxiv.org/abs/2606.24993. arXiv:2606.24993
Pith/arXiv arXiv 2026
-
[40]
Optimizer memory makes shuffle order a first-order source of fine-tuning noise
John Sweeney. Optimizer memory makes shuffle order a first-order source of fine-tuning noise. arXiv preprint arXiv:2606.29554, 2026. URLhttps://arxiv.org/abs/2606.29554
Pith/arXiv arXiv 2026
-
[41]
Anke Tang, Li Shen, Yong Luo, Yibing Zhan, Han Hu, Bo Du, Yixin Chen, and Dacheng Tao. Parameter efficient multi-task model fusion with partial linearization.arXiv preprint arXiv:2310.04742, 2023. URLhttps://arxiv.org/abs/2310.04742
Pith/arXiv arXiv 2023
-
[42]
Lin Tang, Wei Zhang, Jing Li, Hongyu Chen, Ming Zhao, and Yuxuan Wang. Predicting mergeability of parameter-efficient fine-tuning updates.arXiv preprint arXiv:2606.19549, 2026. URLhttps://arxiv.org/abs/2606.19549
Pith/arXiv arXiv 2026
-
[43]
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, et al. 2 OLMo 2 furious. arXiv preprint arXiv:2501.00656, 2024. URLhttps://arxiv.org/abs/2501.00656
Pith/arXiv arXiv 2024
-
[44]
Li, Arnab Sen Sharma, Aaron Mueller, Byron C
Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. Function vectors in large language models. InInternational Conference on Learning Representations (ICLR), 2024. URLhttps://arxiv.org/abs/2310.15213
Pith/arXiv arXiv 2024
-
[45]
Steering language models with activation engineering.arXiv preprint arXiv:2308.10248, 2023
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering.arXiv preprint arXiv:2308.10248, 2023. URLhttps://arxiv.org/abs/2308.10248
Pith/arXiv arXiv 2023
-
[46]
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin. Introduction to the non-asymptotic analysis of random matrices. arXiv:1003.2990 (in: Compressed Sensing, CUP 2012), 2010
Pith/arXiv arXiv 2012
-
[47]
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo- Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. Model soups: Averaging weights of multiple fine-tuned models improves 29 accuracy without increasing inference time. InInternational Conference on Machine Learning...
Pith/arXiv arXiv 2022
-
[48]
Manning, and Christopher Potts
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D. Manning, and Christopher Potts. ReFT: Representation finetuning for language models. In Advances in Neural Information Processing Systems (NeurIPS), 2024. URLhttps://arxiv. org/abs/2404.03592
Pith/arXiv arXiv 2024
-
[49]
TIES-merging: Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. TIES-merging: Resolving interference when merging models. InAdvances in Neural Information Processing Systems (NeurIPS), 2023. URLhttps://arxiv.org/abs/2306.01708
Pith/arXiv arXiv 2023
-
[50]
Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024. URLhttps://arxiv.org/abs/2412.15115
Pith/arXiv arXiv 2024
-
[51]
HellaSwag: Can a machine really finish your sentence? InAnnual Meeting of the Association for Computational Linguistics (ACL), 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? InAnnual Meeting of the Association for Computational Linguistics (ACL), 2019
2019
-
[52]
TinyLlama: An open-source small language model.arXiv preprint arXiv:2401.02385, 2024
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. TinyLlama: An open-source small language model.arXiv preprint arXiv:2401.02385, 2024. URLhttps://arxiv.org/abs/ 2401.02385
Pith/arXiv arXiv 2024
-
[53]
OPT: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. OPT: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068...
Pith/arXiv arXiv 2022
-
[54]
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. Representation engineering: A top- down approach to ...
Pith/arXiv arXiv 2023
-
[2024]
URLhttps://aclanthology.org/2024.acl-long.828/. 28
2024
-
[2025]
URLhttps://arxiv.org/abs/2501.15556
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.