REVIEW 4 major objections 5 minor 43 references
ASLoRA: Adaptive Sharing Low-Rank Adaptation Across Layers
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Sharing low-rank layers beats LoRA at a quarter of the parameters
desk verdict Plausible parameter-sharing trick with directionally supportive experiments, but the adaptive merging criterion is not tested against random pairing and the reported gains are small enough that the central claim is only conditional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the adaptive merging rule: after a shared-training phase, each B matrix is replaced by its time-averaged value (Eq. 3), pairwise L2 distances between these averages (Eq. 4) are computed across layers, and the closest pair is merged by overwriting the lower layer's B with the upper layer's B; the process repeats every m steps until N merges. This is paired with full sharing of the rank-r matrix A across all layers, which removes A's random initialization as a confound in the similarity comparison. The merge rule is what carries the parameter reduction: with 7 of a possible 11 merges on a 12-layer model, per-layer B count drops enough to cut trainable parameters to about 24% of LoRA's.
What would settle it
Run ASLoRA with the same merge budget but pairing B matrices at random instead of by L2 similarity. If random pairing matches the adaptive rule's accuracy at equal parameter counts, then the similarity heuristic is not what drives the gain. A more direct test: keep the merge schedule but replace the overwrite rule with the average of the two B matrices; if performance collapses, the specific direction of overwriting is essential.
Extended reading notes
Core claim
ASLoRA's core proposal is a three-stage training procedure: first, train with a single globally-shared A and per-layer B matrices; second, every m steps after step Ts, compute time-averaged B matrices, rank all pairs by L2 distance, and replace the two most similar B matrices by the upper layer's B; third, continue training with the merged, partially-shared B structure until convergence. The paper reports that on six GLUE tasks with RoBERTa-base, ASLoRA reaches an average score of 85.4 with 0.073M trainable parameters, versus LoRA's 85.2 with 0.3M parameters, and on LLaMA-2-7B instruction tuning it averages 32.33 across MMLU, BBH, DROP, and HumanEval, beating LoRA's 31.40 while using 8.9M parameters to LoRA's 33.6M. It further claims that adaptive merging outperforms fixed every-n-layer sharing at equal parameter budgets.
Load-bearing premise
The load-bearing premise is that the L2 distance between time-averaged B matrices identifies which layers can safely share a B matrix, and that overwriting the lower layer's B with the upper layer's B preserves the useful information in that pair.
Editorial extensions
If this is right
- Adapter storage for fine-tuned LLMs can be cut by roughly 75% relative to LoRA without losing average task performance, on both encoder and decoder models.
- The adaptive merge schedule yields better results than fixed every-n-layer sharing at matched parameter budgets, particularly when the number of merges is small.
- The optimal number of merges is task-dependent; performance on MMLU peaks at 24 merges, on BBH at 20, and on HumanEval at 16, so the schedule can be tuned per downstream task.
- ASLoRA's parameter count shrinks as model depth grows, since more layers allow more merges; the method's advantage widens with model size.
Reading between the lines
- One natural extension the paper leaves implicit: the same merge rule could be applied to intra-layer structure, e.g., merging low-rank factors across the query and value projections, or across heads, to push parameter counts even lower.
- The L2-similarity proxy is a direct, testable hypothesis about representational redundancy; a stronger test would compare it against merge rules based on gradient alignment or Fisher information, which might better preserve task-critical directions.
- If the merge heuristic is validated more broadly, it suggests that much of a fine-tuned adapter's per-layer variation is redundant, and that cross-layer sharing could become a default compression step before quantization in deployment pipelines.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ASLoRA, a parameter-efficient fine-tuning method in which the LoRA matrix A is shared across all layers while each layer initially has its own B matrix. During training, after Ts steps, the method computes the L2 distance between time-averaged B matrices (Eq. 4) every m steps, merges the two most similar B matrices by having the lower layer adopt the upper layer's B, and repeats this N times before a final optimization phase. Experiments on RoBERTa-base with GLUE and on LLaMA-2-7B with instruction-tuning benchmarks report that ASLoRA matches or improves on LoRA while using about 24% of the parameters on GLUE and about 26% on instruction tuning. Additional analyses compare adaptive merging with fixed adjacent sharing, visualize the learned sharing patterns, and study the effect of the number of merges.
Significance. If the empirical claims hold, ASLoRA offers a simple way to reduce per-task adapter storage by roughly 75% while retaining or slightly improving downstream quality, which is practically valuable. The paper's strengths include a clearly specified algorithmic idea, experiments on two model families, and a limitations section that acknowledges hyperparameter sensitivity. The central claim, however, rests on small average margins, an unvalidated similarity heuristic, and per-dataset tuning of hyperparameters, so the current evidence is not fully conclusive. I also note that the only self-citation (MELoRA) appears in related/future work and is not used to support the results, so there is no circularity concern.
major comments (4)
- [§4.3, Table 3] The 'Advantage of Adaptive Sharing' comparison does not isolate the L2-distance criterion. Table 3 compares L2-based adaptive merging against fixed adjacent merging at equal parameter counts, but because the adaptive method can merge non-adjacent layers, the improvement could come from the flexibility to form non-adjacent groups rather than from the L2 proxy. To support the core mechanism, add control experiments such as random non-adjacent pairing, an alternative similarity measure (e.g., cosine similarity or activation-space distance), or an oracle that merges based on validation performance. Without such controls, the paper does not establish that the L2 criterion is what drives the reported gains.
- [§4.1, Table 1, abstract, Section 1] The headline claims are stronger than the numbers support. The abstract says ASLoRA 'outperforms LoRA while using less than 25% of the parameters,' but the instruction-tuning configuration in Table 2 uses 8.9M versus 33.6M, which is 26.5%, not below 25%. On GLUE, LoRA actually achieves higher scores on 4 of 6 datasets (SST-2, CoLA, QNLI, STS-B) and the average gain is 0.2 points (85.4 vs. 85.2) with no error bars or significance tests. Section 1's claim that ASLoRA is 'outperforming the baseline models across all instruction-following datasets' is also contradicted by Table 2 for full fine-tuning on MMLU (47.30 vs. 46.21). Please report repeated-seed means and variances, check the parameter-ratio arithmetic, and qualify the claims accordingly.
- [§4.2, Table 5, Figure 4] The main result is conditioned on per-dataset hyperparameter selection. Table 5 lists different start-step Ts and merge-interval W for every GLUE dataset, and Figure 4 shows that the optimal merge count varies across instruction datasets (N=24 for MMLU, N=20 for BBH, N=16 for HEval). Moreover, Section 4.2 says the maximum number of merges is 16, while Figure 4 reports settings up to 28. Please specify a fixed protocol or an explicit model-selection rule and resolve the inconsistency; otherwise the parameter-efficiency comparison is not for a single method but for per-dataset tuned configurations.
- [Algorithm 1 and Eq. (3)] The pseudocode is ambiguous and internally inconsistent. Line 4 of Algorithm 1 says 'Update Bi by equation (3),' but Eq. (3) defines a running average B^t_i, not an update rule. After two B matrices are merged, it is unclear whether subsequent average-weight computations use the merged parameters or the pre-merge history. The text following Algorithm 1 also says 'we calculate the pairwise similarity between adjacent layers every m steps,' which contradicts §3.3's non-adjacent merging and the non-adjacent patterns shown in Figure 3. Please rewrite the algorithm so the exact merging procedure is reproducible.
minor comments (5)
- [Table 4 vs. §4.2] Table 4 lists Rank r=4 for instruction tuning, while §4.2 states that all methods use r=64. The parameter counts in Table 2 imply r=64, so Table 4 appears incorrect and should be fixed.
- [Table 5] The hyperparameter 'Update Ratio λ 0.5' appears in Table 5 but λ is never defined or used in Algorithm 1 or any equation; either remove it or specify its role in the method.
- [Table 3] The label 'ASLoRA-adp' is used for the fixed-sharing baseline, which is confusing because 'adp' suggests adaptive. Please rename the baseline (e.g., 'ASLoRA-fixed') to avoid ambiguity.
- [Algorithm 1, §3.3] The text says the method merges the B matrices with 'the highest similarity' and the algorithm explanation says 'the lowest similarity'; the two statements are consistent under Eq. (4), where a smaller L2 distance means higher similarity, but this should be stated explicitly to prevent misreading.
- [General] No code or configuration files are released. Given that the method introduces new hyperparameters (Ts, m, N) and the paper reports per-dataset settings, providing code would substantially aid reproducibility.
Circularity Check
No circularity: the central claim is an empirical benchmark comparison, and the only self-citation (MELoRA) is not load-bearing.
full rationale
ASLoRA's central claim is an empirical comparison: Table 1 and Table 2 report accuracy and average scores after training with shared A and adaptively merged B. The adaptive merging rule is a heuristic (Eq. 4 selects pairs of B matrices with the smallest L2 distance between time-averaged weights), and the reported downstream performance is not used to define the similarity criterion or to choose the merging pairs, so no quantity is defined in terms of the quantity it predicts. The per-dataset hyperparameters (Ts, W, N in Table 5) are tuning choices, not fitted parameters renamed as predictions, and they do not make the benchmark outcome equal to an input by construction. The Limitations section concedes that the starting merge step and merge interval affect performance, which is a tuning-sensitivity concern rather than a circularity concern. The only overlapping self-citation, MELoRA (Ren et al., 2024), appears in the related-work description and in a future-work suggestion; it is not cited as evidence for ASLoRA's effectiveness and does not carry any load-bearing step. The paper's claims are therefore self-contained empirical findings rather than circular derivations.
Assumptions & free parameters
free parameters (5)
- Rank r =
8 (GLUE); 4 or 64 (instruction tuning, inconsistent)
- Start merge step Ts =
3000/320/400/1000/200/500 per GLUE dataset; 400 for instruction tuning
- Merge interval m (W) =
2000/240/500/700/100/500 per GLUE dataset; 10 for instruction tuning
- Number of merges N =
7 for GLUE; 16 for instruction tuning
- Update ratio lambda =
0.5
assumptions (4)
- domain assumption LoRA's initialization and update rule: A initialized randomly, B initialized to zero, update is BA.
- ad hoc to paper L2 distance between time-averaged B matrices is a valid proxy for functional redundancy between layers.
- domain assumption Lower layers should adopt the upper layer's B when merging, because upper layers contain more complex information.
- ad hoc to paper Shared A across layers removes initialization interference and allows fair similarity comparison of B.
Cite this review
Pith. "Pith review of ASLoRA: Adaptive Sharing Low-Rank Adaptation Across Layers." pith.science (2026). https://pith.science/paper/J5OTGOTD
@misc{pith2026241210135,
author = {Pith},
title = {Pith review of: ASLoRA: Adaptive Sharing Low-Rank Adaptation Across Layers},
year = {2026},
howpublished = {\url{https://pith.science/paper/J5OTGOTD}},
note = {Machine review of arXiv:2412.10135}
}
read the original abstract
As large language models (LLMs) grow in size, traditional full fine-tuning becomes increasingly impractical due to its high computational and storage costs. Although popular parameter-efficient fine-tuning methods, such as LoRA, have significantly reduced the number of tunable parameters, there is still room for further optimization. In this work, we propose ASLoRA, a cross-layer parameter-sharing strategy combining global sharing with partial adaptive sharing. Specifically, we share the low-rank matrix A across all layers and adaptively merge matrix B during training. This sharing mechanism not only mitigates overfitting effectively but also captures inter-layer dependencies, significantly enhancing the model's representational capability. We conduct extensive experiments on various NLP tasks, showing that ASLoRA outperforms LoRA while using less than 25% of the parameters, highlighting its flexibility and superior parameter efficiency. Furthermore, in-depth analyses of the adaptive sharing strategy confirm its significant advantages in enhancing both model flexibility and task adaptability.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean - Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael ...
-
[2]
Sangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Seungyeon Kim, and Tal Schuster. 2024. https://doi.org/10.48550/ARXIV.2410.20672 Relaxed recursive transformers: Effective parameter sharing with layer-wise lora . CoRR, abs/2410.20672
-
[3]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...
2020
-
[4]
Daniel Cer, Mona Diab, Eneko Agirre, I \ n igo Lopez-Gazpio, and Lucia Specia. 2017. https://doi.org/10.18653/v1/S17-2001 S em E val-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation . Unknown Journal, pages 1--14
-
[5]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pond \' e de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bava...
arXiv 2021
-
[6]
Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022. https://doi.org/10.1145/3485447.3511998 Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction . In WWW '22: The ACM Web Conference 2022, Virtual Event, Lyon, France, April 25 - 29, 2022 , pages 2778--...
arXiv 2022
-
[7]
Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria. 2024. https://aclanthology.org/2024.scalellm-1.4 I nstruct E val: Towards holistic evaluation of instruction-tuned large language models . In Proceedings of the First edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024), pages 35--64, St. Julian ' s, Malta. A...
2024
-
[8]
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019. https://openreview.net/forum?id=HyzdRiR9Y7 Universal transformers . In International Conference on Learning Representations
work page 2019
Show all 43 references
-
[9]
Dolan and Chris Brockett
William B. Dolan and Chris Brockett. 2005. https://aclanthology.org/I05-5002 Automatically constructing a corpus of sentential paraphrases . In Proceedings of the Third International Workshop on Paraphrasing ( IWP 2005)
2005
-
[10]
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. https://doi.org/10.18653/v1/N19-1246 DROP : A reading comprehension benchmark requiring discrete reasoning over paragraphs . Unknown Journal, pages 2368--2378
2019 doi
- [11]
-
[12]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=d7KBjmI3GmQ Measuring massive multitask language understanding . Unknown Journal
2021
-
[13]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. http://proceedings.mlr.press/v97/houlsby19a.html Parameter-efficient transfer learning for NLP . In Proceedings of the 36th Int...
2019
-
[14]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lo RA : Low-rank adaptation of large language models . In International Conference on Learning Representations
2022
-
[15]
Dawid Jan Kopiczko, Tijmen Blankevoort, and Yuki M. Asano. 2024. https://openreview.net/forum?id=NjNfLdxr3A Vera: Vector-based random matrix adaptation . In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net
2024
-
[16]
Xiang Lisa Li and Percy Liang. 2021. https://doi.org/10.18653/v1/2021.acl-long.353 Prefix-tuning: Optimizing continuous prompts for generation . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferen...
2021 doi
-
[17]
Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.41 Exploring versatile generative language model via parameter-efficient transfer learning . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4...
2020 doi
-
[18]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://arxiv.org/abs/1907.11692 Roberta: A robustly optimized BERT pretraining approach . CoRR, abs/1907.11692
2019 arXiv
-
[19]
Xiuqing Lv, Peng Zhang, Sunzhu Li, Guobing Gan, and Yueheng Sun. 2023. https://doi.org/10.18653/v1/2023.findings-acl.656 L ight F ormer: Light-weight transformer using SVD -based weight transfer and parameter sharing . In Findings of the Association for Computational Linguisti...
2023 doi
- [20]
- [21]
-
[22]
Jonas Pfeiffer, Aishwarya Kamath, Andreas R \"u ckl \'e , Kyunghyun Cho, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.eacl-main.39 A dapter F usion: Non-destructive task composition for transfer learning . In Proceedings of the 16th Conference of the European Cha...
2021 doi
-
[23]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. https://doi.org/10.18653/v1/D16-1264 SQ u AD : 100,000+ questions for machine comprehension of text . Unknown Journal, pages 2383--2392
2016 doi
-
[24]
Machel Reid, Edison Marrese-Taylor, and Yutaka Matsuo. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.344 Subformer: Exploring weight sharing for parameter efficiency in generative transformers . In Findings of the Association for Computational Linguistics: EMNLP 2021, ...
2021 doi
-
[25]
Pengjie Ren, Chengshun Shi, Shiguang Wu, Mengqi Zhang, Zhaochun Ren, Maarten de Rijke, Zhumin Chen, and Jiahuan Pei. 2024. https://doi.org/10.18653/v1/2024.acl-long.168 MEL o RA : Mini-ensemble low-rank adapters for parameter-efficient fine-tuning . In Proceedings of the 62nd ...
2024 doi
-
[26]
Adithya Renduchintala, Tugrul Konuk, and Oleksii Kuchaiev. 2024. https://doi.org/10.18653/v1/2024.naacl-long.481 Tied- L o RA : Enhancing parameter efficiency of L o RA with weight tying . In Proceedings of the 2024 Conference of the North American Chapter of the Association f...
2024 doi
-
[27]
Andreas R \"u ckl \'e , Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.626 AdapterDrop : O n the efficiency of adapters in transformers . In Proceedings of the 2021 Conference on Emp...
2021 doi
-
[28]
Logan IV, Eric Wallace, and Sameer Singh
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. https://doi.org/10.18653/V1/2020.EMNLP-MAIN.346 Autoprompt: Eliciting knowledge from language models with automatically generated prompts . In Proceedings of the 2020 Conference on Empirica...
2020 doi
-
[29]
Manning, Andrew Ng, and Christopher Potts
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. https://aclanthology.org/D13-1170 Recursive deep models for semantic compositionality over a sentiment treebank . In Proceedings of the 2013 Conference on Emp...
2013
- [30]
-
[31]
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, et al. 2023. https://openreview.net/forum?id=uyTL5Bvosj Beyond the imitation game: Quantifying and extrapolating t...
2023
-
[32]
Sho Takase and Shun Kiyono. 2023. https://doi.org/10.18653/V1/2023.SUSTAINLP-1.5 Lessons on parameter sharing across layers in transformers . In Proceedings of The Fourth Workshop on Simple and Efficient Natural Language Processing, SustaiNLP 2023, Toronto, Canada (Hybrid), Ju...
2023 doi
-
[33]
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. https://github. com/tatsu-lab/stanford\_alpaca Stanford alpaca: an instruction-following llama model (2023) . URL https://github. com/tatsu-lab/...
2023
-
[34]
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, and Ali Ghodsi. 2023. https://doi.org/10.18653/v1/2023.eacl-main.239 D y L o RA : Parameter-efficient tuning of pre-trained models using dynamic search-free low-rank adaptation . In Proceedings of the 17th Conference of the...
2023 doi
-
[35]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/W18-5446 GLUE : A multi-task benchmark and analysis platform for natural language understanding . In Proceedings of the 2018 EMNLP Workshop B lackbox NLP : A...
2018 doi
-
[36]
Sheng Wang, Boyang Xue, Jiacheng Ye, Jiyue Jiang, Liheng Chen, Lingpeng Kong, and Chuan Wu. 2024. https://doi.org/10.18653/v1/2024.acl-long.156 PR o L o RA : Partial rotation empowers more parameter-efficient L o RA . In Proceedings of the 62nd Annual Meeting of the Associatio...
2024 doi
-
[37]
Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual Mee...
2023 doi
-
[38]
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. https://doi.org/10.1162/tacl_a_00290 Neural network acceptability judgments . Transactions of the Association for Computational Linguistics, 7:625--641
2019 doi
-
[39]
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A broad-coverage challenge corpus for sentence understanding through inference . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computation...
2018 doi
- [40]
-
[41]
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023 b . https://openreview.net/forum?id=lq62uWRJjiY Adaptive budget allocation for parameter-efficient fine-tuning . In The Eleventh International Conference on Learning Representations
2023
-
[42]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[43]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.