REVIEW 4 major objections 5 minor 81 references
A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proposes a three-layer architecture that maps each LLM capability to the layer where it should be implemented.
desk verdict A clear, useful conceptual taxonomy for LLM application architecture, but the efficiency claim rests on underdetermined capability mapping and illustrative cases, not proof. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the layered attribute model: three layers, each with named components and attributes, plus a capability mapping step. An attribute is defined as something inherently responsible for a quality, feature, or operation of the system; a capability is the ability to effectively perform a category of tasks. The mapping procedure works by identifying which attributes a capability requires, marking attributes whose necessity is depreciating when another attribute covers the need, designing a solution at each relevant layer, and downgrading or enabling components in response to access restrictions. The work this does is to turn an open-ended technology-selection problem into a structured design question.
What would settle it
Take ten capabilities not discussed in the paper, ask independent engineering teams to implement each using the layer suggested by the mapping and using a deliberately different layer, and compare on a fixed cost and quality score; the mapping's predictive claim fails if the alternative-layer implementation wins in a majority of cases.
Extended reading notes
Core claim
The central claim is that every capability of an LLM-based system can be aligned with a small set of attributes, each tied primarily to one of three layers, and that this alignment tells the developer which implementation technique to use. For example, generating JSON output spans Knowledge Boundaries and Objectives in the Model Layer, Micro-level Token Generation Control in the Inference Layer, and Probability Amplification, External Information, and Exception Handling in the Application Layer. Because Micro-level Token Generation Control can enforce the format during decoding, the dependence on Exception Handling is depreciating and can be kept lightweight; the mapping therefore directly reduces redundancy. The paper generalizes this into a capability mapping process with steps of attribute identification, solution architecture design, access resolution for gated components, and evaluation.
Load-bearing premise
The framework assumes that a capability's needed attributes can be cleanly identified in advance and assigned to one of the three layers, with the depreciating relationships between attributes visible before implementation.
Editorial extensions
If this is right
- If the mapping is right, developers can decide between fine-tuning and retrieval-augmented generation by checking whether the capability needs knowledge embedded in parameters or supplied as external context.
- A capability like structured output will be implemented most reliably at the Inference Layer via constrained decoding, with prompts and retry parsers demoted to supporting roles.
- Recognizing depreciating attributes lets teams drop or lighten redundant components, reducing engineering complexity and cost without losing effectiveness.
- When a layer's access is gated, such as a commercial model without fine-tuning or low-level decoding control, the framework tells developers which previously depreciated fallback to re-enable.
- The layered view gives a vocabulary for comparing vendor-provided features such as prompt caching with application-level implementations such as hash-based call caching, making the trade-off visible.
Reading between the lines
- The framework's strongest use may be as a capability-planning checklist before implementation, where the absence of any attribute row prompts developers to ask whether the capability is actually being addressed.
- If capability decomposition turns out to be ambiguous in practice, the same mapping could be extended to score confidence across layers rather than assigning each attribute to exactly one layer.
- The depreciation logic suggests a testable hypothesis: for structured-output capabilities, systems that rely primarily on constrained decoding should need fewer retry and repair components than prompt-only systems while achieving the same or better format validity.
- The architecture could be connected to cost models, since each layer has a different cost profile, and mapping a capability to a lower-cost layer could become a cost-optimization rule.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a layered architecture for LLM-based software systems, dividing development into Model, Inference, and Application layers, each with components and attributes. The central claim is that mapping desired capabilities to the appropriate layers/components and identifying 'depreciating' attributes yields effective and efficient implementations. The approach is illustrated with four case studies: JSON output generation, creativity, call caching, and long context.
Significance. If validated, the framework would give developers a structured alternative to ad-hoc technology selection in LLM application design. The paper's internal taxonomies (Tables 1-3) are reasonable and draw on established distinctions (parameterized vs. non-parameterized knowledge, macro/micro token control, prompt/mechanism/tooling/orchestration). The related-work comparison is useful and appropriately positions the contribution. However, the evidence for the central usefulness claim is weak: the evaluation consists of four self-selected examples interpreted through the framework itself, with no metrics, no counterfactual, no external validation, and no falsifiable predictions. The paper also (Section 3.4) includes no decision rule for attribute selection or depreciation, leaving the load-bearing step of the method underdetermined. The manuscript is clearly written and the technical descriptions align with the literature, but the central claim is currently a plausible hypothesis, not a demonstrated result.
major comments (4)
- [§3.4 (Attributes Identification; Tables 4-5)] The load-bearing step of the proposed method, identifying which attributes apply to a capability and which are 'depreciating,' is underdetermined and appears post hoc. In the JSON example, the explanations for why Knowledge Boundaries reduce the need for External Information and why Micro-level Token Generation Control makes extensive Exception Handling less critical are plausibility judgments, not derivations from a stated rule, and Table 5 still lists solutions at every layer (including lightweight exception handling) without any cost-benefit model. Because the framework provides no reproducible decision procedure, the claimed benefit of avoiding 'redundancy and oversophistication' is vulnerable to post-hoc rationalization: any reasonable architecture can be described as aligned, and any omitted component can be labeled depreciating. The paper needs either a concrete decision rule, a testable comparative analysis showing that following the mapping yields measurably more efficient or effective systems, or an honest reframing of the contribution as a descriptive taxonomy.
- [§4 (Use Case Evaluation)] The evaluation does not test the central claim. Each of the four cases (JSON generation, creativity, call caching, long context) presents a table of attributes and solutions, but there is no metric, no comparison against a baseline or an alternative mapping, and no evidence that the framework's guidance improves effectiveness or efficiency over current practice. The paper states that the framework 'was evaluated against several typical use cases' (Section 6), but what is shown is illustration, not evaluation. The authors should either provide empirical evidence (e.g., a developer study, measured cost/latency/robustness comparisons, or an ablation of the depreciation judgments) or revise the claims to state clearly that the contribution is a conceptual framework whose utility remains to be tested.
- [§3.4 (Access Resolution and Evaluation paragraph)] The Access Resolution and Evaluation paragraphs themselves acknowledge the limits of the framework: they state that fallback implementation can be 'cost-effective depending on priorities,' that the framework 'cannot replace the need for thorough evaluation,' and that developers should 'dynamically adjust the sophistication of each component in response to evaluation feedback.' These admissions directly undercut the promised efficiency benefit, which was supposed to come from identifying depreciating attributes a priori. As written, the framework offers no criterion for when to trust the depreciation judgment versus when to keep a fallback, so the 'avoid redundancy and oversophistication' goal is deferred to unspecified future evaluation.
- [§3.4, Table 4 (depreciation examples)] The specific depreciation claims are stated as if unproblematic but are not self-evident. For example, 'Knowledge Boundaries may reduce the necessity for External Information' presumes that the model's internal knowledge of the JSON format is sufficient for semantic correctness, which is a key open issue in structured generation; 'Micro-level Token Generation Control makes extensive Exception Handling less critical' presumes the constrained decoder produces semantically valid JSON, which it does not by itself guarantee. The manuscript would be stronger if it acknowledged these as empirical assumptions and pointed to literature on structured-output failure modes, rather than presenting them as derived consequences of the architecture.
minor comments (5)
- [Abstract and §3.3] The abstract says 'Significant efforts has been made'; this should be 'have been made'. Also, Figure 1 is dense and uses a self-referential legend (Layer, Component, Attribute, Capability, Dependency) that would benefit from an example-labeled annotation.
- [§3.3.2 and Table 2] The text distinguishes 'efficiency' as an Inference Layer attribute but Table 2 lists 'Efficiency' both as an attribute and a developer-access category; clarify whether prompt caching is the only developer-visible efficiency control for commercial models, since the text says inference efficiency is 'centrally managed by the vendors.'
- [§3.3.4] The Intra-layer Dependencies paragraph in the Model Layer says 'the fine-tuning process cannot align the model effectively tune the model' — the phrase 'tune the model' appears to be a duplicated fragment; this sentence needs rewriting.
- [§4 (Creativity, Tables 6-8)] The case-study tables for creativity, call caching, and long context are not accompanied by any discussion of how the framework's mapping changes the decision relative to existing practice; adding a sentence per case linking the table back to the central claim would improve coherence. Also, the creativity case (Table 6) includes 'Less control and alignment over human preferences' under Objectives/Behavior as a positive route to creativity, which readers may find at odds with quality requirements; a clarifying sentence would help.
- [§5 (Related Work)] The related-work section cites several closely related recent frameworks (Zhou et al. analysis of architecture options, Lu et al. layered reference architecture, Liu et al. agent design pattern catalogue) and dismisses them as 'isolated components without a unified perspective.' Given the overlap, the paper should state explicitly what distinguishes its contribution from these works beyond the layer partition and the 'depreciating attributes' notion, especially since no comparison against them is performed.
Circularity Check
No significant circularity: the framework is an under-evaluated mapping heuristic, but its claims do not reduce to their own inputs by construction.
full rationale
I examined the paper's claimed derivation chain: the layered architecture is a conceptual mapping heuristic, not a predictive or fitted model. The central claim is that aligning capabilities with layers and identifying depreciating attributes encourages systematic, effective, and efficient implementation. This is supported by illustrative use cases (JSON output, creativity, call caching, long context) that are authored through the framework's own categories, and the 'depreciating attribute' judgments in Section 3.4 (Table 4) are asserted rather than derived from a reproducible rule. Under the review rules, however, this is an evidentiary weakness (post-hoc or underdetermined validation) rather than circularity: there is no equation, fitted parameter, or self-citation chain whose output is, by construction, identical to an input. The paper itself disclaims that mapping replaces evaluation: 'it cannot replace the need for thorough evaluation of the solutions.' Related-work citations to the authors' own prior work are descriptive and not load-bearing for the central claim. I found no specific circular step that can be quoted and reduced, so the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption LLM-based applications can be decomposed into three layers (Model, Inference, Application) with distinct attributes.
- domain assumption Known implementation techniques (fine-tuning, RAG, constrained decoding, prompting, mechanisms, tools, orchestration) have the characteristics and trade-offs described, e.g., parameterized vs. non-parameterized knowledge.
- ad hoc to paper Depreciating relationships between attributes, where one implementation reduces the necessity of another, can be identified prior to implementation.
Cite this review
Pith. "Pith review of A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems." pith.science (2026). https://pith.science/paper/CRTQLI3H
@misc{pith2026241112357,
author = {Pith},
title = {Pith review of: A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/CRTQLI3H}},
note = {Machine review of arXiv:2411.12357}
}
read the original abstract
Significant efforts has been made to expand the use of Large Language Models (LLMs) beyond basic language tasks. While the generalizability and versatility of LLMs have enabled widespread adoption, evolving demands in application development often exceed their native capabilities. Meeting these demands may involve a diverse set of methods, such as enhancing creativity through either inference temperature adjustments or creativity-provoking prompts. Selecting the right approach is critical, as different methods lead to trade-offs in engineering complexity, scalability, and operational costs. This paper introduces a layered architecture that organizes LLM software system development into distinct layers, each characterized by specific attributes. By aligning capabilities with these layers, the framework encourages the systematic implementation of capabilities in effective and efficient ways that ultimately supports desired functionalities and qualities. Through practical case studies, we illustrate the utility of the framework. This work offers developers actionable insights for selecting suitable technologies in LLM-based software system development, promoting robustness and scalability.
Figures
Reference graph
Works this paper leans on
-
[1]
Code- gen: An open large language model for code with multi-turn program synthesis,
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong, “Code- gen: An open large language model for code with multi-turn program synthesis,” arXiv preprint arXiv:2203.13474, 2022. 16
arXiv 2022
-
[2]
Large language models for software engineering: A systematic literature review,
X. Hou, Y. Zhao, Y. Liu, Z. Yang, K. Wang, L. Li, X. Luo, D. Lo, J. Grundy, and H. Wang, “Large language models for software engineering: A systematic literature review,” ACM Transactions on Software Engineering and Methodology , 2023
2023
-
[3]
Program synthesis with large language models,
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le et al., “Program synthesis with large language models,” arXiv preprint arXiv:2108.07732 , 2021
arXiv 2021
-
[4]
Language models can solve computer tasks,
G. Kim, P. Baldi, and S. McAleer, “Language models can solve computer tasks,” Advances in Neural Information Processing Systems , vol. 36, 2024
work page 2024
-
[5]
A survey on large language model based autonomous agents,
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin et al., “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, p. 186345, 2024
work page 2024
-
[6]
A survey of large language models for financial applications: Progress, prospects and challenges,
Y. Nie, Y. Kong, X. Dong, J. M. Mulvey, H. V. Poor, Q. Wen, and S. Zohren, “A survey of large language models for financial applications: Progress, prospects and challenges,” arXiv preprint arXiv:2406.11903, 2024
arXiv 2024
-
[7]
Bloomberggpt: A large language model for finance,
S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, and G. Mann, “Bloomberggpt: A large language model for finance,” arXiv preprint arXiv:2303.17564, 2023
arXiv 2023
-
[8]
Chatgpt for design, manufacturing, and education,
X. Wang, N. Anwer, Y. Dai, and A. Liu, “Chatgpt for design, manufacturing, and education,” Procedia CIRP, vol. 119, pp. 7–14, 2023
work page 2023
Show all 81 references
-
[9]
Embodied intelligence in manufactur- ing: leveraging large language models for autonomous industrial robotics,
H. Fan, X. Liu, J. Y. H. Fuh, W. F. Lu, and B. Li, “Embodied intelligence in manufactur- ing: leveraging large language models for autonomous industrial robotics,” Journal of Intelligent Manufacturing, pp. 1–17, 2024
2024
-
[10]
Reconceptualizing chatgpt and generative ai as a student-driven innovation in higher education,
Y. Dai, A. Liu, and C. P. Lim, “Reconceptualizing chatgpt and generative ai as a student-driven innovation in higher education,” Procedia CIRP, vol. 119, pp. 84–90, 2023
2023
-
[11]
Automatic generation of programming exercises and code explanations using large language models,
S. Sarsa, P. Denny, A. Hellas, and J. Leinonen, “Automatic generation of programming exercises and code explanations using large language models,” in Proceedings of the 2022 ACM Conference on International Computing Education Research-Volume 1 , 2022, pp. 27–43
2022
-
[12]
Autonomous chemical research with large language models,
D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes, “Autonomous chemical research with large language models,” Nature, vol. 624, no. 7992, pp. 570–578, 2023
2023
-
[13]
Large language models in medicine,
A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting, “Large language models in medicine,” Nature medicine, vol. 29, no. 8, pp. 1930–1940, 2023
1930
-
[14]
Mathematical discoveries from program search with large language models,
B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar, E. Dupont, F. J. Ruiz, J. S. Ellenberg, P. Wang, O. Fawzi et al. , “Mathematical discoveries from program search with large language models,” Nature, vol. 625, no. 7995, pp. 468–475, 2024
2024
-
[15]
Retrieval-augmented generation for knowledge-intensive nlp tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. K¨ uttler, M. Lewis, W.-t. Yih, T. Rockt¨ aschelet al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 9459–9474, 2020
2020
-
[16]
Siren’s song in the ai ocean: a survey on hallucination in large language models,
Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen et al. , “Siren’s song in the ai ocean: a survey on hallucination in large language models,” arXiv preprint arXiv:2309.01219, 2023
2023 arXiv
-
[17]
Improving language understanding by generative pre-training,
A. Radford, “Improving language understanding by generative pre-training,” 2018. 17
2018
-
[18]
Larger and more instructable language models become less reliable,
L. Zhou, W. Schellaert, F. Mart ´ ınez-Plumed, Y. Moros-Daval, C. Ferri, and J. Hern´ andez-Orallo, “Larger and more instructable language models become less reliable,” Nature, pp. 1–8, 2024
2024
-
[19]
Swe- agent: Agent-computer interfaces enable automated software engineering,
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press, “Swe- agent: Agent-computer interfaces enable automated software engineering,” arXiv preprint arXiv:2405.15793, 2024
2024 arXiv
-
[20]
Are chatgpt and gpt-4 general- purpose solvers for financial text analytics? a study on several typical tasks,
X. Li, S. Chan, X. Zhu, Y. Pei, Z. Ma, X. Liu, and S. Shah, “Are chatgpt and gpt-4 general- purpose solvers for financial text analytics? a study on several typical tasks,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Trac...
2023
-
[21]
Struc-bench: Are large language models really good at generating complex structured data?
X. Tang, Y. Zong, J. Phang, Y. Zhao, W. Zhou, A. Cohan, and M. Gerstein, “Struc-bench: Are large language models really good at generating complex structured data?” arXiv preprint arXiv:2309.08963, 2023
2023 arXiv
-
[22]
An explanation of in-context learning as implicit bayesian inference,
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma, “An explanation of in-context learning as implicit bayesian inference,” arXiv preprint arXiv:2111.02080 , 2021
2021 arXiv
-
[23]
Structured information extraction from complex scientific text with fine-tuned large language models,
A. Dunn, J. Dagdelen, N. Walker, S. Lee, A. S. Rosen, G. Ceder, K. Persson, and A. Jain, “Structured information extraction from complex scientific text with fine-tuned large language models,” arXiv preprint arXiv:2212.05238 , 2022
2022 arXiv
-
[24]
Overcoming catastrophic forgetting in neural networks,
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinskaet al., “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences, vol. 114, no. 13, pp. 3521–3526, 2017
2017
-
[25]
On context-free languages,
R. J. Parikh, “On context-free languages,” Journal of the ACM (JACM) , vol. 13, no. 4, pp. 570–581, 1966
1966
-
[26]
Lexically constrained decoding for sequence generation using grid beam search,
C. Hokamp and Q. Liu, “Lexically constrained decoding for sequence generation using grid beam search,” arXiv preprint arXiv:1704.07138 , 2017
2017 arXiv
-
[27]
Foundational challenges in assuring alignment and safety of large language models,
U. Anwar, A. Saparov, J. Rando, D. Paleka, M. Turpin, P. Hase, E. S. Lubana, E. Jenner, S. Casper, O. Sourbut, B. L. Edelman, Z. Zhang, M. G¨ unther, A. Korinek, J. Hernandez-Orallo, L. Hammond, E. J. Bigelow, A. Pan, L. Langosco, T. Korbak, H. C. Zhang, R. Zhong, S. O. hEigea...
2024
-
[28]
Emergent abilities of large language models,
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler et al. , “Emergent abilities of large language models,” Transactions on Ma- chine Learning Research, 2022
2022
-
[29]
Model evaluation for extreme risks,
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt et al., “Model evaluation for extreme risks,” arXiv preprint arXiv:2305.15324, 2023
2023 arXiv
-
[30]
How far are large language models from agents with theory- of-mind?
P. Zhou, A. Madaan, S. P. Potharaju, A. Gupta, K. R. McKee, A. Holtzman, J. Pujara, X. Ren, S. Mishra, A. Nematzadeh et al. , “How far are large language models from agents with theory- of-mind?” arXiv preprint arXiv:2310.03051 , 2023
-
[31]
L. Bass, P. Clements, and R. Kazman, Software Architecture in Practice. Addison-Wesley, 2012. 18
2012
-
[32]
Richards, Software architecture patterns
M. Richards, Software architecture patterns . O’Reilly Media, Incorporated 1005 Gravenstein Highway North, Sebastopol, CA . . . , 2015, vol. 4
2015
-
[33]
TCP/IP tutorial,
C. J. Kale and T. J. Socolofsky, “TCP/IP tutorial,” RFC 1180, Jan. 1991. [Online]. Available: https://www.rfc-editor.org/info/rfc1180
1991
-
[34]
The subjects and stages of ai dataset development: A framework for dataset accountability,
M. Khan and A. Hanna, “The subjects and stages of ai dataset development: A framework for dataset accountability,” Ohio St. Tech. LJ , vol. 19, p. 171, 2022
2022
-
[35]
The art and practice of data science pipelines: A compre- hensive study of data science pipelines in theory, in-the-small, and in-the-large,
S. Biswas, M. Wardat, and H. Rajan, “The art and practice of data science pipelines: A compre- hensive study of data science pipelines in theory, in-the-small, and in-the-large,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 2091–2103
2022
-
[36]
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,
C.-Y. Hsieh, C.-L. Li, C.-k. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, and T. Pfister, “Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,” in Findings of the Association for Computational Linguisti...
2023
-
[37]
Attention is all you need,
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
-
[38]
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in International Conference on Learning Representations, 2016
2016
-
[39]
Mixture-of-experts with expert choice routing,
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon et al. , “Mixture-of-experts with expert choice routing,” Advances in Neural Information Processing Systems, vol. 35, pp. 7103–7114, 2022
2022
-
[40]
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training.”
-
[41]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolo- ...
2019
-
[42]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019
2019
-
[43]
Instruction tuning for large language models: A survey,
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu et al. , “Instruction tuning for large language models: A survey,” arXiv preprint arXiv:2308.10792, 2023
2023
-
[44]
Direct preference optimization: Your language model is secretly a reward model,
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[45]
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
-
[46]
Constitutional ai: Harmlessness from ai feedback,
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirho- seini, C. McKinnon et al. , “Constitutional ai: Harmlessness from ai feedback,” arXiv preprint arXiv:2212.08073, 2022
2022 arXiv
-
[47]
Toolformer: Language models can teach themselves to use tools,
T. Schick, J. Dwivedi-Yu, R. Dess ` ı, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Can- cedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” Ad- vances in Neural Information Processing Systems , vol. 36, 2024. 19
2024
-
[48]
Chain-of- thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al., “Chain-of- thought prompting elicits reasoning in large language models,” Advances in neural information processing systems, vol. 35, pp. 24 824–24 837, 2022
2022
-
[49]
Monte-carlo tree search as regularized policy optimization,
J.-B. Grill, F. Altch´ e, Y. Tang, T. Hubert, M. Valko, I. Antonoglou, and R. Munos, “Monte-carlo tree search as regularized policy optimization,” in International Conference on Machine Learning. PMLR, 2020, pp. 3769–3778
2020
-
[50]
Quiet-STar: Language models can teach themselves to think before speaking,
E. Zelikman, G. R. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. Goodman, “Quiet-STar: Language models can teach themselves to think before speaking,” in First Conference on Language Modeling, 2024. [Online]. Available: https://openreview.net/forum?id=oRXPiSOGH9
2024
-
[51]
Lora: Low-rank adaptation of large language models,
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685 , 2021
2021 arXiv
-
[52]
Hierarchical neural story generation,
A. Fan, M. Lewis, and Y. Dauphin, “Hierarchical neural story generation,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 889–898
2018
-
[53]
The curious case of neural text degen- eration,
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degen- eration,” in International Conference on Learning Representations , 2019
2019
-
[54]
Beam search strategies for neural machine translation,
M. Freitag and Y. Al-Onaizan, “Beam search strategies for neural machine translation,” ACL 2017, p. 56, 2017
2017
-
[55]
Fast inference from transformers via speculative decoding,
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
-
[56]
Accelerating large language model decoding with speculative sampling,
C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper, “Accelerating large language model decoding with speculative sampling,” arXiv preprint arXiv:2302.01318 , 2023
2023 arXiv
-
[57]
Eagle: Speculative sampling requires rethinking feature uncertainty,
Y. Li, F. Wei, C. Zhang, and H. Zhang, “Eagle: Speculative sampling requires rethinking feature uncertainty,” arXiv preprint arXiv:2401.15077 , 2024
2024 arXiv
-
[58]
Efficient memory management for large language model serving with pagedattention,
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Sto- ica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
-
[59]
Flashattention: Fast and memory-efficient exact attention with io-awareness,
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. R´ e, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
-
[60]
Scaling laws for precision,
T. Kumar, Z. Ankner, B. F. Spector, B. Bordelon, N. Muennighoff, M. Paul, C. Pehlevan, C. R´ e, and A. Raghunathan, “Scaling laws for precision,” arXiv preprint arXiv:2411.04330 , 2024
2024 arXiv
-
[61]
Deepspeed-inference: enabling efficient inference of transformer mod- els at unprecedented scale,
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al., “Deepspeed-inference: enabling efficient inference of transformer mod- els at unprecedented scale,” in SC22: International Conference for High Performance Comp...
2022
-
[62]
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys, vol. 55, no. 9, pp. 1–35, 2023
2023
-
[63]
A prompt pattern catalog to enhance prompt engineering with chatgpt,
J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith, and D. C. Schmidt, “A prompt pattern catalog to enhance prompt engineering with chatgpt,” arXiv preprint arXiv:2302.11382, 2023. 20
2023 arXiv
-
[64]
Large language models are zero-shot reasoners,
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in neural information processing systems , vol. 35, pp. 22 199–22 213, 2022
2022
-
[65]
Unsupervised commonsense question answering with self-talk,
V. Shwartz, P. West, R. Le Bras, C. Bhagavatula, and Y. Choi, “Unsupervised commonsense question answering with self-talk,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 4615–4629
2020
-
[66]
Language models are few-shot learners,
T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165 , 2020
2005 arXiv
-
[67]
Active retrieval augmented generation,
Z. Jiang, F. F. Xu, L. Gao, Z. Sun, Q. Liu, J. Dwivedi-Yu, Y. Yang, J. Callan, and G. Neubig, “Active retrieval augmented generation,” arXiv preprint arXiv:2305.06983 , 2023
2023 arXiv
-
[68]
Knowing when to ask–bridging large language models and data,
P. Radhakrishnan, J. Chen, B. Xu, P. Ramaswami, H. Pho, A. Olmos, J. Manyika, and R. Guha, “Knowing when to ask–bridging large language models and data,” arXiv preprint arXiv:2409.13741, 2024
2024 arXiv
-
[69]
React: Synergizing reasoning and acting in language models,
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” in The Eleventh International Conference on Learning Representations, 2023
2023
-
[70]
Reflexion: Language agents with verbal reinforcement learning,
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[71]
Self-consistency improves chain of thought reasoning in language models,
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” arXiv preprint arXiv:2203.11171, 2022
2022 arXiv
-
[72]
Tree of thoughts: Deliberate problem solving with large language models,
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” Advances in Neural Information Pro- cessing Systems, vol. 36, 2024
2024
-
[73]
Benchmarking large language models as ai research agents,
Q. Huang, J. Vora, P. Liang, and J. Leskovec, “Benchmarking large language models as ai research agents,” in NeurIPS 2023 Foundation Models for Decision Making Workshop , 2023
2023
-
[74]
Efficient attention: Attention with linear com- plexities,
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li, “Efficient attention: Attention with linear com- plexities,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 3531–3539
2021
-
[75]
Roformer: Enhanced transformer with rotary position embedding,
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing, vol. 568, p. 127063, 2024
2024
-
[76]
Llmlingua: Compressing prompts for accel- erated inference of large language models,
H. Jiang, Q. Wu, C.-Y. Lin, Y. Yang, and L. Qiu, “Llmlingua: Compressing prompts for accel- erated inference of large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 13 358–13 376
2023
-
[77]
A taxonomy of architec- ture options for foundation model-based agents: Analysis and decision model,
J. Zhou, Q. Lu, J. Chen, L. Zhu, X. Xu, Z. Xing, and S. Harrer, “A taxonomy of architec- ture options for foundation model-based agents: Analysis and decision model,” arXiv preprint arXiv:2408.02920, 2024
2024 arXiv
-
[78]
Towards responsible ai in the era of generative ai: A reference architecture for designing foundation model based systems,
Q. Lu, L. Zhu, X. Xu, Z. Xing, and J. Whittle, “Towards responsible ai in the era of generative ai: A reference architecture for designing foundation model based systems,” IEEE Software, 2024
2024
-
[79]
Agent design pattern catalogue: A collection of architectural patterns for foundation model based agents,
Y. Liu, S. K. Lo, Q. Lu, L. Zhu, D. Zhao, X. Xu, S. Harrer, and J. Whittle, “Agent design pattern catalogue: A collection of architectural patterns for foundation model based agents,” arXiv preprint arXiv:2405.10467 , 2024. 21
2024 arXiv
-
[80]
A taxonomy of multi-layered runtime guardrails for designing foundation model-based agents: Swiss cheese model for ai safety by design,
M. Shamsujjoha, Q. Lu, D. Zhao, and L. Zhu, “A taxonomy of multi-layered runtime guardrails for designing foundation model-based agents: Swiss cheese model for ai safety by design,” arXiv preprint arXiv:2408.02205, 2024
2024 arXiv
-
[81]
From decoding to meta-generation: Inference-time algorithms for large language mod- els,
S. Welleck, A. Bertsch, M. Finlayson, H. Schoelkopf, A. Xie, G. Neubig, I. Kulikov, and Z. Har- chaoui, “From decoding to meta-generation: Inference-time algorithms for large language mod- els,” arXiv preprint arXiv:2406.16838 , 2024. 22
2024 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.