Introduces SAKE benchmark with 2154 questions to assess LLMs on software architectural knowledge, showing high overall accuracy but marked gaps across categories.
In: 2025 IEEE 22nd International Conference on Software Architecture (ICSA)
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.SE 9roles
background 1polarities
background 1representative citing papers
LLMs achieve only modest understanding of HMSC formal semantics at 52 percent accuracy, performing strongly on basic constructs but weakly on abstractions and traces.
A small recency window of 3-5 prior ADRs as context produces higher-fidelity LLM-generated Architecture Decision Records than no context, full history, or retrieval-augmented selection in typical sequential workflows.
LLM approaches ExArch and ArTEMiS reach F1 scores of 0.86 and 0.81 for architecture entity recognition and traceability, matching or approaching baselines that require manual models.
Describes a tool landscape (REST API, TraceView, TraceViz) that makes ARDoCo's four TLR pipelines publicly accessible with a preliminary study showing TraceViz improves developer comprehension.
LLMs achieve 98.22% accuracy answering factual questions about ROS2 software architectures, with top models reaching 100%.
EnergyTrackr detects statistically significant energy regressions in Java commits from 3,232 changes across three projects and identifies recurring code anti-patterns such as missing early exits.
xDECAF is an extensible, open-source framework for architecture-based data flow analysis with a constraint DSL, web editor, and a catalog of 26 example models for information security.
Temporal community detection applied to six releases of the train-ticket microservice benchmark reveals a stable two-community structure aligned with business processes, plus some multi-community services.
citing papers explorer
-
SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models
Introduces SAKE benchmark with 2154 questions to assess LLMs on software architectural knowledge, showing high overall accuracy but marked gaps across categories.
-
(How) Do Large Language Models Understand High-Level Message Sequence Charts?
LLMs achieve only modest understanding of HMSC formal semantics at 52 percent accuracy, performing strongly on basic constructs but weakly on abstractions and traces.
-
Context Matters: Evaluating Context Strategies for Automated ADR Generation Using LLMs
A small recency window of 3-5 prior ADRs as context produces higher-fidelity LLM-generated Architecture Decision Records than no context, full history, or retrieval-augmented selection in typical sequential workflows.
-
Who's Who? LLM-assisted Software Traceability with Architecture Entity Recognition
LLM approaches ExArch and ArTEMiS reach F1 scores of 0.86 and 0.81 for architecture entity recognition and traceability, matching or approaching baselines that require manual models.
-
The ARDoCo Tool Landscape: REST API, TraceView, and TraceViz for Architecture Traceability
Describes a tool landscape (REST API, TraceView, TraceViz) that makes ARDoCo's four TLR pipelines publicly accessible with a preliminary study showing TraceViz improves developer comprehension.
-
Can Large Language Models Assist the Comprehension of ROS2 Software Architectures?
LLMs achieve 98.22% accuracy answering factual questions about ROS2 software architectures, with top models reaching 100%.
-
Systematic Detection of Energy Regression and Corresponding Code Patterns in Java Projects
EnergyTrackr detects statistically significant energy regressions in Java commits from 3,232 changes across three projects and identifies recurring code anti-patterns such as missing early exits.
-
xDECAF: An Extensible Data Flow Diagram Analysis Framework for Information Security
xDECAF is an extensible, open-source framework for architecture-based data flow analysis with a constraint DSL, web editor, and a catalog of 26 example models for information security.
-
Analyzing the Evolution of Structural Communities within Microservice Architecture
Temporal community detection applied to six releases of the train-ticket microservice benchmark reveals a stable two-community structure aligned with business processes, plus some multi-community services.