{"id":"313b1e90-be2c-42da-8a82-d7b07c511799","arxiv_id":"2608.05063","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey uniting chiplet-hardware security and LLM-driven EDA security that identifies a missing bridge: LLM-based security tools are not yet tailored to 2.5D/3D chiplet systems.","lead":"Semiconductor chips are increasingly built from smaller chiplets wired together, and AI language models are now writing parts of chip designs. This paper reviews the security risks created by both trends and argues that AI-based defenses have not yet been applied specifically to chiplet systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gap claim unsubstantiated: no search protocol backs the negative existential assertion that LLM-based security has not been applied to chiplet trust boundaries.","rationale":"The most load-bearing assumption is the completeness of the negative evidence. The reader already identified this, and I agree. The paper's value as a taxonomy and synthesis is not destroyed by this concern, but the central claim that 'current research fails' is an overstrong negative existential. A single omitted prior work would require softening the claim to 'to our knowledge, no prior work' or 'this remains nascent.' The concrete test is a structured search; if none of the queries surface omitted prior art, the original accept verdict stands. I did not find a stronger internal inconsistency: the defense section relies on primary sources with reported numbers, but the central claim is not about defense efficacy, and no formal verification is expected for a survey. The internal ambiguity surrounding reference [11] reinforces, rather than replaces, the completeness concern.","tokens_in":9311,"tokens_out":4151,"duration_ms":42260,"concrete_test":"Perform a structured literature search on DBLP, Scopus, and arXiv (2023–2026) with queries combining (LLM OR 'large language model') AND (chiplet OR '2.5D' OR '3D IC' OR interposer) AND (security OR Trojan OR attack OR trust OR 'root of trust'); screen all hits for work that applies LLM-based analysis or defense to chiplet-level trust boundaries. If any qualifying paper is found that is not discussed in §VI, the central novelty claim should be revised from 'fails' to 'remains nascent,' and the survey should be accepted conditional on this revision.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central contribution is a gap claim: current research 'fails to leverage trustworthy LLMs to secure complex chiplet systems' (Introduction, §VI.6, §VII). This is a negative existential statement about the literature. The only evidence offered is the selected reference list; the survey states no systematic search protocol, inclusion/exclusion criteria, or date range anywhere in §II or the Introduction. As a result, the claim is an argument from silence. The risk is not that the survey is poorly written but that the motivating assertion would be false if even one relevant prior work exists. The paper's own reference [11] (same group's earlier position paper on LLMs for secure hardware design) could already cover part of the intersection, and the text does not explain why [11] does not count. This makes the gap claim both externally unverified and internally ambiguous.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of the intersection of two recent developments in semiconductor hardware: 2.5D/3D chiplet integration and LLM-driven EDA. It reviews attacks on chiplet systems at architectural, logical, and physical levels; defense mechanisms based on 2.5D split manufacturing and active-interposer roots of trust, including transaction monitors and coherence message checkers; threats to LLM-based RTL generation such as backdoor attacks, data contamination, safety misalignment, and IP leakage; and defenses including safe fine-tuning, semantic consensus decoding, dynamic benchmarking, machine unlearning, and training sanitization. It also reviews the use of LLMs for hardware security tasks such as logic locking, side-channel assessment, Trojan detection, red-teaming, and bug detection. The paper concludes that no existing work applies trustworthy LLMs to secure chiplet systems, and it calls for cross-domain research combining LLM agents with chiplet root-of-trust architectures.","tokens_in":9483,"tokens_out":5426,"duration_ms":53034,"significance":"The survey is a useful and timely systematization. Its strongest parts are the organization of chiplet-specific attacks and countermeasures around the active interposer as a physically isolated root of trust, and the organization of LLM-EDA security into backdoor, contamination, misalignment, and IP-leakage categories. It gives concrete attention to mechanisms such as TRANSMONs, CMC-1/CMC-2, SafeTune, SCD, SALAD, and CircuitGuard, and it ends with a testable research direction in Section VII: automated generation of access policies for TRANSMONs and CMCs from architectural descriptions and OS permissions. The paper does not present original measurements or proofs, so its value rests on the accuracy of the literature representation and the validity of its central gap claim. The individual citations are mostly plausible, but the central negative claim is not yet supported by a reproducible selection methodology.","major_comments":[{"comment":"The paper's central claim is the assertion that 'current research fails to leverage trustworthy LLMs to secure complex chiplet systems.' This is a negative existential claim about the literature, but the manuscript does not describe a systematic search protocol, inclusion/exclusion criteria, or a date range anywhere in the Introduction or Section II. The claim is therefore an argument from the selected reference list. Please add a methodology statement and either scope the claim (for example, 'to our knowledge' or 'within the surveyed period') or demonstrate coverage of the intersection.","section":"Section I, third paragraph; Section VI.6"},{"comment":"The gap claim is internally ambiguous because the reference list already contains works that may sit at the claimed intersection. Reference [6] is a survey of hardware security and trust for chiplet-based 2.5D/3D ICs, and Reference [11] is a position paper by the same group on LLMs for secure hardware design and related problems. The text never explains why these works do not count as prior art for 'LLMs for securing chiplet systems.' Please state explicitly what makes the proposed intersection distinct and why those references do not close the gap.","section":"Section VI.6; Refs. [6] and [11]"},{"comment":"The sentence 'structural vulnerabilities of multi-vendor chiplet systems executing distributed acceleration remain a critical and unaddressed security gap' is another negative existential claim used to motivate the survey. It depends on the same unstated completeness assumption as the Introduction. If the authors retain this sentence, they should either supply evidence of coverage or soften it to 'not addressed by the works we surveyed.'","section":"Section III.A"}],"minor_comments":[{"comment":"The description of TrojanLoC is garbled: 'TrojanLoC [70] Trained on the TrojanInS dataset, it uses devises an RTL-adapted transformer...' Please rewrite as a grammatical sentence and clarify whether the training on TrojanInS is part of [70]'s contribution or a baseline setup.","section":"Section VI.3"},{"comment":"The quantitative results (73.7% IR-drop reduction, 18.5% footprint reduction, 2.68% interposer utilization, ~4% monitoring overhead, and 3.2% system power reduction) come from a specific design study in [53]. Please add a citation-context sentence so readers do not generalize these single-design measurements.","section":"Section IV.C"},{"comment":"Reference [9] lists 'arXiv:2601.19908, 2025', but the arXiv identifier implies January 2026; please correct the year. Several other entries ([24], [25], [57]) lack years or version numbers and should be completed.","section":"References"},{"comment":"The paper would benefit from a figure or table laying out the threat-defense taxonomy; without one, the relationships among attack surfaces, defense mechanisms, and the proposed gap are harder to follow than necessary.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The main risk to publication is the unsubstantiated gap claim. I would ask the authors for a short methods addendum and an explicit discussion of why Refs. [6] and [11] do not already occupy the claimed intersection. The manuscript is within scope for a hardware security / computer architecture venue, and the authors are clearly knowledgeable in the area, but the survey currently reads as an expert narrative rather than a verifiable systematic review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful survey, and you should take it seriously as a reference map rather than as a new technical result. The paper does two things well: it pulls together the chiplet security literature (attacks, 2.5D root-of-trust with active interposers, TRANSMON/CMC monitors) and the LLM-driven EDA security literature (backdoors, contamination, IP leakage, defenses), and it highlights an intersection—using trustworthy LLMs to secure chiplet systems—that is plausible and genuinely underexplored. The taxonomy is clear and the citation list is broad. I agree with the reader that this is integrative work, not a novel attack or defense, and that's fine for a venue that wants to shape a community.\n\nThe soft spots are real but not fatal. The central gap claim—that LLM-based security has not been applied to 2.5D/3D chiplet trust boundaries—is a negative existential assertion. The survey gives no systematic search protocol, so the claim rests on the authors' selection of references. That's a legitimate concern, especially since [11] (the same group's earlier position paper on LLMs for secure hardware design) could already cover part of the intersection; the paper doesn't explain why that work doesn't count. I'd like to see that addressed, either by narrowing the claim or by acknowledging the prior position paper. There are also a few editorial slips, like the garbled TrojanLoC sentence in Section VI.3 and unverified quantitative claims (the 73.7% IR-drop reduction) taken from cited papers—normal for a survey, but the gap claim needs a sharper defense.\n\nFor the rest, the paper is well organized, and the citations look appropriate for a survey, with self-citations used as evidence in their own domain rather than to prop up the central claim. It doesn't overreach; the conclusion is a call for future research, not a solved problem.\n\nWho is this for? Researchers entering chiplet security or LLM-based hardware security will get a good map. The paper deserves a serious referee; an editor should send it out rather than desk-reject. I'd suggest the referee ask for a more explicit treatment of the gap claim—either a lightweight search or a clear acknowledgment of [11]—and for cleanup of the prose errors. My verdict: accept with minor revision, and I'd probably cite it as the survey reference for this intersection.","headline":"A solid integrative survey of chiplet security and LLM-driven EDA security, with a plausible but under-defended gap claim at their intersection.","tokens_in":9989,"tokens_out":2871,"would_cite":true,"duration_ms":26193,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Chiplet and LLM hardware security must merge, review argues","keywords":["chiplet security","2.5D integration","large language models","EDA","root of trust","hardware Trojans","data contamination","IP leakage"],"falsifier":"Locate a peer-reviewed publication that already applies a large-language-model-based security agent specifically to 2.5D/3D chiplet trust boundaries, for instance automatically generating TRANSMON access policies or coherence-monitor rules; demonstrating such a tool would undercut the paper's central missing-gap claim.","tokens_in":9145,"feed_emoji":"🛡️","tokens_out":5409,"duration_ms":46039,"temperature":0.7,"pith_summary":"This review paper argues that two simultaneous shifts in semiconductor design—moving to heterogeneous 2.5D chiplet systems and adopting LLM-driven EDA tools—each expand the hardware attack surface, and that the two must be secured together. It surveys attacks on chiplet systems, from coherence-protocol Trojans to physical probing, and reviews defenses built on an active-interposer Root of Trust. It then reviews threats and defenses for LLM-driven EDA, including backdoors, data contamination, prompt injection, and IP leakage. The paper's central claim is that a cross-domain blind spot remains: current research does not yet apply trustworthy LLMs to secure chiplet trust boundaries, nor does it use 2.5D-anchored roots of trust to secure LLM acceleration hardware. If true, the intersection of these two fields is an open frontier with concrete security consequences.","feed_headline":"Chiplet and LLM hardware security must merge, review argues","feed_subtitle":"Two design revolutions each widen the attack surface, and the paper maps the missing link between their defenses.","key_machinery":"The load-bearing elements are two: (1) the 2.5D Root of Trust, a physically isolated security monitor implemented in the active interposer, where Transaction Monitors (TRANSMONs) enforce access control and data masking as hardware shims between untrusted chiplets and the network-on-chip fabric, and Coherence Message Checkers (CMCs) validate cache-coherence flits against OS-managed permissions; (2) the taxonomy of native LLM-EDA threats—backdoor poisoning, data contamination, prompt injection, and IP leakage—with corresponding defenses such as SafeTune, Semantic Consensus Decoding, VeriContaminated, SALAD, and CircuitGuard. These mechanisms are the paper's basis for claiming that chiplet security and LLM-EDA security can be made compatible rather than mutually exclusive.","core_discovery":"On the paper's own account, the central discovery is that the threats to 2.5D chiplet systems and the threats to LLM-driven EDA pipelines are not independent: when chiplets from multiple vendors are integrated on an interposer to accelerate LLM workloads, predictable multi-chiplet memory traffic becomes a side channel for model theft, and when LLMs write hardware, the same models introduce backdoors, leakage, and contamination. The paper's synthesis shows that existing defenses—physically isolated active-interposer roots of trust with transaction monitors and coherence message checkers on the hardware side, and sanitization, unlearning, and consensus decoding on the LLM side—each address one frontier, but neither has been applied to the other. The paper consequently pinpoints a gap: no current LLM-based security approach is semantically aware of multi-vendor chiplet trust boundaries, interposer fabrics, or active interposer configurations.","pith_inferences":["The paper's taxonomy implies that LLM-EDA security and chiplet security could be unified into a single trust model: the interposer root of trust as the physical enforcement point and the LLM as the policy synthesizer.","A testable extension would be to measure whether the 1–2-stage latency of coherence message checkers scales to the higher message rates of LLM inference workloads, since the survey cites minimal overhead for general coherence traffic but does not benchmark LLM-specific patterns.","The paper leaves implicit that the same LLM-driven EDA pipelines that create vulnerabilities could be used to generate the security assertions and monitors that defend chiplet systems, potentially closing the loop.","The review's emphasis on data contamination in LLM-EDA benchmarking suggests that contamination-detection metrics like those in VeriContaminated could be turned into a continuous health check for hardware-security LLM tools."],"forward_implications":["If the gap is real, the immediate research agenda is to build LLM-driven agents that synthesize system-wide security policies for interposer-based systems directly from architectural descriptions and OS permissions.","If the 2.5D Root of Trust with CMCs is applied to heterogeneous LLM stacks, the highly predictable weight and activation traffic of LLM inference becomes monitorable, plausibly closing the model-theft side channel identified in the survey.","If LLM-EDA security measures such as machine unlearning and dynamic benchmarking are adopted, public hardware benchmarks can be made trustworthy for reporting real LLM coding capability.","If domain-aware safety alignment like that explored by HarmChip is extended to hardware-specific guardrails, the false-positive blocking of legitimate engineering tasks can be reduced while still blocking semantic-disguised attacks.","If the proposed synergy is realized, chiplet-based LLM acceleration can ship with a physically anchored root of trust that also guards the LLM pipeline itself."],"supporting_citations":[{"why":"Supplies the 2.5D Root of Trust and TRANSMON monitor design, the central hardware defense object.","marker":"[4]"},{"why":"Defines coherence attacks and the CMC-1/CMC-2 countermeasures, core to the chiplet threat/defense pair.","marker":"[5]"},{"why":"Provides the active-interposer design flow and the physical/performance figures that ground the defense proposal.","marker":"[53]"},{"why":"Demonstrates IP leakage via few-shot fine-tuning, one of the four native LLM-EDA threats.","marker":"[15]"},{"why":"Shows backdoor activation via rare keywords in RTL generation, another native threat.","marker":"[16]"},{"why":"Quantifies near-total benchmark contamination in commercial models, supporting the contamination-threat claim.","marker":"[17]"},{"why":"Supplies the SALAD machine unlearning defense against contamination, backdoors, and IP leakage.","marker":"[58]"},{"why":"Shows a chiplet-based LLM inference architecture, making the chiplet-LLM threat intersection concrete.","marker":"[7]"}],"fun_headline_variants":["Chiplet and LLM security gaps call for unified defenses","Merging chiplet and LLM security is the missing link","Review maps chiplet-LLM attack surfaces and defenses","Hardware security must span chiplets and LLM EDA","Unified defense needed for chiplet and LLM threats"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's central conclusion depends on the assumption that the selected literature is complete enough to establish the gap; the paper does not describe a systematic search protocol, so if relevant prior art already applies LLM-based security to chiplet trust boundaries, the claimed gap would collapse.","fun_headline_variants_meta":{"raw":{"variants":["Chiplet and LLM security gaps call for unified defenses","Merging chiplet and LLM security is the missing link","Review maps chiplet-LLM attack surfaces and defenses","Hardware security must span chiplets and LLM EDA","Unified defense needed for chiplet and LLM threats"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1218,"prompt_tokens":895,"completion_tokens":323,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":237}},"tokens_in":511,"tokens_out":323,"duration_ms":3478,"temperature":1.0,"reasoning_tokens":237,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:09:50.792253+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Locate a peer-reviewed publication that already applies a large-language-model-based security agent specifically to 2.5D/3D chiplet trust boundaries, for instance automatically generating TRANSMON access policies or coherence-monitor rules; demonstrating such a tool would undercut the paper's central missing-gap claim.","supporting_citations":[{"cited_title":"2.5D root of trust: Secure system-level integration of untrusted chiplets,","cited_arxiv_id":null,"evidence_quote":"Supplies the 2.5D Root of Trust and TRANSMON monitor design, the central hardware defense object."},{"cited_title":"Coherence attacks and countermeasures in interposer-based chiplet systems,","cited_arxiv_id":null,"evidence_quote":"Defines coherence attacks and the CMC-1/CMC-2 countermeasures, core to the chiplet threat/defense pair."},{"cited_title":"Design flow for active interposer-based 2.5D ICs and study of RISC-V architecture with secure NoC,","cited_arxiv_id":null,"evidence_quote":"Provides the active-interposer design flow and the physical/performance figures that ground the defense proposal."},{"cited_title":"VeriLeaky: Navigating IP protection vs utility in fine-tuning for LLM-driven Verilog coding,","cited_arxiv_id":null,"evidence_quote":"Demonstrates IP leakage via few-shot fine-tuning, one of the four native LLM-EDA threats."},{"cited_title":"RTL-Breaker: Assessing the security of LLMs against backdoor attacks on HDL code generation,","cited_arxiv_id":null,"evidence_quote":"Shows backdoor activation via rare keywords in RTL generation, another native threat."},{"cited_title":"VeriContaminated: Assessing LLM-driven Verilog coding for data contamination,","cited_arxiv_id":null,"evidence_quote":"Quantifies near-total benchmark contamination in commercial models, supporting the contamination-threat claim."},{"cited_title":"SALAD: Systematic assessment of machine unlearning on LLM- aided hardware design,","cited_arxiv_id":null,"evidence_quote":"Supplies the SALAD machine unlearning defense against contamination, backdoors, and IP leakage."},{"cited_title":"Cambricon-LLM: A chiplet-based hybrid architecture for on-device inference of 70B LLM,","cited_arxiv_id":null,"evidence_quote":"Shows a chiplet-based LLM inference architecture, making the chiplet-LLM threat intersection concrete."}],"review_version":1}