{"id":"c9ad8af6-2897-4dca-a441-46ea1619fcec","arxiv_id":"2509.02053","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Generative AI should be used only as a checked support in technology assessment because persistent structural deficiencies make its outputs unreliable.","lead":"This paper argues that generative AI has structural weaknesses that make it unsafe for technology assessment without careful human checking, and recommends using it only as an idea generator and drafting aid. It catalogues eight causes of these weaknesses, from poor training data to missing continuous learning, and applies them to tasks like horizon scanning.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sec. 2(5)'s claim that LLMs lack a normative 'second perspective' is asserted, not derived; the conclusion that risks are structural rests on this unargued philosophical premise.","rationale":"I read the paper as an argumentative essay in STS whose central claim is that generative AI has structural, persistent weaknesses, so TA should use it only as an idea generator and always check outputs. The practical advice is prudent and supported by common experience and the cited literature. The load-bearing step, however, is the philosophical premise in Section 2(5) that a truth discourse requires a second social perspective and that LLMs lack the corresponding normative component. The reader's weakest_assumption identified exactly this premise, and I agree. The concern is not that the paper disagrees with a consensus; it is that the categorical 'structural' conclusion is not derived from the architecture. The other listed causes are contingent and could improve. Without Section 2(5), the conclusion that the risks are permanent does not follow from Sections 2(1)–(4) and (6)–(7). I therefore recommend CONDITIONAL: the argument holds only if one grants the Brandomian/Habermasian premise and the additional claim that no LLM could ever instantiate it. Since those premises are not defended beyond a brief assertion, the central claim is under-supported. No formal verification or parameter-free derivation is provided, but the paper does synthesize a broad literature and gives concrete TA-relevant examples; those strengths do not fix the missing derivation.","tokens_in":8494,"tokens_out":5603,"duration_ms":70455,"concrete_test":"Attempt to derive from the formal definition of an LLM as an autoregressive next-token predictor trained on a finite corpus a proof that it cannot in principle take a second social perspective as required in Section 2(5). If no such derivation is possible—and the paper provides none—then the 'structural persistence' conclusion is unsupported. This is a single analytical check: test the impossibility claim against the architecture rather than against current training data or benchmark performance.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim—that generative AI's risks are structural and persist despite further development—depends on Section 2(5). There the authors assert that truth-oriented discourse requires taking a second social perspective, which supplies a normative correlate for correct concept use, and that generative AI has only a 'funktional' perspective and lacks this normative component. This is the only candidate in the paper for a genuinely architectural limitation. The other seven causes (data quality, alignment, context-content, reproduction, world model, reasoning) are empirical and potentially remediable: better data filtering, continual learning, improved reasoning, etc. The paper gives no derivation from the architecture of an LLM to the impossibility of such a perspective; 'bisher nicht leisten' in Section 2(4) concedes a current gap rather than a structural one. Nor is it shown that the practical TA recommendation (check outputs) requires the strong claim that AI can never participate in truth discourse; even human work products need checking. If the Habermas/Brandom premise is rejected, or if a future model embedded in social practices is treated as a commitment-bearing participant, the permanence conclusion collapses. The essay therefore moves from a contestable philosophical premise to a categorical conclusion without an argument that the premise is necessary.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper, written in German, examines the use of generative AI in technology assessment (TA). It first characterizes generative AI and formulates requirements for its use in TA (verifiability, traceability, explainability, non-discrimination, etc.). It then identifies eight 'structural causes' of problems in current generative AI: data quality, misalignment, context-content challenges, reproduction, lack of social perspectives, world models, reasoning, and (implicitly) transparency. The central claim, stated in the abstract and Section 2, is that these risks are structural and persist despite constant further development. The paper concludes that, for now, generative AI should be used in TA only as an idea generator and support tool, and that its outputs must always be checked. It also warns against an uncritical 'more information is better' attitude and stresses the importance of human understanding.","tokens_in":8785,"tokens_out":3339,"duration_ms":38957,"significance":"If the structural-persistence claim were established, the paper would have an important message for TA and for AI-assisted scientific work more broadly: no amount of incremental model improvement would remove the fundamental obstacles to using generative AI as an independent epistemic agent. The paper usefully assembles a broad range of well-cited critiques and translates them into a concrete, cautious recommendation for TA practice. The practical advice—treat generative AI as a suggestion generator and always verify outputs—is sensible and robust even without the permanence claim. However, the paper's theoretical contribution, its categorical distinction between 'structural' and merely 'current' limitations, is not adequately argued, and the strongest conclusion goes beyond the evidence presented.","major_comments":[{"comment":"The paper's central claim that the risks are 'structurally induced' and persist rests almost entirely on Sec. 2(5). There the authors assert that truth-oriented discourse requires a 'second social perspective' that supplies a normative correlate for correct concept use, and that generative AI's perspective is merely 'functional' and lacks this normative component. This is presented as a fact, with citations to Habermas and Brandom, but no argument is given for why an LLM, now or in the future, cannot be embedded in social practices that provide such a normative dimension. Moreover, Sec. 2(4) uses 'bisher nicht leisten' ('so far unable to do'), which explicitly characterizes the limitation as contingent. The abstract's categorical 'bleiben jedoch bestehen' is thus unsupported. This is load-bearing: if the philosophical premise is rejected or is merely a current-state observation, the perm","section":"Abstract and Sec. 2(5)"},{"comment":"Seven of the eight listed 'structural causes' are empirical, potentially remediable limitations: data quality can improve; alignment methods are being refined; context-content behavior is being studied; reproduction could change with continual learning (the paper itself cites Shi et al. 2024); world models are under active development; reasoning performance is improving (even if imperfectly). The paper offers no criterion to distinguish 'structural' from 'current technical limitation.' Without such a criterion, calling all eight 'structural' conflates architecture-inherent constraints with engineering challenges, and weakens the central argument. The authors should either define what 'structural' means and show which causes are genuinely invariant, or restrict the claim to 'current' limitations.","section":"Sec. 2(1)-(7)"},{"comment":"The practical recommendation to restrict generative AI to idea generation and support is reasonable and does not require the permanence thesis. Even human-produced work in TA must be checked. As written, however, the recommendation is presented as a consequence of the structural claim ('aufgrund der angeführten Schwächen'). If the structural claim is not established, the 'never use unchecked' recommendation still stands, but the paper's justification shifts from 'impossible in principle' to 'currently inadvisable.' The authors should decouple these two levels or supply the missing argument for why the gap is in-principle unbridgeable.","section":"Sec. 4 / practical recommendation"}],"minor_comments":[{"comment":"The citation grouping has a formatting error: '(Min et al. 2022 , Dai et al. 2023) , Kossen et al. 2024)' — extra comma and unbalanced parenthesis.","section":"Sec. 2(3)"},{"comment":"'ChatGPT-4o1' is likely a typo; either 'ChatGPT-4o' or 'o1' is intended. Also the repeated 'Angeblich' ('allegedly') twice in two sentences is stylistically weak and could be clearer about the source of these claims.","section":"Sec. 2(7)"},{"comment":"Some references are incomplete or inconsistent: 'Grunwald, Nomos (2023): TA' lacks the author's first name and full title; 'Brooks, Tim et al. (2024): Video generation models as world simulators. In.:' is truncated; several arXiv references lack full bibliographic details. These should be cleaned up.","section":"References"},{"comment":"The paper's structure benefits from the numbered list of causes, but the relationship between Sec. 2(5) and Sec. 2(4) is confusing: both discuss social perspectives and learning, and the distinction could be made crisper. Consider merging or cross-referencing more explicitly.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a literature-based essay for a collected volume (NTA11), not a formal empirical or technical study. The practical TA guidance is sound and publishable after modest revision. The main concern is that the abstract and title overclaim by asserting 'structural' risks persist, when the argument only supports current limitations. The authors should either provide a genuine derivation of the in-principle claim (e.g., from LLM architecture) or soften the claim to 'current risks' — the paper's own language in Sec. 2(4) suggests they are aware of this. I would not reject, but the central claim as stated needs rework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a position essay, not a research result. It does a fair job of assembling eight reasons current generative AI is unreliable and maps them onto tasks in technology assessment (TA). The TA-specific application is the real contribution; the recommendation—treat generative AI as an idea generator and checked aid, never an independent analyst—is sensible and well within mainstream practice. The paper is honest, well-cited (Bender & Koller, Shumailov, etc.), and the self-citation (Renftle et al.) is not load-bearing.\n\nThe soft spot is the central claim that these problems are 'structural' and will persist. Seven of the eight causes—data quality, alignment failures, context-content ambiguity, lack of continuous learning, world models, reasoning gaps, reproducibility—are empirical limitations that could in principle be mitigated by better data filtering, continual learning, and improved architectures. The only candidate for a genuinely architectural barrier is Section 2(5), where the authors assert that truth-oriented discourse requires a 'second social perspective' and that LLMs, being merely functional, lack the normative component. But this is asserted, not derived. The Habermas/Brandom premise is contestable, and the paper provides no argument that no future model embedded in social practices could take such a perspective. The text itself says 'bisher nicht leisten' (cannot yet do), which concedes a current gap rather than a permanent one.\n\nThat overreach matters, but it doesn't sink the practical conclusion. The advice to check outputs does not require the permanence claim—even human analysts' work gets checked. So the paper would be stronger if it dropped 'strukturell bedingt' and simply said 'current models have serious, poorly understood limitations; use them with verification.' The philosophical section reads as an attempt to make the conclusion categorical, and it doesn't quite get there.\n\nWho is this for? Readers in TA, STS, and science policy who want a compact, German-language overview of why LLMs shouldn't be trusted unsupervised. AI researchers will find nothing new. It deserves refereeing in its home context (the NTA volume) mainly to tighten the argument in Section 2(5). I'd send it to review, not desk-reject, but I'd ask for a revision that either substantiates the architectural claim or softens it to a current-limitation claim. I wouldn't cite it in my own work—it's a synthesis, not a source of new evidence.","headline":"A competent German-language synthesis of known LLM critiques applied to technology assessment; the strong 'structural permanence' claim rests on an under-argued philosophical premise, but the practical advice is sound.","tokens_in":9235,"tokens_out":2223,"would_cite":false,"duration_ms":23881,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generative AI's persistent problems are structural, so technology assessment should use it only for ideas and drafting, never without human verification.","keywords":["generative AI","large language models","technology assessment","structural risks","truth discourse","second social perspective","grounding problem","AI alignment"],"falsifier":"A concrete falsification would be a reproducible demonstration that a language model, after being corrected on a factual error, stops repeating that error across new sessions (continuous learning), and in an open-ended argument exchange revises its position in response to counterarguments while explicitly taking responsibility for its claims—behavior that satisfies the second-social-perspective standard. If such behavior were observed in a current or future model, the paper's claim that the deficits are structural would fail.","tokens_in":8408,"feed_emoji":"🤖","tokens_out":7827,"duration_ms":76269,"temperature":0.7,"pith_summary":"This paper argues that the problems with generative AI are not temporary growing pains but structural, rooted in how these systems are trained and what they lack. Reviewing eight causes—data quality, misalignment, context-content conflicts, the inability to learn continuously, the absence of a second social perspective, missing world models, and unreliable reasoning—the authors conclude that generative AI cannot participate in the truth-seeking discourse that technology assessment (TA) requires. They therefore recommend that TA use generative AI only as an idea generator and drafting aid, with every output verified by a human. The paper grounds this conclusion in a philosophical claim: taking part in discourse about truth requires being able to adopt a second social perspective, which supplies a normative standard for correct concept use—something a purely functional language model lacks.","feed_headline":"AI's risks are structural and won't fade","feed_subtitle":"Technology assessment should keep generative AI to idea-sparking and drafting, checking every output.","key_machinery":"The load-bearing mechanism is the 'second social perspective' (zweite soziale Perspektive), a concept the paper takes from discursive theories of truth. The paper argues that truth-seeking discourse requires a speaker to take a second perspective, which generalizes an individual opinion, intersubjectivizes it, and provides the normative correlate that confirms the correct use of a concept; speakers thereby take responsibility and enter commitments. Generative AI is characterized as purely functional—it lacks this normative component—and this absence, combined with the absence of continuous learning, is what makes its errors structural rather than fixable by alignment or further scaling. The","core_discovery":"The paper's central claim is that the persistent deficiencies of generative AI have structural causes that further development of the current architecture will not remove. It identifies eight such causes, leading to the conclusion that generative AI should currently be limited to acting as an idea generator and support tool; its outputs must not be used without verification. The deepest cause is the missing normative component: following the discourse-theoretic and inferentialist traditions cited in the paper, truth discourse requires a second social perspective that generalizes and intersubjectivizes a single opinion and provides the normative correlate that confirms the correct use of a co","pith_inferences":["A testable extension would be a benchmark probing whether a model can revise its position after counterargument and explicitly take responsibility for its claims; if any current or future model passes reliably, the paper's permanence claim would need qualification.","The argument implies that continual learning (ongoing training from interaction) is a necessary but not sufficient condition for AI to enter truth discourse; also needed is a normative alignment beyond current reward-modeling approaches.","Applied to AI-assisted democratic deliberation tools, the paper's logic suggests that if AI cannot take a second social perspective, it cannot mediate or arbitrate public discourse without human oversight.","The paper's framing connects to the AI-collapse literature: if synthetic data further pollutes training corpora, the data-quality structural cause will worsen over time, reinforcing the conclusion."],"forward_implications":["TA projects should adopt policies that restrict generative AI to brainstorming, clustering, summarizing, and formatting, with mandatory human verification of all outputs.","Institutional guidance for parliamentary technology assessment bodies should require transparency about training data sources and explicit risk disclosure, since better prompting cannot eliminate the underlying risks.","Research funding in AI for TA should prioritize detection of hallucinations, bias, and proxy effects rather than assuming alignment will eventually resolve the structural deficits.","The same restriction logically extends to other expert domains where verifiable, accountable discourse is the core product, such as regulatory science and legal advice.","If the structural argument is correct, scaling up models or adding alignment will not make generative AI suitable for unsupervised knowledge work; the gating factor is the missing normative-social capacity."],"supporting_citations":[{"why":"Supplies the theory of communicative action and the requirement that truth discourse needs a second social perspective.","marker":"Habermas (1981)"},{"why":"Provides the normative correlate for correct concept use and the notion of discursive commitment.","marker":"Brandom (1994)"},{"why":"Confirms that grounding to meaning is missing in large language models trained only on text.","marker":"Bender und Koller (2020)"},{"why":"Establishes the symbol grounding problem that the paper extends to generative AI.","marker":"Harnad 1990"},{"why":"Shows that training on synthetic data causes model collapse, supporting the data-quality structural cause.","marker":"Shumailov et al. 2024"},{"why":"Demonstrates that a flagship reasoning model fails on variants of tasks in its training set, supporting the reasoning and overfitting-to-data-structure argument.","marker":"Mirzadeh et al. 2024"}],"fun_headline_variants":["AI's structural flaws persist despite progress","Tech assessment keeps AI on a short leash","Generative AI: verify or don't use","Why AI risks are baked in, not bugs","Use AI as spark, not as oracle"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The paper's conclusion rests on the premise that taking part in truthful discourse requires the ability to adopt a second social perspective, which provides a normative standard for correct concept use; generative AI, being purely functional, is asserted to lack this component permanently.","fun_headline_variants_meta":{"raw":{"variants":["AI's structural flaws persist despite progress","Tech assessment keeps AI on a short leash","Generative AI: verify or don't use","Why AI risks are baked in, not bugs","Use AI as spark, not as oracle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000147,"raw_usage":{"total_tokens":949,"prompt_tokens":598,"completion_tokens":351,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":342,"completion_tokens_details":{"reasoning_tokens":284}},"tokens_in":342,"tokens_out":351,"duration_ms":4390,"temperature":1.0,"reasoning_tokens":284,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:54:26.785482+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsification would be a reproducible demonstration that a language model, after being corrected on a factual error, stops repeating that error across new sessions (continuous learning), and in an open-ended argument exchange revises its position in response to counterarguments while explicitly taking responsibility for its claims—behavior that satisfies the second-social-perspective standard. If such behavior were observed in a current or future model, the paper's claim that the deficits are structural would fail.","supporting_citations":[],"review_version":1}