Pith. sign in

REVIEW 3 major objections 4 minor 4 references

Generative KI f\"ur TA

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Generative AI's persistent problems are structural, so technology assessment should use it only for ideas and drafting, never without human verification.

desk verdict A competent German-language synthesis of known LLM critiques applied to technology assessment; the strong 'structural permanence' claim rests on an under-argued philosophical premise, but the practical advice is sound. read the letter →

arxiv 2509.02053 v1 pith:7P7M52I5 submitted 2025-09-02 cs.AI

classification cs.AI
keywords generativeAIlargelanguagemodelstechnologyassessmentstructuralriskstruthdiscoursesecondsocialperspectivegroundingproblemalignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the problems with generative AI are not temporary growing pains but structural, rooted in how these systems are trained and what they lack. Reviewing eight causes—data quality, misalignment, context-content conflicts, the inability to learn continuously, the absence of a second social perspective, missing world models, and unreliable reasoning—the authors conclude that generative AI cannot participate in the truth-seeking discourse that technology assessment (TA) requires. They therefore recommend that TA use generative AI only as an idea generator and drafting aid, with every output verified by a human. The paper grounds this conclusion in a philosophical claim: taking part in discourse about truth requires being able to adopt a second social perspective, which supplies a normative standard for correct concept use—something a purely functional language model lacks.

What carries the argument

The load-bearing mechanism is the 'second social perspective' (zweite soziale Perspektive), a concept the paper takes from discursive theories of truth. The paper argues that truth-seeking discourse requires a speaker to take a second perspective, which generalizes an individual opinion, intersubjectivizes it, and provides the normative correlate that confirms the correct use of a concept; speakers thereby take responsibility and enter commitments. Generative AI is characterized as purely functional—it lacks this normative component—and this absence, combined with the absence of continuous learning, is what makes its errors structural rather than fixable by alignment or further scaling. The

What would settle it

A concrete falsification would be a reproducible demonstration that a language model, after being corrected on a factual error, stops repeating that error across new sessions (continuous learning), and in an open-ended argument exchange revises its position in response to counterarguments while explicitly taking responsibility for its claims—behavior that satisfies the second-social-perspective standard. If such behavior were observed in a current or future model, the paper's claim that the deficits are structural would fail.

Watch

Extended reading notes

Core claim

The paper's central claim is that the persistent deficiencies of generative AI have structural causes that further development of the current architecture will not remove. It identifies eight such causes, leading to the conclusion that generative AI should currently be limited to acting as an idea generator and support tool; its outputs must not be used without verification. The deepest cause is the missing normative component: following the discourse-theoretic and inferentialist traditions cited in the paper, truth discourse requires a second social perspective that generalizes and intersubjectivizes a single opinion and provides the normative correlate that confirms the correct use of a co

Load-bearing premise

The paper's conclusion rests on the premise that taking part in truthful discourse requires the ability to adopt a second social perspective, which provides a normative standard for correct concept use; generative AI, being purely functional, is asserted to lack this component permanently.

Editorial extensions

If this is right

  • TA projects should adopt policies that restrict generative AI to brainstorming, clustering, summarizing, and formatting, with mandatory human verification of all outputs.
  • Institutional guidance for parliamentary technology assessment bodies should require transparency about training data sources and explicit risk disclosure, since better prompting cannot eliminate the underlying risks.
  • Research funding in AI for TA should prioritize detection of hallucinations, bias, and proxy effects rather than assuming alignment will eventually resolve the structural deficits.
  • The same restriction logically extends to other expert domains where verifiable, accountable discourse is the core product, such as regulatory science and legal advice.
  • If the structural argument is correct, scaling up models or adding alignment will not make generative AI suitable for unsupervised knowledge work; the gating factor is the missing normative-social capacity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be a benchmark probing whether a model can revise its position after counterargument and explicitly take responsibility for its claims; if any current or future model passes reliably, the paper's permanence claim would need qualification.
  • The argument implies that continual learning (ongoing training from interaction) is a necessary but not sufficient condition for AI to enter truth discourse; also needed is a normative alignment beyond current reward-modeling approaches.
  • Applied to AI-assisted democratic deliberation tools, the paper's logic suggests that if AI cannot take a second social perspective, it cannot mediate or arbitrate public discourse without human oversight.
  • The paper's framing connects to the AI-collapse literature: if synthetic data further pollutes training corpora, the data-quality structural cause will worsen over time, reinforcing the conclusion.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper, written in German, examines the use of generative AI in technology assessment (TA). It first characterizes generative AI and formulates requirements for its use in TA (verifiability, traceability, explainability, non-discrimination, etc.). It then identifies eight 'structural causes' of problems in current generative AI: data quality, misalignment, context-content challenges, reproduction, lack of social perspectives, world models, reasoning, and (implicitly) transparency. The central claim, stated in the abstract and Section 2, is that these risks are structural and persist despite constant further development. The paper concludes that, for now, generative AI should be used in TA only as an idea generator and support tool, and that its outputs must always be checked. It also warns against an uncritical 'more information is better' attitude and stresses the importance of human understanding.

Significance. If the structural-persistence claim were established, the paper would have an important message for TA and for AI-assisted scientific work more broadly: no amount of incremental model improvement would remove the fundamental obstacles to using generative AI as an independent epistemic agent. The paper usefully assembles a broad range of well-cited critiques and translates them into a concrete, cautious recommendation for TA practice. The practical advice—treat generative AI as a suggestion generator and always verify outputs—is sensible and robust even without the permanence claim. However, the paper's theoretical contribution, its categorical distinction between 'structural' and merely 'current' limitations, is not adequately argued, and the strongest conclusion goes beyond the evidence presented.

major comments (3)
  1. [Abstract and Sec. 2(5)] The paper's central claim that the risks are 'structurally induced' and persist rests almost entirely on Sec. 2(5). There the authors assert that truth-oriented discourse requires a 'second social perspective' that supplies a normative correlate for correct concept use, and that generative AI's perspective is merely 'functional' and lacks this normative component. This is presented as a fact, with citations to Habermas and Brandom, but no argument is given for why an LLM, now or in the future, cannot be embedded in social practices that provide such a normative dimension. Moreover, Sec. 2(4) uses 'bisher nicht leisten' ('so far unable to do'), which explicitly characterizes the limitation as contingent. The abstract's categorical 'bleiben jedoch bestehen' is thus unsupported. This is load-bearing: if the philosophical premise is rejected or is merely a current-state observation, the perm
  2. [Sec. 2(1)-(7)] Seven of the eight listed 'structural causes' are empirical, potentially remediable limitations: data quality can improve; alignment methods are being refined; context-content behavior is being studied; reproduction could change with continual learning (the paper itself cites Shi et al. 2024); world models are under active development; reasoning performance is improving (even if imperfectly). The paper offers no criterion to distinguish 'structural' from 'current technical limitation.' Without such a criterion, calling all eight 'structural' conflates architecture-inherent constraints with engineering challenges, and weakens the central argument. The authors should either define what 'structural' means and show which causes are genuinely invariant, or restrict the claim to 'current' limitations.
  3. [Sec. 4 / practical recommendation] The practical recommendation to restrict generative AI to idea generation and support is reasonable and does not require the permanence thesis. Even human-produced work in TA must be checked. As written, however, the recommendation is presented as a consequence of the structural claim ('aufgrund der angeführten Schwächen'). If the structural claim is not established, the 'never use unchecked' recommendation still stands, but the paper's justification shifts from 'impossible in principle' to 'currently inadvisable.' The authors should decouple these two levels or supply the missing argument for why the gap is in-principle unbridgeable.
minor comments (4)
  1. [Sec. 2(3)] The citation grouping has a formatting error: '(Min et al. 2022 , Dai et al. 2023) , Kossen et al. 2024)' — extra comma and unbalanced parenthesis.
  2. [Sec. 2(7)] 'ChatGPT-4o1' is likely a typo; either 'ChatGPT-4o' or 'o1' is intended. Also the repeated 'Angeblich' ('allegedly') twice in two sentences is stylistically weak and could be clearer about the source of these claims.
  3. [References] Some references are incomplete or inconsistent: 'Grunwald, Nomos (2023): TA' lacks the author's first name and full title; 'Brooks, Tim et al. (2024): Video generation models as world simulators. In.:' is truncated; several arXiv references lack full bibliographic details. These should be cleaned up.
  4. [General] The paper's structure benefits from the numbered list of causes, but the relationship between Sec. 2(5) and Sec. 2(4) is confusing: both discuss social perspectives and learning, and the distinction could be made crisper. Consider merging or cross-referencing more explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity.

full rationale

The paper is a synthesis essay: it aggregates published critiques of generative AI under eight labeled 'structural causes' and derives a conservative practical recommendation for technology assessment. It contains no fitted parameters, no quantitative predictions, and no statistical validation, so the 'fitted input called prediction' and 'self-definitional' patterns do not apply. The single self-citation, Renftle et al. 2022 in Sec. 1, supports the claim that explainability approaches have not met expectations; it is paired with an external citation (Wu et al. 2024) and is not load-bearing for the paper's central argument. Section 2(5) asserts, via Habermas and Brandom, that truth-discourse requires a second social perspective and that generative AI lacks the normative component; this is a contestable philosophical premise, not a circular reduction by construction, and it does not define the conclusion into its inputs. No equation, fitted value, or definition is shown to be equivalent to an output. The paper's main vulnerability is the unsupported nature of the Section 2(5) premise, which is a correctness/evidence concern, not a circularity concern.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper contributes no mathematical derivations, fitted parameters, or invented entities. It rests on a set of domain assumptions about how language models work and about the philosophy of truth discourse. The main assumptions are listed above; none are given formal proof, but they are sourced to the cited literature.

assumptions (3)
  • domain assumption Truth-seeking discourse requires a second-person social perspective that provides a normative correlate for correct concept use.
    Section 2(5) cites Habermas 1981 and Brandom 1994 to establish this. The paper does not defend the premise, but it is load-bearing for the conclusion that LLMs cannot engage in TA truth discourse.
  • domain assumption Current large language models do not learn continuously; they are trained once on static corpora and cannot update from interaction.
    Section 2(4) states this as a structural feature. It is an empirical claim about current systems, sourced to literature, but used to infer that the limitation is permanent.
  • domain assumption LLM outputs are generated from statistical patterns in training data, not retrieved from a knowledge base, so they can be false yet fluent.
    Section 1 and Section 2(1) rely on this to argue that hallucinations and unreliability are inherent. It is a widely accepted description but assumed rather than demonstrated here.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generative KI f\"ur TA." pith.science (2026). https://pith.science/paper/7P7M52I5

@misc{pith2026250902053,
  author       = {Pith},
  title        = {Pith review of: Generative KI f\"ur TA},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7P7M52I5}},
  note         = {Machine review of arXiv:2509.02053}
}
read the original abstract

Many scientists use generative AI in their scientific work. People working in technology assessment (TA) are no exception. TA's approach to generative AI is twofold: on the one hand, generative AI is used for TA work, and on the other hand, generative AI is the subject of TA research. After briefly outlining the phenomenon of generative AI and formulating requirements for its use in TA, the following article discusses in detail the structural causes of the problems associated with it. Although generative AI is constantly being further developed, the structurally induced risks remain. The article concludes with proposed solutions and brief notes on their feasibility, as well as some examples of the use of generative AI in TA work.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 3 canonical work pages

  1. [1]

    Chatbots, wie ChatGPT, Gemini, LLama oder Claude, die auf sogenannten großen Sprachmodellen basieren, lassen sich sehr einfach, ohne Einarbeitung nutzen

    Phänomene der generativen KI Ein herausragendes Merkmal generativer KI (im Folgenden synonym verwendet mit Chatbot und multimodalem Sprachmodell LLM) ist, dass sie es ermöglicht, mit Maschinen schriftlich oder mündlich in natürlicher Sprache zu kommunizieren. Chatbots, wie ChatGPT, Gemini, LLama oder Claude, die auf sogenannten großen Sprachmodellen basie...

  2. [2]

    hidden CoT prompt

    Strukturelle Ursachen von Problemen derzeitiger KI Eine paradigmatische Aufgabe der TA ist die Politikberatung (Grunwald 2023). Sie erfordert verlässliche, kluge, situationsgerechte Lösungen. Um dies zu gewährleisten, müssen die, von generativer KI erzeugten, Ausgaben verständlich, nachvollziehbar, kontrollierbar, kohärent, diskriminierungsfrei und erklär...

  3. [3]

    Für die Technikfolgenabschätzung ist die KI sowohl Untersuchungsgegenstand als auch mögliches Werkzeug

    Einsatzmöglichkeiten generativer KI für die TA Aufgrund der angeführten Schwächen ist es ratsam, generative KI sehr bedacht und nur dann einzusetzen, wenn man in der Lage ist, die Ergebnisse zu überprüfen. Für die Technikfolgenabschätzung ist die KI sowohl Untersuchungsgegenstand als auch mögliches Werkzeug. Im Folgenden werden, ohne Anspruch auf Vollstän...

  4. [4]

    Ungeprüft dürfen die Ergebnisse nicht verwendet werden, man darf sich nicht auf sie verlassen

    Konklusion Grob zusammengefasst sollte derzeit die Anwendung generativer KI in der TA darauf beschränkt werden, sie als Ideengeber und als Unterstützung zu nutzen. Ungeprüft dürfen die Ergebnisse nicht verwendet werden, man darf sich nicht auf sie verlassen. Damit unterscheidet sich der Einsatz generativer KI in der TA kaum von ihrem Einsatz in anderen wi...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.