Pith. sign in

REVIEW 15 cited by

Open Problems in Technical AI Governance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14981 v2 pith:5QVKFUMI submitted 2024-07-20 cs.CY

classification cs.CY
keywords governancetechnicalidentifyopenproblemsactionsaddressanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

AI progress is creating a growing range of risks and opportunities, but it is often unclear how they should be navigated. In many cases, the barriers and uncertainties faced are at least partly technical. Technical AI governance, referring to technical analysis and tools for supporting the effective governance of AI, seeks to address such challenges. It can help to (a) identify areas where intervention is needed, (b) identify and assess the efficacy of potential governance actions, and (c) enhance governance options by designing mechanisms for enforcement, incentivization, or compliance. In this paper, we explain what technical AI governance is, why it is important, and present a taxonomy and incomplete catalog of its open problems. This paper is intended as a resource for technical researchers or research funders looking to contribute to AI governance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences

    cs.LG 2026-06 unverdicted novelty 8.0 of 10

    FLIPS identifies LLM instances with 96% closed-set and 90% open-set accuracy by exploiting biases in generated binary random sequences across 237 instances.

  2. Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms

    cs.CY 2026-05 conditional novelty 7.0 of 10

    VLMs preserve linearly separable visual magnitudes and can compare them, yet collapse at symbolic mapping because visual and textual number spaces remain fractured and disjoint.

  3. No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason

    cs.CY 2026-03 unverdicted novelty 7.0 of 10

    AI systems should output Asserted, Denied, or Undetermined based on the availability of a contestable certificate of entitlement rather than confidence scores alone.

  4. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    On the same 720 replies, scoring exposure versus manifestation shifts the auditor-judge gap by ~0.2 AUROC and can reverse their ranking, so single detection AUROCs are under-specified.

  5. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  6. Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI

    cs.CY 2026-07 conditional novelty 6.0 of 10

    A Basel-III-style two-layer system—coordinated finder-coordinator-defender reporting plus ECAR, CRTH, and ARS buffers—can detect and dampen correlated risk build-up across frontier AI labs’ internal deployments.

  7. The Foreign Policy AI Evaluation Gap

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.

  8. Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms

    cs.CY 2026-05 unverdicted novelty 6.0 of 10

    LM agents' changeable modules prevent persistent identity and sanction sensitivity, making reputation mechanisms structurally inapplicable and requiring protocol-based behavioral harnesses instead.

  9. No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason

    cs.CY 2026-03 conditional novelty 6.0 of 10

    An AI may assert or deny high-stakes claims only when it can exhibit a publicly contestable certificate; otherwise it is obligated to return Undetermined.

  10. AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework

    cs.CR 2026-06 unverdicted novelty 5.0 of 10

    The paper presents a threat model, taxonomy, and six-dimension measurement framework for AI sandboxes to clarify valid testing claims for safety, security, and regulatory assurance.

  11. The Economics of AI Training Data: A Research Agenda

    cs.CY 2025-10 conditional novelty 5.0 of 10

    The paper organizes AI training data into five exchangeable units, documents 24 licensing deals, and argues data should be a separate factor in production functions.

  12. The Economics of AI Training Data: A Research Agenda

    cs.CY 2025-10 unverdicted novelty 4.0 of 10

    The paper synthesizes fragmented research to frame data economics around data's nonrivalry and context dependence, catalogs 2020-2025 AI training data deals, and proposes a hierarchy of data units while listing four f...

  13. Private, Verifiable, and Auditable AI Systems

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.

  14. A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A three-part taxonomy for prompt-based natural language explanations, covering context, generation and presentation, and evaluation with 15 desirable properties.

  15. Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

    cs.CR 2024-09 unverdicted novelty 2.0 of 10

    Survey of harmful fine-tuning attacks on LLMs, their variants, defense strategies, mechanical analysis, and evaluation methodologies.

Pith tools