Pith. sign in

hub Canonical reference

International AI Safety Report

Canonical reference. 83% of citing Pith papers cite this work as background.

16 Pith papers citing it
Background 83% of classified citations
abstract

The first International AI Safety Report comprehensively synthesizes the current evidence on the capabilities, risks, and safety of advanced AI systems. The report was mandated by the nations attending the AI Safety Summit in Bletchley, UK. Thirty nations, the UN, the OECD, and the EU each nominated a representative to the report's Expert Advisory Panel. A total of 100 AI experts contributed, representing diverse perspectives and disciplines. Led by the report's Chair, these independent experts collectively had full discretion over the report's content.

hub tools

citation-role summary

background 6

citation-polarity summary

years

2026 13 2025 3

roles

background 6

polarities

background 5 support 1

representative citing papers

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

cs.CL · 2026-05-31 · unverdicted · novelty 7.0

TukaBench extends JailbreakBench to African languages via human translation, cultural adaptation, curated prompts, and code-switching, finding lower refusal rates for culturally grounded prompts and surfacing comprehension and judging limitations.

Affective AI Safety: The Missing Piece in LLM Safety

cs.CY · 2026-06-22 · unverdicted · novelty 6.0

Proposes affective safety as a distinct class of AI harms with a taxonomy of self-alienation, bias, and relational harms, arguing that existing safety frameworks address it narrowly or not at all and calling for dedicated approaches focused on cumulative and identity-level effects.

Understanding Annotator Safety Policy with Interpretability

cs.AI · 2026-05-06 · unverdicted · novelty 6.0

Annotator Policy Models learn safety policies from labeling behavior alone, accurately predicting responses and revealing sources of disagreement like policy ambiguity and value pluralism.

Evaluating AI Providers' Frontier Safety Frameworks

cs.CY · 2025-12-01 · unverdicted · novelty 6.0

Twelve frontier AI safety frameworks score between 8% and 34% on adapted risk-management criteria, with a median of 18%, leaving them too vague to serve as reliable external accountability mechanisms.

On the Principles of Deep Feedforward ReLU Networks

cs.LG · 2026-07-08 · conditional · novelty 5.0

Deep feedforward ReLU networks generalize two-layer principles via paths, piecewise linear manifolds, and continuity restriction to explain training solutions.

Do LLMs have core beliefs?

cs.LG · 2026-05-05 · unverdicted · novelty 5.0

LLMs generally fail to maintain stable worldviews under adversarial conversational pressure, indicating they lack core beliefs akin to those in human cognition.

Europe and the Geopolitics of AGI: The Need for a Preparedness Plan

cs.CY · 2026-05-13 · unverdicted · novelty 3.0

AGI may arrive by 2030-2040 and reshape global power balances, requiring Europe to close gaps in compute, talent retention, industrial adoption, and unified policy responses through a coordinated preparedness agenda.

citing papers explorer

Showing 16 of 16 citing papers.