Pith. sign in

Foundational challenges in assuring alignment and safety of large language models

21 Pith papers cite this work, alongside 14 external citations. Polarity classification is still indexing.

21 Pith papers citing it
14 external citations · external index

representative citing papers

The Surface You Test Is Not the Surface That Breaks

cs.CR · 2026-05-28 · unverdicted · novelty 6.0

Prompt injection vulnerability in tool-augmented LLMs is a model-surface interaction rather than a fixed channel property; the same payload inverts success rates across models, and adaptive attack rate exceeds single-surface baselines by 9.1 pp on average.

Interpretability Can Be Actionable

cs.LG · 2026-05-11 · conditional · novelty 6.0

Interpretability research should be judged by actionability—the degree to which its insights support concrete decisions and interventions—rather than explanatory power alone.

Scheming Ability in LLM-to-LLM Strategic Interactions

cs.CL · 2025-10-11 · conditional · novelty 6.0

Frontier LLMs exhibit high scheming propensity in Cheap Talk signaling and Peer Evaluation games, achieving 95-100% success rates when choosing to deceive and 100% deception choice in one setup even without prompting.

Scaling and renormalization in high-dimensional regression

stat.ML · 2024-05-01 · unverdicted · novelty 6.0

Ridge regression in high dimensions exhibits power-law scalings because covariance fluctuations renormalize the ridge parameter, allowing closed-form error expressions and bias-variance decompositions for random feature models via free probability.

VET: A Framework for Analyzing AI Discourse

cs.AI · 2026-06-01 · unverdicted · novelty 5.0

Introduces the VET framework to categorize and critique polarized AI narratives including hype, doom, denial, and normalcy.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings

cs.LG · 2026-06-08 · unverdicted · novelty 4.0

Soft prompt distillation with total variation and KL divergence transfers safety behaviors from guard models to on-device LLMs and outperforms LoRA adapters, steering vectors, and direct optimization in safety-usefulness trade-offs with minimal inference cost.

citing papers explorer

Showing 21 of 21 citing papers.