Pith. sign in

hub Canonical reference

Decodingtrust: A comprehensive assessment of trustworthiness in gpt models

Canonical reference. 71% of citing Pith papers cite this work as background.

15 Pith papers citing it
Background 71% of classified citations

hub tools

citation-role summary

background 7

citation-polarity summary

roles

background 7

representative citing papers

BEAVER: An Efficient Deterministic LLM Verifier

cs.AI · 2025-12-05 · unverdicted · novelty 7.0

BEAVER is the first practical deterministic verifier that maintains sound probability bounds on LLM safety properties using token tries and frontier data structures, finding 2-3x more violations than sampling at 1/10 the compute.

Reducing Political Manipulation with Consistency Training

cs.CL · 2026-05-21 · unverdicted · novelty 5.0 · 2 refs

PCT is a reinforcement learning approach that trains LLMs for symmetric sentiment and helpfulness across paired opposing political prompts, reducing covert bias while preserving general performance.

TrustLLM: Trustworthiness in Large Language Models

cs.CL · 2024-01-10 · unverdicted · novelty 5.0

TrustLLM defines eight trustworthiness principles, creates a six-dimension benchmark, and evaluates 16 LLMs showing proprietary models generally lead but some open-source ones are close while over-calibration can hurt utility.

citing papers explorer

Showing 15 of 15 citing papers.