Pith. sign in

hub

BBQ: A hand-built bias benchmark for question answering

19 Pith papers cite this work, alongside 140 external citations. Polarity classification is still indexing.

19 Pith papers citing it
140 external citations · OpenAlex

hub tools

citation-role summary

background 2 dataset 1

citation-polarity summary

representative citing papers

GAIA: a benchmark for General AI Assistants

cs.CL · 2023-11-21 · unverdicted · novelty 7.0

GAIA benchmark shows humans at 92% accuracy on simple real-world questions far outperform current AI systems at 15%, proposing this gap as a key milestone for general AI.

Quality Is Not a Safety Proxy Under Quantization

cs.LG · 2026-06-08 · conditional · novelty 6.0

Across 51 quantized checkpoints, quality metrics fail to predict safety drops in 36 pairings and 10 hidden-danger cases, while a new RTSI screen routes all 10 dangerous rows to testing at matched bucket size.

When AI Says It Feels

cs.AI · 2026-06-04 · unverdicted · novelty 5.0

LLMs trained via rubric-based self-rewarding RL with GRPO enhanced feeling expression and sycophancy robustness but degraded truthful QA performance.

UnBias-Plus: Detect, Explain, and Rewrite Bias

cs.CL · 2026-06-22 · unverdicted · novelty 4.0

UnBias-Plus is an open-source toolkit unifying segment-level multi-class bias classification, biased span localization, neutral text rewriting, and decision reasoning.

citing papers explorer

Showing 19 of 19 citing papers.