Pith. sign in

Freelb: Enhanced adversarial training for natural language understanding.arXiv preprint arXiv:1909.11764

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

fields

cs.LG 4 cs.CL 1

representative citing papers

Cheap Reward Hacking Detection

cs.LG · 2026-06-08 · unverdicted · novelty 5.0

Small transformer encoder with linear probe detects reward hacking at AUC 0.9467 and TPR@5%FPR 0.8296, matching LLM-as-judge accuracy at ~10000x lower per-trajectory cost.

citing papers explorer

Showing 5 of 5 citing papers.