{"paper":{"title":"Relational Preference Encoding in Looped Transformer Internal States","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Looped transformers encode human preferences relationally through comparisons in their internal states.","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Jan Kirin","submitted_at":"2026-04-10T20:00:49Z","abstract_excerpt":"We investigate how looped transformers encode human preference, training lightweight evaluator heads on frozen Ouro-2.6B loop-iteration states on Anthropic HH-RLHF.\n  v2: an erratum is prepended; the original manuscript is unchanged. A post-publication audit found the three headline results inflated by two independent evaluation errors. The 95.2% pairwise evaluator accuracy is a canonical-ordering artifact: the data were correctly split, but the evaluator learned to prefer the first-presented argument; its strict antisymmetrized accuracy on the full 8,552-pair test set is 63.9%. The 84.5% pair"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"loop states encode preference predominantly relationally: a linear probe on pairwise differences achieves 84.5%, the best nonlinear independent evaluator reaches only 65% test accuracy, and linear independent classification scores 21.75%, below chance and with inverted polarity.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the high pairwise accuracy and low independent accuracy reflect genuine relational encoding in the model's learned value system rather than an artifact of the 50% argument-swap protocol, the accidental early stopping, or the specific choice of lightweight heads trained on the same model's representations.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Looped transformer hidden states encode preferences relationally via pairwise differences rather than independent pointwise classification, with the evaluator acting as an internal consistency probe on the model's own value system.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Looped transformers encode human preferences relationally through comparisons in their internal states.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"5d2973ffe2c842584fa0e45d21b06c64425c9a0eae02bef254fc29f714ecd0ef"},"source":{"id":"2604.09870","kind":"arxiv","version":2},"verdict":{"id":"4bf0a652-a8db-4c35-9aa3-faebbfe6d3f5","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-10T16:53:36.083854Z","strongest_claim":"loop states encode preference predominantly relationally: a linear probe on pairwise differences achieves 84.5%, the best nonlinear independent evaluator reaches only 65% test accuracy, and linear independent classification scores 21.75%, below chance and with inverted polarity.","one_line_summary":"Looped transformer hidden states encode preferences relationally via pairwise differences rather than independent pointwise classification, with the evaluator acting as an internal consistency probe on the model's own value system.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the high pairwise accuracy and low independent accuracy reflect genuine relational encoding in the model's learned value system rather than an artifact of the 50% argument-swap protocol, the accidental early stopping, or the specific choice of lightweight heads trained on the same model's representations.","pith_extraction_headline":"Looped transformers encode human preferences relationally through comparisons in their internal states."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.09870/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":2,"snapshot_sha256":"079fbfd2c937727fbd75265f07da373d1df5b5a423a446a52f1a98c7d2ddddac"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}