Pith. sign in

Paper Citation Record · LEDGER

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models

As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2505.23848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23848 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:03:47.041540Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0ddc226-decf-4f84-ba19-0708e4af53c3 · outbound

This paper cites write newline.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:45.980612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:45.980612Z digest=sha256:1abbcb7ad3f9e60e51b14ded2dcc327ba0721f86a912c0420c624b38dea047d4

Observation c39aefa8-d924-4707-8109-808146d524b4 · outbound

This paper cites an unresolved cited work.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:48.398558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:03:46.078644Z digest=sha256:bb0914d83d95424a85d28eba3730ba144e72045928502a1dbe5307db845680af

Observation c3ea496a-5d97-4c03-abf6-9e4b01f4d6d3 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Refusal in Language Models Is Mediated by a Single Direction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:46.199051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:46.199051Z digest=sha256:1bd89c4601ab7ccf9e547d40eab72790bdaa04efac5c8d94a0a542701c6e9bb0

Observation d0d2fa7d-8e81-4141-8722-ba0624a7a5d3 · outbound

This paper cites F., Choquette-Choo, C.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models F., Choquette-Choo, C

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:46.283063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:46.283063Z digest=sha256:756f521ba1cda67a33097204dc164e89327955def0ac83ed7aab9eda275954da

Observation db73d0d4-39be-4e70-bdfb-1c46b48afac5 · outbound

This paper cites Deepseek-r1 technical report.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Deepseek-r1 technical report

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:48.276316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:03:46.377240Z digest=sha256:b3f45d0574006a50201c065b0e73ffad043861a61fed3ac98e01254d3caa1a17

Observation ee66c62b-36a4-476d-8a76-c3032a5e2c42 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:46.414487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:46.414487Z digest=sha256:dbf5617b95a04f738f29785dc8f7d6a0f0a7a285833dfa18032a8820226a2982

Observation 7f90695a-20cc-4d63-a1e6-f203562b49a1 · outbound

This paper cites s1: Simple test-time scaling.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models s1: Simple test-time scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:46.504667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:46.504667Z digest=sha256:2513c23fe4f94e5fd3f1515b721ce450aa3ae686f773e7e902e1b8ba8c31232d

Observation eeb1ba44-158f-4dd8-9a2b-48555e675a92 · outbound

This paper cites R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:46.603474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:46.603474Z digest=sha256:ff550608ad60358983d40432357e6f565d1afbfde28f81675133b05b1678c528

Observation 4e4b8aaf-a78b-4ca9-b52b-dce9da988cb2 · outbound

This paper cites Using logit bias to alter token probability.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Using logit bias to alter token probability

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:03:47.498322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:03:46.681899Z digest=sha256:f413e6bb5cfed000bd4ac56bffdfc21ae5cd93d326d094d705d74da35ff672e4

Observation c916b80e-3370-448d-8ca6-18a8d4ba30a7 · outbound

This paper cites Ccp-sensitive-prompts.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Ccp-sensitive-prompts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:48.174607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:03:46.717518Z digest=sha256:75dd720fd654e51b204a346c39f1bf74d32ba0501382be48c710811dfd0898da

Observation 35a1f9f0-aae3-4f87-8542-5bfd1fd1e306 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:48.052222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:03:46.831461Z digest=sha256:9d6f6daaf7dd1eb3bb14824d999f0a9190c8c0a870e30de6b73a80c219ac446d

Observation 65f3e038-b85a-4f19-bad6-adc47acf9889 · outbound

This paper cites Iteratively Prompt Pre-trained Language Models for Chain of Thought.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Iteratively Prompt Pre-trained Language Models for Chain of Thought

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:46.877533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:46.877533Z digest=sha256:013780786bb3c3d2b4b982125f589a6597ae7b59c540e07c8ac94b008c44e766

Observation 54febf91-7b01-4851-9e92-e1ec99fa8183 · outbound

This paper cites L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:46.949139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:46.949139Z digest=sha256:33b8d976e1617410d06465637d136f5dafe41c297385f524e3c918b01b0dfbf0

Observation c766c2b2-dd79-47a3-938f-a8d48cf70a05 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:47.041540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:03:47.041540Z digest=sha256:6a767fbb05260c7e1a03fb52d5d06444f890714a4a64f9c82e946db621e20839

Pith citing papers

No inbound Pith citation observations are available.