Pith. sign in

Paper Citation Record · LEDGER

Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.16332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.16332 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:12:32.706277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:05:47.060027Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32edb4db-a48c-4551-85f2-5daa3d6165e0 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.155376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:e04bcdf1fbd229011665996b8f52fc989cc9517fb8aedd0ab6e76cfbb927f280

Observation 197cdb0f-11bf-4a8f-b2f3-cdaac7a31d69 · inbound

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs cites this paper.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.706277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.706277Z digest=sha256:4fd38b7d66688b397722b85ca702c837363ad1008356c7e89b3bace9b2095b07

Observation 268c7f87-d909-4ee9-a98d-bf3522936161 · inbound

Selective Safety Steering via Value-Filtered Decoding cites this paper.

Selective Safety Steering via Value-Filtered Decoding Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:03.865596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:03:28.381687Z digest=sha256:1205f2768ca63d9d98d16436d827f620b2639369731acafb712c663812cf5f76

Observation c7f654c8-b951-4a08-a92e-ce281fd065e8 · inbound

Selective Safety Steering via Value-Filtered Decoding cites this paper.

Selective Safety Steering via Value-Filtered Decoding Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T18:59:01.049397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:59:01.049397Z digest=sha256:d7f786fd9e7ebb10deccf7ef193617485d99324b37e59914cf86ba656a78d21b

Observation 1268ed8c-6a33-410f-95e8-367ff9a86958 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.061610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:e890a4d2cfeb686db1ba97820e1f6b0b57e1b5020a0fbc9b731e29db8e6d242f

Observation 37a7ca84-8678-42a5-bbd9-b8ee8a9f137d · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

Reference 115

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:7a5544f670e1adb8d0e79130c5e3e61c04023f65a146be29aaae9ae213d2b0fc