Pith. sign in

Paper Citation Record · LEDGER

Chain of Hindsight Aligns Language Models with Feedback

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2302.02676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.02676 v8

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:49.235474Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

27
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b45d4549-9528-4c2c-affd-c1858f77b22e · inbound

Aligning Text-to-Image Models using Human Feedback cites this paper.

Aligning Text-to-Image Models using Human Feedback Chain of Hindsight Aligns Language Models with Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:15.486004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:39:15.388881Z digest=sha256:57724bec574676afe5a805314c6fb0def4a580d2154102d05b80d2030dcdd513

Observation df7e02a7-8c44-4344-a897-c31194eb80f5 · inbound

Teaching Large Language Models to Self-Debug cites this paper.

Teaching Large Language Models to Self-Debug Chain of Hindsight Aligns Language Models with Feedback

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:24:24.960444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T06:24:24.607354Z digest=sha256:8e583ae1d64ce636b1b88fb141c800a680a5c6c0e3fab5fce113745b7ea087e4

Observation 21d65289-9823-4462-98ee-d142f83cceea · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Chain of Hindsight Aligns Language Models with Feedback

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.552064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:120fa5dcb67f9d80a43a727aea91722153d45720720eb32cfddedf0ca748470d

Observation 888cf924-ce4e-4e37-a18d-8028d9859137 · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey Chain of Hindsight Aligns Language Models with Feedback

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:47.954838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:cea56bc229a43952f5bf761e1dc64d7d23fe3f40521e9a7a1535dfa04c355dbe

Observation 78fe38be-df0c-4001-aa54-18a1e7f11068 · inbound

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation cites this paper.

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation Chain of Hindsight Aligns Language Models with Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:00:51.564458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:00:51.487120Z digest=sha256:fe2ba0f66a57fa19a6a2c964dc50506d5e4c10014335b719c9e35292be1555cc

Observation 51f31bb9-f395-427e-8558-ebef725cee31 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Chain of Hindsight Aligns Language Models with Feedback

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.494101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:55fee7eec106dc427a146877a7e978cd847e6843454f729cd37925dac78d08cd

Observation f059d631-c20d-490d-bdab-2df8213cf7b3 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Chain of Hindsight Aligns Language Models with Feedback

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:34:42.717565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:d785a5ce926abebc84ca0d488486449f908ba370e62d40685134d4032bf34ad8

Observation b82aee17-ec1b-44d3-96ce-91fa6d62a8f4 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Chain of Hindsight Aligns Language Models with Feedback

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.442731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:7f906e2e621f621e68bc96ff984733d712cc483d8d18cd58c834b63455b45239

Observation b34c4e40-7eed-4521-8663-53c83a10bd00 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Chain of Hindsight Aligns Language Models with Feedback

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.001994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:805069302f4c873c0677976f12c11c0268ded3e694817d2ec7310bf61ae8e6de

Observation 0ff99883-5672-4257-aa20-22714700c191 · inbound

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations cites this paper.

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations Chain of Hindsight Aligns Language Models with Feedback

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:47:40.422808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T04:47:40.335477Z digest=sha256:0f38af1060a3648e275ff39e2521db54751c9c7b7e2320eea71e17a5b83a21f6

Observation baa0cbb8-ab8a-4928-9b07-4f8f16043630 · inbound

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions cites this paper.

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions Chain of Hindsight Aligns Language Models with Feedback

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:02:35.944176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:59:36.994758Z digest=sha256:33251377c715dfc2a0daacc409bc3a8df38b96df7a1ad482d906ee6c59efb0ef

Observation e7c2b313-9bb8-48d7-872c-9e44e93999a4 · inbound

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time cites this paper.

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time Chain of Hindsight Aligns Language Models with Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:20.272639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:20.272639Z digest=sha256:5e62735de59aa9112dea900a9a59dc344eb0dd0fd452cd568472b16d8ac6b791

Observation 1fc1ce79-7efe-4436-8321-0ee83f34fb0f · inbound

Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease cites this paper.

Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease Chain of Hindsight Aligns Language Models with Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:49.235474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:49.235474Z digest=sha256:d7a1d7eb3cb8cebb2f7dd5a32ca7fb01dd64546f3a03fa1efcb568c4a59cd655

Observation 3b9a7e87-1844-403c-9371-9a48e56b292e · inbound

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives cites this paper.

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives Chain of Hindsight Aligns Language Models with Feedback

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.446080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.446080Z digest=sha256:fcd4b3e7f5bc13f97872be4c8ab0ae4b769442b147bac46f7863e07494ad1115

Observation 63beea3b-8356-4f0c-96e8-868abdd8ef81 · inbound

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment cites this paper.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Chain of Hindsight Aligns Language Models with Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.730028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.730028Z digest=sha256:6cb8ab67ece664e7464678b1f60e4f74da9120ed4a17cd7c97d89fc70bfcc1d1

Observation c9492d50-6ce5-4b04-bdcf-c9d856e340f1 · inbound

Invariant-based Robust Weights Watermark for Large Language Models cites this paper.

Invariant-based Robust Weights Watermark for Large Language Models Chain of Hindsight Aligns Language Models with Feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:07.896741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:33:07.896741Z digest=sha256:76e214ceddf82b0973c6854190fc903b4e7a40076c9914e5bab9d575cacd57b1

Observation 84a467ab-def9-44bc-904a-b457e82c56f9 · inbound

Feedback-Driven Execution for LLM-Based Binary Analysis cites this paper.

Feedback-Driven Execution for LLM-Based Binary Analysis Chain of Hindsight Aligns Language Models with Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:37.914449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T10:40:32.133423Z digest=sha256:d202fe4ff6e6c9a15b3e76641817629b59016a19d6558bcfdb484f228582d086

Observation 9dd5bbb0-c187-4884-a236-952f529115b9 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Chain of Hindsight Aligns Language Models with Feedback

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:39:00.296891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:d7e07a83e14cac22ac713fd0d0eb057d9651993552b9057abdf03b0f89b14f9c

Observation 4c312351-8c69-4ad3-a390-d8dfa220f5ac · inbound

Reinforcing Human Behavior Simulation via Verbal Feedback cites this paper.

Reinforcing Human Behavior Simulation via Verbal Feedback Chain of Hindsight Aligns Language Models with Feedback

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:24:02.456916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T07:21:48.649289Z digest=sha256:2368c40d519453d3788a3624cc34160e8cb152d160add7ef7e02c9ce7fe2700e

Observation 007cf121-c003-41fe-861a-e03429beac57 · inbound

AI as a Tool for Simulation-Based Experiments in Literary Studies cites this paper.

AI as a Tool for Simulation-Based Experiments in Literary Studies Chain of Hindsight Aligns Language Models with Feedback

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:02:19.091390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T14:54:14.909995Z digest=sha256:f50491e0c51c7a48bf1a44d2955647dad37564f7a8edd86fed046202469386a2

Observation 6c2bb689-008d-4bad-8790-6691437e2c51 · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions Chain of Hindsight Aligns Language Models with Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:20.842111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:20.842111Z digest=sha256:723e4a587c13fc98303de61ae0599a2979c3cec88f3528ecab40d3959fbda8a3

Observation 85eac940-cd36-4287-b81c-573dad37b8fa · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Chain of Hindsight Aligns Language Models with Feedback

Reference 102

Resolution
unresolved
no resolver link, observed 2026-07-31T23:51:58.137701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:51:58.137701Z digest=sha256:de45ff6be0fc8788309ba73d4cffd5f59e3ac1078b3887891e387874bbdaaa44