Pith. sign in

Paper Citation Record · LEDGER

Parameter Efficient Reinforcement Learning from Human Feedback

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2403.10704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.10704 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:18.772325Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T15:35:47.630181Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a58d3885-7902-467a-b295-3112b6003fad · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter Efficient Reinforcement Learning from Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.772325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.772325Z digest=sha256:567accfb3cf35c2416a8a5cad4cf1469e76fa6c00fa23fe2c2bfca3653245c2d

Observation 4ac7c2ab-2fef-4fce-a820-1cde53e3f2d8 · inbound

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning cites this paper.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning Parameter Efficient Reinforcement Learning from Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:07.820412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:07.820412Z digest=sha256:3ee5f92f120d36668b674d831d84e670bc047a4b74f699c82240d719bab339b3

Observation a898b301-55d1-458a-b5c6-098fa3d90597 · inbound

Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation cites this paper.

Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation Parameter Efficient Reinforcement Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T13:31:03.312654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:31:03.312654Z digest=sha256:9837c00872e9d46d6332c1503e6d2ef5e6c12d5cc125a647be4397b2b7ba1e8b

Observation 22b583dd-26fb-47d6-86ec-50dc949c8966 · inbound

Conjunctive Prompt Attacks in Multi-Agent LLM Systems cites this paper.

Conjunctive Prompt Attacks in Multi-Agent LLM Systems Parameter Efficient Reinforcement Learning from Human Feedback

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.694361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T08:13:42.401992Z digest=sha256:a2634da899d19c4c19158f54fbaaae1cf29b52528a1c600724782fdc3044c811

Observation 109df44c-3aec-4c1d-a93d-8b1d2f42adfb · inbound

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs cites this paper.

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs Parameter Efficient Reinforcement Learning from Human Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:42.559347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:09:07.557773Z digest=sha256:7738f400409c1085efee7f810fdf8f017252f6d0ac53c9ce6f0f8361d17e50a0

Observation 8b1eaff2-a7e8-49b4-9d02-ffa6cfc163cf · inbound

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs cites this paper.

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs Parameter Efficient Reinforcement Learning from Human Feedback

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:47.631625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T01:22:16.176398Z digest=sha256:bf60ad3502e54d7914ed94a1416a41d732d4069b3cbaef6b14b6d58dbaba6eeb