Pith. sign in

Paper Citation Record · LEDGER

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

As of 22 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 1 inbound Pith citation observation for arXiv:2605.25850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25850 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T21:17:53.242182Z

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:53:26.206601Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5edcd46b-0dd3-4228-b3f0-3868542a2c12 · outbound

This paper cites MASH: Modeling Abstention via Selective Help-Seeking.

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning MASH: Modeling Abstention via Selective Help-Seeking

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T21:23:59.202807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T21:17:53.242182Z digest=sha256:381fed5ac23ffbad9be9be8446a012da58aa0c83ff7a08499ff58ab97a14a81c

Observation 12066919-c45a-4d6f-bd89-9ae10a0009c6 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T21:23:59.199943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T21:17:53.242182Z digest=sha256:6d8a34b2f58b4a47d96cd5292548212da67a6082acec87c8e3eea37b84dcd97f

Pith citing papers

Observation 1d0ec2be-5923-4e84-8d32-c4160d977077 · inbound

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning cites this paper.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.206601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.206601Z digest=sha256:ab59f4e6b0765ed70002bef202849a84e6209bee47e87285c61cd6cca56f9d8c