Pith. sign in

Paper Citation Record · LEDGER

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

As of 6 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2604.09855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.09855 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:11:15.484711Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T22:26:44.052574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:35:40.643385Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 391fbdef-0764-4a7e-bb02-2fb2a960e055 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Kimi K2: Open Agentic Intelligence

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:26:01.930259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:4085e9317521112be99563e805bc418dfb56f707057dfa8ab3720d3d15e35507

Observation 655c7202-9547-4e29-8c93-d009b353b531 · outbound

This paper cites codename_1.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards codename_1

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.652154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:ee3f9987482595dde5c8f12ed960ad16a038196dbf74ec79fc809c34ebc49435

Observation 8a53b19d-28cc-4b90-8155-9c68803a23b2 · outbound

This paper cites an unresolved cited work.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-17T11:41:46.655409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:04fc226336cb3e468f4b98a5e6fb964c15e98f5a26cf2f027d8d82eb44338d7f

Observation 2dcb90a2-f3c4-4a34-9862-d325c731cc66 · outbound

This paper cites $M (N codename_1) is a exact copy of seller’s previous offer.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards $M (N codename_1) is a exact copy of seller’s previous offer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.644828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:4f2723d5749457745e2a8a8a63e27eebbae39fa560fb3d277fed9f4baaa00794

Observation de725330-876c-47d2-b242-9695dfdb5ebb · outbound

This paper cites Happy By Clinique For Men. Cologne Spray 1.7 Oz.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Happy By Clinique For Men. Cologne Spray 1.7 Oz

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.648565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:64503c8bb5f1550308806403fea666c55161f8a7cbb30461d6c9c0a7cff6a672

Observation a34d4894-f55c-4e2e-bd4d-257c3eab1c59 · outbound

This paper cites codename_1.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards codename_1

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.658549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:4e64dce63d7efbb47e0c92632e310c3673592a394fe728546110dd30e04961c1

Observation 76033cc6-4f99-4fa7-8b9f-5cb5dc201beb · outbound

This paper cites an unresolved cited work.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T11:41:46.638146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:3a228215fa074f9175ccc708e95791c7556d437031eed2e5af37f3629c7cb068

Observation 3e1cd05e-56a4-4491-b652-4ee70e1371b5 · outbound

This paper cites codename_1.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards codename_1

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.634943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:5d9e1d12a66762facd26df9ef3befd4d8b875c2b56d563b68937255a2aecfc61

Observation 3edba574-afe3-4e30-96cb-1f1b72198cf9 · outbound

This paper cites Happy By Clinique For Men. Cologne Spray 1.7 Oz.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Happy By Clinique For Men. Cologne Spray 1.7 Oz

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T11:41:46.641568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:ae4615cf26ee2fff04e240b972d2dbdef1b454fc2d77aadacccf162fef040088

Pith citing papers

Observation d9257adb-84d8-4271-8db2-894109063678 · inbound

Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations cites this paper.

Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.645323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T22:26:44.052574Z digest=sha256:e09b474233fbb40693051f4382e4d630b6fb7a23efd6eac0e29f1b046c2c084b