Pith. sign in

Paper Citation Record · LEDGER

Uncertainty-aware Reward Design Process

As of 23 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.02256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02256 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:39:15.423614Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3993cf90-b9c6-4592-b88f-ee2c2f36d6c1 · outbound

This paper cites grasping.

Uncertainty-aware Reward Design Process grasping

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.020372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:39:15.423614Z digest=sha256:4c849661bf73f33f52613dfae7cf4ebdb925d3beeb1c3a8fa3c20e15cd7b1b27

Observation 701ca3bc-1972-4aff-9005-26a8fc6946d7 · outbound

This paper cites (9) By performing coordinate scaling transformation on the sample points, we obtain new sample points p(i) = (˜p(i) 1 /l1,··· , ˜p(i) d /ld), i= 1, 2,··· ,n.

Uncertainty-aware Reward Design Process (9) By performing coordinate scaling transformation on the sample points, we obtain new sample points p(i) = (˜p(i) 1 /l1,··· , ˜p(i) d /ld), i= 1, 2,··· ,n

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.200589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:39:15.295720Z digest=sha256:05623f06bb1f65f0cddc6543d49c2d9940b0d45e4c029a2ad71c8927f9b544af

Observation 8a56d682-ab05-4f3e-884a-5c185937ea75 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Uncertainty-aware Reward Design Process From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.937399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.937399Z digest=sha256:2795d22037a4b5c411dbb7e6bfedb9c5eabf9f792af4a144b127139ee2c4b089

Observation 74d66903-5814-4b75-bce6-d39ba09089d6 · outbound

This paper cites DeepSeek-V3 Technical Report.

Uncertainty-aware Reward Design Process DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.052660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.052660Z digest=sha256:1d36eecf54233281dcd639c7e1bef44305878809795ef1a6f6cea8b78ac64237

Observation 0365d674-560f-4a63-978f-d41d6024c0c5 · outbound

This paper cites GPT-4 Technical Report.

Uncertainty-aware Reward Design Process GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.232861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.232861Z digest=sha256:d6e1bf893388d5d48250bb64079c162730cbfc223c5313194bd4cb26b4ad3aa9

Observation da250cde-64fe-400e-8ace-fa4420b0013b · outbound

This paper cites Qwen2.5 Technical Report.

Uncertainty-aware Reward Design Process Qwen2.5 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.325053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.325053Z digest=sha256:8b6632afcf4a1ae31b598ab1579935f0331897aa90868726547186c93f6b8788

Observation 3a3d6b55-02b2-4f9b-b2ea-80a85786bd9f · outbound

This paper cites A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions.

Uncertainty-aware Reward Design Process A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.548517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.548517Z digest=sha256:17b2d07597d80ca9b6af032dd7eb2cbf225e89c96579830abb5dd84ff21a2783

Observation 20e09c50-e40b-4152-8b15-49e021975eb7 · outbound

This paper cites Api is enough: Conformal prediction for large language models without logit-access.

Uncertainty-aware Reward Design Process Api is enough: Conformal prediction for large language models without logit-access

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.664937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:39:14.711229Z digest=sha256:c55b87f4e808191d74d3619ee2a3ff545781969fc2c968477c13b0774014ede9

Observation 9af2285d-9255-47db-b961-bd4976b7c8c9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Uncertainty-aware Reward Design Process Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.805208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.805208Z digest=sha256:751c9cb56b21e037d9b25fc8aac1f74045159cdeedaffa3b939306eaec4028da

Observation bb6c000c-b41b-4496-bf78-f820ad8f56e3 · outbound

This paper cites Do phd-level llms truly grasp elementary addition? probing rule learning vs.

Uncertainty-aware Reward Design Process Do phd-level llms truly grasp elementary addition? probing rule learning vs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.898533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.898533Z digest=sha256:d8c2b11eaeac7618fb5fa585f2a91f129394947725f93d2ad8f941168c0534b0

Observation 52545019-4836-4976-b621-651dc3176642 · outbound

This paper cites Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation.

Uncertainty-aware Reward Design Process Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:15.005796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:15.005796Z digest=sha256:7f37aeaaa737a514f3efddaf3928e59a9cff4b6a9b35e43d0b67f7ee6142083f

Observation 15327041-fc96-46a8-b3bf-afecdf26e69a · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Uncertainty-aware Reward Design Process A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:15.150337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:15.150337Z digest=sha256:4c94eda6d05f6f3e103eb37e0c8c6df2236e5ec7c20905b179bd03487345ba8b

Observation 3da67d1c-7edf-438c-b3e6-07fe03250df9 · outbound

This paper cites The robot must push a movable chair from its initial location to a designated target region.

Uncertainty-aware Reward Design Process The robot must push a movable chair from its initial location to a designated target region

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.458326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:39:15.250787Z digest=sha256:7adc2b93e3877f760e7acbe06f5fe9c88605a66e10b2c21438db9eca9d9145e1

Observation 85bba71d-bfa2-4543-b947-d32463ec5e82 · outbound

This paper cites Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics.

Uncertainty-aware Reward Design Process Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.628984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.628984Z digest=sha256:5c764e6223c2ffa79e4c9f97c883b13cf001255da1853930f2788b54b70a7c11

Observation f8f18e13-f230-4c90-a7b1-85703a47297b · outbound

This paper cites A survey of confidence estimation and calibration in large language models.

Uncertainty-aware Reward Design Process A survey of confidence estimation and calibration in large language models

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:17.193725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:39:13.659017Z digest=sha256:329cd8a6a6276997f31caa9f5e49bd492966b0a995dc43359e8ed9fccd1e95ba

Observation 353afe68-687c-4b83-b7af-acde704e0299 · outbound

This paper cites Reasoning with language model is planning with world model.

Uncertainty-aware Reward Design Process Reasoning with language model is planning with world model

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.938341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:39:13.726804Z digest=sha256:817b67765404ca9153c621bdd1445e65ad02ade3de7ed9117b8de8bf86ac0f9f

Observation 38ee5e0b-970b-44c7-9c04-a379d987b780 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Uncertainty-aware Reward Design Process Proximal Policy Optimization Algorithms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.376825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.376825Z digest=sha256:8cd1e2adf41242b4afded18c6ad23b648b27f4081b138a9e597e5dafc3d45e94

Observation 11d557d8-f6ab-4da4-96a2-e659eeccefa3 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Uncertainty-aware Reward Design Process V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.455724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.455724Z digest=sha256:4a7fba46cfff3bdc37aeb61a6a1cc7c1eb15f3015670fd778c4e8274a358a284

Observation e1752300-b526-4ae2-bda8-21ce563c48ec · outbound

This paper cites Empowering LLMs with Logical Reasoning: A Comprehensive Survey.

Uncertainty-aware Reward Design Process Empowering LLMs with Logical Reasoning: A Comprehensive Survey

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.583068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.583068Z digest=sha256:7f5cad556da6eef237f786775f37d3ffa69fd303da57f127d3cae1aab42cbb9d

Observation c063a9bb-689b-4bee-af8f-8e9c0d992ba0 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Uncertainty-aware Reward Design Process Language Models (Mostly) Know What They Know

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.821156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.821156Z digest=sha256:62b4f12555f1decb536f1f509c6455a7ce910776300c95b019ba13404c566fb7

Observation 76f56745-1708-4541-bf02-82be042d9954 · outbound

This paper cites Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey.

Uncertainty-aware Reward Design Process Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.125605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.125605Z digest=sha256:e52f13d642e499e377f38fd571408f0819297f42cbd32d4f78a536c82a5d53d1

Observation d9bca7eb-d4c8-4946-91a0-eb3102484ed8 · outbound

This paper cites an unresolved cited work.

Uncertainty-aware Reward Design Process Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:39:17.365294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:39:13.515501Z digest=sha256:f96d124babd0d485e03dea3b91dee26e6444fd3247f0e3a692cc73628aef6c63

Pith citing papers

No inbound Pith citation observations are available.