Pith. sign in

Paper Citation Record · LEDGER

Uncertainty-aware Reward Design Process

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.02256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02256 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:39:15.423614Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3993cf90-b9c6-4592-b88f-ee2c2f36d6c1 · outbound

This paper cites grasping.

Uncertainty-aware Reward Design Process grasping

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.020372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:15.423614Z digest=sha256:b1caf114a3ba871d649f1e030dbf5fc1378a24075258d1dcc7d75e3450d463f1

Observation 701ca3bc-1972-4aff-9005-26a8fc6946d7 · outbound

This paper cites (9) By performing coordinate scaling transformation on the sample points, we obtain new sample points p(i) = (˜p(i) 1 /l1,··· , ˜p(i) d /ld), i= 1, 2,··· ,n.

Uncertainty-aware Reward Design Process (9) By performing coordinate scaling transformation on the sample points, we obtain new sample points p(i) = (˜p(i) 1 /l1,··· , ˜p(i) d /ld), i= 1, 2,··· ,n

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.200589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:15.295720Z digest=sha256:6aeb98ca1a140e33125ba51e767e67144b541a22a778586016091660c931d517

Observation 8a56d682-ab05-4f3e-884a-5c185937ea75 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Uncertainty-aware Reward Design Process From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.937399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.937399Z digest=sha256:d04e8547ab2e017fbf987f55ff7baa7c83234ca45ee6bd33519d54d212975aca

Observation 74d66903-5814-4b75-bce6-d39ba09089d6 · outbound

This paper cites DeepSeek-V3 Technical Report.

Uncertainty-aware Reward Design Process DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.052660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.052660Z digest=sha256:26ca03c6c4ea162f2cf7c3a12b162e33b5278ec505481c694620a56fe5a5192d

Observation 0365d674-560f-4a63-978f-d41d6024c0c5 · outbound

This paper cites GPT-4 Technical Report.

Uncertainty-aware Reward Design Process GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.232861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.232861Z digest=sha256:6f2d304e9a581ebcadc145c770216896ca1a40571e60181d2560c2080b9b3e87

Observation da250cde-64fe-400e-8ace-fa4420b0013b · outbound

This paper cites Qwen2.5 Technical Report.

Uncertainty-aware Reward Design Process Qwen2.5 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.325053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.325053Z digest=sha256:658fefb28b7e0b16fbd66b8cb11821e341df63ebcbb8cc20de2d6070f2545fdd

Observation 3a3d6b55-02b2-4f9b-b2ea-80a85786bd9f · outbound

This paper cites A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions.

Uncertainty-aware Reward Design Process A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.548517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.548517Z digest=sha256:e646c2ff49ca8f5c98fecdb5bcf477acabd3c0d8c16f0a0adb590944ca856b5e

Observation 20e09c50-e40b-4152-8b15-49e021975eb7 · outbound

This paper cites Api is enough: Conformal prediction for large language models without logit-access.

Uncertainty-aware Reward Design Process Api is enough: Conformal prediction for large language models without logit-access

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.664937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:14.711229Z digest=sha256:7a391a6cf9ede3fdc62eabcff4621b962a68ed9dd415ab3cbd4a65e2664202e8

Observation 9af2285d-9255-47db-b961-bd4976b7c8c9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Uncertainty-aware Reward Design Process Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.805208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.805208Z digest=sha256:f6a7d70e7cf193f4f5d435a25dcb17bfb70721a55ee25ac299a57a02a61d1440

Observation bb6c000c-b41b-4496-bf78-f820ad8f56e3 · outbound

This paper cites Do phd-level llms truly grasp elementary addition? probing rule learning vs.

Uncertainty-aware Reward Design Process Do phd-level llms truly grasp elementary addition? probing rule learning vs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.898533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.898533Z digest=sha256:00a579fa740ce0c4eb9007a79cb01294c55c74409de10434139dacf92c701e3b

Observation 52545019-4836-4976-b621-651dc3176642 · outbound

This paper cites Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation.

Uncertainty-aware Reward Design Process Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:15.005796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:15.005796Z digest=sha256:dfe5181afe98392648e908a0ab743dfcc0f2ac7a0f0e861ad987fdbb4d801378

Observation 15327041-fc96-46a8-b3bf-afecdf26e69a · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Uncertainty-aware Reward Design Process A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:15.150337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:15.150337Z digest=sha256:28109f9984b5a451e7ab730183049fea7016ec519a84cd3fb4449e690ee4dd61

Observation 3da67d1c-7edf-438c-b3e6-07fe03250df9 · outbound

This paper cites The robot must push a movable chair from its initial location to a designated target region.

Uncertainty-aware Reward Design Process The robot must push a movable chair from its initial location to a designated target region

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.458326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:15.250787Z digest=sha256:28a6ba08d598bc12c932d9aec3b5ded54c639611981cee601ddb30d2d9d0216c

Observation 85bba71d-bfa2-4543-b947-d32463ec5e82 · outbound

This paper cites Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics.

Uncertainty-aware Reward Design Process Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.628984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.628984Z digest=sha256:7418b96ff2e36148631305b69b0869af02a5dbf8e44a1dd19d91e472b2bafd69

Observation f8f18e13-f230-4c90-a7b1-85703a47297b · outbound

This paper cites A survey of confidence estimation and calibration in large language models.

Uncertainty-aware Reward Design Process A survey of confidence estimation and calibration in large language models

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:17.193725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:13.659017Z digest=sha256:c38736e7e86786e420f5f52db88b0252674c1073bd9f45a31503ac86c2d6134a

Observation 353afe68-687c-4b83-b7af-acde704e0299 · outbound

This paper cites Reasoning with language model is planning with world model.

Uncertainty-aware Reward Design Process Reasoning with language model is planning with world model

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.938341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:13.726804Z digest=sha256:0af749107d32b135e524b11bb087fdbe092c8321acb2b4f37458dfb387d4dcba

Observation 38ee5e0b-970b-44c7-9c04-a379d987b780 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Uncertainty-aware Reward Design Process Proximal Policy Optimization Algorithms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.376825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.376825Z digest=sha256:cbf1772229eaa691dd9a6b173df755ebfd6d4695d9393f4ca78249bf5444bc48

Observation 11d557d8-f6ab-4da4-96a2-e659eeccefa3 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Uncertainty-aware Reward Design Process V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.455724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.455724Z digest=sha256:971a47058d9a6460176ba22da2c733c38b7b8376be4752fcc5661dfa529ebeed

Observation e1752300-b526-4ae2-bda8-21ce563c48ec · outbound

This paper cites Empowering LLMs with Logical Reasoning: A Comprehensive Survey.

Uncertainty-aware Reward Design Process Empowering LLMs with Logical Reasoning: A Comprehensive Survey

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.583068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.583068Z digest=sha256:d8f163fa8fa4ce8bc8a441da76fe702ff373339c708d1926fa95008e1b934e94

Observation c063a9bb-689b-4bee-af8f-8e9c0d992ba0 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Uncertainty-aware Reward Design Process Language Models (Mostly) Know What They Know

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.821156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.821156Z digest=sha256:d65a10c7d90f1f2bf22001bd94edd0dc36b509e7fc032a5ba8eb3d765af0ba0f

Observation 76f56745-1708-4541-bf02-82be042d9954 · outbound

This paper cites Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey.

Uncertainty-aware Reward Design Process Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.125605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.125605Z digest=sha256:e39301ad7a984c37806be99cb0a94c33592406759e4555eeca14d33640523316

Observation d9bca7eb-d4c8-4946-91a0-eb3102484ed8 · outbound

This paper cites an unresolved cited work.

Uncertainty-aware Reward Design Process Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:39:17.365294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:13.515501Z digest=sha256:7e5509cd400cbdf30f6ae2f44acaa2609c4b91bda511af425367f91828ef90a8

Pith citing papers

No inbound Pith citation observations are available.