Pith. sign in

Paper Citation Record · LEDGER

Towards Cost-Effective Reward Guided Text Generation

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2502.04517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04517 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:31:18.616099Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T07:05:47.150308Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59aafdef-8cf4-4d7d-a14b-6ccacfa308b6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Towards Cost-Effective Reward Guided Text Generation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.450643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.450643Z digest=sha256:2474fa7ecbca0f34cc03b58edb158333b9ef232c1e59fc7980278a75b32561fa

Observation ccf7776c-207c-47e2-a62b-16ee3fa8ef1e · outbound

This paper cites an unresolved cited work.

Towards Cost-Effective Reward Guided Text Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:31:19.279332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.455537Z digest=sha256:8e97be93b4cd625b586b62f18aefddc71999efdaadfe899bdca320347b3d2a38

Observation 8ed0eb1d-e2cf-4133-8fbf-377540eb4a55 · outbound

This paper cites Beyond sparse rewards: Enhancing reinforcement learning with language model critique in text generation, 2024.

Towards Cost-Effective Reward Guided Text Generation Beyond sparse rewards: Enhancing reinforcement learning with language model critique in text generation, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.269384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.459421Z digest=sha256:7f5efdcbe981a905e4f118ae7e757738ec96cdb8d57505f51ce8c09ad67be8c2

Observation 1f347e93-5f45-4724-8feb-98ad11f76c95 · outbound

This paper cites PPL-MCTS : Constrained textual generation through discriminator-guided MCTS decoding.

Towards Cost-Effective Reward Guided Text Generation PPL-MCTS : Constrained textual generation through discriminator-guided MCTS decoding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.259008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.463866Z digest=sha256:984df269e8253e730bbdff027415a38c3d879be8b0c2f23cb1004e28eab25c10

Observation 5a0af50e-0989-478d-9661-7a418d9a41d4 · outbound

This paper cites E., Stoica, I., and Xing, E.

Towards Cost-Effective Reward Guided Text Generation E., Stoica, I., and Xing, E

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.248780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.467925Z digest=sha256:c01a65e2b174b8c634e20e9cc3d224c01278e221effcaaeb0ebdf62b095b4cfd

Observation 75a215f7-23c1-48f2-8136-eb3f0dc4ea72 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Towards Cost-Effective Reward Guided Text Generation F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.238834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.471650Z digest=sha256:7400585f3bb2e5b6523e94a7f0ded9b18bae215689307e9d91df789054288524

Observation 5620212e-1ffe-4e32-b671-bd941a051b9e · outbound

This paper cites Plug and play language models: A simple approach to controlled text generation.

Towards Cost-Effective Reward Guided Text Generation Plug and play language models: A simple approach to controlled text generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.229088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.475588Z digest=sha256:24e73ea63feb6c010a106033c72637851da115ae54d4748b4a8d6e8862ccc07f

Observation fc60ab86-d1d2-41ed-b144-d9c84b1049d7 · outbound

This paper cites and Raffel, C.

Towards Cost-Effective Reward Guided Text Generation and Raffel, C

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.218856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.479221Z digest=sha256:8febf34b6f486ab531ae54c1a8f5d133cba27d6526e41314375a3ebaba99b453

Observation c5d91caa-9cf5-423b-9742-6d38d6a32b96 · outbound

This paper cites RAFT : Reward ranked finetuning for generative foundation model alignment.

Towards Cost-Effective Reward Guided Text Generation RAFT : Reward ranked finetuning for generative foundation model alignment

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.209195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.482892Z digest=sha256:4e5d3da3250d1296151d82643e45e204c1f5bc3ba2537cdb15e2cef9a2a511e8

Observation 509134cd-bc14-46e9-8345-25c4fbfbe2be · outbound

This paper cites Hierarchical neural story generation.

Towards Cost-Effective Reward Guided Text Generation Hierarchical neural story generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.199271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.486379Z digest=sha256:a659f38000bb78771dc41104a1b95d03928dd078fa47c66ef802be8d61909b68

Observation 3090a537-f00b-4fac-8c38-9df9be402e5d · outbound

This paper cites Y., Ding, N., Yao, G., He, B., Zhu, W., Ni, Y., Xie, G., Xie, R., Lin, Y., Liu, Z., and Sun, M.

Towards Cost-Effective Reward Guided Text Generation Y., Ding, N., Yao, G., He, B., Zhu, W., Ni, Y., Xie, G., Xie, R., Lin, Y., Liu, Z., and Sun, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.188991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.489659Z digest=sha256:a497b4ef6cd01b3694ed400c75142c2fc17ed7f8c7a4fdba455b80687550c073

Observation 5960b33a-5423-454c-b3b5-304c808177ca · outbound

This paper cites Value Augmented Sampling for Language Model Alignment and Personalization.

Towards Cost-Effective Reward Guided Text Generation Value Augmented Sampling for Language Model Alignment and Personalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.493155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.493155Z digest=sha256:52614817c270279c59a82205cb9b52a6822688ddae0d7bd12772ac986ac9c77c

Observation 65b9bbd9-f6db-4eb8-a03b-bf6665936af0 · outbound

This paper cites The curious case of neural text degeneration.

Towards Cost-Effective Reward Guided Text Generation The curious case of neural text degeneration

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.179341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.496954Z digest=sha256:f77a76be583d7316330d17ea65c8320988394fc8a8e0e6fe5d76f2ea915e4534

Observation bf01457a-519d-46ca-98af-3e1816fa78fe · outbound

This paper cites Grace: Discriminator-guided chain-of-thought reasoning.

Towards Cost-Effective Reward Guided Text Generation Grace: Discriminator-guided chain-of-thought reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.169062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.500310Z digest=sha256:be96ac04690aacffeda68f1a843bb58ec6dd1b91a8542fe52eccdad32f52dcaa

Observation 21523dc2-3cca-404e-ae86-23ed838c5c98 · outbound

This paper cites Alignment as reward-guided search.

Towards Cost-Effective Reward Guided Text Generation Alignment as reward-guided search

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.158841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.503604Z digest=sha256:1268acc349dd2430115d8c66f86f42809dc4e817a781d9afa558ca893149a980

Observation 25d90cbf-b1bd-418f-bff5-8c1a85066e91 · outbound

This paper cites D., McCann, B., Keskar, N.

Towards Cost-Effective Reward Guided Text Generation D., McCann, B., Keskar, N

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.148664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.507222Z digest=sha256:d15e44339740388feaf4fac865e5ec36916d02157886bef033b45df241047ed1

Observation 547d1eb9-0864-4337-9dcd-a475f1ef69c4 · outbound

This paper cites Rankgen: Improving text generation with large ranking models.

Towards Cost-Effective Reward Guided Text Generation Rankgen: Improving text generation with large ranking models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.138517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.510521Z digest=sha256:6ac2dcd6a38b002c589dc058aabbe65ee064eb9375edfaa06e65422b2bd75794

Observation 80f2c0ed-e2aa-404a-b4f3-e82a0582b0a5 · outbound

This paper cites M., and Abbeel, P.

Towards Cost-Effective Reward Guided Text Generation M., and Abbeel, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.128269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.513941Z digest=sha256:981950312a39bfd17987cd40f62342df5e8b822fe0574704d9624cbbfc28aad4

Observation 49c8c75d-139c-4ae6-9083-4241e8858954 · outbound

This paper cites Cascade Reward Sampling for Efficient Decoding-Time Alignment.

Towards Cost-Effective Reward Guided Text Generation Cascade Reward Sampling for Efficient Decoding-Time Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.517650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.517650Z digest=sha256:047f6933620c242535c57a440237038d5f74d11785c22cf4c4e6936497abe8a9

Observation c8f01739-9745-4df8-9f29-48d9363f98e2 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

Towards Cost-Effective Reward Guided Text Generation Making language models better reasoners with step-aware verifier

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.118181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.522822Z digest=sha256:61f682d698bcbd40b44a4e6580df888c130bcb4b27d83523aaa9d8a604c22ec9

Observation 04a3e171-54e3-4267-b16a-d51ab3671c33 · outbound

This paper cites Let's verify step by step.

Towards Cost-Effective Reward Guided Text Generation Let's verify step by step

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.106760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.526718Z digest=sha256:0d07f5cfdf65b5c5bccba68b1cdca29440d92f7c720427363556f4d189d81d94

Observation f6be779a-9142-40ab-9bd7-d0479da095a1 · outbound

This paper cites Chain of hindsight aligns language models with feedback.

Towards Cost-Effective Reward Guided Text Generation Chain of hindsight aligns language models with feedback

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.096410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.530198Z digest=sha256:2c4d4ff116be1e4c3ce4bb1d3effb39ddfb5bdd7d21dad6ff06940dd071cfa73

Observation ba63e888-674b-477a-90e6-7f9e1771cabb · outbound

This paper cites Attribute controlled dialogue prompting.

Towards Cost-Effective Reward Guided Text Generation Attribute controlled dialogue prompting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.084913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.533949Z digest=sha256:0429ef55e49ac6ddcb20f4e4ed1b1a1db53bd28061203b4bf6eb455345f446aa

Observation 8a0740cc-35c5-44bf-9290-9cd4b77477c3 · outbound

This paper cites Controlled decoding from language models.

Towards Cost-Effective Reward Guided Text Generation Controlled decoding from language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.073861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.538404Z digest=sha256:9ad8f39689ff7942ab27f52df3ea82b85221c2f076edead017dda46f5b55f975

Observation 68234a0f-0b61-4160-90f4-d1ac571ed533 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Towards Cost-Effective Reward Guided Text Generation WebGPT: Browser-assisted question-answering with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.541942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.541942Z digest=sha256:031d3270b9da9156932bd8da0869f49881fc09ad078d7ecc97fcf361db99568b

Observation 9ccd0d22-e5ab-4a45-bba7-13b6c721280f · outbound

This paper cites Training language models to follow instructions with human feedback.

Towards Cost-Effective Reward Guided Text Generation Training language models to follow instructions with human feedback

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.062983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.545658Z digest=sha256:91835947af1d7d0f1b6f42345248f2ba2456290e57bb35913811bfac58ec6219

Observation 7f15150a-e4c5-4c72-a086-b4db42715ec5 · outbound

This paper cites D., Ermon, S., and Finn, C.

Towards Cost-Effective Reward Guided Text Generation D., Ermon, S., and Finn, C

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.548944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.548944Z digest=sha256:5cddd916853bd09432b772e077a36c6269da4c4a3b271f3bfb0fbd312d53a079

Observation 532d3ac3-bd2e-43fc-9a3e-a5673f4f9430 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Towards Cost-Effective Reward Guided Text Generation From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.552875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.552875Z digest=sha256:18cad50bb7f45ac80c60af5f62de8ce72bbfabb5159f7243a134ee8857a3db0f

Observation d48a2270-fadc-45bd-b3de-a0c049828a52 · outbound

This paper cites A critical look at tokenwise reward-guided text generation.

Towards Cost-Effective Reward Guided Text Generation A critical look at tokenwise reward-guided text generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.556893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.556893Z digest=sha256:596bb0803f2dd2a265477da121ab4f2230bc6f032b9a6632c65508e4a75e06f1

Observation bbba2a2a-f5cb-4ce4-bc7c-8912f0f92282 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Towards Cost-Effective Reward Guided Text Generation Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.560438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.560438Z digest=sha256:b747aafd59715b8b729fbb3a3acc41c61a73514c7734b82122bd798539c4522f

Observation 6671bc15-ff5f-4021-9b52-8f235c747b07 · outbound

This paper cites Offline RL for Natural Language Generation with Implicit Language Q Learning.

Towards Cost-Effective Reward Guided Text Generation Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.563950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.563950Z digest=sha256:fae1fddf210af695f41e207d5be0d498b42c5b6a49876ce6b2cf1ec3b7ec6db6

Observation 67ef66ee-a1eb-47fa-8995-f94f18ab9d0d · outbound

This paper cites an unresolved cited work.

Towards Cost-Effective Reward Guided Text Generation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.567671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.567671Z digest=sha256:fe1fcbf10cb048942e270709f6bbbfeb04f81d5b4448933e281b9f997eb6022d

Observation 95675154-baed-4e42-8020-9fa8e19e02eb · outbound

This paper cites an unresolved cited work.

Towards Cost-Effective Reward Guided Text Generation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:31:19.039511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.571054Z digest=sha256:19b3c38a3dec67dcb522cb7948e9300211e351872c0b2dd8d85b19b5f245435f

Observation afa5a19a-35f6-4854-8124-7b1c788343bd · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Towards Cost-Effective Reward Guided Text Generation Solving math word problems with process- and outcome-based feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.574561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.574561Z digest=sha256:68bbee5699dddd96da266107d6109edb5fc9da3f5a76425e28afef73e22e2ff2

Observation b24b40cb-27b7-4804-a2b1-f5218e7a8d0c · outbound

This paper cites Tl; dr: Mining reddit to learn automatic summarization.

Towards Cost-Effective Reward Guided Text Generation Tl; dr: Mining reddit to learn automatic summarization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.028739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.578670Z digest=sha256:6973712da7f0b8b5bf3a9d6c3f109ba4425f287bebf03a6882e2b0dc6dd5de38

Observation f1679c7a-95f9-4e0c-9276-d65f7121ce28 · outbound

This paper cites W., Lester, B., Du, N., Dai, A.

Towards Cost-Effective Reward Guided Text Generation W., Lester, B., Du, N., Dai, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.018156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.582273Z digest=sha256:661c5db10c8c31f9c138a0bf4d93879755532e8c051a49bc1548c4374bb8f906

Observation f694427f-9c12-4ecd-b33f-69b7bb83d72a · outbound

This paper cites Naturalprover: Grounded mathematical proof generation with language models.

Towards Cost-Effective Reward Guided Text Generation Naturalprover: Grounded mathematical proof generation with language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:19.006929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.586079Z digest=sha256:383b73eed30a32052c4cd762e232eb4fabaf227b7c3549a27a867ab80af4c46c

Observation bcefc0ab-d6ba-4ad7-a617-06f04c063e0a · outbound

This paper cites A., Ostendorf, M., and Hajishirzi, H.

Towards Cost-Effective Reward Guided Text Generation A., Ostendorf, M., and Hajishirzi, H

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:18.995635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.589690Z digest=sha256:bebfd4afb52df3051f1e822ffd4b82c53bf763abd6cb01e5c1a58116db75f438

Observation dc2cfb28-bebf-4b09-a968-fccca63a8638 · outbound

This paper cites and Klein, D.

Towards Cost-Effective Reward Guided Text Generation and Klein, D

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:18.984661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.592990Z digest=sha256:05fd7a9d14da702cafb96611a489424f64ded9340828673ca51aa664d4a4b5cb

Observation af7c8150-6ac6-4a99-8705-a7b278833f69 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Towards Cost-Effective Reward Guided Text Generation Tree of thoughts: Deliberate problem solving with large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:18.972858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.597098Z digest=sha256:84b53b9d693cd4eba0290fedd042550bd6c9c14e766e924b1afa7d1784c6e1f9

Observation e9a756bb-819f-4fe1-955f-64a78095ab36 · outbound

This paper cites Probabilistic inference in language models via twisted sequential M onte C arlo.

Towards Cost-Effective Reward Guided Text Generation Probabilistic inference in language models via twisted sequential M onte C arlo

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:18.961625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.601235Z digest=sha256:4fe373ed802d31aa31bcae08a43a0876d703946767e07a81baa7b5e0517f0d0e

Observation 2a6bbb3b-7324-44a9-a687-bcee474741cd · outbound

This paper cites Judging LLM -as-a-judge with MT -bench and chatbot arena.

Towards Cost-Effective Reward Guided Text Generation Judging LLM -as-a-judge with MT -bench and chatbot arena

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:31:18.949850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:31:18.605244Z digest=sha256:f8e367a9e0120219e803803fb979227526c09ac94be13bbf1b39a150bab62b65

Observation 823b3df4-8700-4989-9a3f-4e6f0eb5a569 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Towards Cost-Effective Reward Guided Text Generation Fine-Tuning Language Models from Human Preferences

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.612595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.612595Z digest=sha256:9730513bfda28dae85309f43957a3d7f7afcfb5bee8b92c4f5a88b1eeecc8f36

Observation ef18b4f8-1230-4881-b5db-0136c7ed4abf · outbound

This paper cites write newline.

Towards Cost-Effective Reward Guided Text Generation write newline

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:18.616099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:31:18.616099Z digest=sha256:9f1478f8c679266c2ab3e82111e9451c845d1c942c40fc8f1bc949e2b642c6b3

Pith citing papers

Observation 6c1f8641-65bc-4e7e-ab35-8ea5357eab5e · inbound

Safe Inference-Time Alignment via Lagrangian Reward Augmentation cites this paper.

Safe Inference-Time Alignment via Lagrangian Reward Augmentation Towards Cost-Effective Reward Guided Text Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T07:05:47.150308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:05:47.150308Z digest=sha256:d6981c355740108a66e32fe7cddcc197516faaedb475241416deef3079212405