Pith. sign in

Paper Citation Record · LEDGER

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

As of 2 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2511.07833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.07833 v3

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T23:13:43.754235Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:48:41.262444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact19
  • verified fuzzy2
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78f8feb9-63a0-4407-aab6-7a9673cbd448 · outbound

This paper cites online" 'onlinestring :=.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation online" 'onlinestring :=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:15:27.077451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:2d4c6bd652a1ee24da0341313f9ce71b1dc09c88e62b49269523ff488fcb0543

Observation 0dda1b3c-ec3d-453f-862e-506693063364 · outbound

This paper cites write newline.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation write newline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:15:27.073870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:24b9e73c74363991ac5ac1ffc14743d54f072c70f5a1ad99d39c088d8e92473c

Observation b8fbbeec-10c0-443b-948c-61f774d9dc34 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 3

Resolution
verified exact
doi, observed 2026-05-17T23:15:26.628592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:9f275d38d80a5df514f791548165ddefbfc1ecb8310dc836bc037127eef30ab8

Observation 9af2921c-36e5-4d64-b5b9-26c3390a575c · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.100254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:739e150b92604efac7d08138cf8759e271682300f08ac73f1641228e5c90c0a0

Observation ff021806-6cbc-4c4e-9f76-2040cafb1f6d · outbound

This paper cites Program Synthesis with Large Language Models.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Program Synthesis with Large Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.709803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:c39b0c29c85ca26c1fff510943b5787dde55c98e04d42e65541ebb0bb84977b9

Observation 1732b6aa-d882-4e1a-82a8-01be7729ace3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.704346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T08:08:23.404839+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:9a54a9ef4397b27cb455915ce183693f7e71807d0e560a01f1b8b172841b0e0e

Observation f88ba77c-3ad0-423d-b2a4-0c85c4a8f729 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:15:26.720782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:dee7e1041f382b6062ed51ccaedf80d2c506c06c08ec99853529744bec5832e7

Observation b1eda87f-922b-4c70-8774-9d6a501c5b47 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.687761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:082019c0aba9562cbebe5147b335f27914e7d49708920a4217e01f5707ab8373

Observation 65231ea9-88c8-4882-99e6-313aa18a55cb · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.109445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:97749b02b654fc4db60084c863d68c6d9fbacaf3d6252519e40a002bcce39020

Observation 4722bb41-a52e-45c6-8b49-523a937285b8 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.106174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:fc25913a50f3bfa114880d1f100dd87efc015d40e91d340430f4752b04027d72

Observation f1abe81b-2b82-4b9a-b184-5cf7b69e997e · outbound

This paper cites A survey on large language models for code generation.ACM Transactions on Software Engineering and Methodology, 35 (2):1–72.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation A survey on large language models for code generation.ACM Transactions on Software Engineering and Methodology, 35 (2):1–72

Reference 11

Resolution
verified exact
doi, observed 2026-05-17T23:15:26.622996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:243fa86535335a370b391ec8962c093cca0a0cff8a7fd09e99f40d15cc004434

Observation 8b8e813d-2af0-465f-8e07-a0f2d67932be · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.067389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:90d6eccf9124fe5ac9cae36b5114c9352c9c0579cc67799ff3234a1cee160904

Observation df5a381d-ad51-42b3-9a13-c7d69b8bf89b · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:15:26.729709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:80da4303729941e3009e937adde7141a5a7f9190a33ed095705cc15a5d44e2e0

Observation 8ad6cdf3-0ee1-4abf-895f-66c145ce1fe6 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.070238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:a2a71d09fca911e94c8ae2b2bc88c9ba6496961baeec9332987b11e41098705a

Observation f130e6de-927b-40dc-a66f-41c6705c35da · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.064608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:cb2126ea56f98d5a30318221630aac8f9b8e8fb438034a82b3e01b08daba6c2d

Observation d253d772-7e72-419a-be9a-2a5655fdfa45 · outbound

This paper cites 2 OLMo 2 Furious.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation 2 OLMo 2 Furious

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.681922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:79db227b45518f5ba429d110b8e1b4203842f30888c3b948be84b9b9f9d36872

Observation 74db933c-5edc-4773-95cc-5a5dea52f840 · outbound

This paper cites OpenAI o1 System Card.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation OpenAI o1 System Card

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.739295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:fccdfdbb179ad9893249f86bb4bac820a6e6d10f81e4bb8c2ee88b5469464f03

Observation 0128674f-5ac9-4ba2-b120-789082f92ea8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Proximal Policy Optimization Algorithms

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.715339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:e46b75d6772249630ab0cd806d72fd5c5323441f07f06f851705ec5bc9a20a92

Observation 49832b57-ef4f-4347-9ac6-c76e8f00b508 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.768199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:774dc7c17a02281526800a6a1c44d6602c6d79032597fa2281b7aa1a2be6780e

Observation 587966da-50c4-4cc3-83bd-94e80bf9dd04 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.097267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:0319d9e0be20c780da614f7b2725c28f5976637124bf9cf5ed7741d68b1492ac

Observation e32b2ea2-1a58-457f-979f-3b3d0eeae31a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.675857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:ccea3a4e5b62c9d3b1964f45b15cba456673ab7cb68d7c045ab90822b3c2e676

Observation 088eedf2-c4f2-4838-b9e4-7ee7292faf8c · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.103118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:fa71ddec71245650b895ddd97431054098609b568548cb6a249a02cad661830d

Observation b5332958-03b3-4694-8bc9-a08f8aa574a7 · outbound

This paper cites Demystifying llm-based software engineering agents.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Demystifying llm-based software engineering agents

Reference 23

Resolution
verified exact
doi, observed 2026-05-17T23:15:26.616947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:42ea93a7edd89ea0f15e5f832f99bcd9093e1599a2494152a1d8236a92f27076

Observation 37f2ec39-82a2-48c0-a334-6b26acdf5d34 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.698984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:ad3ea6ef2f6af465a20780c9f3aa47a2560561638396545ed914727475684408

Observation 3c6e2350-a71e-4b38-b8a7-d1f6d650143f · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.757913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:18eb5a05ddf4568aec061f408b8d3e5b9ee57c0c87535ea2d698ce13c96ce7ae

Observation 7d664539-6eb5-49a5-888c-b6feb604fda2 · outbound

This paper cites Qwen3 Technical Report.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Qwen3 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.693481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:719e1b7c28a6805b2db985e2315f1aab47a526a6bf4933a75e342e7b1e52c57b

Observation 07f74694-db5e-4fad-a0f2-5248e4aac6fd · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.089038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:d3013a0e9f1e704046234a42afd9092466c0a40a34092d86f10894db23d324b0

Observation 22249bd6-35ac-4837-870e-f8c1fba5d130 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.091879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:1fbd51f81aa89b4105b9253950be682a39720e573613940dcda4b232c1503ed1

Observation 1831d8ae-812e-4df3-a114-7d67837e9418 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.725203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:fc0f0a10addddb4f8bc381de0d8c6a41f3dd71f7bbe64a852da1019670c50e26

Observation 04809f1f-fea8-4b58-a2c3-8888db5b6d5a · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.747000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:bd1c8e62541e8130771e7f2798227c3529d2ea0b4a9895d326db5ab6425680bc

Observation 99842e8b-521f-4a83-b48a-a42e6b4d92f8 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.734331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:a57016069329016fb3b85d988579ceb37ad341f986e421b7995d97829414fd54

Observation bb5db8fd-4415-4d2c-bb58-9cc652e33de5 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.752887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:f213513d75e1e9b24c8923c2b737d5bb11fe5234f9486e6a1476a731b8ffaa62

Observation 7f5e4ec9-72af-4c30-bd9b-54387848f990 · outbound

This paper cites Group Sequence Policy Optimization.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Group Sequence Policy Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.763421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:96050abd11707d55bf1910998c13b02418a1757efb2560d5a681125fd766f895

Observation bce6996a-dd7c-4530-8e04-81d35c727c3e · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.086126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:546847e8d5b42162a5a036d7416dcb360ba6faa057c8edb6002a5bd5ed4a87fb

Observation a5897ef3-213c-4350-8834-62f414bfaa3b · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.080401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:ee31eb386d8ab224b7173d47367edcca8d3c8b1cedd0f633a7bd56a4d7d06a43

Observation e9b48f81-5e85-4f64-8170-38d6f713ba79 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.094592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:079df6eeaefa27c12ae820f28a9ecea1c14e657a5daa688308fc471c5ab484f2

Observation e0f11304-120d-49c9-ba97-261317bcfb35 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.083288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:c40b48a9071d7b8df32e2592c64f1404cd5539db5dfedaa040c4ec69fe6435c5

Pith citing papers

Observation 7307b0b4-e442-4238-bc34-fa070b144c78 · inbound

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute cites this paper.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.262444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.262444Z digest=sha256:1b7c10bb41f1ef4676188d200d5cb5c9c3791fb67f4331684e06384d9519edf8