Pith. sign in

Paper Citation Record · LEDGER

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop

As of 14 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2601.22900.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22900 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:23:36.558304Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:27:32.645757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T23:27:33.239830Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain2
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cefae43-278d-4f2d-bf9b-26435b37dd92 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:33.604990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:33.604990Z digest=sha256:b0f31b6c36b3378cead2410da272d7c528eab1bbcb8cbeb7c6ba1c190a882793

Observation 39808f75-d686-4975-b68e-79e1dfb4720f · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:33.702476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:33.702476Z digest=sha256:6cb719ff55aa27a32901c696b3bcfe2877b77cf9423b2904c320a64dd59a47a4

Observation d7018b70-7633-4544-b1c0-69a81084532d · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 3

Resolution
parse uncertain
no resolver link, observed 2026-08-03T06:23:33.837001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:33.837001Z digest=sha256:8a2068d878b8286f0db9d4017188b3365f7467c69d8941aaffc8988c26278272

Observation 23a76f03-07a1-4066-a006-8e58f3aa03c0 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.015583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.015583Z digest=sha256:fffc81e1b11c66748f81a4c2e2947f5e75ef7a4609af8406e3207e218720ac01

Observation 9a52a348-46ea-4808-919b-8e4973d06bea · outbound

This paper cites = <final answer>’ inside <feedback>.)\n - You MAY include tiny snippets (a short identity, a one-line correction),\n but avoid long derivations or long equations in <feedback>.\n.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop = <final answer>’ inside <feedback>.)\n - You MAY include tiny snippets (a short identity, a one-line correction),\n but avoid long derivations or long equations in <feedback>.\n

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.113533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.113533Z digest=sha256:c66b3785f4ef7042501f8b76f42b6ea1cd3920d67f815dd968b977e867fbc1db

Observation 2fd936f3-d15e-4e22-9d72-b744531ef22a · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.295625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.295625Z digest=sha256:23926edd5d5a30f935233c81f261b6e85bcb9b8ddb1efc0b429a61cf793f9a80

Observation 641b4ea2-6f17-4862-817b-1b5daa71fd5e · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.451969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.451969Z digest=sha256:4109f8757ee4591746ff8d91aa0073d904f1aad7dad0d461285cdd6a3b5d035f

Observation d2b62e70-85b8-4c26-adcc-d68cd01d0def · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.589801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.589801Z digest=sha256:60791dd2b2ccef5ecf9615d8572a0edece379f88064097645ad49204f37f9e30

Observation a69218a3-0eed-4330-87e6-b1003b24cecb · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.796716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.796716Z digest=sha256:501d4abf8ee4d22d6365ecf3a4a1100aa12aa2766c53cc49db67adf8f078e102

Observation 3ac0aa15-0fac-49c2-89bf-e2bb33accfd3 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.918930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.918930Z digest=sha256:2f5d672a9d577d5141147266c6a1e32537ef6553df02e6e5369c1f28699e94f7

Observation afdb0b6d-24d4-475c-ac52-e69265380952 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.070519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.070519Z digest=sha256:d3a8b7be5f9b3c3d9afde47c2e503969ca96d2d5bf3ca12e8ebfdf9fbbcad241

Observation a23b851c-55e9-452d-8ab8-4b36d0463768 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.470102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.470102Z digest=sha256:777b56e633d6b1864606ad00d0af7cead702575d176bfbef255710a99aebe5e0

Observation be49c4c1-f94c-40f4-9c2e-0d3b87be2e1a · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.615914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.615914Z digest=sha256:85bee4af5bad56abaa28a197210d97c54a241f56cdf739aad5a0b591fa175c0e

Observation bd0a76b0-391f-4067-bb9f-881d475c0ebf · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.679814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.679814Z digest=sha256:b11b016fa1a50fa8300cc1e86c016495b385aaf25bc1196398f2e349cbd408be

Observation 5714db7b-b9eb-41de-a31b-afebac613415 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.766817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.766817Z digest=sha256:68ef75e51a122f48847309db6c2d51cffd155efeee5b4c650b27a4a009695362

Observation 3b1d3473-6fe2-4b39-8013-20ff19281333 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.845672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.845672Z digest=sha256:686998ea4064122afea3137ee53aef28ce8456fb5d217cb7b1ce53ebeca2e5b8

Observation 3c0eb349-cd56-40be-88ee-8b8e95beb06e · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.915756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.915756Z digest=sha256:7ba8fb73a11b71fef2767bee67a103b315cf71141d1ad563b184391342a6a533

Observation 514016a6-2573-49dc-88bc-6b4cb1684cd8 · outbound

This paper cites SFT/RAFT system prompt and CITL-FT initial prompt You are a reasoning assistant.\n Solve the problem step by step.\n\n Output format (must follow exactly):\n.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop SFT/RAFT system prompt and CITL-FT initial prompt You are a reasoning assistant.\n Solve the problem step by step.\n\n Output format (must follow exactly):\n

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.006834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.006834Z digest=sha256:0e18c50b93fbb9ce50c3d618a7f4d53aa4104301bc38ceebbf832820d97834d6

Observation 577a79cf-900d-4da7-8c73-40f546814446 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.071017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.071017Z digest=sha256:23b5cca52e2a956d04f5f54e5f720ffbb41b2266419d5a51e726fc1b98d92496

Observation 6aac5df6-7992-41f9-b3e8-e06468f823f4 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-03T06:23:36.224133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.224133Z digest=sha256:ecbb202b4700b957bed75d693366999602f82ebd71b6a8f9d52e98f6d151a2c5

Observation b98c1b9e-9888-40c3-9cc6-306851cf896e · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.285742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.285742Z digest=sha256:a4141a95b69e7a9b9c6e51e7c52ff80fa1856a1d1391268dd0f370d557acd1b9

Observation 438ab60f-2539-4e78-93bd-ceb78d9abdf8 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 26

Resolution
parse uncertain
no resolver link, observed 2026-08-03T06:23:36.357177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.357177Z digest=sha256:477cb6fe5209f8fddafa1f237c7237ddd98750aaff0fdd04b79cec770da96e44

Observation 8de2b3af-f033-4009-8cfb-b94c3b9c45ba · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.442551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.442551Z digest=sha256:f335f4db96e489c2b6fc60e112ddf4751748d062980f7bad52c8153069916295

Observation c622e5a2-ba65-4ea6-8014-6331d6ba94e9 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.558304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.558304Z digest=sha256:ca173629a19e5306213dfa0bacc9be4c9d32537ff5b085179e7875ade1c3974f

Pith citing papers

Observation 06f62b64-f665-46e8-a1f2-1cfca9c3763d · inbound

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models cites this paper.

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:27:33.244435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T23:27:32.645757Z digest=sha256:dd03c268bbc72ac9bfd96e2ad1f0c21c551dcdf52f9b34b4729ca8ec7e8bdf9e