Pith. sign in

Paper Citation Record · LEDGER

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop

As of 15 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2601.22900.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22900 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:23:36.558304Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:27:32.645757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T23:27:33.239830Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain2
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cefae43-278d-4f2d-bf9b-26435b37dd92 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:33.604990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:33.604990Z digest=sha256:cd43030f1716ea1fb25025134d3ac0b8956c937e1a7822eb76697d523d4166e5

Observation 39808f75-d686-4975-b68e-79e1dfb4720f · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:33.702476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:33.702476Z digest=sha256:5cdc905367e4e922670a8f79bf9ca91aaacac9e1f0676459aaed1777c2da5d8e

Observation d7018b70-7633-4544-b1c0-69a81084532d · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 3

Resolution
parse uncertain
no resolver link, observed 2026-08-03T06:23:33.837001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:33.837001Z digest=sha256:2abf5d85fab2bd73ef67cfa04fa865401d1f24db123807fe848f45e04ad298d8

Observation 23a76f03-07a1-4066-a006-8e58f3aa03c0 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.015583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.015583Z digest=sha256:3a21195ec276ab68e3b86e320766e4b56f9e3d4f35edb59ee84a3471964c554a

Observation 9a52a348-46ea-4808-919b-8e4973d06bea · outbound

This paper cites = <final answer>’ inside <feedback>.)\n - You MAY include tiny snippets (a short identity, a one-line correction),\n but avoid long derivations or long equations in <feedback>.\n.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop = <final answer>’ inside <feedback>.)\n - You MAY include tiny snippets (a short identity, a one-line correction),\n but avoid long derivations or long equations in <feedback>.\n

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.113533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.113533Z digest=sha256:bbf3070443dde1c3f95bbfdee0a09c3c9189be618be921dd829532d85584be38

Observation 2fd936f3-d15e-4e22-9d72-b744531ef22a · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.295625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.295625Z digest=sha256:671a8c367e102de55dbc313eb00d37faae22ae4c71414e07b99c98e94ae34ceb

Observation 641b4ea2-6f17-4862-817b-1b5daa71fd5e · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.451969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.451969Z digest=sha256:fb4c492fe7c57067e2bcfcdd25f525da4bb46437f802bb6ae05d1f0487af74a4

Observation d2b62e70-85b8-4c26-adcc-d68cd01d0def · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.589801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.589801Z digest=sha256:86357a6f0ad5a21206065c548a711f7fdeb732ed3691ec556287c93e0948321a

Observation a69218a3-0eed-4330-87e6-b1003b24cecb · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.796716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.796716Z digest=sha256:d8b09fb484c2f92c36079f6478ef0ed92148dc165922ae1bcff27bd04734e49a

Observation 3ac0aa15-0fac-49c2-89bf-e2bb33accfd3 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:34.918930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:34.918930Z digest=sha256:f1654cbc9b5e39b85cebb6782c595c9ebc5e3a9f6c15019b8d251601fa950c17

Observation afdb0b6d-24d4-475c-ac52-e69265380952 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.070519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.070519Z digest=sha256:0afa51d456781e8f11248859752a9dedbfaef289412e7ce5e7de3911a8324ac4

Observation a23b851c-55e9-452d-8ab8-4b36d0463768 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.470102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.470102Z digest=sha256:a5f412f452d7ef43ac4a76771778577524c3d83ecde8709b881247075c3de5ca

Observation be49c4c1-f94c-40f4-9c2e-0d3b87be2e1a · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.615914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.615914Z digest=sha256:aa395db70d2d75953a6fd0aea09937c49e86804a67bf91e85108f50aa68e01b2

Observation bd0a76b0-391f-4067-bb9f-881d475c0ebf · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.679814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.679814Z digest=sha256:f92b35b369182212281370bfb5912c519340087d0310e3f78e796e544359c056

Observation 5714db7b-b9eb-41de-a31b-afebac613415 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.766817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.766817Z digest=sha256:159e5ac6413b9807e7b3077c1bc841150ef4cda1b3cfd89d5afaa97c30789368

Observation 3b1d3473-6fe2-4b39-8013-20ff19281333 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.845672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.845672Z digest=sha256:c0b19e79b3b2009c108e5af74e021e6f69caa136eb2d9095a1f597f79910147d

Observation 3c0eb349-cd56-40be-88ee-8b8e95beb06e · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:35.915756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:35.915756Z digest=sha256:a2b0c1ba65e0c395cafb88d957eec39cbeb4de8f8a72a7ccc819994f05a1de42

Observation 514016a6-2573-49dc-88bc-6b4cb1684cd8 · outbound

This paper cites SFT/RAFT system prompt and CITL-FT initial prompt You are a reasoning assistant.\n Solve the problem step by step.\n\n Output format (must follow exactly):\n.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop SFT/RAFT system prompt and CITL-FT initial prompt You are a reasoning assistant.\n Solve the problem step by step.\n\n Output format (must follow exactly):\n

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.006834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.006834Z digest=sha256:7d98f11a66806e1a39436fd396482fd0a1ac0787de088717cfa60f79b33e284c

Observation 577a79cf-900d-4da7-8c73-40f546814446 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.071017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.071017Z digest=sha256:25ed614fcf5727bd1535cf70248ec4326ab340ccc53fb50942614e68b43ac2fd

Observation 6aac5df6-7992-41f9-b3e8-e06468f823f4 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-03T06:23:36.224133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.224133Z digest=sha256:8ddc248fb53934d305d7dcbc28a36078ef0113294772dfe0070b9a3331ed2329

Observation b98c1b9e-9888-40c3-9cc6-306851cf896e · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.285742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.285742Z digest=sha256:807cbbad9063b2cf3dc94adc790130a3a914234098cf2a19843f854d14aaaa1c

Observation 438ab60f-2539-4e78-93bd-ceb78d9abdf8 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 26

Resolution
parse uncertain
no resolver link, observed 2026-08-03T06:23:36.357177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.357177Z digest=sha256:321ccc740d69cef7022b8c482330a9d4c08d86abffa4019d37622b2cb152cce7

Observation 8de2b3af-f033-4009-8cfb-b94c3b9c45ba · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.442551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.442551Z digest=sha256:e895593a82a1bc766c1655066c59121559bb2c8916840b84cd241eae88dc82e6

Observation c622e5a2-ba65-4ea6-8014-6331d6ba94e9 · outbound

This paper cites an unresolved cited work.

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:36.558304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:36.558304Z digest=sha256:3ccacf1a47d9946e5ada666808b0a8c16c1a1925aefd4f9737c243006410478b

Pith citing papers

Observation 06f62b64-f665-46e8-a1f2-1cfca9c3763d · inbound

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models cites this paper.

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:27:33.244435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T23:27:32.645757Z digest=sha256:765fbadbdd677ddb8713bac51241705a33376dd6cd1ded6970cef6365eff9cdd