Pith. sign in

Paper Citation Record · LEDGER

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2508.18642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18642 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:24:52.802595Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:38:50.313305Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.225457Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f8721a2b-bd86-4f7e-a0d9-bb699f7d0df8 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.969774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:50.825058Z digest=sha256:8ae6d230aba3726ec0163a2ae15efe65b5b0204c0fdba812001fbbed567d0b38

Observation 0d22fe9f-74eb-4f4a-85a9-b96fdaf95e82 · outbound

This paper cites [System] You are an answer quality assessment expert.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing [System] You are an answer quality assessment expert

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:55.901772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:50.977092Z digest=sha256:1eeb303ac416741819446224d702dbfc66d58f7c19655e8b2d13301bbc80290f

Observation 296e50b2-f124-4d35-918a-7e3a8813e238 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.484808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:51.389246Z digest=sha256:8c867d2d6b111c45fd0f673bb4d0865d72206ec5e5d0ac7992b4dcb6a7d2873c

Observation c15c6627-a112-4e83-b2d6-f5cbba1a9096 · outbound

This paper cites - Conclusion: Correct/Incorrect.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing - Conclusion: Correct/Incorrect

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:55.234743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:51.494733Z digest=sha256:67de4d829acad9d2b7442482be611ddc8866da8c71515061fb58621771297f74

Observation 56238734-a64b-439c-bae8-20432c8c7ded · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.754742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:51.117105Z digest=sha256:d8d255c72a2d4a228dcb66dfc436c6aceac848288969171859d8184850be88a4

Observation 3432a930-0580-48bc-a107-53702569a92b · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.649394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:51.232908Z digest=sha256:1b10a3c96613bee0342d37a483f8d091248e480ac76cc5c975170782bc60c706

Observation 2f44dcde-9f93-41e9-9214-056e2e024254 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 9

Resolution
parse uncertain
raw_fallback, observed 2026-08-05T16:24:55.080322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:51.594739Z digest=sha256:9a030c4bf5065af26ecb39819efc7fbe7a620de1d94996d29236ec0da331ce6d

Observation 678a2c23-aed3-4092-987a-c2bd8b9edb16 · outbound

This paper cites Write a script for a modern history video group assignment.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Write a script for a modern history video group assignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:54.910547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:51.776086Z digest=sha256:a337d719d25bc9fa0a3495951d9b695b1d7280fe9487fc7bf3a8e6f27f460e44

Observation 2e1cc6c4-06e7-493d-88d3-50f9322d3b87 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.724737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:51.909221Z digest=sha256:e189ea1df41a48d399d8e5844390e34b6ad62f7c0a1e2f44c01db4e17af298f9

Observation 6b7fd8a9-6f28-409b-99ff-b89d623f7ab7 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.530654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:52.024731Z digest=sha256:660acd7027fc390b19634dbe811f57d999332010352af9bcbf244c9afc190e3d

Observation 9212d3d3-1a8f-45a5-b5b9-3eb0f4a7bf66 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.313759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:52.164758Z digest=sha256:530cd22490bf584eedff71f61e2e44d3d3e828e51112b3c50d755de5fde39b22

Observation c485604b-9e36-4c4e-b4cc-fddbcdf531af · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.024945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:52.273448Z digest=sha256:c8fca1aaa0be6dcd52f967f1824a10e65963f2195fc760306f7cc367c1dd452d

Observation 8e0c6504-2b28-445b-9af7-f246677ecb54 · outbound

This paper cites My mother.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing My mother

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:53.644926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:52.574748Z digest=sha256:8741b83ff1ea7206c4fb2309718c429989e1383b6102a72b14691530add16d77

Observation 56091e6d-dbc0-41f2-8462-4e7b70d157f9 · outbound

This paper cites come up with some three-character sword names.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing come up with some three-character sword names

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:53.314458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:24:52.802595Z digest=sha256:d8ebcebb42261b1a692b4751aee20f70dd68e06d915bb64507070d8b9dabcebf

Observation 9e8fb600-23f1-449c-88d1-27803be1d3aa · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:50.583302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:50.583302Z digest=sha256:241cf6b809b17e271a15c1d7f7d51e2c003a580714797b08e412f740f9c34da4

Observation 87ace097-1856-44ad-a58a-9406b4e8e721 · outbound

This paper cites LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:50.674749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:50.674749Z digest=sha256:58d71d7068a40499bf1ee61c8fa3093ebe4cec3b49edd606818abb86ca361534

Pith citing papers

Observation 96345c8f-51f8-4537-b756-e25287d8d992 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.227008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:a10fd5db59ad568cc9f79d005e232ac9282569aa26022c5941468a8509f322b4