Pith. sign in

Paper Citation Record · LEDGER

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

As of 9 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2508.18642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18642 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:24:52.802595Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:38:50.313305Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.225457Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f8721a2b-bd86-4f7e-a0d9-bb699f7d0df8 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.969774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:50.825058Z digest=sha256:418f249c4a5e36be118a220a4253e07fcf1464cf32daa4ec1d311ab10dc340aa

Observation 0d22fe9f-74eb-4f4a-85a9-b96fdaf95e82 · outbound

This paper cites [System] You are an answer quality assessment expert.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing [System] You are an answer quality assessment expert

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:55.901772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:50.977092Z digest=sha256:57294b6e241e8bd648f1e91d9ecb13005f13ba7a9b5cadedb2fe917a45323e65

Observation 296e50b2-f124-4d35-918a-7e3a8813e238 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.484808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.389246Z digest=sha256:03415bf9b53750fc4e1b493cf91acf94847255c78a1c225af84e0c527ea11fdf

Observation c15c6627-a112-4e83-b2d6-f5cbba1a9096 · outbound

This paper cites - Conclusion: Correct/Incorrect.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing - Conclusion: Correct/Incorrect

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:55.234743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.494733Z digest=sha256:31fe6bdca485e654776fb676b3f8b956d0961a7734a3b06966a44c77b2b36cfc

Observation 56238734-a64b-439c-bae8-20432c8c7ded · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.754742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.117105Z digest=sha256:a2abb35d595c20d7fa519d501aa6086446f98a928d0e699faaf201e0a876db88

Observation 3432a930-0580-48bc-a107-53702569a92b · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.649394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.232908Z digest=sha256:251138d1701bd8faf56dcf29d459c272549147c3449c1c769e95e1653aca727c

Observation 2f44dcde-9f93-41e9-9214-056e2e024254 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 9

Resolution
parse uncertain
raw_fallback, observed 2026-08-05T16:24:55.080322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.594739Z digest=sha256:b313ce44b63772a218d74016e13ad5532deab4708f0cfff3726e85442b11e8cd

Observation 678a2c23-aed3-4092-987a-c2bd8b9edb16 · outbound

This paper cites Write a script for a modern history video group assignment.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Write a script for a modern history video group assignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:54.910547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.776086Z digest=sha256:ea20be0939d8db7341711be91abeff9aa1704a83944a0d2154d816a2e6a01b70

Observation 2e1cc6c4-06e7-493d-88d3-50f9322d3b87 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.724737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.909221Z digest=sha256:f70b465377b20f202f09002446606a81b0e528a4f2228bc6ca50102208c058e8

Observation 6b7fd8a9-6f28-409b-99ff-b89d623f7ab7 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.530654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.024731Z digest=sha256:56437b5991ef6651155f2b10681881bd82388ce10c2716c16b84be1aeefb349a

Observation 9212d3d3-1a8f-45a5-b5b9-3eb0f4a7bf66 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.313759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.164758Z digest=sha256:6fa78a329132a586ed3820dd57947006ab2fd94069c2e96d70272e7b9d154149

Observation c485604b-9e36-4c4e-b4cc-fddbcdf531af · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.024945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.273448Z digest=sha256:9662bd61c65b0efd5eb6057d23eb4d533526500c146fa32eae20a4c5d718b4e7

Observation 8e0c6504-2b28-445b-9af7-f246677ecb54 · outbound

This paper cites My mother.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing My mother

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:53.644926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.574748Z digest=sha256:ec82384fe621a49bad3d3ad260a61622b8eff23ca339404b0853f6d07bc823e0

Observation 56091e6d-dbc0-41f2-8462-4e7b70d157f9 · outbound

This paper cites come up with some three-character sword names.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing come up with some three-character sword names

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:53.314458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.802595Z digest=sha256:c4ac23a2f3f7a44eb4c864f89a6f47247d43d85edeabc50fe0aa52daf1ad465e

Observation 9e8fb600-23f1-449c-88d1-27803be1d3aa · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:50.583302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:50.583302Z digest=sha256:6ce3ae48733613ae80b4359c9fb9bf8e8dbde37af234ac77604456b2d5688748

Observation 87ace097-1856-44ad-a58a-9406b4e8e721 · outbound

This paper cites LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:50.674749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:50.674749Z digest=sha256:7cd6a4b86cb0b85f023d2b8826c9bdaff41e16a241f35b81cdb1fda75938b69b

Pith citing papers

Observation 96345c8f-51f8-4537-b756-e25287d8d992 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.227008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:785c6789b4a274d931116bf015fbe18c8d5eb5245a3c08fdb246af244ef54821