Pith. sign in

Paper Citation Record · LEDGER

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following

As of 23 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2508.02150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02150 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:14:36.480639Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30ef7ae3-2cc4-46d8-8669-388d24235a35 · outbound

This paper cites To make the instructions more complex, I want you to identify and return five atomic constraints that can be added to the seed question.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following To make the instructions more complex, I want you to identify and return five atomic constraints that can be added to the seed question

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.131823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.131823Z digest=sha256:d58beaa732fe8345b69ed1f9090d5e936bcb12b61855294ad5c6ce61b05e6dcc

Observation 36fea6dd-fb2d-48b9-9067-968e99c5cdf4 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.168195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.168195Z digest=sha256:7d93ee9c1f6f64cf0514c9bb0e11a130e760d4ae89aebdaabe9cba1dc643dc30

Observation 88453eb8-30b2-4c03-8bcd-bcf6610ad178 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.226387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:35.964825Z digest=sha256:ab5aaf2687496bfc6f7d9ac48f6b5d06d001caec519129343deecbc229914a22

Observation 070800b4-ace1-497f-b440-6ef268daf33c · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.208220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:35.985981Z digest=sha256:54ef949f4827e1135b333e23c11643d46bfe67093deea6c7ed99890a80d60bef

Observation 9f26d81b-e226-42fa-809b-e6c91f2e289c · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.190858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.009417Z digest=sha256:10776a29ff6a23a6d02d43d7e39782d2ea709a4e3389b9499f58c2557be842a6

Observation a037e7a0-6edb-40ab-9a61-51cae0ebd1bf · outbound

This paper cites Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.026653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.026653Z digest=sha256:f91b91e3b9da19f9c4e8f3e3d3a4871f17e7b0d4a79f0603cd7c025b8a05822f

Observation 67020480-fe39-4fbd-813e-01d0dd2054e7 · outbound

This paper cites 5-thinking: Advancing superb rea- soning models with reinforcement learning.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following 5-thinking: Advancing superb rea- soning models with reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.169330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.032798Z digest=sha256:fc1f2b0ba31e5f201aff1cd7ac0f893bdd578c269dd60f3e9c35a5c9288fc6f9

Observation 63de9e48-05d5-465a-9c44-e346864dee00 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.149630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.044328Z digest=sha256:5a06f44f0d26f024dba844fe757cb197b892c3f3eb0cecf4af6cdcf952440e6b

Observation 2c1eee46-b55f-402a-bea0-bc133bc518f5 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.127153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.060611Z digest=sha256:570106f04a3c6bf6d6b0d00a0b119188d71bebabb8373ef36f2fb063bb491bb9

Observation c35bf93f-2f11-48ec-8320-bd251713d45a · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.100189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.086975Z digest=sha256:51b99568341ba56db611a1ba36f60548b0aeea35846808b9aa66eb2f1131494e

Observation f10955ec-320c-4dcf-88b3-31dc5b3ad4b1 · outbound

This paper cites joy,” “anger,.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following joy,” “anger,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.053649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.104630Z digest=sha256:782a22ccfcb887d492af938d92b8756df090bfd68e80c07e67d3b4d43efa7419

Observation 80807db4-1b09-4088-9b74-8b1f1dce4a6e · outbound

This paper cites You may choose one or more constraints from the list or propose new ones if needed.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following You may choose one or more constraints from the list or propose new ones if needed

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.185917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.185917Z digest=sha256:a46942d19e8027f8d67b2a6f133c4d563eb30f123d531fc21647db17badcc418

Observation f0fe6964-45f9-4668-b253-2224dcc20fbc · outbound

This paper cites Your task is only to generate new constraints that can be added to it.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Your task is only to generate new constraints that can be added to it

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.214321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.214321Z digest=sha256:9ac6de9bcbec48f9a9acbf851911afa3e0d0170dd0c9ecd6b5939b5184d23a4f

Observation f92f8690-123b-4b0e-be3b-87b7688341dd · outbound

This paper cites c1": "<first constraint>.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following c1": "<first constraint>

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.232092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.232092Z digest=sha256:bde2ff129aa2677ffc53fad86e14ba5e17ba647b2941aebb4ba2914b9435ab41

Observation 9361bf8c-c861-48cb-bb46-f8bf52b01c33 · outbound

This paper cites No explanation, no reformulated question, no analysis—only the JSON structure.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following No explanation, no reformulated question, no analysis—only the JSON structure

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.256221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.256221Z digest=sha256:456dfd11f6d9d8982efeb2169a5590f36fa0fbfb4a3d8a26f60e7480cf38b943

Observation 1e7f2873-a75e-42d1-8e88-bdb640d3c879 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.274838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.274838Z digest=sha256:4d17796ac3ae0be550f9e2f36ce6146f0800fdee0fffca4474f6fa3d6f80d328

Observation bf2fcc35-c469-47f5-824d-79fce6f97408 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.865813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.314901Z digest=sha256:645af7750cf91926e5fb0a9a4ca898bb8fdb862e088cc115deef48514efa1318

Observation b66d7156-b68c-4fe7-81ac-19e21775824d · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.759569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.366593Z digest=sha256:1c08b0beb4312659df12450ca863c88bb58e2090fe528b59ee9c9499f48467f6

Observation e9c347f7-e45b-4334-83c9-30fd4bc1f17a · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.729126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.390194Z digest=sha256:ce2f5e06d96e3566feedfd6ca0ac55069093d7833cda7a95c5096c25c9146b9d

Observation b174ff10-11e9-43f2-ad0e-d195c6ee65df · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.706982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.417236Z digest=sha256:e1281286bc61f0bd8b206535a37d3bc2b631cd4ce60c1377eb5e1b786cddb0e6

Observation 2e865080-540f-4fbe-8c8a-f2986f6aa140 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.671862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.455905Z digest=sha256:81059697ab741c17fab718430ba8e1c8835a93c821546b677844cf41795132a6

Observation 3bae2fd9-3e61-43c5-9537-737eb96254dc · outbound

This paper cites You are a meticulous assistant who precisely adheres to all explicit and implicit constraints in user instructions.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following You are a meticulous assistant who precisely adheres to all explicit and implicit constraints in user instructions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:36.788801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.344774Z digest=sha256:68aba800fac3e76ef9f47a3b6d22a27e410c3b1ffcc485ed5d3c388687031361

Observation 26ba5337-378b-4f70-81c4-be680506db97 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.632653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.465633Z digest=sha256:2ca94354101596cb2c8c9d570be2a7b47ed5433572c9e9a4fa6353d1980313d4

Observation bbec9262-9820-4269-a73a-bdcc7c5e36f2 · outbound

This paper cites Whisker’s Quest.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Whisker’s Quest

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T05:14:36.609633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:36.480639Z digest=sha256:2f26b5274a41747b9fb6bd664f8550240ebf5b060fb439ab62d151344c0af4e7

Observation de110678-2512-4b55-bcfb-39ab272fa84a · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.243055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:35.823310Z digest=sha256:9d6a3f0d8a11e9ea033e2995ead77a652a7a7d7cf0aca04d3db66e824e4cca6f

Observation 8919734b-a4e7-414d-a5ab-e7bf419868f0 · outbound

This paper cites In �������� �� ��� ����������� ��� ������������� ������������ ��� ����, pages 18632–18702.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following In �������� �� ��� ����������� ��� ������������� ������������ ��� ����, pages 18632–18702

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.259997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T05:14:35.700324Z digest=sha256:83c777bc816fe8599e21e76c491942c62d705029c41a577eae9745cac879017b

Pith citing papers

No inbound Pith citation observations are available.