Pith. sign in

Paper Citation Record · LEDGER

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

As of 23 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.08224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08224 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:23:27.192432Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a055507b-9cc2-4844-a9cf-fb1774a5aa21 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.220225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.220225Z digest=sha256:07da6e14c1e05b5794bd3bfcd759ea514a8f0411ba51230e0fb9631650fdf39a

Observation 09c720e4-4280-4d93-a8f4-76a21bb0f01c · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.300358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.300358Z digest=sha256:08995371ef0faf4256baf3b28913f5bb4f56661b814b824f220fb8960dec3bb0

Observation 5772d3d2-5b68-4017-a57a-0ea646db5f5c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.325949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.325949Z digest=sha256:61f39bac969a6f2c6f35e2a2bf4257c8fc24b5312cb4e65969cac9f180f390ea

Observation 9c0a9770-474b-407e-9c5c-14dab82bd794 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.354692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.354692Z digest=sha256:4d67be35ff18e9cf263391b2f1e7a61d4a725c91d559c2cfe85569a8d1650d6e

Observation eb2c0d45-5c10-44cf-a228-dfb3972a0ee7 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.373998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.373998Z digest=sha256:37790ad4d584835e1e49aa7de7c268f58ea6e489c78c47686bdfaf3a1bcfd43e

Observation 22bcd684-565c-4e41-8d61-2288361cbb08 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:29.161951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.387998Z digest=sha256:61abe88cb00f087301c468fdf6bf42f629d7a3fa7785dc4eadbd0455a804fe4b

Observation 0b41890b-103d-4d18-bc7f-3367627a6fef · outbound

This paper cites and Liu, Alisa and Dziri, Nouha and Lyu, Shane and others , journal=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training and Liu, Alisa and Dziri, Nouha and Lyu, Shane and others , journal=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:29.121004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.394372Z digest=sha256:41fd6454b60d6d520e8a50632c9c1f0123e1110b7986159b2197ae3412bbec4d

Observation f722d4af-9da7-437c-8511-7ce41877fd62 · outbound

This paper cites 2025 , howpublished=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2025 , howpublished=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:29.066668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.399464Z digest=sha256:b022cbe7573b3ad27774fffbe5120941cd1b2e74e02c2715449719bff797d54b

Observation 9dbc91ec-8d30-4a85-b75b-e4f0c38f2c53 · outbound

This paper cites 2025 , howpublished=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2025 , howpublished=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.405038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.405038Z digest=sha256:fd5fea561771be0cb8a6599adec7a26faf0f68b2333e573f581840081a2e44ad

Observation 6c499c69-1e6a-4d53-a4d7-96b1c097a1c8 · outbound

This paper cites Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.430231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.430231Z digest=sha256:b778da58764f442c1be043bfe7fbc7494573439906290dacbd4f8b55881aefac

Observation 0d8c7570-6afc-4257-a1ac-c80131a63690 · outbound

This paper cites Multi-Task GRPO: Reliable LLM Reasoning Across Tasks.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.510719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.510719Z digest=sha256:d7edcca45c9fec99ee9cb768d65b8ab9868faf042d98c05e11c482c920c8be06

Observation 8a5937d4-b8bf-4ee2-aa04-5d674c99f2c2 · outbound

This paper cites Proceedings of the 7th BlackboxNLP Workshop (EMNLP) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Proceedings of the 7th BlackboxNLP Workshop (EMNLP) , year=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:29.039496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.566066Z digest=sha256:b48777c948585b0b9b3363a16d17649120419b38e3c9e579e9ddb767dec89d1c

Observation fed65094-a888-4d86-afe9-656b5468259c · outbound

This paper cites 2021 , howpublished=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2021 , howpublished=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.601514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.601514Z digest=sha256:579b73a586eaf3c177d0664e848ef2fa45bd6c45872b26323db698ca799e1ced

Observation b8a92781-348e-4b16-ab7a-85198e2ba2b4 · outbound

This paper cites 2023 , howpublished=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2023 , howpublished=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:29.013927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.633873Z digest=sha256:42c919d2e004e7a277347548b3792cd0a36b6cbf0c0b28c2da48de47ef3f250a

Observation 1ff3eea1-771f-4546-982b-10489d92c7f5 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.682339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.682339Z digest=sha256:1509ed1da9204500cadda75d87ef462601fa2ac302e5aaa9bfb95edc0eb2ea61

Observation ed7e3211-8d88-4d3a-a7ad-dc791b4dd4e6 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Interpretability in the Wild: a Circuit for Indirect Object Identification in

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.691452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.691452Z digest=sha256:70dec98bba2caf7549da0b4561afc6e7fd2f7bdf78969b9e34a57e3e3d43c787

Observation 99ffb5c0-1972-4a93-9f01-27fb72ec9e53 · outbound

This paper cites AtP*: An efficient and scalable method for localizing LLM behaviour to components.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training AtP*: An efficient and scalable method for localizing LLM behaviour to components

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.696774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.696774Z digest=sha256:a4c3a867a2d59f82e5053c26afe7511df98f11aa8b7e251c4fb8f5ef1c3c2b75

Observation 929fb0ac-c46e-42ab-b4af-5b9dbcef26c6 · outbound

This paper cites International Conference on Machine Learning (ICML) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Machine Learning (ICML) , year=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.976791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.708725Z digest=sha256:b8e1f506b7682051aeb7d4a4b9026eed7ef8d2d26b6c7da87455ff20716d5a40

Observation 0da3dfd1-fdf3-449d-8e36-71e9348a3295 · outbound

This paper cites Conference on Language Modeling (COLM) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Conference on Language Modeling (COLM) , year=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.960239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.713646Z digest=sha256:4c92c8f4ba77077238b6eeb97d14f9014c10836a979a8e3106f246ca83cc89b4

Observation c1ae0d81-43fa-47ad-b972-961ca6787fdc · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.905324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.719130Z digest=sha256:9ed9d9a18bd7b38236264c0ac32a33084fc4c38972d3cbe9b7223b2c1517a0c2

Observation 40aa5ad5-69f2-477b-b54b-c4e866773d13 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.720643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.724014Z digest=sha256:e078769889027f56600bd00b09536be4e00f766ed931da76b567d32620053f8d

Observation 95492037-5286-4823-8135-0a378bec20da · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.702207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.729024Z digest=sha256:6bbe8dc475b5c0d83655af97aef7c18f736f96d5e9da601a1c65a67f27b3089e

Observation 058f235b-cfc3-438f-83cf-8f2350d2fd1d · outbound

This paper cites Biochemical Journal , volume=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Biochemical Journal , volume=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.554302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.753658Z digest=sha256:a1e6e30f9a04c1f254ee197bbd0d53d1c5737091b8af94fe7f64bf575b4999c9

Observation e13722e8-4788-4939-99e7-6d1074321f8b · outbound

This paper cites Symposia of the Society for Experimental Biology , volume=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Symposia of the Society for Experimental Biology , volume=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.454334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.812215Z digest=sha256:78a2f53108022feae8d4e50453826f28ea81015b0c41b935ac59b18aad09d48e

Observation 53fea1a3-e560-45cb-96f3-d204d2cfea1a · outbound

This paper cites European Journal of Biochemistry , volume=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training European Journal of Biochemistry , volume=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.437656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.858760Z digest=sha256:07f614fe2552cbe84805c4a656a1be28fcf4fe2353a029e147508610704ac7cf

Observation 224f42e4-896b-4004-aaae-8cac7f066e53 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.326734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.878863Z digest=sha256:ec0c9cf76da8d4a42d228d1e883fc8f2fad5e4d5bf6b1934b29840199fa2aac8

Observation fe57f195-1991-42d7-8e57-dddc10f14c5b · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.887253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.887253Z digest=sha256:2138e17f9ec1b57571ea53d1523e0d7f72e0d39799cd8def6bb9ba57cf11f8b8

Observation afcf2de7-4c10-4195-9f03-ad7185b12932 · outbound

This paper cites , journal=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training , journal=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:28.186855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.895086Z digest=sha256:ce97a62a3e1ca104b5583939878d27d32e26fff855ef515f324344390182936b

Observation 01d8e8d1-7353-4217-bbae-6cf437d6c40a · outbound

This paper cites an unresolved cited work.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:23:27.974523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.900837Z digest=sha256:7f72123c56cb87752e633c35559bf3fcf8b1169a1892b91a2ae25134865c9ddc

Observation 220f296b-cbe5-4803-8af7-48e389d3ba3d · outbound

This paper cites an unresolved cited work.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.905757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.905757Z digest=sha256:b1a72c2bf4515b3cbff776ff151ac1e83762c8c66ddb52fe3aa357f01e7163a7

Observation 075ce667-b95c-41ae-bc44-1f86246c8002 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Evaluating Large Language Models Trained on Code

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:26.910603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:26.910603Z digest=sha256:5fe848997add24b4aec5c0424ff3f2a79ec28eaf3f9908179d0ce4bc800d99eb

Observation 37a31f4b-bfc3-48aa-9dc2-10cd029b6215 · outbound

This paper cites Measuring Mathematical Problem Solving with the.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Measuring Mathematical Problem Solving with the

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:27.877469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:26.983319Z digest=sha256:62fc35352d614e8d75b6a017468cc34f41ef1d789be57718ff4a84c149affe54

Observation a94e6974-09dc-4b80-9d13-f51a57224350 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year=.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:27.861633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:27.076333Z digest=sha256:25df2d688cd666b7d82fcc1d732c9a260007d95ec2c38dd08fc3dc63ca7338bc

Observation 9fdf7d35-340a-4d6a-8cf9-a938314e233e · outbound

This paper cites Program Synthesis with Large Language Models.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Program Synthesis with Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:27.156108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:27.156108Z digest=sha256:5cb7a7fc6a4a922a24c230e784c6b8edab25527b1ba3571cc28766255b747798

Observation de6ecf71-5e42-44d0-a878-ea08ad374705 · outbound

This paper cites Is Your Code Generated by.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Is Your Code Generated by

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:27.742907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:27.162185Z digest=sha256:06363aacda86f324acef054d4722605e2c2a6eea18c6479fa9956475cf3132c0

Observation 5ebbaaee-44c4-4e00-85b0-e1a2c16fc801 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:27.167588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:27.167588Z digest=sha256:ff5c99e7a6b8d795ba2423296917a955e53b9eb2c3ce5519464a3f261cb6da40

Observation 4815e24e-f3ed-4b71-85e6-06cec19a1fdc · outbound

This paper cites and Ermon, Stefano and Rudra, Atri and R.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training and Ermon, Stefano and Rudra, Atri and R

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:23:27.499733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:27.173533Z digest=sha256:a0e99849691bc6d10475db97be3166216dbef6fb939e6e8fec4747aaa4bc8aad

Observation e1b7bc2a-7a62-4734-a5f1-e4910240fe1f · outbound

This paper cites an unresolved cited work.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:23:27.482296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:23:27.178674Z digest=sha256:211f77e8bd18abaadac3c8fb65edcfc72b62bc3c39a6720aad417a2829b62a92

Observation 596ff04d-1672-46f6-aabb-f8cd0c9bef7b · outbound

This paper cites an unresolved cited work.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:27.183472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:27.183472Z digest=sha256:e056252a8ba4325f096f315d461d090b1c32f719a8657808e8e20da0cd8169cd

Observation 1eaa5a27-8834-4c94-9767-111c41ff01fa · outbound

This paper cites Qwen2.5 Technical Report.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:27.188092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:27.188092Z digest=sha256:e70333d268ac2b63a04fc75137f2819ad69356708b0cb9f4b6029673d8772589

Observation 283f2359-6c66-4766-8286-f628810b91a9 · outbound

This paper cites The Llama 3 Herd of Models.

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training The Llama 3 Herd of Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T00:23:27.192432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:23:27.192432Z digest=sha256:d4df0282b407dd9d31b9964f58a4cee1a662f06df8aeb1856f9bdaea845a79fe

Pith citing papers

No inbound Pith citation observations are available.