Pith. sign in

Paper Citation Record · LEDGER

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

As of 5 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.03972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03972 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:53:13.195296Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact6
  • verified fuzzy22
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71316cb8-5dfe-4eca-aeb6-84068d6591a4 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.524151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.932970Z digest=sha256:5cf27f5f013ba5b75097ecc37ab38cd96e3469a34e3cf258c5683d1f379d8fb2

Observation a6cd4b45-4225-435a-9331-cba37575e76f · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning DAPO: An open-source LLM reinforcement learning system at scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.503525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.938125Z digest=sha256:ac39b06930974c378df6553d60818d7ac9cb7a44779294af44e2f0e7bb14792c

Observation 8dd9ea9a-6b98-4d28-9433-78731f14513b · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.942968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.942968Z digest=sha256:fe9310cbbf1166270444277e476f6a0d54ce6f9ce51fa7a86adedba611c26a6d

Observation b9f173e5-5435-4412-8d60-8680b15edc49 · outbound

This paper cites Loong: Synthesize long chain-of-thoughts at scale through verifiers, 2025.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Loong: Synthesize long chain-of-thoughts at scale through verifiers, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.484558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.947974Z digest=sha256:344d5ce6905903df7ac48e31962c02b559ccae6683b9222ec297710301338f29

Observation 676f668a-52ec-46bb-8205-8d92a83901b9 · outbound

This paper cites Reinforcement mid-training, 2025.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Reinforcement mid-training, 2025

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-05T04:53:13.931848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.952698Z digest=sha256:c5865d8a77a03a076a71ff8d4a597582f33877c02040ad3ee40059beca127c77

Observation 44f0d7fa-e63b-4d3f-95ab-b22290f112a8 · outbound

This paper cites Self-evolving multi-agent systems via textual backpropagation.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Self-evolving multi-agent systems via textual backpropagation

Reference 6

Resolution
verified exact
doi, observed 2026-08-05T04:53:13.269181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.957456Z digest=sha256:84af89b73728cecf83291e68e6e14a6afcae39a3ccf19f606413372185c9a62d

Observation 67a31706-7647-441c-8dc1-7555dd841db5 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Connectionism,.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning On-policy distillation.Thinking Machines Lab: Connectionism,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.967025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.967025Z digest=sha256:6752d74137bd6b15377c0f30de2e7a396781260413901730658f7216490fc5e9

Observation 08623710-4d5f-4471-9ed1-b85223d145b3 · outbound

This paper cites EchoRL: Reinforcement learning via rollout echoing.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning EchoRL: Reinforcement learning via rollout echoing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.431786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.977027Z digest=sha256:ed60f97181333722ac594d1cf9dc661bbb2a3de736ebd3960eb10a7373d2c0e1

Observation 52a20fe1-d4db-4942-bd10-d5997539c958 · outbound

This paper cites Learning to reason under off-policy guidance.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Learning to reason under off-policy guidance

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.403651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.981914Z digest=sha256:bb97e41468c703f8374a853081708efc8515a9656d40d64720f385f33252f34b

Observation b0ef0794-26c1-4934-9d36-ef04fc58c1de · outbound

This paper cites Ozdaglar.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Ozdaglar

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.384926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.986788Z digest=sha256:940cfb461bb18db45154d5b5bdd1e8ad5f856b9b9e52952edd01a8feba15dea2

Observation b85e76e7-0456-4d29-bc9a-c56812d5dcb0 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.991139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.991139Z digest=sha256:2f398dc4b5a29558382b84bf9dc818e4c116d8147eadab9026695827c91eed03

Observation 3e10b2e6-8c8f-44a6-aa55-41da810147f7 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.995759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.995759Z digest=sha256:351a83018bb4d564594922e5ac6edb388fb2ede7244e98d4e9a3cd1b5aaf4052

Observation ed2e493a-78a5-41cf-b3d3-c6115491da56 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Measuring mathematical problem solving with the MATH dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.000441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.000441Z digest=sha256:c8e110eaf22e656bfb4b4dc45143f2f084462650b9b9aa761e7e3c5c39bf4087

Observation f3f5af37-516a-4256-a71d-ad4ed9033892 · outbound

This paper cites Solving quantitative reasoning problems with lan- guagemodels.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Solving quantitative reasoning problems with lan- guagemodels

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.356660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.004741Z digest=sha256:1e1c8e992aa140b0f9c559337cecea97822db5f774dcaffedcb72cbbbfc88aef

Observation 9a4ab40e-1c4d-42b2-abc4-82df9bd4935a · outbound

This paper cites OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.009142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.009142Z digest=sha256:892118304b26f846a30ac1b4a6248f5ed76257baef583c439ece0dc498ddbdd0

Observation 459ab9fd-1d12-4713-a675-b2c08aa7424b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.014357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.014357Z digest=sha256:d4f53686da158046136412752040b06bf0569271a800ea6c3ec1a55a8ad95a7f

Observation 1112e77e-80b2-45ad-b511-b36b7e5fab67 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.019460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.019460Z digest=sha256:dcb3d5cca215252696bc3e2333cead5c06193c8ebece462cae61c4ca98ae438e

Observation efe2d354-b221-4331-b963-3c3aa75c560c · outbound

This paper cites MMLU-pro: A more robust and challenging multi-task language understanding 12 benchmark.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning MMLU-pro: A more robust and challenging multi-task language understanding 12 benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.319942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.024162Z digest=sha256:a9335ef88b3360f08b39b38277f1ae1ab74499e6410533cd96b0e22a423224d7

Observation 8f65f649-835a-4863-9ab0-ad4026eaee28 · outbound

This paper cites SimpleRL-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning SimpleRL-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.303448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.028821Z digest=sha256:154663c383ad25b0084140d66a639cadfaf468e7293aa9b051c3b06c203a535f

Observation 6ebc4b3a-56c6-45b5-84ad-3dcc8aa093ca · outbound

This paper cites Open- reasoner-zero: An open source approach to scaling up reinforcement learning on the base model.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Open- reasoner-zero: An open source approach to scaling up reinforcement learning on the base model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.286659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.033764Z digest=sha256:6eff3e96e1fc9bb8b5f46f086493e42400e6179526f3a994663322d75857c5d2

Observation b6cf11d9-1ffb-4e6d-b7ca-4fd5826cd031 · outbound

This paper cites Process reinforcement through implicit rewards.Transactions on Machine Learning Research,.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Process reinforcement through implicit rewards.Transactions on Machine Learning Research,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.268959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.038050Z digest=sha256:49f4ac3760bc33e1df99abc9af59accaff7ed55e335ca0bdd9babc2ddf86a137

Observation db9f4ef3-9ecd-428d-9fed-a5df98748910 · outbound

This paper cites Qwen2.5 Technical Report.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.046874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.046874Z digest=sha256:53f2e4ba4911363c83fd9bf3ffcd6e383d96e6173d83f4325f5935f8e453196e

Observation 8e59e571-48aa-4a9a-8545-fbb844e7488b · outbound

This paper cites The Llama 3 Herd of Models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.051552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.051552Z digest=sha256:27348b252089682bbc74d7dda71b5d8e17b150308ee69dd8c65106a2972b851f

Observation c97db0fe-842d-4390-928a-c1a0eee4a97a · outbound

This paper cites Think outside the policy: In-context steered policy optimization.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Think outside the policy: In-context steered policy optimization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.237674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.057291Z digest=sha256:88e03664703bfbad3c04f90706e7ccf67e12a4042bc9d31c257c30778a5dde0f

Observation f4754bd8-ebb5-4826-81cc-08f8df5b9e12 · outbound

This paper cites SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.062069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.062069Z digest=sha256:42d5b245a97b967be756607c758a0e4e195bc7b62e42c13ef907c8fd7b187d79

Observation c6f71012-67b5-4f5f-8950-493008b635b8 · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.066820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.066820Z digest=sha256:318c9e1bbb209810dc4feeff8f938a3f645cb46201d1f50341769a752a8e3c9b

Observation 4113cf82-1aa0-4f83-a1e1-30163a41d3b3 · outbound

This paper cites Don’t overthink it.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Don’t overthink it

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.220813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.072079Z digest=sha256:5cfba109f49c6ade4810af4f5333db77266ad5ce36a0686559a272f2861a68fb

Observation aa62ebab-5ecd-4e67-913b-e590d9786438 · outbound

This paper cites LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation- steering.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation- steering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.199622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.077044Z digest=sha256:5a27050d9856e720a3a2c0c4727f19aa067baaafe39b9ab3b3b030a602fdccf5

Observation 00d862b3-3b99-465b-8966-3f520b032d86 · outbound

This paper cites PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.081608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.081608Z digest=sha256:afb04c53a613ef6ee4958f1e014789c6264dd3b5c9ed80a35ce0a743bfda497d

Observation 11290c47-93f1-4ba8-a121-b8c902397195 · outbound

This paper cites Beyond nl2code: A structured survey of multimodal code intelligence,.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Beyond nl2code: A structured survey of multimodal code intelligence,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.169785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.091539Z digest=sha256:3f0e242231ef3637a2ece693f78461219094f21aee5860d5b69a392c7107e95b

Observation 3cf0b7a4-e5fd-4080-a050-efd75a89652f · outbound

This paper cites SPOT! Revisiting Video-Language Models for Event Understanding.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning SPOT! Revisiting Video-Language Models for Event Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.101410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.101410Z digest=sha256:23d90b2db29d86ff02ce40dc09d9afd58079d80452d2dc3a30fc8426c8348a6d

Observation d02da753-95da-47eb-86c2-22a74beb1c38 · outbound

This paper cites Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:53:13.646604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.106954Z digest=sha256:a1f0c93cee617b3ce9f81a14c322264483f7cc0a863aad7c4b556b21dcbf1c5b

Observation 614957f0-49a5-4284-9ca1-8e36fb0678d9 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.086535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.086535Z digest=sha256:f20c759a7ffff2bdb6b7322cdd5a44dd3337221d14bb0328449597a518b3565c

Observation caefa43e-66b5-448b-9654-1b652cf39833 · outbound

This paper cites KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.116166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.116166Z digest=sha256:22ec2901db83e7a2fd57bc27b179e1f7774e0548fecd935353d2a14fe2ff6c10

Observation 641dc76e-2364-4908-a5ea-84cbcf94c2f3 · outbound

This paper cites Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:53:13.690777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.095917Z digest=sha256:8a1884e491d38f3c29eda2f36f63392aadf5308210d71bc803613e60f0eaccbf

Observation f9a3aeed-c64c-4e64-a40c-744ce16e7669 · outbound

This paper cites Aditya Prakash, Yizhou Sun, and Wei Wang.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Aditya Prakash, Yizhou Sun, and Wei Wang

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.125098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.125098Z digest=sha256:1d001cc5d99ac2c63121cc90aebcb9db739597c6918bad51c05f4c096907873d

Observation cd29fcac-7c1b-45c6-bf24-4b2c68e8208f · outbound

This paper cites HYPERION: Fine-grained hypersphere alignment for robust federated graph learning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning HYPERION: Fine-grained hypersphere alignment for robust federated graph learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.154663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.129286Z digest=sha256:735480e8a001b3819685e752350ca30b2ad968126ed67b8757dfa66d93cac513

Observation 86e77089-5ef2-4b20-ad58-6bfb328ad5cc · outbound

This paper cites Alignsae: Concept-aligned sparse autoencoders, 2026.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Alignsae: Concept-aligned sparse autoencoders, 2026

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-05T04:53:13.621870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.111636Z digest=sha256:1703e3036365267605212002ab3802dc13cad90f829d85f030bb8716f2d96cb6

Observation bf1166f2-560f-4ee7-86c6-b287bfd71e1e · outbound

This paper cites Can visual input be compressed? a visual token compression benchmark for large multimodal models, 2025.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Can visual input be compressed? a visual token compression benchmark for large multimodal models, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.138245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.138245Z digest=sha256:c3ede6be494dbc42b842da2040011f1d7c3e517192dafaab073e20d4ae895ac1

Observation 33ee0171-b0d2-45d6-a8c8-35bc4f89def1 · outbound

This paper cites Ascd: Attention-steerable contrastive decoding for reducing hallucination in mllm.Proceedings of the AAAI Conference on Artificial Intelligence, 40(12): 10306–10314, Mar.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Ascd: Attention-steerable contrastive decoding for reducing hallucination in mllm.Proceedings of the AAAI Conference on Artificial Intelligence, 40(12): 10306–10314, Mar

Reference 40

Resolution
verified exact
doi, observed 2026-08-05T04:53:13.238886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.120867Z digest=sha256:00cfdfe656b0a370628642452c6407b598347cef498c445cb2475dd4039a3c3a

Observation 9f59fc7c-c5ff-4cb9-9a09-f4498dc40694 · outbound

This paper cites Let’s verify step by step.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Let’s verify step by step

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.121834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.148046Z digest=sha256:922cd1be60b151def8cb0e37c9396c407b2487369f1d36166cd9d01534244d08

Observation 7480c158-ceb8-4a12-b726-2534a5dde6e2 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Star: Bootstrapping reasoning with reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.152614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.152614Z digest=sha256:8e00e5996635eae8908fff5063e015489be34a04b5ecf9d24064ccf6050e8743

Observation f23d5009-9e53-4424-922c-c96b54579004 · outbound

This paper cites Backdoor cleaning without external guidance in MLLM fine-tuning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Backdoor cleaning without external guidance in MLLM fine-tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.138677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.133448Z digest=sha256:4c128c40c351afada16ea616c64243a3c1cdd011dddea56557881bde70d26fed

Observation d32c9b7a-0a3b-45a1-8917-2b78a3643fa0 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Minillm: Knowledge distillation of large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.074366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.161276Z digest=sha256:081ad2f029794ef2256376a65c5eddf14474b33a8567978b17b9aff28dded2fb

Observation a51bb7c4-78cd-4b2f-99bc-40b975e53ee3 · outbound

This paper cites MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.142470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.142470Z digest=sha256:60032f4a1cefe6029418228975f4793c180901be97bf5ad108e17a9c9cbfcc6c

Observation 74b080b8-944d-4d8a-a72b-534a39c68219 · outbound

This paper cites Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.092400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.156977Z digest=sha256:d3556cd17979218e44cf5da5a66fbd41ce8f31caccf1f85e09ae3f4d95859f42

Observation 76c6f394-d6e1-4557-8a99-8089f7dcf0cb · outbound

This paper cites Generating Sequences by Learning to Self-Correct.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Generating Sequences by Learning to Self-Correct

Reference 50

Resolution
malformed identifier
no resolver link, observed 2026-08-05T04:53:13.165423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.165423Z digest=sha256:a7cadc9f32eacce1afaffb992133027883b1cdef7ee7d011b04f21fa40644219

Observation 4a368c62-a29d-4e6e-9f93-ecdb5d9c6a58 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.057496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.170212Z digest=sha256:ba317e529373419f56fd7847822f6609daa2ac04fa87e5fb53f243c259952a71

Observation cc02cac2-914a-4afa-830f-3f0e1a453564 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.039323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.175101Z digest=sha256:77a1a872873e4efc88fe59e08afc6eb2af1e662cb1432464e8686d37a62e9b66

Observation e82dd163-107b-494f-8a7d-0843a8f6677a · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.021601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.179774Z digest=sha256:de99e786010be59b8ce9c52f2dc21f040057a3d8916aecfdb8905fd44d4b4f38

Observation 615e3fea-8876-41e4-b695-975684456f7a · outbound

This paper cites 17 Figure 10Case Study: Reflective Reasoning (Success).

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning 17 Figure 10Case Study: Reflective Reasoning (Success)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.002582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.184559Z digest=sha256:a427598a673546eb078cc2cb8e4150829f430b55814be6880934bca93682e1b8

Observation 165e5c50-9541-4901-b645-34bffda28b24 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:13.985398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.190617Z digest=sha256:ec43dac58c87468df6f2fc1054a4d6124660a16a1a8102c8424676ae89e60dff

Observation 185dcd77-df9f-48aa-b89b-249267a1f57f · outbound

This paper cites The reference solution suggests there were some errors in the previous attempts. Recomputing:.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning The reference solution suggests there were some errors in the previous attempts. Recomputing:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:13.968802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.195296Z digest=sha256:da32182f426673a7cab63ea3c63ead55682700ed131b46ab5c2d42f7edbe9974

Observation c1872307-fff5-456f-8650-4dfb8e02e1e9 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 483

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.461761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:12.962737Z digest=sha256:389d9185f64450e20f35a8b5145f8147a371105cdfec93c3f5a462117c281e3d

Observation edd914fa-d6f8-45f0-8048-8e58b2befa49 · outbound

This paper cites https://thinkingmachines.ai/blog/on-policy-distillation.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning https://thinkingmachines.ai/blog/on-policy-distillation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.971795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.971795Z digest=sha256:f46a446d9b09f95a8032eae41ac34bcaffe2dbd173b8c148f799d5189037cfff

Observation 3fa1c871-1e68-4fbe-a32c-6e00ee3212e2 · outbound

This paper cites URLhttps://openreview.net/forum?id=9SkkifLopZ.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning URLhttps://openreview.net/forum?id=9SkkifLopZ

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.253295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:53:13.042559Z digest=sha256:3f8f963ece36b44e782a2c77190f854568a7d1c0cd06370207cb4eeb0b5a415d

Pith citing papers

No inbound Pith citation observations are available.