Pith. sign in

Paper Citation Record · LEDGER

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 17 inbound Pith citation observations for arXiv:2511.07317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.07317 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:08:10.236337Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:49:38.226910Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T18:40:03.168966Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 711a9378-d43d-44ae-85ac-9eaee17f53f4 · outbound

This paper cites write newline.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.359917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.359917Z digest=sha256:b4810b29b3e5eea209f5ce47a2dd8033b6f567fd35682e54efbfc05fd67de649

Observation 41507d03-bcbb-42eb-bab7-f069ab16b6ad · outbound

This paper cites Aime problems and solutions.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Aime problems and solutions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.445615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.445615Z digest=sha256:d825ad6b2a66c0f573f5da67f947724f45940b4dc120aa3a95e3064a7ad8786c

Observation 7c049207-1df0-446a-aa0e-d0a5684a3135 · outbound

This paper cites M., Wu, Y., Powell, G., McGrew, B., and Mordatch, I.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments M., Wu, Y., Powell, G., McGrew, B., and Mordatch, I

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.518494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.518494Z digest=sha256:bb8a513bd429a11fc0218197ec71555fe52c462491cb0a30773ff7199be0b1fa

Observation d0f403d2-1bb2-4db0-a725-8eeb63a00929 · outbound

This paper cites Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.572067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.572067Z digest=sha256:623922d1285cfcccf2e1f6b05bd3eb6ba0493ba0f96cddf3e99aa6f16ade442a

Observation b04e6ac3-fb27-4529-a655-615ad5de2d78 · outbound

This paper cites Self-evolving curriculum for llm reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Self-evolving curriculum for llm reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.630272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.630272Z digest=sha256:4fbccd7a26d205023ecbeb4c07934162a0600d0acdd66cc872be9fb174dca8dd

Observation c29c4463-2279-4cd2-901a-de145fdfcccd · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Leveraging procedural generation to benchmark reinforcement learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.748584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.748584Z digest=sha256:81d6931f521ee367cc26e205979d93125f8e72ef73a4d06e99ff87fc800bc781

Observation 9535e49d-2159-455a-a9de-101bbf0ad069 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.802059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.802059Z digest=sha256:ad5ed4a00290fbd79363afd709502ba7b7766fa05d2cbe4a06bd107c01732a5e

Observation c0ebb816-2aed-4d66-8ac4-8c34f78ff225 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Process Reinforcement through Implicit Rewards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.849138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.849138Z digest=sha256:1ac524f847a8ab5e1e4d49679d072f65c8b1a2630850d9870b26b260d6e7be1e

Observation b4f9eeb1-52fa-43cf-84bf-786875497883 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.933968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.933968Z digest=sha256:a9c0b0b92f5c8981896dac2eba8f21bef880312ec3ac61356f188e89f0d93c42

Observation 83c99156-8ba9-4c91-aa52-838e1cac10e4 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.021984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.021984Z digest=sha256:f0b893f0ce04a4dddac1355392bddc5b11ea5421114d27430f97c64020b7121f

Observation dd6500fe-364f-47cb-a18f-034c9a757fcd · outbound

This paper cites Y., and Tan, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Y., and Tan, L

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.083453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.083453Z digest=sha256:0d80728f193e8e9ef97c9a53f7c23a74096c172c9e2e3872ef724cc4510da853

Observation b88fc6d6-9a7c-4967-99de-c508fd59c761 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.189419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.189419Z digest=sha256:f30fc5e0ffa444a17f295c5f62db8e35d35d6e6820a49b73ad8c0cc0d090f2d1

Observation 6241cec8-4ab9-4249-ac86-d3ee2cc9e701 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.247460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.247460Z digest=sha256:07104dab6ac8d6e62adb273643b559e4c2b9cb1250fed3789f7097855034d0da

Observation bfe257a2-376d-4514-a611-19c0b2631ef3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.337250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.337250Z digest=sha256:6f89d58fd130dcf00605f5dc57b7ad52d29f06b503fd0b32c3e7c0c2fc1422a9

Observation 63845372-c90a-463d-83d9-a1ae4dedb9cb · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OpenThoughts: Data Recipes for Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.394862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.394862Z digest=sha256:303917d6f871b07d0c1207dbdee12e8ab8910e6a4a988308aa13069aa7979163

Observation c0722e1e-dee5-4d8b-9c9c-7859f02264a0 · outbound

This paper cites L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.472364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.472364Z digest=sha256:36acad459daf49c6b2f576d5170d95cd804af75781e691c11fc7f039c5db0fbf

Observation 7eff8bab-4ce0-4f8b-b813-c430af938f21 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.560042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.560042Z digest=sha256:d500d2d9848f50cff762c1bda382e393598944c33412ee3929f2fe88d46006ed

Observation 5ba791f7-8638-45e5-ba25-680772fc0b65 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.673978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.673978Z digest=sha256:9d10b2682981bd3a3a9cc78f14f810e1328c8d45cea6148c6e3b63b15bfcffb7

Observation 42f9225f-c8f4-4bc3-837a-6100450466b2 · outbound

This paper cites Prorl v2: Prolonged training validates rl scaling laws.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Prorl v2: Prolonged training validates rl scaling laws

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.744103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.744103Z digest=sha256:daf3c2f7a0d8a713d3caea4b17aedc7cf160035524e77a430850b09dbb0f4bfa

Observation 54ca0bb3-93fc-4c14-aab0-4a2359dd51bb · outbound

This paper cites Brorl: Scaling reinforcement learning via broadened exploration.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Brorl: Scaling reinforcement learning via broadened exploration

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.823057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.823057Z digest=sha256:0413a585f7a0dc85c81bd08a7508d676c2d70115c8c55ff183c0f49b23e1072f

Observation 38fba38e-0463-4c0c-82b2-238b28506783 · outbound

This paper cites Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.895267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.895267Z digest=sha256:4f97e7b6380c937cd0075370400054a391da680a1dbaa1297703e295e4446a9e

Observation 68b9e502-9711-4063-87ce-6832ab832e08 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.978760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.978760Z digest=sha256:01d77a552ff82c62803d210f95f75d55dc96809561c88cdb04377b25a9282e07

Observation b1b10ddd-5362-49f6-9954-0259c47ac4f9 · outbound

This paper cites Prioritized level replay.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Prioritized level replay

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.061589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.061589Z digest=sha256:8ff75e228d26883db5e3b4b8b4d39fe94369c30d12d26fee03407c7892ac050a

Observation 59a2cb9f-857c-4e43-8d08-a06b88f9b181 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.138272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.138272Z digest=sha256:35425a4639eb1ed5ff8fe05c30ce738764dd160bc4d4b620450a7863f250a3f3

Observation c9b3361c-8da1-4f76-9532-371b05ce2414 · outbound

This paper cites V., Jain, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments V., Jain, L

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.221024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.221024Z digest=sha256:adf820dec2bbfc52a7df99b64eecdb023c9846686b59bcf37c3199d94066bb81

Observation 7132b63c-3286-4b58-8493-a9a9e8218e2f · outbound

This paper cites The Art of Scaling Reinforcement Learning Compute for LLMs.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.327269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.327269Z digest=sha256:e76a6ac84baa0d814ea99415eefe1c2678f20a91b1d9193016930a2ccd327e7b

Observation b1bbae63-0a3f-429b-b99a-05baa91c1571 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Kimi K2: Open Agentic Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.409104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.409104Z digest=sha256:af79e6fd75f9c47578a8f26469764b4c1a00010d2701d6ebbf3ad57da95fa5a7

Observation b0dc3d5f-5cf6-4b9b-ba7f-6fd593e0749e · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.494187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.494187Z digest=sha256:98635305d4e8eb63c19f72942049e41d4a3341b515f55903a4ebb999908e5151

Observation 84a955eb-c0bc-432a-89c3-a50398bfaba5 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.605289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.605289Z digest=sha256:b44d14d0f319cbeb8f653c2a442235be230d397a6a26ce80e5904fc07bec5b74

Observation 8b87b70c-90e3-459e-ad40-839f969c78e5 · outbound

This paper cites The Need for a Big World Simulator: A Scientific Challenge for Continual Learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments The Need for a Big World Simulator: A Scientific Challenge for Continual Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.689110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.689110Z digest=sha256:bb348734a31b81a0b92c9ee99c1c3fd28efda1e19c9f970066fbc3af5083bf5f

Observation 0c65452e-c59e-4c8d-a26f-34f68ec3cdec · outbound

This paper cites D., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments D., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.771218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.771218Z digest=sha256:55e2161edc2c49b22d8bd24b5dce837afdc876f793c3d94c2a67fa24e1b727c8

Observation 0c8e6dbc-9ba9-4cfc-b8db-a4b698fc7381 · outbound

This paper cites InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.836334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.836334Z digest=sha256:8f85f37cd669a640d071e4f923cd5897ce8466642512a10244346fc4f7006f8f

Observation ca4b7185-7809-4763-8279-75bd5f975fb0 · outbound

This paper cites S., and Jaques, N.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments S., and Jaques, N

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.902982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.902982Z digest=sha256:67ca7aebd70a73499d37853f09d87d77ab715d1379d515aefff870c4a697b987

Observation eb63c31e-e793-413c-b0f4-ef3d588965c7 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.071246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.071246Z digest=sha256:5b79b5dc9d6acb37177234425185f1cf7a1473ad9491bff76cb8c9b15bf359a8

Observation 75ec6c54-8332-439a-acf0-24e5e12fed8c · outbound

This paper cites Saturn: Sat-based reinforcement learning to unleash language model reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Saturn: Sat-based reinforcement learning to unleash language model reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.185149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.185149Z digest=sha256:285740763634e05e4c883f28ec46ceff5397c34f1ca4967c1da90cffd2839b05

Observation e2683b59-d82d-4645-8929-4126cc34d343 · outbound

This paper cites Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.393807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.393807Z digest=sha256:3fe8bafdc6d7c27bb2775ff427f0be90e19222becb693d032d989cade0b949d5

Observation f8b28bf1-c0ac-4f36-805a-8d86fc864981 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.481047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.481047Z digest=sha256:d72c99878c672f3f770409201e1927ef6d211abf61c2d9fa36c78ef99b74f4a3

Observation bb1abdb8-287c-4a71-9ec2-a779e459fed4 · outbound

This paper cites Y., Roongta, M., Cai, C., Luo, J., Li, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Y., Roongta, M., Cai, C., Luo, J., Li, L

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.569271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.569271Z digest=sha256:37d81f3abb563d2fa18dd0e77186856ac24e3ca2bd891b7a690ec0f0660626de

Observation 9f0efe4a-2ac9-4b6e-b7b6-fc352bbd9358 · outbound

This paper cites OpenAI o1 System Card.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OpenAI o1 System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.683868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.683868Z digest=sha256:37245c770d877c7d4d292da0def8a4c8ae13e7b849d9f167c68f6c4bc486829d

Observation 500fd292-5a3c-466b-b6e0-e7e7133f8384 · outbound

This paper cites Deep research system card.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Deep research system card

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.806197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.806197Z digest=sha256:726aa17f8f3f9ab0d0aa80ac3150c97e0ebd61194a7282ea5370bbd83813138f

Observation a58544e5-1854-424d-9ed9-b49781722c25 · outbound

This paper cites L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.929035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.929035Z digest=sha256:555cbdaa4f08b875bf2d3a78db6e5cffee8b7148390c86de7bec0841c1c664a8

Observation 12d60c90-3bd8-4418-8b28-8628cd7c5e9d · outbound

This paper cites Tinyzero.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Tinyzero

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.053532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.053532Z digest=sha256:c9341d1cdf16ff924a4267b1543d5b8cd330b9aeda3d432f8650bb2abeaac218

Observation 76dc2670-35e2-4c5b-b56d-266d55562eb5 · outbound

This paper cites Automatic curriculum learning for deep rl: A short survey.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Automatic curriculum learning for deep rl: A short survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.135012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.135012Z digest=sha256:47307f4b6403f17d1c77cd4075badf2a778a1ac523b8daa3fb4b95bdc1d27de2

Observation 181a8df9-d81f-457a-96c0-fecd3eda50ff · outbound

This paper cites Qwen2.5 Technical Report.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Qwen2.5 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.255858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.255858Z digest=sha256:088faf8cd84e692001ca124557045225082cef9e8afac6083ca4d9b9b45daf15

Observation dfdb21fd-fd72-49dc-99d0-2698d935ad28 · outbound

This paper cites Qwen3 Technical Report.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Qwen3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.418042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.418042Z digest=sha256:49c7d7cc5befc5023965978612ab9c17811103692f3100c95358644b5a848764

Observation 34c31c20-9cfa-44c2-b1d6-5420b9cae465 · outbound

This paper cites M., and Littwin, E.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments M., and Littwin, E

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.537091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.537091Z digest=sha256:4a4c14be31d2f92453cc2cbe4d0afb1c49a0b59a67a60a451926697a8536147b

Observation 5608e14a-e46e-4d74-8c30-9b628f50b09c · outbound

This paper cites D., and Arora, S.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments D., and Arora, S

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.623207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.623207Z digest=sha256:60fc920a596b86cd7601642d18ad1afdd60241d4da834dcb0695034dd9b6dee8

Observation 40eaaeff-fdc9-462d-ad16-c0d6842b87f3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.747991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.747991Z digest=sha256:18e8e80cf2b2b7e90553b50fecac31bf88624e1dd78cfa0a35f7210e1b82a104

Observation 315a5eb3-8e26-4444-9d11-3368db7147eb · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.872364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.872364Z digest=sha256:94cd2182fdef2478c179ae883145d6256f62134fab95be76fcdb8046f2baf445

Observation ece12c58-8a02-44be-a5bd-3f495b3650a7 · outbound

This paper cites Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.957659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.957659Z digest=sha256:8736d0302a763dcfa581a0cf84e50feb303e6c328f5c0f55be3542d3aa92bef5

Observation cde4d9cd-cad7-49f7-9f5a-6b546170dfe9 · outbound

This paper cites A., Zettlemoyer, L., and Yu, T.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments A., Zettlemoyer, L., and Yu, T

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.054674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.054674Z digest=sha256:7ff08cd3ab60bfeb215f50124eee30a6671f793cea370709be0ddd3bdd89436e

Observation ff53744e-1a34-46c0-a3cf-94671faf3c08 · outbound

This paper cites OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.137256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.137256Z digest=sha256:3873d868d1887bcec9296645204edb2d1bdc2815f8aa09060d791c0009c8a5f0

Observation a106d9d0-8c57-4f76-89e9-cca574c655b8 · outbound

This paper cites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.253737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.253737Z digest=sha256:cf3ed5b29a774731e416d050957daa3605d2295176a7bfe5424b818aafa73480

Observation 1cd2ef87-67bb-48cf-9a7e-6e047fbd4d9b · outbound

This paper cites S., Arunkumar, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments S., Arunkumar, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.362784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.362784Z digest=sha256:e0751c2fb553d6bd0c798c30d0a535332036c313c3d1af9a30cf255dfe16bedd

Observation 0f352196-bd35-47c5-9741-09d84abd32fc · outbound

This paper cites A., Khashabi, D., and Hajishirzi, H.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments A., Khashabi, D., and Hajishirzi, H

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.480977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.480977Z digest=sha256:8ae830ae8290010b518e2be55886df2f9b4cea3139dc9ae0047902713484680f

Observation e2b3afbe-65f8-42f1-941a-0635d7f30ff1 · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments On Memorization of Large Language Models in Logical Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.608509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.608509Z digest=sha256:53ae27844072361c67a48da3b0d119486ae14803da1729d9ddbd5e69db579fe1

Observation fc43979e-56b5-47c1-857e-606e250eaff7 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Your efficient rl framework secretly brings you off-policy rl training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.724972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.724972Z digest=sha256:841e6f5dbb63af95e6d2829292f4ae6f84dc8822088c6cb850c213b805adede2

Observation 32b4327d-65f7-4ff6-a194-9c1365e4c8fc · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.847499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.847499Z digest=sha256:2d514d579414354c88dfb9543c47551510f3ddf982296f90beb2e21b4caa022d

Observation c1794073-6282-49e4-814f-6ae36e646daf · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.963111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.963111Z digest=sha256:de2509ef3a3e3fa928f0bff9aae6aadc5d598a6db17fb4517ea2594902df9f7d

Observation 8cd0af95-2e55-49a1-8e49-7722f2dc7d11 · outbound

This paper cites Absolute zero: Reinforced self-play reasoning with zero data.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Absolute zero: Reinforced self-play reasoning with zero data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:10.022415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:10.022415Z digest=sha256:0c0751e1e9b9ea3cc6dca40090a56448b4a609db43b9b84076cd057bf10612f1

Observation ce03bf9c-b9ee-4b57-b1a2-cd0e63d90554 · outbound

This paper cites H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:10.134897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:10.134897Z digest=sha256:0da3faa60ea4b12aeb18eae512b7ee0fa3801549bfaade398475065ecfb16a43

Observation 316f22f8-a7ad-42db-9f99-f2ee44fe808d · outbound

This paper cites APRIL: active partial rollouts in reinforcement learning to tame long-tail generation.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments APRIL: active partial rollouts in reinforcement learning to tame long-tail generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:10.236337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:10.236337Z digest=sha256:28d30a08e14c34ebcd704c1ba83de71a4a4fa999ef9f346498e855b436e7bd0d

Pith citing papers

Observation 73b8833c-a07a-46ce-bdce-b308209207d5 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:1560f9ca92d6db75eaa9e5ac53399d9a98a942911c4332e637c8bad2adb99e64

Observation c38355a6-2f37-4df5-b1ea-47e7e3e27f18 · inbound

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration? cites this paper.

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration? RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T19:12:23.254065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:12:23.254065Z digest=sha256:cb3a045c7c064dbf5d2e1d94476776979802711c98aa0788bad7a7c81b9931e9

Observation 1a62abc8-ae43-41e0-a18b-09106001c2f1 · inbound

Gym-V: A Unified Vision Environment System for Agentic Vision Research cites this paper.

Gym-V: A Unified Vision Environment System for Agentic Vision Research RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T10:05:09.049846Z digest=sha256:37d73eb6391ef1d496e15479923b98dc2a9b831b4070ee399f28376aafec38df

Observation 54ad51cf-ae04-47f6-bc8b-ee32a8f4f506 · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:08:53.731480Z digest=sha256:2b4bec80b89fa5505b4143dc48587b5afe7450fe043f90f1564e9c47a21aa8f5

Observation d762091a-30e5-4ab7-b095-96ba5a8a2c6e · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:55:12.135512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T00:48:54.797750Z digest=sha256:65de9b4367fdb6b8cba994c4e43db677c79df4df12c7cde872de92bf6fb6189e

Observation 5141a38d-f1c6-4c9f-bf84-84ca7f3dad6a · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 222

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:5a2954d29f653268816fffea0ab5e15c790b0a76255f24c3e8015ef359bb4650

Observation 34638b7f-9db4-4b4f-a216-9f243aa5a349 · inbound

ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenes cites this paper.

ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenes RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:04.708277Z digest=sha256:bc883a4dacdc13d13d2bd90efffdc948a1ca53cdca95418fc0dfd21389fc2cfa

Observation a31563ec-0660-463c-bd50-6a643bd37c69 · inbound

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis cites this paper.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.654889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:739a0534c65cd9782f803bd0d4678cd8c04c3be252c80b240d3cfd6e35515be4

Observation def3f0a8-eae3-476a-8705-e4276688909e · inbound

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL cites this paper.

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:19.981261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:55:07.045625Z digest=sha256:d8a9e2573ae376244c37f99c7980594c83cd549bfe8fd66f35dfe7056152c851

Observation 1f8654e5-305c-46fe-9781-a682b495bb4e · inbound

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning cites this paper.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.080522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:ded2b7eab9aedc712739ca7c7408685491f8f6622fea25fbbc1d55292cc556dd

Observation c9047f27-7620-46b8-acb0-88b701d01ba7 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 263

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:40:03.170224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:6e96498a9bec6a0fe12d4a90b8e7ab995a43ef0a9817b2b60ef624848980efde

Observation 9fafa30a-1157-4851-84e3-bfeb04c14e04 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 263

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T18:15:59.080655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:6770c4c86967a12c88f4cab2005c824ac53dabccabe4b3541f3f16184659338c

Observation b9405cd7-b621-4d4b-aa7d-e2fa6c9684ed · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:53:26.544967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:fb53295a597dafc5b00db2264385bafe2a17f76231aff54ad897f63a1ce29dd6

Observation 55d954bc-7085-45db-83f6-904ce94409e6 · inbound

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs cites this paper.

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:59:51.886210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T04:47:47.691913Z digest=sha256:7e3efd64ea6e3d67ecb48503201c68ed6be5f10c59abdeb82f04b81a76bfb49d

Observation 88541c9e-7c6b-43db-ac95-2fa12382a28d · inbound

SETA: Scaling Environments for Terminal Agents cites this paper.

SETA: Scaling Environments for Terminal Agents RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:39.248762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:39.248762Z digest=sha256:fe0695d717dff0d362b7ec9db4cb1b9503412b28935ed1efa2daf25c023c5d72

Observation 7d008aa5-dedb-4519-bd8b-fe0a4c2c5431 · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:10.993520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:10.993520Z digest=sha256:f70ea33029cb15804e6fa09a78d631f5ef3ca24855fb9d694eaa118e8ac890c0

Observation 03046ac1-6a00-4528-a8a9-70b08d86f949 · inbound

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning cites this paper.

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:49:38.226910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:49:38.226910Z digest=sha256:eed7580b66c8d7fa7112903a1521bfbf845c35005e3ec60d6646f0ad1cc261e0