Pith. sign in

Paper Citation Record · LEDGER

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.18830.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18830 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:34.013794Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:34:06.517190Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T15:16:09.090199Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2ffe78-a13b-4ec9-8192-4f7d15c71d18 · outbound

This paper cites Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.082862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.082862Z digest=sha256:476bc31319a4544d48d5620b62bb65bcf3de0777305b1f17c8676679c9503eef

Observation 3519a9e7-04f0-4dbb-8464-61e670ba6ace · outbound

This paper cites DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:29:34.477945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:30.182581Z digest=sha256:597010b9048961cda3484de7f2ee120e4c91e2a9cd42a1133fa495833998097b

Observation b8da91cd-c47c-4176-8984-193b2cd7c071 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.352756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.352756Z digest=sha256:88d82ed5663e5f5933f559120af4aea0708380323ac023cbcfaf68f1eac0ee6f

Observation 0086d934-0735-4c00-b015-e426cf4b1793 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.472429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.472429Z digest=sha256:73f8c5adf05586e2f9229859af9553f4395e663e6064eed40df5b252e69db734

Observation ac04d9fc-c2c3-43dd-8890-51ecedb29804 · outbound

This paper cites The local elasticity of neural networks.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization The local elasticity of neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.731705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:30.598230Z digest=sha256:43e06290812ef9f2ed44ce4d3e7ea83e4d9b99ed83ec3a9e137313d9a4804fda

Observation 4e9a462d-48e0-493d-a2e0-c32b91e59503 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Measuring mathematical problem solving with the math dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.787987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.787987Z digest=sha256:8c30fdb7bcb76bb6f7fc2be8b6c6031f7fcbbcadcb458101b046fb83e85a6624

Observation 01ee2673-2467-4663-a4dc-fe0a0a5ef0a9 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.979686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.979686Z digest=sha256:c3ae130f283ccc0363ab93b0de0dc6f77cdea817251cf8a28d61e32a4b9c665e

Observation e5e4e96a-22cc-4bcf-a91a-88ad472cf7c5 · outbound

This paper cites OpenAI o1 System Card.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.111041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.111041Z digest=sha256:8f5f91f3282cfc48204e74f67c1a5e27190053b073f04e1337c578ed9ece01fa

Observation 2e95c004-02d1-453c-956a-8e17c1eb7a31 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.259139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.259139Z digest=sha256:69d881887db5150f9d3026a225bc738a5d0ce0ce6bd9723423360d9722796397

Observation 0213d9c1-098d-47b5-98e7-bc3263b303e9 · outbound

This paper cites Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.370676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.370676Z digest=sha256:64cfdcb55a5f7a37039fe51aebced91b102dd427fc517323cdbebe90abe07172

Observation 287758ef-d131-493f-a2a3-6fab4b88f346 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.497301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.497301Z digest=sha256:c0de8ae57ce2af4861eb1c6a1e954fe5a3b205438159364f191e9fc12c069fa2

Observation 9138a9b8-c697-454e-9c8b-b2c04ae9b2f8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.672136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.672136Z digest=sha256:8eeccfd9f8149829b43601fc560f694b77d08e031aecb6c095b965ceeab17f55

Observation 4bbda4f8-5219-4a63-b4d3-aca75d7097e0 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.844838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.844838Z digest=sha256:67edd6f35741ecc2e6a7e35332ff3134b11168b2337dfc5279f2834cd8ec707b

Observation 65cf08cb-4b81-423b-9bfe-765f73129ae8 · outbound

This paper cites Neural collapse with unconstrained features.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Neural collapse with unconstrained features

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.585539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:32.001012Z digest=sha256:8781b055834dc1ee0fd757ba554d31e4dbdd9396dc3d4bf34da580bf03e9172e

Observation cc9af7d2-8533-41df-8f1b-6f9a67183a9e · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.198096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.198096Z digest=sha256:68b7e982f434836b444551b1756ad71a6522f7aeabbc84de24db535df2c80dbf

Observation fd6b3d69-c9a3-4268-8dc1-340694b26bfb · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.334464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.334464Z digest=sha256:247a10bbf95182be4ba08f549c3d14b7af6d0fbe8adc6b8cb0b73833e755bc9e

Observation d3f472fa-05a6-4984-b680-13d8fa5b52de · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.444781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.444781Z digest=sha256:858abac3f817d94687e0e3b6b4099ffefb590cbb1779c64f9a169c68635fbd42

Observation cb0aa861-d8ce-4d6a-9d29-a536550589d8 · outbound

This paper cites Learning Dynamics of LLM Finetuning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Learning Dynamics of LLM Finetuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.572317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.572317Z digest=sha256:f09ace80d43da3ef59fb57483e5dfb371e2eca739054a5e65bc8cd6a2c8026de

Observation 41608882-f242-4ed3-8191-4bc586c69a70 · outbound

This paper cites Proximal Policy Optimization Algorithms.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.710722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.710722Z digest=sha256:8dad55a161443eb9ea453fdb633d0b81f4d237dc0faf03b906f1c3e61de8257f

Observation 726ede61-1ee1-403e-9235-575fa2c83438 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.815490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.815490Z digest=sha256:ae5935bd8d343fbd3886807ca5b9e3d73774b9c7f28190f99f922e1f5165e85e

Observation 01b333f3-1f0b-4378-a7d7-4d48e3286691 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.921726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.921726Z digest=sha256:3e07ddeb5625be8b1adcf00e33558203f3707a1e3253e571f3858aa03b8845aa

Observation 96dff3bd-5d64-4f4f-ae3c-5372b0de278a · outbound

This paper cites Aime problem set 1983-2024, 2023.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Aime problem set 1983-2024, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.509164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:33.000764Z digest=sha256:7f550d2f406fa671c57b8729ad513a9bcf6c7293aa8533a4eb8dbb36f9eb1ae7

Observation 21575429-1287-40a5-b8cd-f7afa7856141 · outbound

This paper cites Qwen2.5 Technical Report.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.073364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.073364Z digest=sha256:1b0bed87f3d8b0e2f1af1f88229a5643e49417b8f400d4c96af1ac0bf2807d38

Observation a15f090f-5325-4091-b01e-7fe9210c3d57 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.120361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.120361Z digest=sha256:437cde8bd207e59239ba600a47fe85cccaafdde5272723d29942e355ab416f42

Observation bd1349c0-d83a-4598-8710-ddbb685d6d35 · outbound

This paper cites Breaking the Softmax Bottleneck: A High-Rank RNN Language Model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.191566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.191566Z digest=sha256:039cf2973fd5d4cda34a5f0619e4b508108e3e30d61c229740abf986f04eaf61

Observation 86e8509a-9d1f-4541-8271-238cb2c324a4 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.267870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.267870Z digest=sha256:d65a51c39e001a932ce120816470e1bdedc7ce182224f75afbc0db022f242c5c

Observation 1e7a4dd1-a25f-4783-b71d-bac8905ba860 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.323315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.323315Z digest=sha256:f97b3b03abb028fc0b4c1fe1bce569379f9578ba8e39570e636d4fb5ebb41372

Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.431718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.431718Z digest=sha256:f446f771dd2f44fd8e9c1690f46abd70c4e9ed50d28ddaa018adddb7af26e6dd

Observation b04807af-9df9-4bc3-b30c-5558edf2ef4f · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.492988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.492988Z digest=sha256:bc5595c6eddddd8db480fd29b15fed6a04a46fd2c396926fe538a66813e5a358

Observation 4177ec77-2d36-4819-af22-815479f91d08 · outbound

This paper cites Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.608339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.608339Z digest=sha256:98e59b34811a0a2e2cba05cf7819d5229851a1db3b72db1f6208eb7e4fc8084e

Observation 93de0474-6ead-4896-9f95-87061801dad4 · outbound

This paper cites one commercially available ten-button lock may be opened by pressing – in any order – the correct five buttons.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization one commercially available ten-button lock may be opened by pressing – in any order – the correct five buttons

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:29:35.313069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:33.696512Z digest=sha256:adfe0080a4bf5df39861353c917ac6eecac0562e755d0e2b21009bcfc4205091

Observation 37e0f6f9-5bc9-40c9-95f9-8c882721387e · outbound

This paper cites From the given graph, we can observe the following: - The roots off(x)are atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization From the given graph, we can observe the following: - The roots off(x)are atx= 1 andx= 3

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.094259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:33.803416Z digest=sha256:a9cdbf40147e84faa8e157de5cfe20a4a5ad7631dfa13c7820308d1f9b35c692

Observation 1ad0abfb-49ad-4c55-80ac-4b4a33f33c88 · outbound

This paper cites Therefore: - The roots of g(x)are also atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Therefore: - The roots of g(x)are also atx= 1 andx= 3

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.956361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:33.875782Z digest=sha256:459978b78cc0bc5b1aa1ed7881b174b3b0d9cb5874dcabfc2aa6f315c5491ea7

Observation 7ba4801a-f4b9-485c-8328-5cb5b74f5928 · outbound

This paper cites This simplifies to: f(x) = 0 The roots off(x)are atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization This simplifies to: f(x) = 0 The roots off(x)are atx= 1 andx= 3

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.870775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:33.935421Z digest=sha256:d1c818442849e9ad342c6408712b7c42399f555fe7891c1c655056e47f2eb806

Observation d34fc82a-47eb-4b63-883c-158d98591680 · outbound

This paper cites This implies thatf(x)is an even function, and its graph is symmetric about the y-axis.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization This implies thatf(x)is an even function, and its graph is symmetric about the y-axis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.738641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:29:34.013794Z digest=sha256:76760cd3019381347a296b44ac4c8d0d1e527b2dc08cb97b2813f4189146c117

Pith citing papers

Observation 151aac32-085e-47ea-a121-e760e2719af9 · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:06.517190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:06.517190Z digest=sha256:7ae2309da05d3c77c7ddb0885cc9f2365a0047e8d871ea6c857d9443743e65da

Observation 89a73c6b-0b37-423f-a8fd-f01398c0874b · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:09.093200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:117c2daf0ea04aa6712b4661044da64138f09ac1f953ffff6b541d99c31becf2