Pith. sign in

Paper Citation Record · LEDGER

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control

As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2605.11775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11775 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T06:04:32.640299Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact30
  • verified fuzzy36
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 557f233a-4fee-4e72-8364-09c3b847d187 · outbound

This paper cites author=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control author=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.178340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:b017924e15f1e67f7c433f5a7d1a3249e5aa9dca631ab3ebab43b01aa45c68ab

Observation ce5fddca-35e7-41b5-8be1-4d5731ef2474 · outbound

This paper cites 2026 , url=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , url=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.298290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:4e06439fff4b0f5a35407e7687898b30eed57a295d965cb078de32336355b20d

Observation 313e936e-be72-4dfb-bbb5-d37f5199f9bc · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:05:06.381914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:565c81053b78376155050361767a74682e7fe74062c183e6f36807bc0b63ac4d

Observation f873a061-5c55-4445-81bc-85c975667337 · outbound

This paper cites POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models , url =.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models , url =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.355452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:1fafb76366be3b777fe82e92300c5727289c764d5c542dcec9557ea5cdb69694

Observation 0a881924-055b-4ef7-8808-56063c53e986 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.308112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:178d6e5158d3ab94c596622187e91aab3420a1296ea55287bc4cfd9cfbfe0b87

Observation 182b54fa-f5fa-411b-a846-b5f255b77e46 · outbound

This paper cites Entropic: Towards stable long-term training of llms via entropy stabilization with proportional-integral control.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Entropic: Towards stable long-term training of llms via entropy stabilization with proportional-integral control

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.683882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:6ab900fa243ae17f13d3923788800f8ebf84d989aac43565c82d56e7b99d14b4

Observation c06b9d68-3e1f-4e47-91e1-e9bd88e40a56 · outbound

This paper cites 2026 , url=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , url=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.322782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:c650b35d6e6b86fe6f78460535f9c4f60ba2113f0baf44254c5add0f0d12b3f9

Observation a4d55725-1f74-4e37-bb09-f1f4a64a2498 · outbound

This paper cites Beyond Magnitude: Leveraging Direction of.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond Magnitude: Leveraging Direction of

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.236941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:32f505da90ed022ec9c92ffddd32423bf6a08ececdb7b765f47b92daf91d5c1a

Observation 37c4a159-06da-4d24-ae81-5085b1039e86 · outbound

This paper cites Sparse but Critical: A Token-Level Analysis of Distributional Shifts in.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Sparse but Critical: A Token-Level Analysis of Distributional Shifts in

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.313005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:d04d7d227ea0c0c42334eead40a2568e5cd02c0dfe85ca04a924784aeb0db519

Observation f438d893-6709-4658-b4c7-667725a7f3e3 · outbound

This paper cites 2026 , url=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , url=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.262613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:64593723ca0117844e3d72c428a1d0cdda383b11a6f7fc854802aa39584a6458

Observation de73cd91-e00d-45aa-9b59-b679f5d5985e · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.205264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:4211be9fd0aa51ba610c655f001ab33061db0b498c037f00a7e10ee5c92366d1

Observation 050d202a-b668-465f-9724-5a4ab47c1c86 · outbound

This paper cites The Twelfth International Conference on Learning Representations.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The Twelfth International Conference on Learning Representations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.221413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:711ccbe2394d8deb5dd2201b7cb741c8d465b3881a577fb9e0672003310c1a0a

Observation 3cc92413-f495-4bb0-8abd-edd46d9e0635 · outbound

This paper cites Hugging Face repository , volume=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Hugging Face repository , volume=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.210876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:18be3b98205e9d3ecf781f01ffa4cb7a661609d0531f61c2817d8c61b12f433c

Observation 9896f1ff-195d-422b-a268-ca62a66e184c · outbound

This paper cites 2024 , note=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2024 , note=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.272200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:7b8a6c59a73d3c1caa3fd2c4b63f70b42796b51eecb088bc878278276aa109cf

Observation eee99b74-a6f2-4a57-82ac-4634fb304340 · outbound

This paper cites 2025 , note=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2025 , note=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.328451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:872790db9f9d36ef3cb21962daf6e753e0c7617502752e161ef3e78c1a92c746

Observation 99306081-5c8e-40e0-9783-22c0df46b672 · outbound

This paper cites Advances in neural information processing systems , volume=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Advances in neural information processing systems , volume=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.195031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:acc3aee66ac91a9a0518bc03e0e9277b81a30bcd72c4d8f52e0cd273b05d7f85

Observation a18c04ee-9b41-4c31-b9e3-55586f324673 · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Forty-second International Conference on Machine Learning , year=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.318024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:c61cee60da892d5a9a5d94f30220af2d2a78c52483c45d1e0554ff1a742ea8cc

Observation 2bf86b94-26f7-4b77-9f91-79cd7260367e · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.257769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:c3b0ea66b5b4e47d3ff32439bde9addcecdc7784dd04822ed74507ee55547258

Observation 1a138f03-af49-4f6f-965d-7b6285561c74 · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The Thirteenth International Conference on Learning Representations , year=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.199878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:8cb280ab0f8b4df269869e61ceaad53ac8a43c021b0ec941e33dc27898f4d86e

Observation 92fb0707-0513-4087-9512-cfdcd102e3a5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Advances in Neural Information Processing Systems , volume=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.251816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:2d8db3dd41f32d5658b098289a3aec199b9f475b15759ca1f83691a97d49fc79

Observation 2ba9d8b1-cb94-4556-91c5-386517159698 · outbound

This paper cites On the direction of rlvr updates for llm reasoning: Identification and exploitation.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control On the direction of rlvr updates for llm reasoning: Identification and exploitation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.676421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:0f6068d0fe4844171ec071100761653d09b391f3d0b851d1a1156b051a4df74a

Observation 684f3e3d-be52-4585-89f5-ef1409597fbb · outbound

This paper cites Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices , booktitle =.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices , booktitle =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.189826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:e00b581f34c924fa9a419e35b5edd2d03962866e9b6de075eb4744b034e03033

Observation 5ea532d5-47c1-4dcb-bce0-9768d4bcfb0d · outbound

This paper cites 2026 , eprint=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , eprint=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.246248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:a7846dd2eb1fa762e72859f01cbc021d4fc9ccd26925406af974d7d567857fed

Observation c407257f-d74f-49ac-ba54-04b7e24e2575 · outbound

This paper cites Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.216231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:f7625a63b8dbc8a6693393359685e6ce51141b7a1e2357f3f8984e32792f1798

Observation 1920804f-4cc4-4079-8e3f-4b55f205bb76 · outbound

This paper cites Claude code.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Claude code

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.302944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:6b4df6c25cb1189831bdb98fa9d22c4e470381ec968f25e0e1f4dcc90f9797e5

Observation 677e0d37-8cf0-4496-8131-0cc646a28049 · outbound

This paper cites TTVS: Boosting Self-Exploring Reinforcement Learning via Test-time Variational Synthesis.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control TTVS: Boosting Self-Exploring Reinforcement Learning via Test-time Variational Synthesis

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.375698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:e4eaeb2f56d0075aeaf519a05d28a4dfd444ba3c18e0f0aa0f8b7df26ab6cc7d

Observation 927c14cb-bf21-4c27-a24b-109b5a3833f4 · outbound

This paper cites Acebench: Who wins the match point in tool usage? arXiv preprint arXiv:2501.12851, 2025 a.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Acebench: Who wins the match point in tool usage? arXiv preprint arXiv:2501.12851, 2025 a

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.405209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:b7c65ff1da69f06a16b929810e757376bec56133a22acab42e26d15f8f855445

Observation 5843447a-5e6c-4832-907c-262a029d902e · outbound

This paper cites Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:05.721920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:1d050756dc1d581799148a12a4efe13f474422cc8984b08b230d8bd41a9ac8b3

Observation 5643cca9-9766-46b1-a244-d459dea31216 · outbound

This paper cites Exploration vs exploitation: Rethinking RLVR through clipping, entropy, and spurious reward.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Exploration vs exploitation: Rethinking RLVR through clipping, entropy, and spurious reward

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.599968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:8672fecd79dc4d8620e749b0ac6fa10c0b83943489d23c74818ff17666916cf0

Observation 904c25f1-aa07-44cd-abb0-ccf7694504e7 · outbound

This paper cites Beyond high-entropy exploration: Correctness- aware low-entropy segment-based advantage shaping for reasoning llms.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond high-entropy exploration: Correctness- aware low-entropy segment-based advantage shaping for reasoning llms

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.715670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:2012c95d190c71850b7211b74793c493a04f7bba4a05851d4d8edcd0962a5a1f

Observation 2fd17d16-dcca-478d-9018-e3d43b12d25d · outbound

This paper cites Reasoning with exploration: An entropy perspective.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Reasoning with exploration: An entropy perspective

Reference 59

Resolution
verified exact
doi, observed 2026-05-15T06:05:05.561384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:15abc00334ab3b33a7fcb8347e223dc1ce098be63e776670e76cc763f9cecf1f

Observation 3596d672-72f2-4c79-b135-e472c245fe98 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:05.642690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:aaeb3f687eb1708b8f12a103a1ec4a960e23686f3c575353f0cea7e4bb3bb225

Observation e5f3f6be-1f8d-4b43-b24e-4c2a57f78860 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Group-in-Group Policy Optimization for LLM Agent Training

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:05.659613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:a09702ea2fdc34bbc06d4b88d9d376c81a54e7d69c47cd2282272d3d0792786f

Observation 8856dd9e-94e2-4b81-9eb3-f3c951952dca · outbound

This paper cites Soft Adaptive Policy Optimization.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Soft Adaptive Policy Optimization

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:14:32.897338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:b5146fc0edb77af5a1fbdef1641e69dd09b85a59357f198bcc463af8b2c9e600

Observation 6ba870ad-b434-4d29-96d7-550254fd19fe · outbound

This paper cites Cruxeval: a benchmark for code reasoning, understanding and execution.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Cruxeval: a benchmark for code reasoning, understanding and execution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.241583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:268627a3461d5b5fe79451dc594adfa9fb1d9cd68e792ff808acf56ce2324591

Observation 5e39f9a5-9ee6-4017-843e-11a10abceef8 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.341060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:518dc0b4fec887cb994a49852e53bd697b074565288ae45c191da3d0192853c1

Observation 6e75603b-2526-4ebc-b4cd-a2d248c8d701 · outbound

This paper cites Justrl: Scaling a 1.5 b llm with a simple rl recipe.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Justrl: Scaling a 1.5 b llm with a simple rl recipe

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.411153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:6b15558bc37db9ff29ff2c0d7a8bc6eabcec67ac09d634cf4ce014538220156e

Observation fba3d225-0820-4d67-b561-0af838b82cf1 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.232094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:704cade419c94c7da3af9bfbe2fa085c46aac4e322f6f28628d5e6aa83df2f8a

Observation 9c7ad0fa-9c35-4b1b-8695-6f7a878a8d24 · outbound

This paper cites Beyond magnitude: Leveraging direction of RLVR updates for LLM reasoning.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond magnitude: Leveraging direction of RLVR updates for LLM reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.345882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:4684d897c9be54de096611d9004a0d027ca60e7c414cddd5a130faeb77357234

Observation 24ae3fb6-e3be-4267-9a71-6dd8dae4e587 · outbound

This paper cites OpenAI o1 System Card.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control OpenAI o1 System Card

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.358962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:78c0d305f00db8b44313be59e8d6cdb665e722938735ab88267d807c9012a20c

Observation b0d0c371-6857-4216-a62a-ab18fd029ed8 · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.282461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:864fba43d3a9be3eea60cb32344c179171d040f2b26cb54e9bd5f7686f92648b

Observation a2ef35dd-7d2d-48ad-8050-9688232c2bd3 · outbound

This paper cites Let's verify step by step.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Let's verify step by step

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.336982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:863e04cac63378093803daa7a1a30e65f5469a3a28fb7fdaba58b4e3be41756c

Observation 2dde0b84-2aea-48ca-9a2d-70b6c24df2ee · outbound

This paper cites Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.287566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:1e375ee3b1e7283beea74b73907544a929f2ab77af12b11f010ba9051c6c9c26

Observation 36da2de1-db94-4260-a5e6-412e8fd3eabe · outbound

This paper cites Sparse but critical: A token-level analysis of distributional shifts in RLVR fine-tuning of LLM s.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Sparse but critical: A token-level analysis of distributional shifts in RLVR fine-tuning of LLM s

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.350658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:5ae8db23fb8b8698b831665d35fffecc4b99436be4a636a29dd31d28e0bdd428

Observation b0dff27b-85c0-4c12-85c9-957b900868f5 · outbound

This paper cites The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.226641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:b7dcb6e2a3d9aa34ae1e3bcd471897e49bdbab09076f419867b4ad9acde07f17

Observation 070ce4c6-bb11-42bc-89fc-9f0f16a608a5 · outbound

This paper cites arXiv preprint arXiv:2603.11682 , year=.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control arXiv preprint arXiv:2603.11682 , year=

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.708407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:e4e662fe6376ab3fa4ed0851976c9c59d211a5f4020e012048f2a70cdbf55bd2

Observation b73e19ba-45b9-4726-b24c-e3606efe6291 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Maximizing Confidence Alone Improves Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.692197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:3e0e0c37eade237abba043e324cfa167c5fb8fb0f6051eeadfb7d2ff17af21e7

Observation 79a78249-4165-4de3-a3d8-78d1945ba278 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.367627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:7249eea5f391c0537c1656832c8abc23ffe950f341ee25c3309493968a74f511

Observation 64eb57fc-08ca-427f-b6b9-b07d867ba27d · outbound

This paper cites On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.699864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:f7398d52863adcec81ffc1c055a1600a23ce88b3128e3c8e51183e6a2a83c90f

Observation b0deea67-a24b-4311-924c-526b4531eaaf · outbound

This paper cites CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:05.608396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:bfe24e8efe9d0052e07aee01078dbc616941278051d7b10d5726b3438a4ce3c7

Observation 528c41a9-d157-4d40-be88-acf15de1cbbe · outbound

This paper cites MIT press, ??? (2018).

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control MIT press, ??? (2018)

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.617258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:dc1f3f2b659350127f3ebe7fd67d97915f02d14c762a880a612f9b8346706baf

Observation 006a1428-f998-4e51-99ab-92ef524468ff · outbound

This paper cites Rethinking sample polarity in reinforcement learning with verifiable rewards.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Rethinking sample polarity in reinforcement learning with verifiable rewards

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.634180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:847559a1713b019fe2ea29228d634cf27a65d530cbffb58614f84844f55c5b9a

Observation c4e907c9-0b5b-466c-ad64-da9b26aaa512 · outbound

This paper cites Skip-Connected Policy Optimization for Implicit Advantage.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Skip-Connected Policy Optimization for Implicit Advantage

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.343713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:10a020d846290d8d012b2a77133f16ad5e7418d6a3c459450d45fce0fd25c0e2

Observation 297a523d-7f92-48ad-a280-3973a6a0d88c · outbound

This paper cites Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.277297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:ef33c428ec83b738291dd64a221e95ee018f8b6405e2d3bfdbaa946e9a1cbe0d

Observation 111f4477-1712-4f52-80f1-8eda29f0efdd · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.293189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:e961f618e3a0d5f4ba185abbb150052acca22e34741ed46b529352427bef0669

Observation 199a2e9b-eb9a-4a3a-940b-4e6a27353c78 · outbound

This paper cites AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.669079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:ceb55590c8eaa77cd47ba2cc68814c9292f439f43a14b420fff5cbb4a584b3eb

Observation 2526582d-bd22-45fd-a8ca-d41b717612f9 · outbound

This paper cites Can RL improve generalization of LLM agents? an empirical study.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Can RL improve generalization of LLM agents? an empirical study

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.625527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:1a2a8c726cab61dfa6d9714c5ff525f807d7840f2e71f386d464417d85a70579

Observation 74cbe1d4-6a43-45fa-ac36-e21311350669 · outbound

This paper cites BAPO : Stabilizing off-policy reinforcement learning for LLM s via balanced policy optimization with adaptive clipping.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control BAPO : Stabilizing off-policy reinforcement learning for LLM s via balanced policy optimization with adaptive clipping

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.332968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:494ee43eda4ce7dc24859a31d61c0d9b342684a95a945dfa534b550785775c28

Observation ec936a71-24ac-4c62-bad9-098338dd98df · outbound

This paper cites Qwen2 Technical Report.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Qwen2 Technical Report

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.418576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:8b1033e55443c2188a1ba881fbe1c6dbc49b10c5200073e8bca7346c9d7c794f

Observation 69d64e5d-9cc7-4ce0-b5fc-3631a6d2fb5e · outbound

This paper cites DAPO : An open-source LLM reinforcement learning system at scale.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control DAPO : An open-source LLM reinforcement learning system at scale

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.267547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:b863fcbbcb8e57407a78001160c5ceb1bb3e1190d15b7eae6ad110880aa8a4f4

Observation 456abcf3-3977-4592-8a7f-568bbf7d12d0 · outbound

This paper cites AgentV-RL: Scaling Reward Modeling with Agentic Verifier.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentV-RL: Scaling Reward Modeling with Agentic Verifier

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.388468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:28e8bf4235f1cf5ff3f52f496411d42be2312343ba96a8c2502b6e5ceef375ad

Observation 783afc32-a3c2-4eb0-b8e0-1f14ff5ed753 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control A Survey of Reinforcement Learning for Large Reasoning Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.511734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:42f7d2f94fbb133a9bf53ad885aca2ce73dee84e6a9e0cd0cde1f252b70c89ea

Observation 43e4fe24-08dc-4499-af74-b9b8f36dbb38 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.427969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:2828f66fd8c3e5c6bc9bd715db4046c9591489fca27412e64f3fdd1328de32dd

Observation 1af3f861-a842-4766-98b6-41912ba6e535 · outbound

This paper cites American invitational mathematics examination (aime) 2024.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control American invitational mathematics examination (aime) 2024

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:05:07.183500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:95cfa88c5af0ef02aeb6e237eaf194e5b48579898a34b635fa7fa1260de7701c

Observation 2d6da1e6-b7fd-4619-afe4-228f2c821d11 · outbound

This paper cites Why reinforcement fine-tuning enables mllms preserve prior knowledge better: A data perspective.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Why reinforcement fine-tuning enables mllms preserve prior knowledge better: A data perspective

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.435798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:7c11db0ffa49b88600e1acc59d42f2971fb70553bb77883e8315fbd9bb014e82

Observation f4af5695-b46f-47ec-a2b8-c376b1329923 · outbound

This paper cites Group Sequence Policy Optimization.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Group Sequence Policy Optimization

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.397134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:94c1ac54c49ee5f1ce9bd7915b74d3d4b1f1e6f2ff454136bedc51a1c1060487

Observation d4440e23-b9ac-4779-a490-10a0565f32c3 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Instruction-Following Evaluation for Large Language Models

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.350493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:cbdb232c8587160221a0736251e99603ff444e6c3ab4e1ecb73e440e39f222cc

Pith citing papers

No inbound Pith citation observations are available.