Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:04:32.640299Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2605.11775.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:04:32.640299Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 557f233a-4fee-4e72-8364-09c3b847d187 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control author=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce5fddca-35e7-41b5-8be1-4d5731ef2474 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , url=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 313e936e-be72-4dfb-bbb5-d37f5199f9bc · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f873a061-5c55-4445-81bc-85c975667337 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models , url =
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a881924-055b-4ef7-8808-56063c53e986 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 182b54fa-f5fa-411b-a846-b5f255b77e46 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Entropic: Towards stable long-term training of llms via entropy stabilization with proportional-integral control
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c06b9d68-3e1f-4e47-91e1-e9bd88e40a56 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , url=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4d55725-1f74-4e37-bb09-f1f4a64a2498 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond Magnitude: Leveraging Direction of
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37c4a159-06da-4d24-ae81-5085b1039e86 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Sparse but Critical: A Token-Level Analysis of Distributional Shifts in
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f438d893-6709-4658-b4c7-667725a7f3e3 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , url=
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de73cd91-e00d-45aa-9b59-b679f5d5985e · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 050d202a-b668-465f-9724-5a4ab47c1c86 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The Twelfth International Conference on Learning Representations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cc92413-f495-4bb0-8abd-edd46d9e0635 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Hugging Face repository , volume=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9896f1ff-195d-422b-a268-ca62a66e184c · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2024 , note=
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eee99b74-a6f2-4a57-82ac-4634fb304340 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2025 , note=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99306081-5c8e-40e0-9783-22c0df46b672 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Advances in neural information processing systems , volume=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a18c04ee-9b41-4c31-b9e3-55586f324673 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Forty-second International Conference on Machine Learning , year=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bf86b94-26f7-4b77-9f91-79cd7260367e · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Proceedings of the 41st International Conference on Machine Learning , pages=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a138f03-af49-4f6f-965d-7b6285561c74 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The Thirteenth International Conference on Learning Representations , year=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92fb0707-0513-4087-9512-cfdcd102e3a5 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Advances in Neural Information Processing Systems , volume=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ba9d8b1-cb94-4556-91c5-386517159698 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control On the direction of rlvr updates for llm reasoning: Identification and exploitation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 684f3e3d-be52-4585-89f5-ef1409597fbb · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices , booktitle =
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ea532d5-47c1-4dcb-bce0-9768d4bcfb0d · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control 2026 , eprint=
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c407257f-d74f-49ac-ba54-04b7e24e2575 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1920804f-4cc4-4079-8e3f-4b55f205bb76 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Claude code
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 677e0d37-8cf0-4496-8131-0cc646a28049 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control TTVS: Boosting Self-Exploring Reinforcement Learning via Test-time Variational Synthesis
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 927c14cb-bf21-4c27-a24b-109b5a3833f4 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Acebench: Who wins the match point in tool usage? arXiv preprint arXiv:2501.12851, 2025 a
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5843447a-5e6c-4832-907c-262a029d902e · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5643cca9-9766-46b1-a244-d459dea31216 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Exploration vs exploitation: Rethinking RLVR through clipping, entropy, and spurious reward
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 904c25f1-aa07-44cd-abb0-ccf7694504e7 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond high-entropy exploration: Correctness- aware low-entropy segment-based advantage shaping for reasoning llms
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2fd17d16-dcca-478d-9018-e3d43b12d25d · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Reasoning with exploration: An entropy perspective
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3596d672-72f2-4c79-b135-e472c245fe98 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5f3f6be-1f8d-4b43-b24e-4c2a57f78860 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Group-in-Group Policy Optimization for LLM Agent Training
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8856dd9e-94e2-4b81-9eb3-f3c951952dca · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Soft Adaptive Policy Optimization
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ba870ad-b434-4d29-96d7-550254fd19fe · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Cruxeval: a benchmark for code reasoning, understanding and execution
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e39f9a5-9ee6-4017-843e-11a10abceef8 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Deepseek-r1 incentivizes reasoning in llms through reinforcement learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e75603b-2526-4ebc-b4cd-a2d248c8d701 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Justrl: Scaling a 1.5 b llm with a simple rl recipe
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fba3d225-0820-4d67-b561-0af838b82cf1 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c7ad0fa-9c35-4b1b-8695-6f7a878a8d24 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond magnitude: Leveraging direction of RLVR updates for LLM reasoning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24ae3fb6-e3be-4267-9a71-6dd8dae4e587 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control OpenAI o1 System Card
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0d0c371-6857-4216-a62a-ab18fd029ed8 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2ef35dd-7d2d-48ad-8050-9688232c2bd3 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Let's verify step by step
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2dde0b84-2aea-48ca-9a2d-70b6c24df2ee · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36da2de1-db94-4260-a5e6-412e8fd3eabe · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Sparse but critical: A token-level analysis of distributional shifts in RLVR fine-tuning of LLM s
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0dff27b-85c0-4c12-85c9-957b900868f5 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 070ce4c6-bb11-42bc-89fc-9f0f16a608a5 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control arXiv preprint arXiv:2603.11682 , year=
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b73e19ba-45b9-4726-b24c-e3606efe6291 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Maximizing Confidence Alone Improves Reasoning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79a78249-4165-4de3-a3d8-78d1945ba278 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64eb57fc-08ca-427f-b6b9-b07d867ba27d · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0deea67-a24b-4311-924c-526b4531eaaf · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 528c41a9-d157-4d40-be88-acf15de1cbbe · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control MIT press, ??? (2018)
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 006a1428-f998-4e51-99ab-92ef524468ff · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Rethinking sample polarity in reinforcement learning with verifiable rewards
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4e907c9-0b5b-466c-ad64-da9b26aaa512 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Skip-Connected Policy Optimization for Implicit Advantage
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 297a523d-7f92-48ad-a280-3973a6a0d88c · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 111f4477-1712-4f52-80f1-8eda29f0efdd · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 199a2e9b-eb9a-4a3a-940b-4e6a27353c78 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2526582d-bd22-45fd-a8ca-d41b717612f9 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Can RL improve generalization of LLM agents? an empirical study
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74cbe1d4-6a43-45fa-ac36-e21311350669 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control BAPO : Stabilizing off-policy reinforcement learning for LLM s via balanced policy optimization with adaptive clipping
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec936a71-24ac-4c62-bad9-098338dd98df · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Qwen2 Technical Report
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69d64e5d-9cc7-4ce0-b5fc-3631a6d2fb5e · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control DAPO : An open-source LLM reinforcement learning system at scale
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 456abcf3-3977-4592-8a7f-568bbf7d12d0 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentV-RL: Scaling Reward Modeling with Agentic Verifier
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 783afc32-a3c2-4eb0-b8e0-1f14ff5ed753 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control A Survey of Reinforcement Learning for Large Reasoning Models
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43e4fe24-08dc-4499-af74-b9b8f36dbb38 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1af3f861-a842-4766-98b6-41912ba6e535 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control American invitational mathematics examination (aime) 2024
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d6da1e6-b7fd-4619-afe4-228f2c821d11 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Why reinforcement fine-tuning enables mllms preserve prior knowledge better: A data perspective
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4af5695-b46f-47ec-a2b8-c376b1329923 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Group Sequence Policy Optimization
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4440e23-b9ac-4779-a490-10a0565f32c3 · outbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Instruction-Following Evaluation for Large Language Models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.