Pith. sign in

Paper Citation Record · LEDGER

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2507.01327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01327 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:38.886258Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31a208ec-737b-4732-823d-902ac8dc57a8 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:36.942746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:36.942746Z digest=sha256:261c76e8193a3041f531049c4169817d9e19c1b8d8aefe7828cbd844dd8e6658

Observation 5f9e17e1-3cc5-483a-b922-d22bd53228f7 · outbound

This paper cites o3-mini vs DeepSeek-R1: Which One is Safer?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy o3-mini vs DeepSeek-R1: Which One is Safer?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.032293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.032293Z digest=sha256:ab533042e0b88b8e131b79a1a7dc012656f3e78cab7ba8c520e81d0c5935413b

Observation 14b9f2ce-14ee-4cf7-a291-fb1d8309115d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.183211Z digest=sha256:9f15259d524284277cd89b1bebe467a27a3af1c8f96f89ba8ab333097a139fe7

Observation 0d5c8270-2ff8-49f1-b33a-2dbbc9d82af1 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.234899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:37.319463Z digest=sha256:45b2fcb352a52b66a8c2c4c0b9a0c3dcd143f6c68305b34881552f0cf14a148e

Observation 023fce87-47f5-49f7-8100-39b5698eaf4c · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.502465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.502465Z digest=sha256:5c432e6ce5adade2b5f6457117cce051a2b4fe8d82ebc220525252e567a14efc

Observation 0aa882ab-7b75-4db8-8aef-e46eea1ddc3a · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.699455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.699455Z digest=sha256:ed721508e1879d882537aaee107bcc4ef82fc1fee86cca96f4332cd9164f64f2

Observation 86e7978e-5dd8-4787-9961-5830e10eec17 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.820378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.820378Z digest=sha256:fe9aed6a4b78aede96203c862374cf437a1df4b5bf189be928b40d466ae57e54

Observation ce74fb56-f1ee-4738-a1c1-2e2a673cc6c3 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.220572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:37.955406Z digest=sha256:4e80dab99015a79ce0656253e34ee005a248f52a4917b1c1f183b9c2b7848a0a

Observation c657feba-84a2-4a22-8038-9a9c4d28439b · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.205589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.044785Z digest=sha256:0d0988c7d35599cd212859d36ace07b64f02e762153570a020740d19da7afb33

Observation b4835359-7fc4-40df-b9fd-6224fa5dd951 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.151830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.151830Z digest=sha256:e06c8b91b596aa6ff56e3859b2ffb245e37d6f0d6b299a08463f3ea10cd768b4

Observation 2c2ad01c-1d6e-4ccf-abb3-9329f1992d0e · outbound

This paper cites Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.200430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.200430Z digest=sha256:afbad86b664f17bd022e0d123028b0ef709923fde96e1ba162ed856c16943b65

Observation a1327e2c-c0e8-4beb-85d4-dfe89e595765 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.180524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.263726Z digest=sha256:6a0353b22aaf12feb6f4bcd4541da3053b8500f0522c5ef37ed32ab75c4f873d

Observation e57f1145-b8fc-42c4-bed9-231ca0008c03 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.448293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.448293Z digest=sha256:7a136237efeb0cbbf1fe61e8b5cdcda266934957d45f2d5faa77d56ee87b8c66

Observation 8878cbaa-47b4-4a76-80a2-7262b7f49da0 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.165410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.528388Z digest=sha256:39e4ae40f3f3a3a2087e5de3487c36f68cb5477e8f4e19d17da055c082412f45

Observation dd28b2ec-b78f-4368-bd84-ba7e374f707d · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.646820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.646820Z digest=sha256:c8f1ce51cee5457dbe62018e49d65d5f1043b59952d3cf2e02ee328288662dbf

Observation 9c259246-689b-4490-928f-5ad5de36cd71 · outbound

This paper cites OpenAI o1 System Card.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.664449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.664449Z digest=sha256:4db99de711f11881fb2f74c6aa1cf1749a135c426bfc065200ec05badcaf50a9

Observation b7e29413-02af-4dc8-8d42-a5e0a05c3671 · outbound

This paper cites Scaling Laws for Neural Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling Laws for Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.671165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.671165Z digest=sha256:c0d59ca8199599204f35a90b2ff853ecb61bde36f2a3c02302b2e4e40bbf0d1a

Observation f43a6cd5-3164-47a8-a521-720e6188e685 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.676052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.676052Z digest=sha256:e8471a9e0171b58500574c86786d602644532f3460c331f9e293da64d5af1062

Observation 12f06d24-1b9c-4673-bf13-d373f5c1b33a · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.681209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.681209Z digest=sha256:eb2c9c5102759d0c301ee4fbca095f07d0bcd889d1fa8812d92c376a218198b8

Observation 3b8e6642-fd7f-4db2-babd-ce61d14f1753 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.685458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.685458Z digest=sha256:d0352e9e4cd315b58dd50dfa8c02c0e0868bd2dabd79535e1c9d901215c99762

Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.690014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.690014Z digest=sha256:7a9d846b800ee62c55daf8bcb608d627fce53ed6860a5936d704afd8b3d30e21

Observation a3d2dc45-2930-4ced-9b96-f57ef3b601f6 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.694881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.694881Z digest=sha256:fcdd63f70329e2f7ebe0bf10ca9acad793185764f497ca59561dc0e879e2f65a

Observation 806618c4-0cc6-4a35-9b9a-34319bb0c842 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Understanding R1-Zero-Like Training: A Critical Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.699466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.699466Z digest=sha256:048cf53b14ec86d010e8c98744dedd7fa6b390da32f0765eaed809cb6fb628a4

Observation 7586d958-0dc2-4586-9fd6-44af33d8ac51 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.704508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.704508Z digest=sha256:8a5849aceb2a386e05d714df6617ed4a98bd8ef4ad874bee75869235da2ada78

Observation 07392453-e916-4d8d-8ef8-bc9b199f4f63 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.709251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.709251Z digest=sha256:4cc1e68f2f2a692bce33deefe9ae200e8fedd6352363f9fe7f7170b0074115c7

Observation 9aeb3252-57f0-4b23-9cf4-20a64e488caa · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.129593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.713899Z digest=sha256:6b55e5f41823eaea7d73b0b68b12b246b2af74c0b6809aa13ffc0087777a6dd1

Observation b96b7442-c81b-4c59-80d4-041fdf07e4dd · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.115473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.718779Z digest=sha256:6b316a169860dbd65a3026e4329b569847c20d5d224c63b82c127ac3bca7cae2

Observation 8d6d9956-b807-470b-ad2e-2f02d026acf9 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.100897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.723190Z digest=sha256:7831a8a0760e14e2a16fc05c4335286ec3728829309577307e429c00649d798e

Observation ceacf498-c68d-439d-9547-2842c0e27f5d · outbound

This paper cites GPT-4 Technical Report.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.727798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.727798Z digest=sha256:e232b81a60f5012f93cd9d7cdfaf34ff776f37ed62321434897ecb7f85438bca

Observation f1d61937-ff6d-4540-bd52-7cfa4ed879ee · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.085501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.732475Z digest=sha256:60156613fe401867b475507e5474c1f75895896886ce44515c8afcd8d1b0e6ab

Observation 04ea2fdc-426e-4521-8c9c-92d6778a9918 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.736901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.736901Z digest=sha256:c099be195288dc75470050babce53b7900c2b2eca5c0c29a7da452ed8548641b

Observation a00037ce-f4da-4fb2-ba28-2a7ffbd0fbe1 · outbound

This paper cites Multi-class Text Classification using BERT-based Active Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Multi-class Text Classification using BERT-based Active Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.432882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.741316Z digest=sha256:69ea536b991726b70792ae9e97f1f2b69c8c6f53b221c82556b48b8b8eecd98c

Observation 32ffb3d0-ec8e-4c7b-b520-c4910cf7dc8b · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.059976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.746100Z digest=sha256:e1c71b912b8de45f53a54312613937bd3ff73407070770418e35fe41ec2e3ef6

Observation 0591341e-520c-4e6c-af89-90d0f5dffecf · outbound

This paper cites Qwen2.5 Technical Report.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.750400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.750400Z digest=sha256:b7c477eb8ab393808e5e12df7a9fbebea049236d990ae53255bb7db9a1c37a1e

Observation c7ace076-2e60-4f14-8da7-2890c9818394 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.754928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.754928Z digest=sha256:737d7a0f4b32c6972e6e69716662e359bbc02cd76c065fd840d4b0b46c30a796

Observation d9babff3-870f-4bcd-b9bc-6ed2395b8610 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.035515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.758816Z digest=sha256:8dc78c94cddb12a493d94378e4bec35916a94143bd748d8316305eb3015c6d66

Observation 2715de2e-6c1d-45f2-ac73-649341fc1bea · outbound

This paper cites Ray Interference: a Source of Plateaus in Deep Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Ray Interference: a Source of Plateaus in Deep Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.763176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.763176Z digest=sha256:dcb4337ebe1090b1ed2c7b9e196566854f62835620fe427458fefddad4638926

Observation 7b522326-3a75-4857-915a-6ccf1c508f45 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.767702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.767702Z digest=sha256:20b7d4e5870418ed7f5a9be5ee15eea6c0b5f7f4724cfc7a3db71e287903f6d3

Observation fdb864e8-2e6b-4ad6-92e9-26720e25ec49 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.776906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.776906Z digest=sha256:41fa8e2ebb4e9fac2cd5c1ec8cf04040e915e59045bd3049db44b046dc0729dc

Observation 4571d06d-73c8-4ff6-83e7-161d0e3e36d0 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.781944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.781944Z digest=sha256:635393f2dda347a2d7ab9d40914bf9544026f8835b914d7053ace3935880f3e4

Observation 042cd7d2-6082-4890-bce1-bb2480a4d2e9 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.786360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.786360Z digest=sha256:1376ce5377a233bfd19a1d7cc77f42bdb77e20beb4410266d8b3748dfdd12c2c

Observation e05d6af9-440d-4d16-bb6d-5ce7cc0b3662 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.021093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.790628Z digest=sha256:967631510b5abfc375ac68b1ede4ea6a9a858326613a34c3daa174bf65e8a8d7

Observation 11905553-2ea7-4744-ad24-d71fc44b1561 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.794746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.794746Z digest=sha256:737edcd94b58e2551857936d915c8e75a8c568ff313a62e994c2bf8133c468c1

Observation 9e72aea4-7d5f-474f-be6f-67a07e0a119e · outbound

This paper cites Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.295404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.799491Z digest=sha256:d84053e8c3d950f3060ed2ee20f9f3a8942ba432e9698b43cee6d7e99d8094b3

Observation 8a02cf4b-ebfd-4335-aac8-44d35fa87c26 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.804100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.804100Z digest=sha256:1ebd57831e01eb9e8f45c316add9cd6e359dd4efe4de2674c39b58f43700b288

Observation df9f09b6-8183-466e-98d8-68bd965f2c22 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.808953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.808953Z digest=sha256:8f8d57b93d57551f3667199db411dd0dae36711ce28cb3e8e1b926d2c1d3e124

Observation 33e7ef3a-5699-49e0-8229-fa15229b2143 · outbound

This paper cites ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.239638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.813932Z digest=sha256:ec62b5958ec40e2b61e2e2dec5fc01a30c7c280565a1f6306a12893a3bc41f4f

Observation 5765637a-d77e-4077-9058-a0f4022bc6d1 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Finetuned Language Models Are Zero-Shot Learners

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.818681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.818681Z digest=sha256:bc8ae80120d7e6ca5a05dbc57682e07eb044802620c34e6ac527c576be8c0988

Observation 8e8d32ef-5428-4420-81a6-7d0ec150fb46 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.006437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.823905Z digest=sha256:2e55be73db9cc3af7e14ba55126553353a0c15d113c50a64a6dfad3222524592

Observation dd18e72e-1c9a-4bbb-92b7-fee9f47c2141 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.828256Z digest=sha256:873b16b1d1a14e33311c9fead4f99f0d60c150d26976a9d59a00d8a055a786ea

Observation 20842445-d568-402d-b50f-703292704157 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.833031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.833031Z digest=sha256:79bfa008812c69dd4f3eefa239fea91971909e8dbf24ee049d99773487b03bcf

Observation 4dc20609-ed91-4ef5-a216-9aa1bdba487c · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.981747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.837355Z digest=sha256:1e6a33f043a958bb451834cf75fbbe3dbc65a9a1e624028722c147eeca1beabf

Observation 777fff92-f2f1-415f-9946-188530af05c8 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.841898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.841898Z digest=sha256:5894b334785e01de5145b498f7a43398ae2a860ac2f301572f47c6b2b00b3957

Observation c1c42d31-9ed7-48d7-8c42-8cefef7bf7ad · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.846079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.846079Z digest=sha256:ad1eb09f566f784f709efbc569c0ded2778f3fa2994c8eda8955cdb293052716

Observation 1bb0e20d-945b-41d7-a4a4-30fc039d6810 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.957748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.851156Z digest=sha256:c41119879a3926a1ed6a59b97e615b61009ffeaceebd39f6d665a2be26796184

Observation 48d27688-5e9e-45c2-a2d9-51bd55bfa345 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.855807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.855807Z digest=sha256:74468d1648746ad5c2a438d54b48a0d47ee2a920981a86af1869c2f167c8879c

Observation 71613651-96da-44e0-bf60-01d6a1a7ebf9 · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.860058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.860058Z digest=sha256:396433e894744985bae2c2118cddc15d2d47f6d51d7a99d9b1b93fcc476a285a

Observation c0df93c6-3af7-4ce3-ad93-beb0d070b678 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.864617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.864617Z digest=sha256:a158f1c400b80c0074c8361c7231d3d674303357c8dfc51cf50b95f5a74548a4

Observation 6c168912-83aa-46d6-94cc-af00a997d0dd · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.868488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.868488Z digest=sha256:fbbc116d2d4b3ed98fdfaf990311057220dfc8d56452e83b9516682b7f0dfaac

Observation c0586836-fa80-412a-920f-13e07fa9f098 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.873273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.873273Z digest=sha256:8b4e6066456c0ebb8aeed48edc9b22e637ab271cc6d76caeaaba69910226c58b

Observation 656c26ef-ca14-4248-8233-0421db02e172 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.942268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.877329Z digest=sha256:598ad655166104e00dd450302ba67601afe27854220ff2f4faa8f4738e901969

Observation 1d90440a-2dd7-43ce-8ece-082d594f8f6d · outbound

This paper cites online" 'onlinestring :=.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy online" 'onlinestring :=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.881524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.881524Z digest=sha256:8019eed890f3d500323027555336f649109668c7e920ee1c9232dbb75dade168

Observation 16f3e57d-a978-4b96-9d35-2180688f1449 · outbound

This paper cites write newline.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy write newline

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.886258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.886258Z digest=sha256:e63d3c117b2fb4daffcfed0718e8fd864664b8c4f992dbc4ccd861b243cfbe82

Pith citing papers

No inbound Pith citation observations are available.