Pith. sign in

Paper Citation Record · LEDGER

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2507.01327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01327 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:38.886258Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31a208ec-737b-4732-823d-902ac8dc57a8 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:36.942746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:36.942746Z digest=sha256:ad1ea2c16d8b64f3461c349db292808d528ad066c76d0c4879772e6ec2ac25b3

Observation 5f9e17e1-3cc5-483a-b922-d22bd53228f7 · outbound

This paper cites o3-mini vs DeepSeek-R1: Which One is Safer?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy o3-mini vs DeepSeek-R1: Which One is Safer?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.032293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.032293Z digest=sha256:3746ccc02dfc60e09a034a5fb16576af2dbee031214af21e9862d3920c9cdd63

Observation 14b9f2ce-14ee-4cf7-a291-fb1d8309115d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.183211Z digest=sha256:74f4fa91168b7e90f4cdb4001c173c9fc24fd3144169147576d899082cdfbdef

Observation 0d5c8270-2ff8-49f1-b33a-2dbbc9d82af1 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.234899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:37.319463Z digest=sha256:8e57df17d2aa22ea0f4e3d5357ce3dbe51fb40834d3f81717464e233bff01d14

Observation 023fce87-47f5-49f7-8100-39b5698eaf4c · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.502465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.502465Z digest=sha256:b8559be637e248d456113706c938cfec1fe3dff024fb4d3849ebea22bc80de25

Observation 0aa882ab-7b75-4db8-8aef-e46eea1ddc3a · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.699455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.699455Z digest=sha256:2f52e1b315c95110caefd4170544eb111560a059cff26ac1ef01d4ec8294d619

Observation 86e7978e-5dd8-4787-9961-5830e10eec17 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.820378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.820378Z digest=sha256:1067756b2bee200ddc3d184671de556dfb9a9b8de01565d17debcc53221ad4af

Observation ce74fb56-f1ee-4738-a1c1-2e2a673cc6c3 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.220572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:37.955406Z digest=sha256:fda64360ff0683ca0d885609dbcf4033c21237fc74a151d08a10f7870c3d9912

Observation c657feba-84a2-4a22-8038-9a9c4d28439b · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.205589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.044785Z digest=sha256:30344db9d1b6809ec54483dcbccb9cd6193e97c41ea90640cee406362ea1cc4b

Observation b4835359-7fc4-40df-b9fd-6224fa5dd951 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.151830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.151830Z digest=sha256:b4d0d13413a1d15d5de80dfa40d7a8cb6e0f28d9e1adbca0ad76441f25e8260c

Observation 2c2ad01c-1d6e-4ccf-abb3-9329f1992d0e · outbound

This paper cites Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.200430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.200430Z digest=sha256:14b8bc06a119e16119cf770625c5d881cfaeda0e9643f8d8a604a0aba2a99b3c

Observation a1327e2c-c0e8-4beb-85d4-dfe89e595765 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.180524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.263726Z digest=sha256:5bdcf382b05e125ec9f95ebdefaf5fe4f30106184a0b272c2d3dc0edf16e93cd

Observation e57f1145-b8fc-42c4-bed9-231ca0008c03 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.448293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.448293Z digest=sha256:8d21258590e24d056144925478c3793365317bc27c544059ae317cb4094ed6ad

Observation 8878cbaa-47b4-4a76-80a2-7262b7f49da0 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.165410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.528388Z digest=sha256:5802e9028a14b739b40a3585da904347d10edbf3c21648f5843048c3db79f16a

Observation dd28b2ec-b78f-4368-bd84-ba7e374f707d · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.646820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.646820Z digest=sha256:c8bc6308aa7514fc694803a9647559d7e37cfe4ef21f29312f367bba6b48afe8

Observation 9c259246-689b-4490-928f-5ad5de36cd71 · outbound

This paper cites OpenAI o1 System Card.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.664449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.664449Z digest=sha256:8a8b0bc082ecad6e672bbf1dde5223077cf99e8e279058fd0c18a12658356dd9

Observation b7e29413-02af-4dc8-8d42-a5e0a05c3671 · outbound

This paper cites Scaling Laws for Neural Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling Laws for Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.671165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.671165Z digest=sha256:b6bd68eef81c5b46b18924f97fc5862da4cc839f685b4063a48e2f45663271cd

Observation f43a6cd5-3164-47a8-a521-720e6188e685 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.676052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.676052Z digest=sha256:95758e74f5086b6530f60a648411b78d86fc679ca520dd82781136666036c0e8

Observation 12f06d24-1b9c-4673-bf13-d373f5c1b33a · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.681209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.681209Z digest=sha256:1dd88155f9f279ac66e247104498a8f952bbf24aef098537a24aac359c4c3eb4

Observation 3b8e6642-fd7f-4db2-babd-ce61d14f1753 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.685458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.685458Z digest=sha256:1b5df20fc028c3a1e7853e36ab6b688fdc856ac44d0407211432582d5e8d4ccb

Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.690014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.690014Z digest=sha256:0e7f96e0b317108a8e2bf6e17214de08efad39ae142296559b587daa4ded10d6

Observation a3d2dc45-2930-4ced-9b96-f57ef3b601f6 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.694881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.694881Z digest=sha256:4ebbbdcab8bcef6f2202f509b66f295c3e60435f0b429cf76d9e65a4857cbe84

Observation 806618c4-0cc6-4a35-9b9a-34319bb0c842 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Understanding R1-Zero-Like Training: A Critical Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.699466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.699466Z digest=sha256:d511d0d6a376f01258ad1f1c3317dd5d0f5d726fd23c0da43b529df2f6b6d9ca

Observation 7586d958-0dc2-4586-9fd6-44af33d8ac51 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.704508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.704508Z digest=sha256:9e7f9892d75b92c7e98523a5c9b33c57b33e0599c896fad91e7de0d393de7ff9

Observation 07392453-e916-4d8d-8ef8-bc9b199f4f63 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.709251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.709251Z digest=sha256:7ef417b9f100b06d0f23388c79c7b72e4a68d8e96cb1d34b3574c730e6f2c920

Observation 9aeb3252-57f0-4b23-9cf4-20a64e488caa · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.129593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.713899Z digest=sha256:6128b7168fd36a0b2468b2e287192ceb420e73138c2ad24d5d6aab699f330b0a

Observation b96b7442-c81b-4c59-80d4-041fdf07e4dd · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.115473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.718779Z digest=sha256:9750c69940b684069de58ac41d76adb1a08f6671e5e2014a7e01999841a071fc

Observation 8d6d9956-b807-470b-ad2e-2f02d026acf9 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.100897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.723190Z digest=sha256:ac9c5934d77b28e808c9bbb2193219d3f63c8ed8ab0e6ac35f17b55621a97d03

Observation ceacf498-c68d-439d-9547-2842c0e27f5d · outbound

This paper cites GPT-4 Technical Report.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.727798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.727798Z digest=sha256:12e2f3ac9fda3a6028e7f6489624f9d9c62f927e06d25fba113c3a6c7909ed2f

Observation f1d61937-ff6d-4540-bd52-7cfa4ed879ee · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.085501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.732475Z digest=sha256:72fa5633aef3c37de01ab929a436de48c281b1536eee35227d8192516900b718

Observation 04ea2fdc-426e-4521-8c9c-92d6778a9918 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.736901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.736901Z digest=sha256:81f54d1e14944dd4bd23a80d2aa317c43e9ff8cb80f5b71894daa42655e85cd6

Observation a00037ce-f4da-4fb2-ba28-2a7ffbd0fbe1 · outbound

This paper cites Multi-class Text Classification using BERT-based Active Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Multi-class Text Classification using BERT-based Active Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.432882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.741316Z digest=sha256:1aa85a99fe6d6e9937fd7c21c68ab166ee84f0c99878e3f518a137d9cc79978f

Observation 32ffb3d0-ec8e-4c7b-b520-c4910cf7dc8b · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.059976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.746100Z digest=sha256:d66da2b77d5d64a0c3afe8c529ded9d8384af0aaa5ea4ecc60980354eb76bdee

Observation 0591341e-520c-4e6c-af89-90d0f5dffecf · outbound

This paper cites Qwen2.5 Technical Report.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.750400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.750400Z digest=sha256:2a909ab52aa41d7e3ed470a0ed36066023404fc8c4fc1f81fe5151b51edc07ec

Observation c7ace076-2e60-4f14-8da7-2890c9818394 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.754928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.754928Z digest=sha256:aa77e3296559b551f6dde9987926d837cb72614eb6558afe10ac125580da8cfc

Observation d9babff3-870f-4bcd-b9bc-6ed2395b8610 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.035515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.758816Z digest=sha256:ea2de304d7fb8b7f06217d12a4ba7dacf45e60c2b0bda66e62d7928b1cf8abc0

Observation 2715de2e-6c1d-45f2-ac73-649341fc1bea · outbound

This paper cites Ray Interference: a Source of Plateaus in Deep Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Ray Interference: a Source of Plateaus in Deep Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.763176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.763176Z digest=sha256:45ab1534c91e92af8eebda54b94da12b2506809fe69fc33f4e08e463144a467a

Observation 7b522326-3a75-4857-915a-6ccf1c508f45 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.767702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.767702Z digest=sha256:320ee9be6f422eaf04bd93ba7b21c6b16e2f0134a14a011a38bb4e2dc7d5b553

Observation fdb864e8-2e6b-4ad6-92e9-26720e25ec49 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.776906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.776906Z digest=sha256:418891b385a1b17f26027414e09e9028d0cf2aed2345253ec5d7ac4c3fe0fa3b

Observation 4571d06d-73c8-4ff6-83e7-161d0e3e36d0 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.781944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.781944Z digest=sha256:872cb5cb1d996c61074eb2f8f2db5f129c7728642f3ce454cca8dea868a8561a

Observation 042cd7d2-6082-4890-bce1-bb2480a4d2e9 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.786360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.786360Z digest=sha256:24568b1f1810509c5dfcf01531514c853d1227ec592b60a8262f01d71f0e99a3

Observation e05d6af9-440d-4d16-bb6d-5ce7cc0b3662 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.021093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.790628Z digest=sha256:b9e26c37df5363ee6d8eaf24fde3990d9e013ec17bf8c16becbea90ef545886d

Observation 11905553-2ea7-4744-ad24-d71fc44b1561 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.794746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.794746Z digest=sha256:9be42f63baba6d9a1618d7c2bf1c26a52274dceaf73948d423c10e0e165c9f2f

Observation 9e72aea4-7d5f-474f-be6f-67a07e0a119e · outbound

This paper cites Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.295404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.799491Z digest=sha256:ebdfda3d436a24a0d628cc8ffd40ed471814c1b66504a55bfd1e4a8ae7c2bb5a

Observation 8a02cf4b-ebfd-4335-aac8-44d35fa87c26 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.804100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.804100Z digest=sha256:e51f7afa6e22f96e9a6ae4e3f3d59b6f3f97cde29f53d5fb3e13b12efde35b4b

Observation df9f09b6-8183-466e-98d8-68bd965f2c22 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.808953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.808953Z digest=sha256:8af9e0efce730a8c2ae31a4f7f0fff6b527aa0e6ef346104c051a335ad5affb1

Observation 33e7ef3a-5699-49e0-8229-fa15229b2143 · outbound

This paper cites ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.239638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.813932Z digest=sha256:7805f0dd9d3af3c1bc4e0faa7bd3dd28be5e9353281acd870c5efae464f098e4

Observation 5765637a-d77e-4077-9058-a0f4022bc6d1 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Finetuned Language Models Are Zero-Shot Learners

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.818681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.818681Z digest=sha256:5c7f6a59b9100634a244c7ae3cae8be4956fc990165792c6a227992443ac0a44

Observation 8e8d32ef-5428-4420-81a6-7d0ec150fb46 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.006437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.823905Z digest=sha256:65beec501e87d109feec11622f4845a3c604203f9456416b32f3c40cdd6620d8

Observation dd18e72e-1c9a-4bbb-92b7-fee9f47c2141 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.828256Z digest=sha256:658dfed00fc27da624141198cd0a3ba280d022bd0fb2ad4e278e793b48567369

Observation 20842445-d568-402d-b50f-703292704157 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.833031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.833031Z digest=sha256:7b203b02bc1573767ba67169d4279b0173cc04fe80fbb5a17f38293bcec45ad1

Observation 4dc20609-ed91-4ef5-a216-9aa1bdba487c · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.981747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.837355Z digest=sha256:2a88449bde1d5eb04ca850d0a867b778ecdc69369f77a63b3468e50ccd4a95da

Observation 777fff92-f2f1-415f-9946-188530af05c8 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.841898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.841898Z digest=sha256:a472bc6dbfc12377de3108c12724271f20fa1bd7753d564ae874029b3612f07e

Observation c1c42d31-9ed7-48d7-8c42-8cefef7bf7ad · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.846079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.846079Z digest=sha256:76532fe36fba8d61502f94f084825e64554d5dab896371e92f42de46b1361e4d

Observation 1bb0e20d-945b-41d7-a4a4-30fc039d6810 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.957748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.851156Z digest=sha256:a3c8823d75ca26720bc83c77df6386847a9c28b2600c31a2f2d45df0951fb8c7

Observation 48d27688-5e9e-45c2-a2d9-51bd55bfa345 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.855807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.855807Z digest=sha256:af8b8a08183e42a6c7a5cf43d0b86e78ddfca9c37c00d0903cd0bb88acd0ef4c

Observation 71613651-96da-44e0-bf60-01d6a1a7ebf9 · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.860058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.860058Z digest=sha256:857e8ab5b63c9f600a7edc0d52fa07f0cb14cedd1171c614106ea53938e59395

Observation c0df93c6-3af7-4ce3-ad93-beb0d070b678 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.864617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.864617Z digest=sha256:a320783a609f67003ba2f5a5bbf71414cb55f174d863dd78054d3a334f00e146

Observation 6c168912-83aa-46d6-94cc-af00a997d0dd · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.868488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.868488Z digest=sha256:2282f46211140593526dbe3f8a114d11223867628a2a239ca86a956c7572b476

Observation c0586836-fa80-412a-920f-13e07fa9f098 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.873273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.873273Z digest=sha256:1fd12bf0d77dacbdc1594a52641d5d6eac79dfa762ef9f2f909138adda23ad00

Observation 656c26ef-ca14-4248-8233-0421db02e172 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.942268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.877329Z digest=sha256:90bdff8f27463aa1e5c2af4459b927c5a7b02233e79779d774e1839d51f8ffea

Observation 1d90440a-2dd7-43ce-8ece-082d594f8f6d · outbound

This paper cites online" 'onlinestring :=.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy online" 'onlinestring :=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.881524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.881524Z digest=sha256:a6ee8e9ac039a936ad6555e42369e09503c85e62e287f19b13a55629fc3a82c7

Observation 16f3e57d-a978-4b96-9d35-2180688f1449 · outbound

This paper cites write newline.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy write newline

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.886258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.886258Z digest=sha256:472a570eda012fb11b3893614a986f25d667a4d6fef3e0dda3a2acc8ccec4ffc

Pith citing papers

No inbound Pith citation observations are available.