Pith. sign in

Paper Citation Record · LEDGER

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

As of 19 August 2026, this Paper Citation Record lists 100 of 295 outbound references and 0 inbound Pith citation observations for arXiv:2607.21793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21793 v1

Coverage vector

measured 100 of 295 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:48:31.213384Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 295 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 730dca30-1ac0-4e30-ae2b-17eb3ba55269 · outbound

This paper cites DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:16.678850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:16.678850Z digest=sha256:9093a7e122582e4572d699f463e890e5008e8a14481f36ed43d57266a1ae5f67

Observation 2dae9d61-fa00-4fa7-ad58-5d5af9c2354e · outbound

This paper cites Advances in neural information processing systems , volume=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Advances in neural information processing systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:16.733341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:16.733341Z digest=sha256:545d19103baf93a8d77fa61f96b82593e51c6393ecdbf5c544a17fc8be1ab4a6

Observation 3071401e-40f4-45d2-a28d-7531fd8fb165 · outbound

This paper cites 2023 , eprint=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2023 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:16.826143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:16.826143Z digest=sha256:2c5617fbcf2efed721dae01ccdedef052cbc8ce37b5f65a2427c340e5c4a26ff

Observation c98d662d-f60e-4985-81e5-0a7a94553563 · outbound

This paper cites 2025 , eprint=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2025 , eprint=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:16.952557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:16.952557Z digest=sha256:8cb0dc9b109ca804c2357f2c494c45de2068e26829571a8e25b6b5aa9a674f05

Observation a006505f-6d53-4019-8b16-a95198a2106d · outbound

This paper cites arXiv preprint arXiv:2509.18521 , year=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization arXiv preprint arXiv:2509.18521 , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.095474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.095474Z digest=sha256:d1aa6e8bf59253915d26f80925834157f989d071de1b4753514d3634848a382d

Observation 5df58ff4-7bb4-48ce-af43-85fb27c606db · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.150050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.150050Z digest=sha256:a9bad1a4c936cd137642eac3cbcb4c55ac45390a1589e11291bb89671226797f

Observation 8dcd30dd-5f01-4957-979c-a3a7d152910c · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.223313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.223313Z digest=sha256:24e36616207c2e05e220b9688743532f7f42c7ec96067614e8f3fd66258a8b92

Observation 4cd5dede-04e4-47e8-a00a-115009c85765 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.311342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.311342Z digest=sha256:e3c167fd3dfed5a39d865cdb5c24bd477f1889626672ff53ea78ee3771bc4258

Observation ad6a37c0-ab62-44c1-9b77-863eed331d99 · outbound

This paper cites arXiv preprint arXiv:2509.21009 , year=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization arXiv preprint arXiv:2509.21009 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.455645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.455645Z digest=sha256:ffee547b43ce55c2cd9dc2ee9acd661bff9804972f4d9305a1a3660ac912dc70

Observation 60ec856e-94c0-4350-bf9f-4bfffd6cd99e · outbound

This paper cites arXiv preprint arXiv:2510.11345 , year=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization arXiv preprint arXiv:2510.11345 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.573665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.573665Z digest=sha256:0dc9f27787f0e886fc301e0a2a9a1e1d61ebc19ca7e97d01141904b35ee83c19

Observation bc987d12-9c0f-4ba4-b4f1-b067647d1df2 · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.765327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.765327Z digest=sha256:a9104225537918e4f8e725f58ac78e2c67a438b0e2d6e29256af4c6e528f7be2

Observation ecca540d-4d77-4e58-825e-2fe83203497b · outbound

This paper cites AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.927596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.927596Z digest=sha256:ebf689560208254e6949861aee7e33cf0e89d4edefaee37a6a460211a650d835

Observation 3f176268-ebbb-498f-a953-8ba9718194a7 · outbound

This paper cites History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.067189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.067189Z digest=sha256:8b44bdbf6435f5aa3707236b642a98019769f4e5916a489ade9567c712ee5c88

Observation 2ecf6c05-00e6-4f71-925e-c9666b6c47ba · outbound

This paper cites arXiv preprint arXiv:2510.12633 , year=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization arXiv preprint arXiv:2510.12633 , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.186367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.186367Z digest=sha256:612d2976d27f1174574107e89ae8d813a208bee395633e0a41f9bc5af01cfbca

Observation e2bdb299-5d36-422f-9ada-e8d8f3cbf42f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.278039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.278039Z digest=sha256:6ab37ac7cd58d4389e4638e90e5d6be9253682e1ed34f5c64c13559f7869c4b9

Observation cc8c3976-a21a-4dce-b251-ff605ce939a4 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Kimi K2: Open Agentic Intelligence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.387148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.387148Z digest=sha256:2a7a61a4eeff81936e60b3c0f973543fe08c996ae4390d387e788903280263da

Observation 3069c028-7ba3-41a3-82d7-662d311e9b8f · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.460683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.460683Z digest=sha256:46b7dd9d13323b80cbf7740067d3abaaccfb50b21991924506ac155cb020b926

Observation 9a823554-4551-44fb-b8b6-a66682107a08 · outbound

This paper cites 2025 , author =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2025 , author =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.589656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.589656Z digest=sha256:e2500ba8ce1336f01e2f4ea8f3bd27855fc95206cb7a4a0281fdba92216e69a8

Observation 97326a05-8a72-4959-9f8b-b91a05796733 · outbound

This paper cites 5-thinking: Advancing superb reasoning models with reinforcement learning , author=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 5-thinking: Advancing superb reasoning models with reinforcement learning , author=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.682757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.682757Z digest=sha256:e6c9268b026a389d51080ddaf29c80a1835fb4e51520f060bce81336868d28cc

Observation 89f6d3aa-e164-49d6-bc60-72ee3c4b8db4 · outbound

This paper cites OpenAI o1 System Card.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.811803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.811803Z digest=sha256:d9411f310241b1ed447d20180d02e62358525064dee346aa627df44f9f0667b8

Observation ebca9c4d-3346-425c-977b-5075038ec0c4 · outbound

This paper cites 2025 , author =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2025 , author =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.894432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.894432Z digest=sha256:ee2ffd7a83ba6e1092107b0327294cf39a31b74f016d0d785e6faee2f464b975

Observation 6893356d-40d9-4c82-b43b-a8e7514d04a3 · outbound

This paper cites 2025 , author =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2025 , author =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:19.060983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:19.060983Z digest=sha256:011a1c1780b682e7faff335398a9c004c3aa0007cb40cfc19f6228de233c5d26

Observation 6468e8b6-a33d-496d-b2cf-5a9cd94cb397 · outbound

This paper cites 2025 , author =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2025 , author =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:19.192118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:19.192118Z digest=sha256:33e9a0cecdfcc3c4d7245419f8758f8f4599797fac3563f9b84041d7a06a412b

Observation fef6798f-043f-46dd-a766-2c883a648f71 · outbound

This paper cites 2025 , author =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2025 , author =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:19.346898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:19.346898Z digest=sha256:e682adbecb4fd7d166a1d873730543d257c67b22f6b162226ca24feb522820fc

Observation 9d10e571-88bf-46ba-8bc6-9843a60311c8 · outbound

This paper cites 2025 , author =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2025 , author =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:19.520246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:19.520246Z digest=sha256:b1398f34c8ed6c0fd411a4955cf673d2494a7465e142fc2d04500fb706c37021

Observation f009fecf-ce6b-44c1-8297-2464b743c73e · outbound

This paper cites Qwen3 Technical Report.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:19.652409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:19.652409Z digest=sha256:684341b03f763d3cf364f0b96427841c4a0636040694070a411f07ffdf1e8aad

Observation 05c132e5-9a79-4f71-9ae7-ffe57a3f4bf5 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:19.850653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:19.850653Z digest=sha256:d07bf7766ef9c9617ea6458fc8cd9c3a1b948c96a15b689f341f77ad1bf4daa2

Observation a84f644d-1d01-4c94-b618-fc75e3aaa21d · outbound

This paper cites , howpublished =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization , howpublished =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:19.986483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:19.986483Z digest=sha256:b8a276bdb76a008f7dc501096f0253c15118a80e45b7b1bc89e31c0f5cd3deff

Observation 39cea982-f41a-4213-bd56-5292db9e36b6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proximal Policy Optimization Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:20.108214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:20.108214Z digest=sha256:448a569351192d3482c8f6be005d624a14372ba330a1f33ffe04f01348d92ac7

Observation d14ca402-6671-4899-9e44-88c6f5bb392e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:20.278182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:20.278182Z digest=sha256:c8abad267e7f03ad5da418b5c6ca76f76cc89687c6c891554779a95f803fac35

Observation 3c9f35f7-8376-4959-b82a-baba6ba5c515 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:20.440300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:20.440300Z digest=sha256:9b2f21cdb0ee222487f3090130c4c8aaee7d3cfc72f9a185397ebf7e3585fe9a

Observation 53419a6a-e742-48f6-add0-5e7a470e019c · outbound

This paper cites Notion Blog , year=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Notion Blog , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:20.565573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:20.565573Z digest=sha256:cc1f281feb5cdef1b8f41fb3fde77c4720f18b18526881d62e8e7c060d17a23c

Observation bdfbd23b-38b7-44c2-bf9c-99ff829ae9c6 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:20.696913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:20.696913Z digest=sha256:e70a036e605cb72d0621b50d28bee8a86973ad835880853089ff5317a4674009

Observation 946b68e5-08df-4dbd-a21b-737f24eceb68 · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 34

Resolution
verified exact
doi, observed 2026-08-01T06:53:45.059650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:20.793420Z digest=sha256:ccae71221ce8bfc205b3da343767494a8bf94b07ad22be6a4c1bdecf5c712c10

Observation 8aebadd1-4328-4c92-8dc7-dcb14391ffa6 · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:20.928710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:20.928710Z digest=sha256:3f421713d94ee7a139bffd0272cfd251700cd50136a2253045e6920293ed8a65

Observation 5213734d-b2cb-4a2a-ae08-32685fefd028 · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:21.055312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:21.055312Z digest=sha256:b096d2ad6498e02d18d9220765378d6be8364adb4aa05e4ed438bc5d1c6ec062

Observation 8d0665fe-1f71-42fc-ae9a-91433af58654 · outbound

This paper cites Data Sci.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Data Sci

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:21.189664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:21.189664Z digest=sha256:4f5cc0301c8fd2e40213e41253ce21dba8d063a8a8c2399c1ec7a675724c6153

Observation 2aa8d52d-edf6-46de-b007-6086a482d32d · outbound

This paper cites 2024 , issn =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2024 , issn =

Reference 38

Resolution
verified exact
doi, observed 2026-08-01T06:53:45.039697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:21.374831Z digest=sha256:0dea8b792bd63378679d0edec0dd9139ef6f009f6439204e652b2f1025d2dcd3

Observation 8428b6db-1160-4059-9dc2-2e239a58650f · outbound

This paper cites Advances of Pipeline Model Parallelism for Deep Learning Training: An Overview , journal =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Advances of Pipeline Model Parallelism for Deep Learning Training: An Overview , journal =

Reference 39

Resolution
verified exact
doi, observed 2026-08-01T06:53:45.027440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:21.525119Z digest=sha256:0af13009d04fe95dde136c0b0124fedfbe768d9c02b5996c2866b2fa988cccdb

Observation cf32b25b-7e98-4067-9363-541be3aa67fe · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:21.639547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:21.639547Z digest=sha256:77afed3efdb7dea0d99f4ff00faeb5f428600e50e9d5442db11f00423ba53d07

Observation 972c8d87-f9e4-4004-b7c2-d27ddf053f33 · outbound

This paper cites Big Bird: Transformers for Longer Sequences , booktitle =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Big Bird: Transformers for Longer Sequences , booktitle =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:21.807410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:21.807410Z digest=sha256:842b4a9d35d5d90f1cf28e24945d0deba70ae472d0448a86672892a77fd91b66

Observation 73ae6671-9688-4ef3-8473-eb002fdbbf01 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Baichuan 2: Open Large-scale Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:22.009226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:22.009226Z digest=sha256:c7c800f855f9c407ca7b8661f0e8ea5a3df68acbe8e1b04e9f18b198275f9e70

Observation fa685c76-f8fb-4ad9-ace7-2c8fcb2be258 · outbound

This paper cites The Llama 3 Herd of Models.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization The Llama 3 Herd of Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:22.134379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:22.134379Z digest=sha256:dc072d407ff2f534346dd221e3e02329d076aa88eaedb7f2f8de0cab21736ae4

Observation aa52407b-62a8-4b85-8a4f-74b747b76e3a · outbound

This paper cites 2023 , eprint=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2023 , eprint=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:22.456520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:22.456520Z digest=sha256:ce112441bf0b5de9f4ecbddd5def9a1f07ff55d4e8680d5beedb5f2cabc63d73

Observation 8cc9a333-5718-4ff1-aef4-591522e358cc · outbound

This paper cites Proceedings of the 2022 International Conference on Management of Data , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proceedings of the 2022 International Conference on Management of Data , pages =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:22.817150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:22.817150Z digest=sha256:16323d0f3446b614d1ee25ac8641d39d0815e3f595326972aa88c6abeb768b34

Observation af8e665c-bd7c-45ec-a012-2d9155bcac3d · outbound

This paper cites Proceedings of the 2021 International Conference on Management of Data , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proceedings of the 2021 International Conference on Management of Data , pages =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:22.999024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:22.999024Z digest=sha256:7250b954a8d467c0116e51531a1b5607c2c481b420cefb24b39fb28bd71722ac

Observation ebadc26b-5beb-444c-8020-42e37a0bad1d · outbound

This paper cites Forty-first International Conference on Machine Learning,.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Forty-first International Conference on Machine Learning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:23.172019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:23.172019Z digest=sha256:bb95a14ff96ca9fc77c0c91d409220c38eb139253bb3989f88cabdf03afb6f07

Observation 34e710bc-2a0a-42b0-8975-2b94628390dc · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 50

Resolution
verified exact
doi, observed 2026-08-01T06:53:45.015569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:23.292403Z digest=sha256:9969a2737b40f2d76cac7c5bec51963bb8355ca4c4df568c00437ec3b03b6c4a

Observation b3c89438-2031-4ec8-9aed-fbc0aa5e7c0e · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 51

Resolution
verified exact
doi, observed 2026-08-01T06:53:45.006116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:23.476245Z digest=sha256:34f3fd5c3ce191aa992c7d9d6560f8fa4703ca377e066255b29893b7e26f93ad

Observation 46195bd6-c61b-4c76-af21-dc2fffbe4ae5 · outbound

This paper cites Proceedings of the 2022 International Conference on Management of Data , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proceedings of the 2022 International Conference on Management of Data , pages =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:23.652852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:23.652852Z digest=sha256:d2967dcf6d1db4ad6381a84fd9783c3316216c4fab5e5708ec89ad20c85d73c6

Observation 604b023d-6c69-4e93-b42e-06ca0e1ecc98 · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 53

Resolution
verified exact
doi, observed 2026-08-01T06:53:44.995372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:23.792234Z digest=sha256:50a163d45a45debe86c8089fb8cb389a26326fe7ce91ef52722b62296dd4a8a1

Observation e4071d2c-5035-47d2-a93c-92e803e18c4a · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:23.933805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:23.933805Z digest=sha256:a1e86e6cc268bfa4e3e2e2c0e1774e7832ac6ada7db2272feded9f994e262ae9

Observation 2ce3e0db-ec3d-4dd4-a5c9-bcf591e8e843 · outbound

This paper cites 2009 , isbn =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2009 , isbn =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:24.103766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:24.103766Z digest=sha256:ae578d4382364f1c764592727d4d3dacb02e4c62eb3a14c00f1e9ec32c2a71c5

Observation ff7e6146-e46f-4aaf-bd98-d12152445f7f · outbound

This paper cites Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , pages=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:24.241588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:24.241588Z digest=sha256:70fa7d227ac219b373768c8ee9e820486116be005b706e265fde0b22cd0acbb8

Observation c3cdc54f-0b07-4a50-8da2-15c8ee6a0860 · outbound

This paper cites Proceedings of the Eighteenth European Conference on Computer Systems , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proceedings of the Eighteenth European Conference on Computer Systems , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:24.387022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:24.387022Z digest=sha256:ec4c8a1a5ed7fba1d5f24bc2312e1994b401d73771647ab9ba748beac173b492

Observation 3cfa6997-b49a-44ea-9811-125e03e2f02c · outbound

This paper cites 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , pages=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:24.510917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:24.510917Z digest=sha256:3558cb6b508f95a20d600c7e322f2052c4379a718c4cb32de0e7fa080d83cfe4

Observation 44b7b8e9-2396-4869-a7c0-fe7ab5475b37 · outbound

This paper cites Qwen Technical Report.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Qwen Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:24.689017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:24.689017Z digest=sha256:f88b175ec1c0fe6910561abc5ead2d3208ac2f49b6e54d28c8423061d8ff4bf6

Observation 194838b9-db7b-47ec-876a-d76b29b232f5 · outbound

This paper cites Gehringer and Daniel P.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Gehringer and Daniel P

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:24.788593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:24.788593Z digest=sha256:ea77a06c59ee63fd77c9ae22df70735913e4624942d04473a0caee5af6e89c32

Observation adabdd42-3354-486e-82f6-b152a7422f7f · outbound

This paper cites 1992 , isbn =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 1992 , isbn =

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:24.974725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:24.974725Z digest=sha256:2671c137ab244815a4794a1b4bacc312d75097196f409027368652fb06cee6a9

Observation 8991d225-1c38-4467-9a54-b0e78760e51b · outbound

This paper cites Kovalyov and Maciej Machowiak.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Kovalyov and Maciej Machowiak

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:25.125782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:25.125782Z digest=sha256:6635dbbe55760318721b7a43bcf52a227f63099a9aaaa7a21911f1a0046fd115

Observation 71d54759-8ecb-4665-afbb-348a07a13158 · outbound

This paper cites Annals of Operations Research , year=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Annals of Operations Research , year=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:25.199734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:25.199734Z digest=sha256:920f0dcc932ac4ce5a33ed2ab720d50aaf462ec503f4152e18c5692c5105ef5f

Observation 4f9dfeb7-667e-4abe-971a-06dd5ceb9796 · outbound

This paper cites Optimization and Control of Dynamic Operational Research Models , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Optimization and Control of Dynamic Operational Research Models , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:25.316083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:25.316083Z digest=sha256:f80e7e10350e9425dae248617cd7a0b86a805b2e9f097c5915cfe669d182b454

Observation 55a0ae03-3a2c-46fd-b44b-47b912bc3550 · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 65

Resolution
verified exact
doi, observed 2026-08-01T06:53:44.967511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:25.491619Z digest=sha256:4e139f91b67c574d9d5508c1aa244a248157de60933e5542b481930b6085a37e

Observation 7494cd29-ff52-4641-8c41-62244ea16002 · outbound

This paper cites Gomez and Lukasz Kaiser and Illia Polosukhin , title =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Gomez and Lukasz Kaiser and Illia Polosukhin , title =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:25.671198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:25.671198Z digest=sha256:03c60af90968392ca1dd4f2d18e350a909dc945b9f1102e656d9eec93a839b2e

Observation 9767258a-d542-46f2-9d4b-644b7a7f1537 · outbound

This paper cites wav2vec 2.0:.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization wav2vec 2.0:

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:25.827498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:25.827498Z digest=sha256:37746e26722ba4ba4e569c965d17c60ca2cad68134441343b1ab714aaf545ced

Observation db24f521-4b52-4d44-af97-c1c7f884a9ff · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:25.998800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:25.998800Z digest=sha256:92ee5ed7744115cff6dce32e5ed936cd9013235a2733aa4092d3c9c48f31b9eb

Observation b0d89cf1-f9e4-4e5b-829b-a7310da19006 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training , booktitle =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training , booktitle =

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:26.122269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:26.122269Z digest=sha256:417829f30fa7040b04f5899d162c76fb4b66a5154522c3231082f2869bdbf6c8

Observation 2b21a471-863f-4e51-8355-b25d958f7726 · outbound

This paper cites Large-Scale Self- and Semi-Supervised Learning for Speech Translation , booktitle =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Large-Scale Self- and Semi-Supervised Learning for Speech Translation , booktitle =

Reference 70

Resolution
verified exact
doi, observed 2026-08-01T06:53:44.957144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T06:48:26.252524Z digest=sha256:b18b3d5acb6e8b034899f29752a7ec869a72f37fc221e3decbc1197c3c45750d

Observation 0933a20c-d9d2-4fdd-9edb-2780946baa5c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision , booktitle =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Robust Speech Recognition via Large-Scale Weak Supervision , booktitle =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:26.382795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:26.382795Z digest=sha256:d7e8d3bc8a99f5a1468e8acf07bad1d96b7190c5ef2955d6c1c86983c0afd22e

Observation 889c2fc9-ad58-4cdd-9b47-9dff8769d8e4 · outbound

This paper cites The Tenth International Conference on Learning Representations,.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization The Tenth International Conference on Learning Representations,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:26.511229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:26.511229Z digest=sha256:144b75378dd12117a55441358abd130d52c7875c5532ec548912d33610adb718

Observation 9e5b5479-cfe7-415d-ae15-0f4966110247 · outbound

This paper cites Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , pages=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:26.652841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:26.652841Z digest=sha256:5086b9f9c603dd9f04eb04e5fd1a5cd9dce2f173b7389ff78398b1ef3a160a54

Observation 809e6696-f35a-4bce-93e2-8b5b67599f79 · outbound

This paper cites Optimization Methods for Large-Scale Machine Learning , journal =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Optimization Methods for Large-Scale Machine Learning , journal =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:26.840288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:26.840288Z digest=sha256:93c76ab273365556226ef08de66bc871c3a24c5d9f6f3605ae8102b29bb33a93

Observation 1a117779-5100-430d-91eb-497fc1f32260 · outbound

This paper cites ICML , volume =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization ICML , volume =

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:27.028317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:27.028317Z digest=sha256:e225d9d8e478aec34648293d5f549bdf45d2cdb70d92bf451ca979241e8f8d2a

Observation 37983ffb-a50b-4e68-8af5-df44feea2556 · outbound

This paper cites ICCC , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization ICCC , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:27.178611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:27.178611Z digest=sha256:2ffd4aa12acdd88ef7314c84e08fee5f0eb9d595d24aeee1bc71d610aab92fc4

Observation 855859ae-48c3-4212-a9c9-0a6bc493c26b · outbound

This paper cites Mechanical Systems and Signal Processing , volume=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Mechanical Systems and Signal Processing , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:27.375003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:27.375003Z digest=sha256:cd9272e069e50d7bf5ec9ec77936108047c2d9d344dbf6262da8488174bf635e

Observation d188cc07-0f75-40c3-8092-cf9a02deb696 · outbound

This paper cites NeurIPS , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization NeurIPS , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:27.513786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:27.513786Z digest=sha256:bf6da808738afccc3f489011e11dfe97be6f14680cd9e5049dff8932ff0958bb

Observation 59e81f45-3b42-43ef-8aca-a71f5b1adc9e · outbound

This paper cites ICML , pages=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization ICML , pages=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:27.714608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:27.714608Z digest=sha256:40d7103a8e72e98bfb2eeaaad52b3f39a5f43f2abcd14512b81023d8eeae6dad

Observation 243179c4-64ce-4de4-b319-fcd70e86b21e · outbound

This paper cites Andersen and Jun Woo Park and Alexander J.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Andersen and Jun Woo Park and Alexander J

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:27.862289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:27.862289Z digest=sha256:1bb4124cc439b92098d6a138e828f4b276b94bcf21aba462df34ffd589a8e7ed

Observation 0753f9a1-5fe5-4b31-b815-4dedc3770ff1 · outbound

This paper cites Taming unbalanced training workloads in deep learning with partial collective operations , booktitle =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Taming unbalanced training workloads in deep learning with partial collective operations , booktitle =

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:28.061742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:28.061742Z digest=sha256:893bfdd2e297ff44df01fe6da5d9b57457dc31ea4a7c2430305428d97c9c2530

Observation 9fb47e1d-6851-49b9-9550-9efe17c0c1d8 · outbound

This paper cites ICLR , year =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization ICLR , year =

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:28.212237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:28.212237Z digest=sha256:118306474bf1fd384b594c847ace793be1ae67bb2d1fa4d9e2f7e271f6110685

Observation d657bb5d-061f-4630-884d-48e881f4c224 · outbound

This paper cites ImageNet:.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization ImageNet:

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:28.382045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:28.382045Z digest=sha256:cc673f0393cbda84f79ec04733c8d87d04f5d58b405f79843ef9639ed3e8ebab

Observation 482c5b41-2c73-485f-a4e4-370b1e3b4761 · outbound

This paper cites Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis , journal =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis , journal =

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:28.464126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:28.464126Z digest=sha256:81e931e863b8317e1fadad3abdf1a3fe6c2b41d2b16b394d5af1339b8086557b

Observation de7604b1-df12-47d0-9061-586cac9673cd · outbound

This paper cites NAACL-HLT , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization NAACL-HLT , pages =

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:28.655173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:28.655173Z digest=sha256:5dc5cbd91959b8388137628e694d7befee234f1d268dc779efb1028c05cd4032

Observation 9f6246a3-f105-4ea9-aa20-18531db9e8b3 · outbound

This paper cites SIGMOD , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization SIGMOD , pages =

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:28.811138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:28.811138Z digest=sha256:53a06d3eab970cbe2fa30744b5220ddb3fe7fb16b45d3b91858148d219fc92d9

Observation 9501a609-af27-431c-9d9c-bdd019d77f22 · outbound

This paper cites SIGMOD , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization SIGMOD , pages =

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:28.999778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:28.999778Z digest=sha256:83845f6295190599c524d5154c4242aefe49dba74486db11c505a0a78402410e

Observation f86578c3-52a9-4c71-a1ec-c90f34d6a292 · outbound

This paper cites DimmWitted:.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization DimmWitted:

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:29.135349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:29.135349Z digest=sha256:18226adb68c286cef381c769d79daba545c9e575933cb811b3061593fcb985fa

Observation 075f639d-7e21-42b1-9309-331f440e8f30 · outbound

This paper cites SIGMOD , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization SIGMOD , pages =

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:29.280545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:29.280545Z digest=sha256:2d30c84aff258af995c32a16343de148646c1cc10e58f559a46557107c6a833f

Observation cf75ecd1-e47a-47db-ae12-c9562b25199b · outbound

This paper cites TensorFlow:.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization TensorFlow:

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:29.417956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:29.417956Z digest=sha256:e64c671c466d3fe3d51daabc991fa06a3e7276784e05d9d209089c1d4689be51

Observation f117424c-8f3d-4fed-a377-9e0dc843f0b7 · outbound

This paper cites SIGMOD , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization SIGMOD , pages =

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:29.586098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:29.586098Z digest=sha256:a7ce7332c468fb019968f7c2b6cbcd2e1dd858408730a34e8fc56359369da4d0

Observation f5f24593-9dc5-426f-bdfc-206e7906f3d5 · outbound

This paper cites Gibbons and Garth A.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Gibbons and Garth A

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:29.750868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:29.750868Z digest=sha256:97559e4b305398ad63654860e8bcb3e97663060b857bdcda58261c1f35c8b212

Observation 22ec1fe5-53a2-457b-8ec1-bc234f38b606 · outbound

This paper cites Ganger and Phillip B.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Ganger and Phillip B

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:29.881896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:29.881896Z digest=sha256:3b82255c286bb4c160019920d9aeb700da749c866ada000e0f661938639b9356

Observation c5595b54-dc8c-4e34-a563-ea432d492613 · outbound

This paper cites Brewer and John Wilkes , title =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Brewer and John Wilkes , title =

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.052849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.052849Z digest=sha256:231f5186aa0324597e374d8a9333e2b305bfa98ea0899ad5409890dafc7548c1

Observation 7dc3e747-8c97-4ac5-bb07-e03a6e26d475 · outbound

This paper cites OSDI , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization OSDI , pages =

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.154090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.154090Z digest=sha256:2e29c1aaf7bedd3fdbb25e65b53c544c3483b6b8c5882e65cd25a7a7017d622b

Observation 83824867-0535-4fd7-9770-41b2e3dbf70e · outbound

This paper cites NeurIPS , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization NeurIPS , pages =

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.211586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.211586Z digest=sha256:3f5f967729438fefcb099c72ca13a940c4f28552b2b6c803732997dd7f7d43ce

Observation 51fe7cbf-fef2-42cc-a0df-09c260ce33d4 · outbound

This paper cites an unresolved cited work.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.349214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.349214Z digest=sha256:5c242a6e212b19977d782bef172a32221cae73b8409776153c4defab65632811

Observation c11d2c1b-1629-4eb8-a08d-5808c242e80c · outbound

This paper cites CoRR , volume =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization CoRR , volume =

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.508705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.508705Z digest=sha256:eaa9ea393dd727838e1fa6633e015536966e0e19842c4250354c923e954b53c9

Observation 2853d193-fbd4-42d6-b1b6-340539ec93e8 · outbound

This paper cites Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks , booktitle =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks , booktitle =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.672121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.672121Z digest=sha256:52f405025330c76898c1c4a59553811aae3480327cbcf8840f75cca0bc4eb0d0

Observation 787d79dc-7118-4c7d-bae6-a986d90e989c · outbound

This paper cites NeurIPS , year=.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization NeurIPS , year=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.847372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.847372Z digest=sha256:08b9445d02883df2bb45612b4a3225d914520ac888a9c8ee4abb806792242cb8

Observation 6354e144-ba4f-445e-a293-f76a8b106cda · outbound

This paper cites CVPR , pages =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization CVPR , pages =

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:30.967573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:30.967573Z digest=sha256:ec9e06f9d2156babdfe0d46594f31932bd94a751f748259ff4b283bfa5b62f73

Observation dbfdc495-44cc-4719-8533-fd1acd75edb5 · outbound

This paper cites Weinberger , title =.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Weinberger , title =

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:31.213384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:31.213384Z digest=sha256:490b84c8ce21e36e6e87d89b109f37da9110d12b1af4288e40092a371ea20c8c

Pith citing papers

No inbound Pith citation observations are available.