Pith. sign in

Paper Citation Record · LEDGER

L0: Reinforcement Learning to Become General Agents

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2506.23667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23667 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:00.896630Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T21:51:10.972744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:51:40.816217Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bef74cd5-63f2-434e-a12e-c43090396fe9 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

L0: Reinforcement Learning to Become General Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.706854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.706854Z digest=sha256:94f6396415945ba3a0f27cc3443bc9d68beec1198d75e754be79ae35bbc1ea04

Observation 8608f17b-f4a3-47d9-a94b-ba8cc8b87810 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.778161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.778161Z digest=sha256:004f2d3d883b026fe9cb3a4fb3ba837efab13464fd82da671e3398e10de9b9f1

Observation 1ff586b0-bb82-4514-b5a2-ad612c94b987 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.773356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.884891Z digest=sha256:7d8e0419d8c2f029b4a89df0a8d199a10c776a2f4e0ce03e62f251fd5d80cb62

Observation 899f4bf4-0f05-44c5-bd4e-30cd78ef346f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.597391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.933455Z digest=sha256:f0570d2ad948bf76654d0dd389e9f80c1da36f523b458953dd685a11210c953c

Observation 72fa6a89-92e5-439a-b933-439aa24f0a6a · outbound

This paper cites RAGent: Retrieval-based Access Control Policy Generation.

L0: Reinforcement Learning to Become General Agents RAGent: Retrieval-based Access Control Policy Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.007761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.007761Z digest=sha256:821961276511c6aed78e9c4492298cf41832606fab250da4bed8743e32e317c5

Observation 7acf283a-68bb-4849-8f3f-ee404e39634e · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.084961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.084961Z digest=sha256:be589a026c7705f53208160c4e21ba089aa742bdcf341b9b86449d467df0d9ba

Observation fa7a83f0-e58a-4379-86a5-f54943bd6301 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.398920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.190054Z digest=sha256:2c6953c1290db4f0a1e62a431a43c8f76c39c42d6c6a7078e7afa7204e9c3a87

Observation 0371d9fe-77f0-4a83-8b55-39be7cf64e4a · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.224600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.264906Z digest=sha256:8c37340004e8de44fc96fefa1abad54e0bd1e81d958c64eaab756298ff7b098f

Observation 9f9a0550-db7d-461b-b08a-5dabcfa900b1 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

L0: Reinforcement Learning to Become General Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.349407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.349407Z digest=sha256:6333fac6ac02dae1f5835111fb1db8541754c3455870e40b2adce016b22bd56f

Observation 8063a002-7b35-4e2b-bd16-40aca3581f7f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.072346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.408556Z digest=sha256:46583348c4358485d2e36f217b8014b54d41534c37b074e7e8afa419c5055be0

Observation 10349242-37f1-4f2e-b007-f0b20d7ab42a · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

L0: Reinforcement Learning to Become General Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.471283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.471283Z digest=sha256:97f446bb5be9e6f6603e3ca0b5db8142d9f2087f35717af3b2f3e616d948bc08

Observation 6394dd20-57cb-42e8-b990-3c0703ba3894 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.893036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.537973Z digest=sha256:8875b9b41b14cfd51e5ab9e920a0f5a8b21c946c1c4f00fd161939dd8556d250

Observation b553d063-6f8b-41cf-a758-fb45b1b0bfd6 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.720030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.588082Z digest=sha256:5acf59c8fbf15e9ab9f183a629178c310175b6b3340709fca98b9e3cb8fc54a1

Observation f7fbddeb-c4cf-4e61-afac-b28e34b7a3a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.551841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.635927Z digest=sha256:f9912f1ebe57c043524f462a10729d721e363f2c4195623d060b9e4059ab03a2

Observation e091268f-41eb-4008-bc86-6a39117b640e · outbound

This paper cites Measuring short-form factuality in large language models.

L0: Reinforcement Learning to Become General Agents Measuring short-form factuality in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.718465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.718465Z digest=sha256:c2aa7e2eedfdaaf6dfe579a476c78ecd4896b56fdd2bdd2aaa337701bb98455e

Observation 8a587171-c7c8-4a4d-b714-114a72425a2e · outbound

This paper cites Qwen3 Technical Report.

L0: Reinforcement Learning to Become General Agents Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.782069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.782069Z digest=sha256:af45ab882df8f40011c0a84eb414bf19c7cd9b5b0e9088d4ab6c0a1ac7b3626c

Observation 51f948a6-9df1-4632-a191-5819d4b31832 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.382091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.834498Z digest=sha256:cc534ba98d6a6b2c0d0500ebbdcde4951ca51a1d4a24287eedb18333b02919e8

Observation 9caf3835-4d4d-403e-ba6c-4f0e747b91a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.188222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.896630Z digest=sha256:63285ab496f49cf86c9d7f1dab9fa8f26c2ac135d3a445aa2b75a474910d0bae

Pith citing papers

Observation 7f6b2a6f-fb70-41df-906c-053903ef2397 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation L0: Reinforcement Learning to Become General Agents

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.820434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:bf5af88b8a04c7542f66583eab3b0a9d1b2b8b7a0ec7a1c57e299851e483b037

Observation 1309c335-5c4e-4801-9ba6-78bcd1cdb231 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning L0: Reinforcement Learning to Become General Agents

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.725122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:2a22d61cc75681782c6cc3f23564a5f540d03933f20876de1aa7916de332ee39