Pith. sign in

Paper Citation Record · LEDGER

L0: Reinforcement Learning to Become General Agents

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2506.23667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23667 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:00.896630Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T21:51:10.972744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:51:40.816217Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bef74cd5-63f2-434e-a12e-c43090396fe9 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

L0: Reinforcement Learning to Become General Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.706854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.706854Z digest=sha256:94f6396415945ba3a0f27cc3443bc9d68beec1198d75e754be79ae35bbc1ea04

Observation 8608f17b-f4a3-47d9-a94b-ba8cc8b87810 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.778161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.778161Z digest=sha256:004f2d3d883b026fe9cb3a4fb3ba837efab13464fd82da671e3398e10de9b9f1

Observation 1ff586b0-bb82-4514-b5a2-ad612c94b987 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.773356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.884891Z digest=sha256:2b621a3e2751917130084e8e7aeca23111a395fc9a88552639cee0bb4eb7380c

Observation 899f4bf4-0f05-44c5-bd4e-30cd78ef346f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.597391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.933455Z digest=sha256:ba9bf4a87d46acf7c1dc5e306645d62ac9645434ecc23f45ee2221839ab6d1b5

Observation 72fa6a89-92e5-439a-b933-439aa24f0a6a · outbound

This paper cites RAGent: Retrieval-based Access Control Policy Generation.

L0: Reinforcement Learning to Become General Agents RAGent: Retrieval-based Access Control Policy Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.007761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.007761Z digest=sha256:821961276511c6aed78e9c4492298cf41832606fab250da4bed8743e32e317c5

Observation 7acf283a-68bb-4849-8f3f-ee404e39634e · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.084961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.084961Z digest=sha256:be589a026c7705f53208160c4e21ba089aa742bdcf341b9b86449d467df0d9ba

Observation fa7a83f0-e58a-4379-86a5-f54943bd6301 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.398920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.190054Z digest=sha256:4efc14f3b6f36e69cfd518576dd8d49c1cf14329408abf9897cf074f250c929a

Observation 0371d9fe-77f0-4a83-8b55-39be7cf64e4a · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.224600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.264906Z digest=sha256:79bf4d90b667c6e59740bee4e429fa35569106dbcf86111a023a3532c1fa29d0

Observation 9f9a0550-db7d-461b-b08a-5dabcfa900b1 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

L0: Reinforcement Learning to Become General Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.349407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.349407Z digest=sha256:6333fac6ac02dae1f5835111fb1db8541754c3455870e40b2adce016b22bd56f

Observation 8063a002-7b35-4e2b-bd16-40aca3581f7f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.072346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.408556Z digest=sha256:00d50554f431bf620c0fb4db491999632c7cb731825d1975b71dbe83e210f83c

Observation 10349242-37f1-4f2e-b007-f0b20d7ab42a · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

L0: Reinforcement Learning to Become General Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.471283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.471283Z digest=sha256:97f446bb5be9e6f6603e3ca0b5db8142d9f2087f35717af3b2f3e616d948bc08

Observation 6394dd20-57cb-42e8-b990-3c0703ba3894 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.893036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.537973Z digest=sha256:99df83cd6953e8e83a0945ed6fbdde5650659c208535d0b4aecf59eaee1054e2

Observation b553d063-6f8b-41cf-a758-fb45b1b0bfd6 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.720030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.588082Z digest=sha256:c6f63365f1e3c2ec860ab41bad96268758585652e3d370c600a4623b04236075

Observation f7fbddeb-c4cf-4e61-afac-b28e34b7a3a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.551841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.635927Z digest=sha256:8c5711f1a7ce443b9f1062144f6d9086f583e2c33a95963f506c24bfa7fc2a48

Observation e091268f-41eb-4008-bc86-6a39117b640e · outbound

This paper cites Measuring short-form factuality in large language models.

L0: Reinforcement Learning to Become General Agents Measuring short-form factuality in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.718465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.718465Z digest=sha256:c2aa7e2eedfdaaf6dfe579a476c78ecd4896b56fdd2bdd2aaa337701bb98455e

Observation 8a587171-c7c8-4a4d-b714-114a72425a2e · outbound

This paper cites Qwen3 Technical Report.

L0: Reinforcement Learning to Become General Agents Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.782069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.782069Z digest=sha256:af45ab882df8f40011c0a84eb414bf19c7cd9b5b0e9088d4ab6c0a1ac7b3626c

Observation 51f948a6-9df1-4632-a191-5819d4b31832 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.382091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.834498Z digest=sha256:058fba7c13ae370429e6083c3cc5fd015f3ebe87af2f4dbb0b4e611a83e2a1e8

Observation 9caf3835-4d4d-403e-ba6c-4f0e747b91a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.188222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.896630Z digest=sha256:7321dcfadcfa5a1492c5198c29ab5aff4120c8a41d6403577bee6141f444024f

Pith citing papers

Observation 7f6b2a6f-fb70-41df-906c-053903ef2397 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation L0: Reinforcement Learning to Become General Agents

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.820434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:2858f0a3be46bcbacc6b1663c8f0c3bf77fa2863cf818d22550946a7f2cb1261

Observation 1309c335-5c4e-4801-9ba6-78bcd1cdb231 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning L0: Reinforcement Learning to Become General Agents

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.725122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:1a950f0b6144bf6c88e17452ec9f830a26454cd3fd75cb924fef6e00e184480c