Pith. sign in

Paper Citation Record · LEDGER

Training Agents with Weakly Supervised Feedback from Large Language Models

As of 18 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2411.19547.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19547 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:08:13.270416Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0164537-341b-4af9-9018-a9ad87aa6898 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Training Agents with Weakly Supervised Feedback from Large Language Models Dota 2 with Large Scale Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.103779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.103779Z digest=sha256:d087f26098bc1d87da99e2a4ecc74c57047f7462e46f8e15755b9815283fc652

Observation 3911caca-315b-48dc-ab0f-151a45476f58 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Training Agents with Weakly Supervised Feedback from Large Language Models FireAct: Toward Language Agent Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.110427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.110427Z digest=sha256:c57d3d00b07cce6cca112571b28deb4579f3859ab15e393a75347becb89d771f

Observation ef04eb93-7dd7-4bca-bd77-6e1efd7c23fd · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Training Agents with Weakly Supervised Feedback from Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.115910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.115910Z digest=sha256:0dcf61b57a3f5bb4c017fb94e6ae44889735275860ab2334ade3057f85d94d76

Observation 6f793bbd-7632-4028-b1b5-75ba93f4de6d · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.122617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.122617Z digest=sha256:3b098623b86ad0b112eda1cc6537ebc3ee43fece64c46efdd6a9ee9824b89ccb

Observation 28587567-42b0-4924-987c-a60391b68760 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:08:13.766325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.128743Z digest=sha256:b83d8989acd751588d2d9a6d57684c7cbad00f411858c67eac90232e869aac10

Observation 718e2d56-b39d-4f36-be58-647422554ef5 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

Training Agents with Weakly Supervised Feedback from Large Language Models A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.133623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.133623Z digest=sha256:794434d497b1b34cce1af88186015ee76b587f9390857b615200ee169276d422

Observation db37de30-47ce-433c-a044-113fc4eb932f · outbound

This paper cites Large Language Models Can Self-Improve.

Training Agents with Weakly Supervised Feedback from Large Language Models Large Language Models Can Self-Improve

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.139576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.139576Z digest=sha256:7ad9922e13e35a9f503ed85b535eb28eeae159f94e9b9b8e5836fcee65f44b68

Observation 4d8aa783-376a-4892-b90d-fe4e86b880c8 · outbound

This paper cites SelfEvolve: A Code Evolution Framework via Large Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models SelfEvolve: A Code Evolution Framework via Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.144370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.144370Z digest=sha256:4fba5a81ffefb88a426a86898a9dca43fe952f38a826570819c6386039454e00

Observation 9fd38620-f1e9-4e65-b8e7-b22ecbfcf6a4 · outbound

This paper cites Self-training language models in arithmetic reasoning.

Training Agents with Weakly Supervised Feedback from Large Language Models Self-training language models in arithmetic reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:08:13.751304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.149088Z digest=sha256:63fd575aec3fa6c9adecc32be48535ed46bb8fe02c375a3ce9005a6998898feb

Observation 8b844dc2-c8b8-4dc0-9edb-81bc36c5c626 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:08:13.734739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.154116Z digest=sha256:507fe611b235347c4c738b43c5a77d643cf15b0750457a20d9bda907e41f4698

Observation 28d9802b-9719-4342-a050-5e0f7222eac1 · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

Training Agents with Weakly Supervised Feedback from Large Language Models MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.158944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.158944Z digest=sha256:4c5efdfeef7aa49b1cac8e8a3e2b7894ac4af9fb9df03e458020beddf1cfea55

Observation 7d6a8eb0-7b47-468f-a64e-8bce11b4c6ed · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.164434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.164434Z digest=sha256:da67cf8b75d7110472da22b730978b81efba3088a3c270e53d16fe6cc0796581

Observation b8711026-5546-4613-a603-0378ef86e67a · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Training Agents with Weakly Supervised Feedback from Large Language Models WebGPT: Browser-assisted question-answering with human feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.169686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.169686Z digest=sha256:f39c467a21bdd920941e223961670b7611c11541361b46776ad1329f29a0c6cf

Observation 22770d27-8490-46bf-b7e4-ddb966e6a351 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.175010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.175010Z digest=sha256:39e3dbeaf5d0578521802c54afc79198548192be55fcf39c1361a88341d7b8b5

Observation fc839b6f-707d-4c4d-b329-ff278a5b0229 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Training Agents with Weakly Supervised Feedback from Large Language Models Gorilla: Large Language Model Connected with Massive APIs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.179868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.179868Z digest=sha256:ecdc891c984c1569b56cab66eced01a4c1d8e5613b8c85141b8439a11ffa2556

Observation 8342b52c-e514-4a8a-9361-2ac05e32dcba · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.184620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.184620Z digest=sha256:5e985733fef3477eaa525e7949b4b392ad189e29b71754faee08d3a93bac41e0

Observation 464f3854-a11b-4c8a-9b4f-92286fd17e73 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.189463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.189463Z digest=sha256:8774e575b89c797f29c0a6622d4fda7df7a302b80d8a3b98c759081bdc0e6643

Observation 3e5d3a35-5ba9-467f-afd8-0d04ad5bd38a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.194394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.194394Z digest=sha256:e9a4b9a0737d8408ec79c069f6f425267014c36d465ae0a57959de8fd8759e0f

Observation 93d9b338-0bdf-4cd7-ae47-22bb95f431e7 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.199373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.199373Z digest=sha256:876cdf69e866125c1659f1f9c7296523ee3b759a68bd813a938c8857b34c4f3e

Observation ca4bf90c-b752-47e8-ab71-18f14afc5c18 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.204838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.204838Z digest=sha256:9f972f932ea9d7abf2464e535ac319f7966608dfd1c576d5c3defa54d672cec3

Observation 30f9efe6-4235-41ab-80fe-fae95410b525 · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

Training Agents with Weakly Supervised Feedback from Large Language Models ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.209522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.209522Z digest=sha256:1ad5d227dd1fb2b0e161d2ee7d37b818c0de8f6965fea39ee78d5e005c8fc8df

Observation 961843a0-f1d2-4a48-a538-4833fad57d09 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.214607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.214607Z digest=sha256:fca8998da6472e07ead9f4f8962e31372040f7a826b85a92c23394ee43b9fdb9

Observation 10a9cc14-2510-4922-9bd9-38ad76b28956 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Training Agents with Weakly Supervised Feedback from Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.219817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.219817Z digest=sha256:28332ca50536b5f6462fad540368ffac20eda3a15229c0308cf2a405eff09e4b

Observation d4bca0c1-da40-4e88-a251-c4d224b5cad8 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:08:13.662093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.225241Z digest=sha256:bbe7530a98d7952d9fd578e48c1e40dedd71e07215a466c0a9b4414450f89d0e

Observation 9ee9c83f-1709-4c35-82f4-122a2c3981aa · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Training Agents with Weakly Supervised Feedback from Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.230991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.230991Z digest=sha256:5c7cf54ef05d82abcb8191e745f304c9f6ec25b83ca4596b7f48f6f8eaef8123

Observation 66349724-564e-4ced-804f-1e789f21684b · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Training Agents with Weakly Supervised Feedback from Large Language Models The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.235856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.235856Z digest=sha256:fbf0d5711eba471b20ac20404f759ba78d47e6cc7baff7ef866a04c3b53c3cca

Observation c40f4bc0-521e-4a6c-aaa8-3033dc969e04 · outbound

This paper cites Lemur: Harmonizing Natural Language and Code for Language Agents.

Training Agents with Weakly Supervised Feedback from Large Language Models Lemur: Harmonizing Natural Language and Code for Language Agents

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:08:13.380613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.240119Z digest=sha256:96a7b18c749a440f03cf182974a08b6ed142c2b52b14c7b479029b811b9af611

Observation 56c74e4a-5805-45d7-af8c-af24eb062a9e · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.244892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.244892Z digest=sha256:320d14aeda6972509ad2566b255e0ee472ee82c855d7bac72d2d3b1122e34a44

Observation 2894738e-c67c-4bff-8d49-7d942bb8df60 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Training Agents with Weakly Supervised Feedback from Large Language Models Yi: Open Foundation Models by 01.AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.249386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.249386Z digest=sha256:bb7dbe7d51a8f8df56b1a323548cb186069bb360a1a412a583e4f4ea8e7ec1b9

Observation 8030f6ea-7901-4745-9615-cce2475dd0d9 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Training Agents with Weakly Supervised Feedback from Large Language Models AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.254686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.254686Z digest=sha256:ed3ad37978ceca7f96721d9ab9b4e1bbaba1b7242d0b2ffc361129bdec51b934

Observation af16f172-6a01-4b14-89ea-2613ee381586 · outbound

This paper cites Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization.

Training Agents with Weakly Supervised Feedback from Large Language Models Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.259699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.259699Z digest=sha256:08868d9c925cd450bdfcbd15d1a15888571787d272f2cf08f243715bc62c0cce

Observation 2395cfca-8a4e-477b-8f87-0948d7c215ed · outbound

This paper cites online" 'onlinestring :=.

Training Agents with Weakly Supervised Feedback from Large Language Models online" 'onlinestring :=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.264817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.264817Z digest=sha256:3602806a2b7a3d61d98749277d6c53ecbcf118a655de1658c659bcddd5348622

Observation 159d06ae-b93f-4975-ad0e-a008cba38128 · outbound

This paper cites write newline.

Training Agents with Weakly Supervised Feedback from Large Language Models write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.270416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.270416Z digest=sha256:6a518997d2e1285043554831c81806ed4b5c9ce85c420e9cdeb0ad77d16a235c

Pith citing papers

No inbound Pith citation observations are available.