Pith. sign in

Paper Citation Record · LEDGER

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

As of 18 August 2026, this Paper Citation Record lists 97 of 97 outbound references and 0 inbound Pith citation observations for arXiv:2608.11772.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11772 v1

Coverage vector

measured 97 of 97 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:33:33.662364Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

97 of 97 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09d97e4f-ced6-45d7-b2c7-f740760deaba · outbound

This paper cites Many-shot in-context learning.Advances in Neural Information Processing Systems, 37:76930–76966, 2024.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Many-shot in-context learning.Advances in Neural Information Processing Systems, 37:76930–76966, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:31.971243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:31.971243Z digest=sha256:b10b6c1eaaa57621d073cc8a0bb7f0f890e1742322a3e0e69544e9f67c5da1c7

Observation 4cd93409-239c-459f-aabe-a81841f4d6e6 · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction AutoMix: Automatically Mixing Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.022958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.022958Z digest=sha256:341ceababab17771589d4b8ff62657f4ba7428db65ba150172401584a1d601e6

Observation 0b72f677-7202-4924-9558-e84a1d72653a · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.077850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.077850Z digest=sha256:430c7f632cea72a29f64409a266bdf2168b1b3be61323dea97a4263f28c97cd8

Observation a557d41f-20aa-4b00-817f-a22960f39aba · outbound

This paper cites InProceedings of the 62nd Annual Meeting of the Association for Computational Linguis- tics (Volume 1: Long Papers), pages 12248–12267, 2024.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction InProceedings of the 62nd Annual Meeting of the Association for Computational Linguis- tics (Volume 1: Long Papers), pages 12248–12267, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.137007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.137007Z digest=sha256:40e7eba89666aa8ce707f32ccca6217a90309e4aa1549954fe23118d8c0d571d

Observation 189ba2c5-240b-47dd-8b9b-3248219af6b5 · outbound

This paper cites Self-RAG: Learning to retrieve, generate, and critique through self-reflection.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Self-RAG: Learning to retrieve, generate, and critique through self-reflection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.140785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.140785Z digest=sha256:9b050989c72f8ebc2d4fea227b7ff6b5c2cbd32466ac608e3a636bf8e7952fa0

Observation 80c5210e-e431-49a4-bc98-06596e564ebc · outbound

This paper cites Program Synthesis with Large Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.144679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.144679Z digest=sha256:c0ba6a1fd0b03985ff2b08f5dee4fb5f24ad6d52be626145dc72d584c7781036

Observation a9a07f0d-e26c-4c80-aaa5-58032f277c49 · outbound

This paper cites Digirl: Train- ingin-the-wilddevice-controlagentswithautonomous reinforcement learning.Advances in Neural Informa- tion Processing Systems, 37:12461–12495, 2024.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Digirl: Train- ingin-the-wilddevice-controlagentswithautonomous reinforcement learning.Advances in Neural Informa- tion Processing Systems, 37:12461–12495, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.148827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.148827Z digest=sha256:a739a7d388239cec46728d28589a7445f7815c5e6930a7f48b9519c4f269bdf1

Observation 32950d6a-7f69-4603-b4ed-89caaa41b9e3 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.153242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.153242Z digest=sha256:4f8ee3c4360865085c662851d40774d1e86d9192512f2a7e4901950a3d61c5e8

Observation 47acc963-bb6c-4f64-9f31-3004eb87f8b5 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Longbench: A bilingual, multitask benchmark for long context understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.157190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.157190Z digest=sha256:4fb744b42868018d5102324170e7b8d040f7ff97c7f6fe011eb1226bb20b6c3a

Observation 5cbd9b77-0cb8-445d-be94-2dc0f90a379c · outbound

This paper cites CodeT: Code Generation with Generated Tests.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction CodeT: Code Generation with Generated Tests

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.160580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.160580Z digest=sha256:06ea4bc315f1103c3d8e422cd222b2693f4ca74a5230f16937abde170520c276

Observation d078e36f-47f6-44cb-a93b-60be7de38413 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.164506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.164506Z digest=sha256:f1b8050906a24e2265d6f54683a11054840ff1941c6dc69872e972a03764cb1a

Observation 68ecc026-c855-4dd6-b247-1378d728b4e4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Evaluating Large Language Models Trained on Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.168031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.168031Z digest=sha256:bee583e40593f73cae91bb9d718ded87b8ed3ba421a120b8aa5a5831bb45be3d

Observation c874b937-257f-4ef4-bd2f-47a0a4774bc9 · outbound

This paper cites Teaching large language models to self-debug.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Teaching large language models to self-debug

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.171640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.171640Z digest=sha256:1a3efcc89fc00bc083f9b5aeb03addc09e35f69f2060c3121d812720b4d36bb0

Observation 2b073000-76d2-40ae-9712-a8f8872eb97f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.175234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.175234Z digest=sha256:34d9c27efda3c3563328f1cefcf1123c7fd1f12128be0ba139a5b1d841a028f7

Observation 0f48e516-0c7c-4e28-bc03-41dc99ab008b · outbound

This paper cites Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.179467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.179467Z digest=sha256:893c242bf6d4bc748611949f76e3e54f153296a04f4f63b975ed86ff2e4b0881

Observation e2fb1888-51f0-48fd-8924-dfa3ddd58756 · outbound

This paper cites Precise zero-shot dense retrieval without rel- evance labels.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Precise zero-shot dense retrieval without rel- evance labels

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.182753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.182753Z digest=sha256:6d2a7ac5006103f20e562ddeffe2c36167b3cd2596d2e63e791dd5f8fb260e7e

Observation 96df189d-b461-4611-9233-1430c2a4a05d · outbound

This paper cites REALM: Retrieval- augmented language model pre-training.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction REALM: Retrieval- augmented language model pre-training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.186272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.186272Z digest=sha256:d8353c9f70a7d011593e8abbe73ec892d172194d9061e843a466c655c2e5c2f7

Observation af93ce11-40a2-4316-929a-f29a1c26016b · outbound

This paper cites Measuring massive multitask language under- standing.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Measuring massive multitask language under- standing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.189460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.189460Z digest=sha256:ac25b1bed985dc021e85d160aa230ee178db924173f25950b5b77a1202fbb322

Observation fd1acf69-ee37-487d-823e-9bc27a45faca · outbound

This paper cites MetaGPT: Meta programming for a multi-agent col- laborative framework.International Conference on Learning Representations, 2024.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction MetaGPT: Meta programming for a multi-agent col- laborative framework.International Conference on Learning Representations, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.193279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.193279Z digest=sha256:89d4079288c8ca104ef095dd73c2e513af24e2c0df6ef2e4e546b50f41077d9b

Observation fa226199-b19d-48e6-add2-59c7a12eb9c0 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.196910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.196910Z digest=sha256:d620aedfdcf80998967f25b6bd0ec9b7c602573045fbe80f6609af43ee59af9e

Observation e841e3bc-1a57-44dc-959a-0d6ebac2a071 · outbound

This paper cites SEAL: Synergistic Co-Evolution of Agents and Learning Environments.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction SEAL: Synergistic Co-Evolution of Agents and Learning Environments

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:33:34.206650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.200281Z digest=sha256:446bdb64e0508268d14471e31baf2236d50960eb3a5dde1e7394a54c277c89f5

Observation 51eb4269-139a-4e06-b618-13f06e4d30c0 · outbound

This paper cites Leveraging pas- sage retrieval with generative models for open domain question answering.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Leveraging pas- sage retrieval with generative models for open domain question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.204786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.204786Z digest=sha256:9a53e6ef8a1f00ccc6c5abaf74fbc7387ccd1bca4a306188151c5664dcad0a64

Observation 002df34c-bb15-4c7b-84a3-eed606ef1358 · outbound

This paper cites Atlas: Few-shot Learning with Retrieval Augmented Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Atlas: Few-shot Learning with Retrieval Augmented Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.208555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.208555Z digest=sha256:f7e7550f917550e0eebc9abd20d067906e3755e232a32c3b073f16357271c898

Observation c1426c6e-e4a0-4500-87fd-c8064ca820d1 · outbound

This paper cites Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.228347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.228347Z digest=sha256:3cbac4f353ef743ef0455aac653557ad438622a819e31df361191f552ce8d84d

Observation 75256c16-6b1b-4bf2-9c60-344eb5af9e3c · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.288079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.288079Z digest=sha256:20c319c32d6ec38bb187e72dd576d1c46cdecd23b006b6e9ad9f1aa58f24d19b

Observation f22a8f96-6674-4744-90b3-77beeff398b5 · outbound

This paper cites MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.416509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.416509Z digest=sha256:a3385af1304b41f097f63866e7fe064f44cabf24ef55fc74e56bd83e67b4ee5b

Observation 8d91ba6c-99f5-49aa-a9cb-63cd2ee6690a · outbound

This paper cites Densepassageretrievalforopen-domain question answering.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Densepassageretrievalforopen-domain question answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.467329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.467329Z digest=sha256:ba52a20d7ce122453a22850ad65ff9ec8b60bc7b0d64a15d37eb347e10a01cae

Observation 620349e2-6004-4fe6-a454-7177b30781fe · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.474145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.474145Z digest=sha256:6663dbafd186e8b0d7366442b6faf35cde50dd63e6f3d58d2b9cd55afe8db2f0

Observation 6be8f683-fa64-4f16-92c4-505ba26e9f0c · outbound

This paper cites Decomposedprompting: Amodular approach for solving complex tasks.International Conference on Learning Representations, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Decomposedprompting: Amodular approach for solving complex tasks.International Conference on Learning Representations, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.478752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.478752Z digest=sha256:2e88cb1ff0bca75583a539c9cedad5effce1c64609f457da87b5f498bf178d2a

Observation 2ae42f66-aeb2-46db-9777-42c899ab93d3 · outbound

This paper cites Language models can solve computer tasks.Advances in Neural Information Processing Systems, 36:39648– 39677, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Language models can solve computer tasks.Advances in Neural Information Processing Systems, 36:39648– 39677, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.481838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.481838Z digest=sha256:6b4e3af980cb090e239ea4006236220de36d2703c558b6b4087ac896d537a8b4

Observation 8515de51-d582-453a-bb2f-e1e49c602329 · outbound

This paper cites Large language modelsare zero-shotreasoners.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Large language modelsare zero-shotreasoners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.484935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.484935Z digest=sha256:ffa4603f8b02cb661c30ab85fb7cfe82cb416f03838cbdc3a185818a67e9d391

Observation b40e6beb-36f6-425f-bd6f-e4328a11da63 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.488403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.488403Z digest=sha256:b836a0f598e0cf0e9806b66c01b2fa05725388d3ee22154a498f306348b17c61

Observation 08a108b0-374a-4483-a78c-86ff582a0635 · outbound

This paper cites Retrieval-augmented generation for knowledge- intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Retrieval-augmented generation for knowledge- intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.492546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.492546Z digest=sha256:6b8df40f5f005707da06140c7eaaeb4474b673fe5b7c58a73d44ec8712d07fde

Observation 5211ec49-0bb9-4d1e-bded-90916a296193 · outbound

This paper cites CAMEL: Communicative agents for mind exploration of large language model society.Advances in Neu- ral Information Processing Systems, 36:51991–52008, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction CAMEL: Communicative agents for mind exploration of large language model society.Advances in Neu- ral Information Processing Systems, 36:51991–52008, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.581738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.495468Z digest=sha256:7988cd12d776cbb13c2f8166f3b9cd3d83fd909b4f7d44258bbec9e53b1e02fa

Observation ffa740c5-05f8-4e55-b9c9-ffb58f9373f5 · outbound

This paper cites LooGLE: Can Long-Context Language Models Understand Long Contexts?.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction LooGLE: Can Long-Context Language Models Understand Long Contexts?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.498179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.498179Z digest=sha256:08fc58e0e811c1814bfe5afbae320b389d72cca9fcd802339105028b7bdc46f2

Observation 48ae8444-11cf-4995-945a-6bdd9cbe490f · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.555235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.555235Z digest=sha256:2510a1315a559fc011c2738a5cc41861ada52ed3ac3dbb5cfd74f4b3837fcc68

Observation 7339dda7-b780-4ef4-9a17-47c79564b0b5 · outbound

This paper cites Rethinking the role of entropy in optimizing tool-use behaviors for large language model agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Rethinking the role of entropy in optimizing tool-use behaviors for large language model agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.571104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.622026Z digest=sha256:118ea3e2e00d64e3e4a177dd788945e1d9259159b20c6631a59f2e726a63d84d

Observation 82d101be-2327-4503-97b4-9f1b2144bb32 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.695930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.695930Z digest=sha256:e4f6b6e8afd138f4dfe4be36d3f06734467cc1ff73237bf1e878f60a97b5ae92

Observation 8c273b47-1aa3-4ddb-a50d-f7eb068852f8 · outbound

This paper cites Holisticevaluationoflanguagemodels.InTransactions on Machine Learning Research, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Holisticevaluationoflanguagemodels.InTransactions on Machine Learning Research, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.499331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.761395Z digest=sha256:3838c74c30917e3bd093e2d965fa7d50843e87ccd57be6fec4c3521c3c3c3622

Observation 02fa0e63-0548-47d8-86b0-1d655b8ffab3 · outbound

This paper cites Let’s verify step by step.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Let’s verify step by step

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.399743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.765307Z digest=sha256:d5c798d11aecbb4ae7307ef851d2782cb75f2c210fba3e97102211c33109d93c

Observation b560e841-de80-490f-b16d-3224831a7d2e · outbound

This paper cites Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.390237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.769066Z digest=sha256:505c347ac6c7da27439650d839165949492f32c2043ce1ecf6a837cb8d29f3f0

Observation 86a963b1-50dd-4db6-a4a5-f74944028721 · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paran- jape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Liu, Kevin Lin, John Hewitt, Ashwin Paran- jape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.379798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.773447Z digest=sha256:e0eab9e27a672051044c7ede281c79cbb018c8e10f6d6113351f3b2f937614b4

Observation d224acb1-c0af-4393-9c86-7685f3ca1f9b · outbound

This paper cites Agentbench: Evaluating llms as agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Agentbench: Evaluating llms as agents

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.369200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.777679Z digest=sha256:3b89acbe7925c5a43c7166d4575521f775809308800cc889c62a7a061b663838

Observation c12d4d46-ca23-48da-b358-38ff6cbbf340 · outbound

This paper cites AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.781745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.781745Z digest=sha256:368fadb6c099cf80a738cdfb6377f67a68b67c2abafd1530d530d65f5fc1bdf5

Observation fa9fc889-98e2-46f7-a0f0-14f1e6633a05 · outbound

This paper cites Finer: Finan- cial numeric entity recognition for xbrl tagging.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Finer: Finan- cial numeric entity recognition for xbrl tagging

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.359186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.863716Z digest=sha256:f9d89bc887caf6ab12a279c8a1dc117985c35bc5d33b10ceb445bb399df459ee

Observation 8e82d994-dcbb-4bea-9c43-bb0db391b830 · outbound

This paper cites Self- refine: Iterative refinement with self-feedback.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Self- refine: Iterative refinement with self-feedback

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.348917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.921545Z digest=sha256:00618ef723590263d3b4f9142978c003a3f137f075aa9b88d9ab80c277f51a2c

Observation 0d30f7b4-3722-4297-8b49-91ede3b6da8c · outbound

This paper cites Lever: Learning to verify language-to-code genera- tion with execution.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Lever: Learning to verify language-to-code genera- tion with execution

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.338188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.925203Z digest=sha256:b9ac02954c46df5f70d48752a05ed377ee9cf328c714c7c59e57b7d97a98615c

Observation 2dba78d6-0612-4f7a-b24f-bc3841774a86 · outbound

This paper cites Optimizing instructions and demon- strations for multi-stage language model programs.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Optimizing instructions and demon- strations for multi-stage language model programs

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.287312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.928763Z digest=sha256:5f28dd2fa0fc927a9f4dbe713e2c9c5e12e21d55c0d270b53e473135dfb48c24

Observation b4ceec2c-27ba-43fe-a8f3-3a57918d247e · outbound

This paper cites Training language models to follow instructions with human feedback.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Training language models to follow instructions with human feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.932441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.932441Z digest=sha256:bb8c0caf543b53aa30b02814e77d2a4e2fa803602c8c014e8100dad9c4d0d4b5

Observation 3a935459-b786-4a8a-8438-c2408fe2c0c1 · outbound

This paper cites Understanding and miti- gatingoverrefusalinllmsfromanunveilingperspective of safety decision boundary.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Understanding and miti- gatingoverrefusalinllmsfromanunveilingperspective of safety decision boundary

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.205605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.937025Z digest=sha256:377fcc4833b38ded1694c95a73219778cb1d8bebb7451014996e5eda9642d0db

Observation 9613c332-f4ba-43a0-8f25-58e6c8b3ad20 · outbound

This paper cites Optimal Transport for LLM Reward Modeling from Noisy Preference.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Optimal Transport for LLM Reward Modeling from Noisy Preference

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:33:34.071641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.940440Z digest=sha256:8080f7ed5b984d7a7e8024928cbbce1a6173bdb815d0a46b1b174091338f836c

Observation 51a50405-d900-4128-ac77-8a3e5d0687de · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction ART: Automatic multi-step reasoning and tool-use for large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.944284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.944284Z digest=sha256:df62c772fd0185097ff059da6b165a4e2361c2467c0763b8e3516943255fb7fd

Observation 724498f6-bcc6-4194-b5e7-ebd3be301342 · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Generative Agents: Interactive Simulacra of Human Behavior

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.949504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.949504Z digest=sha256:f5bf15587a34551138adb0dd854da4f370cef5140b8e66b7fb2a430590b51a72

Observation 7aecd48e-25eb-439e-b06b-123a90b728ba · outbound

This paper cites Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.952720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.952720Z digest=sha256:19e7d61894fd9766fca8aaacd4a03c4ec4fb1c8e810bc4a1ea429ddc2918bc77

Observation b7ca47cf-09ef-4cbe-8e99-15c781a63853 · outbound

This paper cites Webrl: Training llm web agents via self- evolving online curriculum reinforcement learning.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Webrl: Training llm web agents via self- evolving online curriculum reinforcement learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.128534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.956425Z digest=sha256:443bd3e3e7629bbb80c4b1ff03432962c3b1bb9120fc0ba6439bad0f9c76db90

Observation 926f485e-c26d-47d7-8e51-ae9a50e2d937 · outbound

This paper cites Toolllm: Facilitating large language modelstomaster16000+real-worldapis.InThe twelfth international conference on learning representations, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Toolllm: Facilitating large language modelstomaster16000+real-worldapis.InThe twelfth international conference on learning representations, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.094626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.959681Z digest=sha256:04f857045ffb017e2a2b7e6b1a27ce1c118f2b249b88b4bdda14bc395f224988

Observation 53834e41-48b2-4bb5-9a3e-10e333d704a2 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Qwen3.5: Towards native multimodal agents, February 2026

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:32.963369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:32.963369Z digest=sha256:24ab8e289fdec56d0b7147e7dba64a005fd56387ff3984764764d3726d69b981

Observation 843c4dc3-8dc5-4edd-8626-97b439ac488f · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.078721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:32.996014Z digest=sha256:e612a5fbdb35fb7487a27f8f30aeda5a962e6cbc388f9ccc62f1a840df8a596e

Observation 5979b046-04a3-4fda-a452-1aae76c2bd1f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural infor- mation processing systems, 36:53728–53741, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Direct preference optimization: Your language model is secretly a reward model.Advances in neural infor- mation processing systems, 36:53728–53741, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.068292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.072226Z digest=sha256:7116ca4cdcae31e6f85bc15cde929d6984035b060963be2e080cf3c2627f6068

Observation 98de1b21-8a3c-443d-86eb-1b51b29ff57b · outbound

This paper cites Tool- former: Language models can teach themselves to use tools.Advances in neural information processing systems, 36:68539–68551, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Tool- former: Language models can teach themselves to use tools.Advances in neural information processing systems, 36:68539–68551, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:35.056321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.076577Z digest=sha256:d41ee5497a65bf902bcf7eeb304bc292aaae2adf68b95d3684c08bbdc4fcefc8

Observation 6c3a2dc9-614e-4199-b5ce-391f0d27dc1a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.080607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.080607Z digest=sha256:4a4dc5407d098afa5ab9bf65810e8a65fda7551b047574cd2765921de0f2d01d

Observation 66175715-5efd-44f8-b949-db35c5e00be9 · outbound

This paper cites HuggingGPT: Solvingaitaskswithchatgptanditsfriendsinhugging face.Advances in Neural Information Processing Systems, 36:38154–38180, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction HuggingGPT: Solvingaitaskswithchatgptanditsfriendsinhugging face.Advances in Neural Information Processing Systems, 36:38154–38180, 2023

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.944393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.084094Z digest=sha256:a7633e4a72ec795cd614afd4ad2af1601a09770e819ac69c3cd77204aff8312f

Observation 76a573c4-808a-4277-b3aa-fdc10cefd3f6 · outbound

This paper cites Chi, Nathanael Schärli, andDennyZhou.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Chi, Nathanael Schärli, andDennyZhou

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.896116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.088514Z digest=sha256:afd7da80a27c470e4d55baa1fa072082104a7da344da47652339ae3c9a324e71

Observation bac1e0a3-91ec-4b3c-b897-c8d5004b4344 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Reflexion: Language agents with verbal reinforcement learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.091708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.091708Z digest=sha256:cb35b47e4e6b326e2ea5a67f3ed6264c5f4b80884a2db2184f1756969148aeb9

Observation e47806f5-7490-4a44-b27e-aedce987065a · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.095593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.095593Z digest=sha256:92f36a973c5eb03ce09640f39d5fc0d6d41fac85390d900b954c8db8e0ca0501

Observation 24c2a113-bd2e-4d10-874d-acd65cf0e319 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of lan- guage models.Transactions on Machine Learning Research, 2023.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Beyond the imitation game: Quantifying and extrapolating the capabilities of lan- guage models.Transactions on Machine Learning Research, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.880215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.099964Z digest=sha256:8bf5ccc768cd957edfe750df708769d0f251ed8e1ad7e9d229a9f5259e8bf7cc

Observation 7f3c1aee-d14e-4347-83bb-b07b6977ca5f · outbound

This paper cites Cognitive Architectures for Language Agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Cognitive Architectures for Language Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.103535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.103535Z digest=sha256:2d587352eb1d7cd357b0f23ccb2468576ccd503c77559192fcd8ef8990125b91

Observation 7d41055d-f86f-4679-b786-b1159d401579 · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Challenging big-bench tasks and whether chain-of-thought can solve them

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.870263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.184604Z digest=sha256:1e255ff278c28a331f683fabdeacef56058ab8e9499150c4911d94ee257642fb

Observation 3b250c3d-05f6-4c53-8346-29ba93471cb1 · outbound

This paper cites Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.224105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.224105Z digest=sha256:cdf398bc26f16cacc9fd98a118a793bf12b15a71d036786f67b38579af7e561d

Observation 26f07466-fff7-4e29-85ae-6ab6f60cf66a · outbound

This paper cites Eliminating Reasoning via Inferring with Planning: A New Framework to Guide LLMs' Non-linear Thinking.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Eliminating Reasoning via Inferring with Planning: A New Framework to Guide LLMs' Non-linear Thinking

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.228103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.228103Z digest=sha256:f5fb5d7101d55c97a6a299d086fc3ac921a48d467b67f88a57733d7154f23137

Observation c547dd5e-94a3-4c2a-9a6e-fa6183c73134 · outbound

This paper cites Canllmslearnfromprevious mistakes? investigating llms’ errors to boost for rea- soning.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Canllmslearnfromprevious mistakes? investigating llms’ errors to boost for rea- soning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.860646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.231626Z digest=sha256:19a9824807ef7c48374bf0174e6cf2d2433477e2021361b4926d1ce65e47de64

Observation 578b6c1e-9374-4af1-9efb-6502704c0d7f · outbound

This paper cites Optimizing Language Model's Reasoning Abilities with Weak Supervision.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Optimizing Language Model's Reasoning Abilities with Weak Supervision

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.235031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.235031Z digest=sha256:3fa272c90e0a9c53939137ccc23967201eafb3e19b343d870c6ca2f0489be30f

Observation 2789af54-40cf-46f6-ae27-64ea6b5d4cb1 · outbound

This paper cites Appworld: A controllable world of apps and people for benchmarking interactive coding agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Appworld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.850747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.238494Z digest=sha256:1c7c849979def478d8a36476189255e466ae35be6f2b6fe36e6bf618508ab193

Observation 32516c05-0ee4-47f5-922e-0d1c91c624a6 · outbound

This paper cites FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.242741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.242741Z digest=sha256:eb704df17b1f50e0693dd8a349c8f6ed62ebe6e55e122338df3321928c7cc8d5

Observation 58f28120-81b0-4846-b256-acf8819992f7 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.285570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.285570Z digest=sha256:143296bfeeac16246a94266baced044c82247e3ea5a57901d540ceec3a9aee22

Observation e92c1da4-0770-4dba-bcf1-7c5d81521fe3 · outbound

This paper cites REFLEX: Reflective evolution from LLM experience.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction REFLEX: Reflective evolution from LLM experience

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.840862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.342258Z digest=sha256:34e1d173182e5b20310814eb8521bef9029a6b5993098e808c286649626d340d

Observation 4e5eadef-055e-4428-8a26-08917106bcd0 · outbound

This paper cites AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.435925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.435925Z digest=sha256:2e4d3e7d56db171164fd86479d2ba7c7db60099398b9786ee8eb14de8c01df6c

Observation 896cc7b9-c227-48b8-b059-40c9543f532e · outbound

This paper cites Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.490521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.490521Z digest=sha256:c2a161466baf955feae6ada42eaa3dda4ea490a316dbf547b989e624cd880d30

Observation 88ca43e4-4628-4468-b315-fa7325705753 · outbound

This paper cites BPO: Towards balanced preference optimization between knowledge breadth and depth in alignment.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction BPO: Towards balanced preference optimization between knowledge breadth and depth in alignment

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.830847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.495167Z digest=sha256:805b7ffc36330acfce68ee654d78e897ccdc54ee07af4d7c9306ded7461f0ae8

Observation 088094cc-c51b-4956-9ecf-7f9145f63a4d · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.498815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.498815Z digest=sha256:604c96a57b6f7ae11c521589ef56ede45ed16956312037ed29931a19ff679e48

Observation 1ae4ec73-b70d-4cfd-90d6-ffd9223b2f86 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.503441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.503441Z digest=sha256:12a41203bd0ec44835d8d151337dd4a79c9b3ec3d069e30e78ba544b19b050bd

Observation 82300874-c092-48f0-9e21-94999a2fcfed · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Chain-of-thought prompting elicits reasoning in large language models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.814090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.506774Z digest=sha256:1c4702031a26bf9b449f70f6b607bd2443ae8146b61f2ba79c592f819efc9818

Observation 9e6b5135-869a-43ef-8e63-a87a5a065628 · outbound

This paper cites Generating sequences by learning to self- correct.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Generating sequences by learning to self- correct

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.661793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.509673Z digest=sha256:d88a11f3cb26475f1cff135cd50d26620d7c7dc5c2d3381a61eeb4cc775d0ac8

Observation 189255af-02e9-4150-b435-18a40aa37400 · outbound

This paper cites Autogen: Enabling next-genllmapplicationsviamulti-agentconversations.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Autogen: Enabling next-genllmapplicationsviamulti-agentconversations

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.558483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.512348Z digest=sha256:36a54684393f88d28ce85569834c13760c7f7497812f1f7d33ed2e70c3f9c66a

Observation 80b0ac03-dfb2-4ef9-a65b-af2f0ad7feb6 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Agentless: Demystifying LLM-based Software Engineering Agents

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.515805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.515805Z digest=sha256:e8bdef0aecd978242826f21a93d42b2bb150a92fa5bf6083ab00a7f5eb0a30f7

Observation fcc01372-0216-406d-b5a0-4c595510e4f1 · outbound

This paper cites Understanding conflicts in multi- objective alignment through reward consistency.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Understanding conflicts in multi- objective alignment through reward consistency

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.545677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.519209Z digest=sha256:a8e7afbe3ad58c5da75e7e55dec4bc162b9e78a5fe5fd6a8b920a013be92e19b

Observation cc1f9429-1855-4ebc-88aa-e9db5d0afafa · outbound

This paper cites Corrective Retrieval Augmented Generation.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Corrective Retrieval Augmented Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.522211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.522211Z digest=sha256:02361f4e1687833ee55ca7ad31035922bc088d101ff659c54515d2e496170780

Observation e443eff3-8741-4a37-9111-932790c0bef7 · outbound

This paper cites SWE-agent: Agent-computer interfaces enable automatedsoftwareengineering.InAdvances in Neural Information Processing Systems, volume 37, 2024.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction SWE-agent: Agent-computer interfaces enable automatedsoftwareengineering.InAdvances in Neural Information Processing Systems, volume 37, 2024

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.533251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.525916Z digest=sha256:d892f10846737504397ddf596cbce67e8049c32cc00c70ae87e35c363811d77a

Observation 2ea8998c-607c-4368-8d9e-2f37126e16fc · outbound

This paper cites How is llm reasoning distracted by irrelevant context? an analysis using a controlled benchmark.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction How is llm reasoning distracted by irrelevant context? an analysis using a controlled benchmark

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.523071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.553434Z digest=sha256:a692d275ccb85c176309e293ca54763e98fed7da6099d55fc6dc646e0337e656

Observation 0d040aec-434f-4b5a-a34a-b9441f54389d · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction ReAct: Synergizing Reasoning and Acting in Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.588385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.588385Z digest=sha256:815257d830fb0a9a562dd2190a6c6176f3d0abbcb514cc86ab5fd64ec263f4b6

Observation 34d85fb8-9456-4f94-bf2e-1fd169af05e9 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik Narasimhan.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Griffiths, Yuan Cao, and Karthik Narasimhan

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.511755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.641615Z digest=sha256:a732266eaf6180efb71bcc00eebe0b08e0a09fab93d0e11f49ae204928f2f3dc

Observation 1aa9ee3c-da5e-4dc4-8dc7-733ba2c82a06 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.645802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.645802Z digest=sha256:b696eac3767b9c07fd57b0b7c96f3ac1b51382f2814e037cd5c81353be494abf

Observation e72c3e34-d083-45f0-a36b-91aad86262f4 · outbound

This paper cites an unresolved cited work.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:33:34.502260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.649136Z digest=sha256:c509b8156a4af136e1aacddf145557ea80f5e5c6b01a683ae7222aaf5d408e6f

Observation 3f749eb0-d9a9-4aae-be9c-94aebded2811 · outbound

This paper cites Agenticcon- text engineering: Evolving contexts for self-improving language models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Agenticcon- text engineering: Evolving contexts for self-improving language models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.492208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.652423Z digest=sha256:4d9034d4a06e0df3ffb0921bbbefa110e9b93b0b9112ef2eb91001af30013387

Observation 49112e5b-16e2-488a-b499-b839b9e8c802 · outbound

This paper cites Debug like a human: A large language model debugger via verifying runtime execution step-by-step.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Debug like a human: A large language model debugger via verifying runtime execution step-by-step

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.481415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.655858Z digest=sha256:79f280bc15bca4a57c63db1ded8637ba1d9f4920efb878fc7fb03eb8b3a4cd1f

Observation 089374cc-2037-440c-a863-59c4f7bc4ff1 · outbound

This paper cites Least-to- most prompting enables complex reasoning in large language models.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction Least-to- most prompting enables complex reasoning in large language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:33:34.384144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:33:33.658664Z digest=sha256:c9ae6d421d359396d8d178dec1411885320f977b3b8b85cbdd805932745da98f

Observation a6063bd7-fd89-4283-bd7b-34fe3c9f1b14 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T00:33:33.662364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:33:33.662364Z digest=sha256:ac3e32936402ee6beb7cebc6626670af2a22aedc009d1196bdc231b2d387c4fc

Pith citing papers

No inbound Pith citation observations are available.