Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:44:06.111125Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2506.13977.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:44:06.111125Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:44:28.494145Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T09:55:41.045115Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 217bbc1a-257d-4b69-9794-05b4d991820e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c0cfbec-79ff-4061-9b7b-227bdb64af60 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f0377f4-e565-4569-89f6-490bbdc4321c · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7fecc16b-e1d9-4ccc-a44f-8c7668226c5e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Learning From Mistakes Makes LLM Better Reasoner
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fadd0eb-838a-4d0c-8203-639c4e8f39d3 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c47a9e-9490-41e8-bc0f-24251e672eb8 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 845a9e89-81e7-41c7-a4e2-7771fdf4b6f8 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1654c7bc-fb7d-4a53-a740-21d386f76742 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689c701e-040c-47bc-b1a0-7613609aef61 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72fdc844-aeb9-4214-adb9-c7bfd31a1ef9 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e759c2be-f0f9-4544-99a3-e67b35b0900e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ffd358cf-dc7e-4423-8913-b22c9e5d2a6e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eeb95397-334c-442f-8fbb-74addd9d7918 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7931df2-f944-4b98-b455-0e7d77241325 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e7d13f1-3ad6-41a8-9b1b-52245361afbc · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d2444bff-9f8a-4a7e-a33c-c5eaed1e5a6d · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66880987-ac1e-4fc5-aa42-4bb68f4f11c0 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios GPT-4o System Card
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac4134f-cf69-46af-9fee-3fc533a43831 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios A Survey on Large Language Models for Code Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c834917b-d2d1-4eb4-8348-e13f40af22a9 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7c21ec4-5f7a-43fd-b22e-34bb9c2be634 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f3cafb9a-70c5-4a84-bdef-16b74ec4736e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios CriticEval: Evaluating Large Language Model as Critic
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c78a4a2-24f3-47c4-9a50-4634594879d9 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ce7f19c-92ea-4f12-b406-aa7f1353fe1a · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation acc7b5f5-65db-484e-bba3-d81f28142f86 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolACE: Winning the Points of LLM Function Calling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02bc1dfe-46d6-417a-bbef-fef24ef561b9 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LLM Critics Help Catch LLM Bugs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7883b501-a81c-4491-9c33-e6e412c7962d · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fbd3ae0-f63f-492b-a458-819abe786dac · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d1f1c625-1a6f-4184-94ef-08a192067740 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Gorilla: Large Language Model Connected with Massive APIs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93279016-7dd9-445e-adf8-3b5823caf82c · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172fa7a8-5323-4f0a-bd87-b9e72adeaf4c · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffde937-60a1-4432-ac8b-4277c4bf7c0d · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a54ff64-1f1d-4520-a5f0-b99007a8a92c · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios TaskBench: Benchmarking Large Language Models for Task Automation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5587e5-b660-4644-a920-f09aac92215f · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c48ce5-9139-484e-8c75-e579e826ac76 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 76a0d6b7-153b-412d-9ca2-33d62202f73e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0647b2b7-efd5-4d81-97fd-78b1d48a8f1e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Qwen2 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a6ee20-8ae0-4209-9e94-6ba84077b99c · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780b2f58-8832-4486-844d-86d6bf3a3b8b · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae17871-d694-4766-815c-5583c4ebcb13 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e8efb481-ee2b-4016-ad62-e2f495fb91d5 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd377a5-55a9-47e0-8e70-c03d14bd779f · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9a216b23-d2b7-4c7b-a3e1-e26f67b2ee8d · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios On the Tool Manipulation Capability of Open-source Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a6544b-3b5d-4ac5-802f-efc3b7c78a1e · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Patil, Ion Stoica, and Joseph E
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation be33c64d-3ebd-4ac8-af3b-46dcf181c08c · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b304eab5-2ffb-4dca-acf6-3cbcc9d960da · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92ede64-81df-43ec-805d-2bc53fecf758 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ef958cb6-c84d-4d1c-9d3d-11da26766dbf · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6049b07-f49e-4e96-9aed-d247befb0e8b · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c6ef08-dc2b-47df-8e58-ae41965ded89 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios AgentTuning: Enabling Generalized Agent Abilities for LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c75468e2-0c74-43f0-aaa8-0b4cb451c5ef · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation abf54ff6-b669-4a22-bb08-2eb7a92e59db · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios A Survey of Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b177e4ee-7bfb-496a-b3b8-d879c4c037b6 · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios online" 'onlinestring :=
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76269b9-64a7-456d-af9e-e20f598c35dd · outbound
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios write newline
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568c73ec-2278-4e24-ac49-9a81921e4d2b · inbound
Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4904fd2f-b77a-4838-b57f-5cab4c0c216a · inbound
Recursive Self-Evolving Agents via Held-Out Selection CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1a7a4e19-1268-4d10-a232-8b4d46724ac3 · inbound
SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation da4b10f7-2315-4881-8542-3bc14fb3835f · inbound
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.