Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2310.03128.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:45.870083Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T21:18:59.715072Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation bfdcb510-4117-4c31-bdb4-3897e0a75dd1 · inbound
TrustLLM: Trustworthiness in Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15ec5b6d-4d56-46f8-a061-989c9f6bc2bc · inbound
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55495220-d2ad-4deb-8d59-2736b19284fe · inbound
$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c18c7bb9-1a8c-4176-9085-053352f1ff88 · inbound
Prompt Injection Attack to Tool Selection in LLM Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e16af64-e263-45dc-b0d4-27ecf863a911 · inbound
$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b59b1997-5a74-45d7-8a59-a4d8437cf4cc · inbound
The Curious Language Model: Strategic Test-Time Information Acquisition MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20756517-6596-4bec-9837-d139ff5bdd0b · inbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac1df20-9361-4c43-82c9-f4898d7e2d0e · inbound
Teaching a Language Model to Speak the Language of Tools MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337c765f-ae53-452a-8937-39caf61300d2 · inbound
MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1ca4a9-c1c7-4390-8c31-1e095790d3e2 · inbound
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22550fcd-f160-4d2f-bbd9-269ad5bc33d1 · inbound
Towards Compute-Optimal Many-Shot In-Context Learning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51babeb4-227a-4516-95b0-2f094cd24194 · inbound
GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1576480d-e593-4139-8f1a-353f60d92f52 · inbound
Evaluation and Benchmarking of LLM Agents: A Survey MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67a0b2f9-4580-4f30-9e34-68085d9089bd · inbound
UserBench: An Interactive Gym Environment for User-Centric Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 915671c3-e5cc-4ec2-9679-5d17a1a45d50 · inbound
Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03ba34e-3ea1-4c44-a50e-849e2df3cda7 · inbound
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e9d9af4-cf15-48a8-824e-b99be0ce07c2 · inbound
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 523d99e1-82e0-4b96-84b8-3501b8c110a9 · inbound
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0700dca1-1a6d-46f8-8cd4-7a0dabd4525f · inbound
Memory in the Age of AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 261
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad97d443-4774-437f-b967-c37f48d6ce05 · inbound
Toward Efficient Agents: Memory, Tool learning, and Planning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ec2ede-71a8-4033-86a9-01bfcebf3c9f · inbound
Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b853e52-1d2d-404e-83a7-6d6775680c58 · inbound
Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2bdf27-338b-487a-8c00-d32a047c8fd4 · inbound
To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a5476e6-df85-4c89-9155-3d38d16c87f8 · inbound
FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d00cbb03-51f7-491b-9c00-d4a541116888 · inbound
FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f87b296a-6676-48f1-9d59-9f30edbfe7a2 · inbound
FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7709fb7c-2ad0-4ef3-886e-0853e6383a01 · inbound
Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 684b6e09-e7cb-4db7-9017-2bb70f976afd · inbound
From Intent to Execution: Composing Agentic Workflows with Agent Recommendation MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a86b0777-73cc-4fea-906b-4dfb1a7f260f · inbound
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b053461d-5763-4d1d-8f70-7b549a050ee9 · inbound
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e545d30-a3b9-44e7-b915-75cc50001943 · inbound
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d473db4c-3e82-4cd3-9767-bcb8063c5b44 · inbound
Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 164
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 749d96c9-e863-41ee-9a88-6548c8b301dd · inbound
Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b36e908b-d4c6-4335-a441-73f48c1b2f6c · inbound
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7301589-2913-4ce2-9e68-39c437d74e07 · inbound
MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df94b234-2959-4db7-a73e-2469787000e9 · inbound
Capability Self-Assessment: Teaching LLMs to Know Their Limits MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cbf862b-fdcf-46a9-adff-cdc6cc2f28a2 · inbound
NTILC: Neural Tool Invocation via Learned Compression MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f6a9ba2-b68d-4f41-8fe9-02eb8f016c6e · inbound
The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c152ec48-6f8b-4634-82ed-6532c4cc7091 · inbound
The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b490de2a-c3a7-4ef9-8b45-7932e70decb2 · inbound
Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf78261e-76a9-4938-8b66-59b2e68a263f · inbound
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5bd2c1c-a64a-4a8f-bc81-35180af2a342 · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 156
Source-reported events for the cited work
Unavailable: canonical work link unavailable.