Pith. sign in

Paper Citation Record · LEDGER

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2505.16944.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16944 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:21:00.066780Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:18:33.541374Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fc2f4517-1872-4a63-960a-5c081147e530 · inbound

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following cites this paper.

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.546219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T06:54:21.052326Z digest=sha256:9b1af0634451f3f0a5a6599a749740454783ef9f51c338d811478f10a501541f

Observation 0c9fa60a-ca55-4835-8ce9-21d885d6028f · inbound

Controllable LLM Reasoning via Sparse Autoencoder-Based Steering cites this paper.

Controllable LLM Reasoning via Sparse Autoencoder-Based Steering AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T12:21:00.066780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:21:00.066780Z digest=sha256:5ac858052f3c0289b9a85c5dabedf9595e82ff96845a903a3fe07b61a07ff580

Observation e2c3ea5d-9691-4407-a4fb-898ca0a234d5 · inbound

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation cites this paper.

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:07:44.406753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:07:44.406753Z digest=sha256:d328ecfa96e5a01c4ba2e177fa29eadf7857b27baded03cf542851fb842207f4

Observation c0fe5729-db7a-426c-a6a4-399768cd8de1 · inbound

SEIF: Self-Evolving Reinforcement Learning for Instruction Following cites this paper.

SEIF: Self-Evolving Reinforcement Learning for Instruction Following AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.896543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:51:09.514927Z digest=sha256:597c067e6c275ed8cf6c7f06462f943d9039d2b0a3f000050082a31e9e6e5bc4

Observation 8c58bbf3-a6ce-454e-a644-574f5df195d0 · inbound

RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems cites this paper.

RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.515059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:04:08.454422Z digest=sha256:c1ce2ee210dddb9732f1feb7f8b979012a584077c978dfc3d98e1446a23aba86

Observation 1a87ecee-0c1a-413e-9efb-2a6f6f06ac2e · inbound

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback cites this paper.

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:23:10.615029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T09:22:06.285118Z digest=sha256:e741cd5edd0fb277b6462c9bbd382ab8726e5f4839c8bc06db7a0d7033e5ed21

Observation 7d556ebd-4c18-473b-8be4-ed6554e9953e · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.609311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:dba70aab99c20a1732cb155061804e6a5112888526c98a7384b7977a72fa6b31

Observation 01b7c6f0-63e1-42c7-a09c-b2b95494937e · inbound

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison cites this paper.

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.341187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:51:09.668951Z digest=sha256:e2dcf7b19ac7e854b4f4129df625d6cab8130074b1f752351d86725f171370ec

Observation 2a8a97d4-cb4c-4ae7-991b-5956fc61d87a · inbound

Uncertainty-Aware Clarification in LLM Agents with Information Gain cites this paper.

Uncertainty-Aware Clarification in LLM Agents with Information Gain AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.606851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:27:18.135089Z digest=sha256:b796c879c105ef643c6defbe3f2e334a34b9759387913bf62b9a569c86f2e218

Observation 2d5e20ff-65c2-4c5d-b127-97d92cb185db · inbound

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments cites this paper.

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.907822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T01:25:51.788544Z digest=sha256:1e9c82a77b1ba1df2814dbffb9fbdb3303ebce4c53b0c56036dbf612d7ddbae4

Observation ad1aed7f-9dbc-4336-8c32-51dd99256bdc · inbound

EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery cites this paper.

EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.543087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:35:11.702124Z digest=sha256:f60e9b0da7f4bfc2cd9ac19cc20096b2d4da3029eb527ecc305bc0267a0eff14

Observation 92056536-b256-4f56-8023-b70188111921 · inbound

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse cites this paper.

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T17:45:46.872944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:45:46.872944Z digest=sha256:b1ba46b1e6c82ecc4036c91733405fac48eb870b2c50a942de788e081126a8dc

Observation 543eb8a8-7eae-49d6-88b4-e0cbab2b4dfc · inbound

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios cites this paper.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.848292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.848292Z digest=sha256:9974fa15e2477aec128c5a925953adc2f343a28430713f4c43620319f7370ae1