Pith. sign in

Paper Citation Record · LEDGER

InFoBench: Evaluating Instruction Following Ability in Large Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2401.03601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.03601 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:25.195307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T10:27:02.414584Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 929b7e06-34b1-45db-b993-85ccfa2ad7df · inbound

Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era cites this paper.

Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:25.195307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:25.195307Z digest=sha256:f58b304b7d668f56e86cd980dc04eeedfe5ad5b149d15715c0c7cfee6c108f7b

Observation d1750171-cdc5-41f6-bd96-f1c6d3e70c9c · inbound

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models cites this paper.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.641907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.641907Z digest=sha256:875d8d5626503e8bbd39e952a88a56d02eedb7d3c661726446d2e6f5395549ba

Observation 278ea87f-4455-44f5-9d49-8c03e8c98561 · inbound

LearnLM: Improving Gemini for Learning cites this paper.

LearnLM: Improving Gemini for Learning InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:40:18.495020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:40:18.495020Z digest=sha256:a9f0d3c9953bf9c9e0a3f077bca5e2a4cec71ead980c50246f405b001e6efbd5

Observation 406503e4-4d61-4bf2-89fa-0ef09f8ba446 · inbound

Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models cites this paper.

Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:39:45.745269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:39:45.745269Z digest=sha256:6f0e58abedf1ba720852dcdd809c5eccea4cf2d0a5cc77a84eeb847000da48d9

Observation 146b0000-1fe2-4c00-92bd-13c45b35aba3 · inbound

Atla Selene Mini: A General Purpose Evaluation Model cites this paper.

Atla Selene Mini: A General Purpose Evaluation Model InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.877232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.877232Z digest=sha256:7e6850db4d004bda168f258b2326ef339ec0b425362f1e9be89ee3986fb1efdf

Observation e80dd5c7-9aa1-4733-a722-e1501b9e443c · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:31.007005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:93eb06a2c4565da62a3ddcf79c60f302ee258a1f7005a33bb6234e5e9d3017c1

Observation 94735af9-e993-47f0-9f2a-0734165c463f · inbound

Shuttle Between the Instructions and the Parameters of Large Language Models cites this paper.

Shuttle Between the Instructions and the Parameters of Large Language Models InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T12:41:02.357079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:41:02.357079Z digest=sha256:d953c93cd2f1db7b24323c57bddedb106872c91137de1555b93e6e12a965b672

Observation 5a2c937c-c775-4abb-b282-826e32733246 · inbound

LLMs can be easily Confused by Instructional Distractions cites this paper.

LLMs can be easily Confused by Instructional Distractions InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T10:50:12.464076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:50:12.464076Z digest=sha256:4374e83e252e0c0b544d63e413f5159c836b521773a49c8bf7c2905fb2e55b26

Observation 21265f56-1a53-4bf6-898d-c558e87f9219 · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.309327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.309327Z digest=sha256:b3f435165512751b5846aeded7f79c2d4ca7401069d0a76f89f1fa59e25a6216

Observation 4e913d6f-c7e5-46c3-ba35-28801db497c4 · inbound

IHEval: Evaluating Language Models on Following the Instruction Hierarchy cites this paper.

IHEval: Evaluating Language Models on Following the Instruction Hierarchy InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:56:00.088954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:56:00.088954Z digest=sha256:5d4ffac81b1e531197d2b7bebdafb1dcc949599b08543ed222d777668be023cf

Observation 2969fc50-390e-4da1-b5a6-eae3c06f5de8 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.811415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.811415Z digest=sha256:339bdec981cd6bb47d177aebe51f8fc9a7e08bbbc0656b0b4abacfb3ff2cb120

Observation e71064a1-1795-472e-b8d0-ce017204634d · inbound

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios cites this paper.

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:07.702430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:07.702430Z digest=sha256:0fba971088b966078674c9de4c92588c9ff045f47258f159c32820c38f2b9a58

Observation 1ee71b22-1db2-4efc-9878-d7e5e85d0a4c · inbound

The Price of Format: Diversity Collapse in LLMs cites this paper.

The Price of Format: Diversity Collapse in LLMs InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:54.459160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:54.459160Z digest=sha256:52f3a96f50817ea80ec8dcd54eebcf32ab2463f6cbb3c95207d4389df0b15c72

Observation 63bfd505-16a8-4a7a-927d-ac111e0c5158 · inbound

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring cites this paper.

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:27.234774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:16:27.234774Z digest=sha256:b1652eff9fed28a0c69aa718e70f60e7185b9d2000957cac92339952f55a9a92

Observation 0ffed680-fcb8-4413-9c16-0d51fb04d1fe · inbound

Beyond Facts: Evaluating Intent Hallucination in Large Language Models cites this paper.

Beyond Facts: Evaluating Intent Hallucination in Large Language Models InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:44.620268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:44.620268Z digest=sha256:d2be0628216820d3e5cb2ff932c2c183b2265c905e175a9a0eacd58db7c54472

Observation da751ba6-c1d1-47df-985f-5a1f713dda1c · inbound

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback cites this paper.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.489193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.489193Z digest=sha256:be9ab0c74decfeede7b125e3aaf40a205233b2d4dff57c49dacffa6539cab1b6

Observation f09dd33e-e5d6-4e26-9d47-ec3e91cb9b24 · inbound

MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning cites this paper.

MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:18.719781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:18.719781Z digest=sha256:152db96fd51ba653e01fbb5681a404ac02b276b5ae4bdabf35a3e7cf4fc58dd6

Observation 1af21bce-07ba-4364-b4ea-b1fbbe0452e2 · inbound

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training cites this paper.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.631107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.631107Z digest=sha256:f29bdcf2b1b1010758bc8459290a7ab59fd56eaa2528124e9daef7ed2752c2d0

Observation 5f945864-5747-43ed-a98a-04bcd6a3e7df · inbound

Token-Level LLM Collaboration via FusionRoute cites this paper.

Token-Level LLM Collaboration via FusionRoute InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:26:31.539200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T12:25:59.747665Z digest=sha256:3eb86e1bafddf19b3bc024f22b0a5863bc2402ec7df960aa36fe42c2be70f020

Observation 0c7a68c3-7c11-407d-9a6b-ccae2181552f · inbound

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users cites this paper.

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:35:27.401956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T10:34:30.152982Z digest=sha256:d78d5dc9c70851460d546333578ac1e87ee684a00fbe92829755baf2e0793866

Observation 8a015d1a-4e1a-4f49-b01c-f306e0c4c63d · inbound

SAGE: A Service Agent Graph-guided Evaluation Benchmark cites this paper.

SAGE: A Service Agent Graph-guided Evaluation Benchmark InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:21:00.981457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:41:23.956104Z digest=sha256:9635762ba259ac4d4121091cfe54ddc1c4375c42bee8b8a0a1e5aa74a544d26f

Observation c165e069-7c20-4359-9d21-21d1f684c6ba · inbound

Intelligent Drill-Down: Large Language Model-Driven Drill-Down Technique for Human-AI Collaborative Visual Exploration cites this paper.

Intelligent Drill-Down: Large Language Model-Driven Drill-Down Technique for Human-AI Collaborative Visual Exploration InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:41:36.870761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:37:21.309366Z digest=sha256:4b046b335daa76d509295b5c3f3f16eccc34093a3c217b6629458ef2e8faf3a6

Observation e7d903ac-52b6-4a6b-afd0-babd513714a1 · inbound

SpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generation cites this paper.

SpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generation InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:28.336851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T07:58:39.106432Z digest=sha256:9ce81eee7a4bc10a28e64e30ad6f6a10969c54ce2e2c33a4e09af08b20527c39

Observation 97e9ca44-6223-4166-814d-30555baad599 · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.699202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:02d13789cc7c166ad6ab1835c78a7fba220c1389fee024d9ccb3e051148e7e1d

Observation 20183586-043a-4750-8fe2-e5c55bfdeff8 · inbound

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs cites this paper.

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:34:02.846609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T07:30:27.297971Z digest=sha256:cb61103dcae87531e562f6864e845322832b48b7ca6f73588906721a399181a6

Observation 7b264fa5-9b8a-41ae-8398-1ca14e1ff122 · inbound

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs cites this paper.

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:04:58.150119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:00:29.237972Z digest=sha256:7d9aaf0d0dabe74b5a84abe17c69cb2802577f3eacdfe4bfe39935c88fa2531e

Observation 7d284006-9e52-4787-a5c5-cc1cb7ffa5ab · inbound

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following cites this paper.

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:33:24.266472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T12:32:00.055500Z digest=sha256:7520c9f6988d860b2e59857d2294f4f167cac2efff0cd613d1414637edfb8f65

Observation 47f6b657-6b4a-40c4-bcc8-e542d59266e9 · inbound

ComplexConstraints and Beyond: Expert Rubrics for RLVR cites this paper.

ComplexConstraints and Beyond: Expert Rubrics for RLVR InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:31.369399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:37:11.141846Z digest=sha256:f1283a2ea7cbe749e549e3ba8df1e2af5d69eeacb2d6b4b07c858db251eca312

Observation 1a05813a-c35e-4da0-bced-a4d65cbd79c3 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.795430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:333c2f1f84b74672898cf0c97bb7be9902ea7b0104bf47e163d6211a8d07188d

Observation 2a668a69-0949-40a9-886f-f4c7ee3b9daf · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.748737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:63e14d7bf47b37867732b59fa89f7c6329db074f66a7f3730ea7db5d7ccabcbe

Observation 59047028-7964-4c12-9bf9-c7957d67eaaf · inbound

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment cites this paper.

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-10T10:27:02.415936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T10:20:46.103409Z digest=sha256:b6736f97a6c1b1aa555730a85f74c6c7dac9f2b3024d41758efb33fd178d8110