Pith. sign in

Paper Citation Record · LEDGER

FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2310.20410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.20410 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:10.940404Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.890478Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03b59bb3-e6c4-428e-bdeb-90f0d24ae605 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:23:31.015188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:4198a84020ea578c98a83c97909c67976293bba068693e0bebb89963a41ba7c5

Observation 6e794753-fe77-449e-8d01-208d32f78d11 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.940404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.940404Z digest=sha256:eb809748ae73a5c6f881703d822e540b39e23c50d43a3ecd9685af4534e6d46c

Observation 78054f9a-c2f4-46be-9c4d-6224a5874124 · inbound

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios cites this paper.

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:06.978691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:06.978691Z digest=sha256:73fd18c48b07a308957b0b288f3570f4e6bf802a152d9097fd74da18d8f041a7

Observation d99caefb-5ec6-40dc-bab6-a222343171bf · inbound

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components cites this paper.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.709513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.709513Z digest=sha256:0e267bb9225fc4c89c4f166573462e2508fafc2ee842faab1990b438b2632013

Observation cea143ab-1230-4175-9893-9bb48f9b74f8 · inbound

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback cites this paper.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.694748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.694748Z digest=sha256:03e644155b9ac24ee80eea14f9fc811e91778d5cdebc5508ac960f6ecccef54a

Observation 3a2b176a-8bcf-4207-b9d0-f2b39abd00f9 · inbound

How Many Instructions Can LLMs Follow at Once? cites this paper.

How Many Instructions Can LLMs Follow at Once? FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:47.092552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:13:47.092552Z digest=sha256:71bac2fac81c35e298661c1f8561b572012785e6ad2662bfc2fe318db357572c

Observation a50026f2-3946-4563-a8c4-4e0bca440f0f · inbound

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting cites this paper.

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:39:22.252419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:39:22.252419Z digest=sha256:f94d8513ad8a29908d24466730b0c2d4d750576235dae2d6d22b22d15611769d

Observation e5b2b43c-60c3-4bcc-b384-4614c96dfddc · inbound

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization cites this paper.

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:52:06.021759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:50:53.673037Z digest=sha256:cf831bc8e5f2af998290e3909567d7a24c491f8253da94e4cc73519ece3233b1

Observation 7b785a7e-2495-4583-9e12-2a48d583ac09 · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:02.499731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:ddc6cefae02c503dd235c808558d1d1c1a1f87e6de9f0af37325345beda18158

Observation d863c62e-c4c5-4539-932c-7380e5556d66 · inbound

ComplexConstraints and Beyond: Expert Rubrics for RLVR cites this paper.

ComplexConstraints and Beyond: Expert Rubrics for RLVR FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:31.354635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:37:11.141846Z digest=sha256:82db6d0f37106dd4c8e8ae9cfe6085ccb2fffbc369c781b04cd0b01ebe9aff9a

Observation 776a9f8e-bbe3-420f-8cb0-85c8cf6d1e8c · inbound

LLM-as-Code: Agentic Programming for Agent Harness cites this paper.

LLM-as-Code: Agentic Programming for Agent Harness FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:44.164149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:16:58.964046Z digest=sha256:06db77fcf6ffc4991e1da616b14628afb48f65a4e1c821afba978bb5609216e0

Observation a1f6e6a3-bd3a-4068-8ad4-f8288a904160 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 232

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.891870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:bf5185c33b0d6164ff8abb0aa9f1ac0e767ff1686decb018ae6519ee0e157699

Observation 008130e9-02a3-4996-b4dc-c1ed6818f878 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.626380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:7eba831edbaec54b6e9d022a2ea76bebf87c7d9ca929c2f53c00f53c7b828058

Observation e08387bd-acb7-4937-b3e6-d36ba320e29e · inbound

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following cites this paper.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:36:07.970657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:36:07.970657Z digest=sha256:52f1eaf805f21b0cfc8b01eba695aaf9da7bf0b99b3b949e5d4f655c77faadaa