Pith. sign in

Paper Citation Record · LEDGER

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs

As of 1 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2604.22937.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.22937 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T11:33:21.391661Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T19:01:21.650333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T19:07:35.241005Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact20
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c46bb44e-3c1a-4045-b196-572805592936 · outbound

This paper cites online" 'onlinestring :=.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs online" 'onlinestring :=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T14:07:51.807847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:a0c75041d810b499a7a3bff0ed2947fcaa69f442f64bb38eb5763775c75e5e41

Observation 36aafd43-bd3d-4c22-830f-d6076901e1b1 · outbound

This paper cites write newline.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs write newline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T14:07:51.829409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:fa2d41029c7d29069a2485e20adb57047f16a64da645ce914eac84dfcbf73ce3

Observation 44cdc814-62bf-4033-9e5f-a12af7dc4de3 · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:10:14.938194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:1823ea05abedd2cbfb571bd5584332afbfb5d81d3ae3978598193554e326a686

Observation 7407d28d-99e4-4b14-b14a-af2bcd5a30c8 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.750346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:d7ff62fc1ef2e060a5c21a78a267c362254f156ccd1ed1248d97ba8c62127897

Observation 1ebfef7f-f252-4e21-aae1-29f8a5981718 · outbound

This paper cites Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.787691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:860e7e5cf90fc3a50ca1214b2495dffd08799ca24e1edc53a5a21417aa42bb84

Observation d8ea1117-5b30-456d-8dac-147a32ff1d04 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.832178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:68b03b40cf76d4b74291f26650313a132cd55f830e18b659c62d3be7bf5052cb

Observation 7ecbb288-1d78-4a7c-89a9-b47f84a32448 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.820152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:7849186e3eaf3146043756d62645896f0f5a3debc0209d75bbf130bc971d1378

Observation 21cfe2ab-2b31-4a46-97d2-828433491170 · outbound

This paper cites Beyond oracle: Verifier-supervision for instruction hierarchy in reasoning and instruction-tuned llms.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Beyond oracle: Verifier-supervision for instruction hierarchy in reasoning and instruction-tuned llms

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T14:07:51.813014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:04608258c8f69d746f291c575fbe29fe3dc7b0400c926f5d38e4c27396cf7b44

Observation 48029d7e-f467-4632-9309-499a21e736a7 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.731528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:1c675a92c8ecf7125c41d0e545c927944d9d3c974d1273a9f7838d9eff8cb795

Observation 2ac75d33-6d5b-4bf5-b07b-3fe8b882cb8c · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.835109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:e2c379544dda084f6ad9eb2e320ccf8ac3d208ff83424b37ba37a5c812c4158e

Observation 45661ab3-f3cb-4ea9-a9c9-efe335a0f535 · outbound

This paper cites Process reward models that think.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Process reward models that think

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.762894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:5656b5b937d6c81a8cec7038e6978c6d9856f35baf81572bc317ca71973834f6

Observation f128a48d-731f-46f1-8d57-728379350875 · outbound

This paper cites How to Correctly Report LLM-as-a-Judge Evaluations.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs How to Correctly Report LLM-as-a-Judge Evaluations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:02.584259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b7cdc3b853cdbd24daec8a0c61eb8065364905aaf884efaa28aa2b730d220969

Observation e763ad97-8017-4944-a0c3-90190430f9c1 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.815193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:ae4c629f2fc519dd630928fd7b02913f851cb89fce8271f0183243b13a65850e

Observation ec19ef6e-2772-44da-aae4-b49188f278c2 · outbound

This paper cites LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.714408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:af71d843fc3eca537b665ca4dec2d33c05eee8eeae79a9aa82a1ee110d0d0017

Observation 8ad72153-e9aa-4dd1-a290-1e3c637781a2 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.805284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:453d848edfdad94f24da3edcf91f5fec1038ad038b6c3f161af709087bc5f876

Observation 768b5491-8865-4ee8-963c-b45f37ef4ee0 · outbound

This paper cites Autoharness: improving llm agents by automatically synthesizing a code harness.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Autoharness: improving llm agents by automatically synthesizing a code harness

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.775751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:38332efa3fb9c700285e2ff13c6b23580d3552af82e536654bb485b1652e883c

Observation a75a7856-c5f8-409c-b872-eaec0a77c1f8 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.799298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b33ae5429afe40b95be623e3a4520a46d7f6d9056731ef280c63785841be9814

Observation 3a597ab0-7276-4732-b31c-d285b5cfc3a8 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.810638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:5aebc67ddc8c52f3fbae83fcfc6a108bf6f5749cc17b7a0b3eb9a49308dd30fc

Observation 91b8ae1e-4a0e-402a-a76c-98d6d5cb03de · outbound

This paper cites Natural-Language Agent Harnesses.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Natural-Language Agent Harnesses

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:30.082400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:ec91e65d62026cee760dc6f5f4cfe00cecc86fa2c8d9b6c30b8351d1059101a3

Observation 89ed779a-d700-45f5-963a-a3e25281d343 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.822974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b3a9bda1d97b525a08e8f701367f0998b53e98ff29da86f6bf3c1472cac73dcd

Observation 67002705-f036-44ce-95a4-3a7e5675ba47 · outbound

This paper cites Beyond outcome verification: Verifiable process reward models for structured reasoning.arXiv preprint arXiv:2601.17223.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Beyond outcome verification: Verifiable process reward models for structured reasoning.arXiv preprint arXiv:2601.17223

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.725542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:c8ec85e0df97a2301791f800d07af5d70f48141132c2dce74c383230d257af65

Observation bc17e442-b7b8-4bf7-a2e3-e26fa3ebcc59 · outbound

This paper cites Generalizing Verifiable Instruction Following.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Generalizing Verifiable Instruction Following

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.699446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:0417cff4f2a8acabb374c0f9b51c07dac43c98fb9835333a06d584e6ed90ed27

Observation 85b937d4-e414-4c45-83ff-e793b734385e · outbound

This paper cites OpenAI GPT-5 System Card.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs OpenAI GPT-5 System Card

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.679536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:ceace05a8b40046b57af80784f7ecc375e9496eea3f24bf213ba7119f10e84c1

Observation 119e44ff-95f2-4253-a83a-9b506b9a545c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.660771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:747c5c23bf3f11e8602989c2387a9b734851c65dd4a1793395543c0c83f6b335

Observation 145e1f44-a3ec-43c5-a78b-146fd6307c3c · outbound

This paper cites Trust- judge: Inconsistencies of LLM-as-a-judge and how to alleviate them.arXiv preprint arXiv:2509.21117,.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Trust- judge: Inconsistencies of LLM-as-a-judge and how to alleviate them.arXiv preprint arXiv:2509.21117,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.649003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:2ebc6029316a912059dada6c9ae0b75a6ba35d5c9077ec135e0be9311cc7a9d7

Observation 51f8e7e6-b35a-42e7-a489-d3107768fd20 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.826464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:75d3f8423880579362aa05648685abcd75a69da3dc655e01c77e15ad41a48276

Observation 1017f8af-d607-4e1b-b968-50192fc8f6da · outbound

This paper cites StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.704509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:71aaa0dc7fa978c7aae7cf95483460799eb6ce305efa26bc8b81e0cd48779d7b

Observation 7875a516-346b-4f3a-8e20-5c64e9321aca · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.802944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:947810363fe9e2509102fabefd1f64e00c4a559c8e48a27550a93c841edd32f8

Observation 099524b1-63dd-425c-b722-87a3f0b40fdd · outbound

This paper cites MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:10:22.017517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:1915defc4319429bf55726019219720b0869c56635c7a94cbb989bcd1c4cb832

Observation 92666ef6-494a-4a0a-be5b-9eb7e62098bd · outbound

This paper cites AgentV-RL: Scaling Reward Modeling with Agentic Verifier.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs AgentV-RL: Scaling Reward Modeling with Agentic Verifier

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.694509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:df6b11597e1bd1f8be61a6748285897af48640c76ed7456ea0e2f7aba86fe343

Observation fd5e532e-81b4-4da3-aca3-7a62b4d7a490 · outbound

This paper cites Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:44:08.223014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:ada3cb0e74ca3cbfeae53bd2dfc397c25135503944258e1077da9b2c0bb6bd91

Observation 00d1a3ea-39d5-4299-b112-6b2d5ca445a7 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.817477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:0fb7f98f4219f6b7668250e8cef5ec9bf7d35666ab57c67ac80d0cbad6605fc2

Observation f3f8c785-c41e-426c-b7d3-cbd8aa881af9 · outbound

This paper cites Available: https://arxiv.org/abs/2603.11445.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Available: https://arxiv.org/abs/2603.11445

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.668075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:63d5c9f2666abe581c5bf9e4c0b6b9f736e323da9fdffbbc110a4abaf85e722d

Observation c90841e7-ff2f-4029-b6d1-7a368a59c826 · outbound

This paper cites ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.672811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b965e03a9f5ba9510323da3f89f9877b6ad4dabb3d5992b6026676365d479966

Pith citing papers

Observation 939fb746-b159-48e3-a0cc-4d79b29405b3 · inbound

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents cites this paper.

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T19:07:35.244603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-07-10T19:01:21.650333Z digest=sha256:e81cfda69385eb0f2c664c5e84b252be3e62b595609ee9bfa5454ccab9dc78eb