Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T11:33:21.391661Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2604.22937.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T11:33:21.391661Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-10T19:01:21.650333Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T19:07:35.241005Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c46bb44e-3c1a-4045-b196-572805592936 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 36aafd43-bd3d-4c22-830f-d6076901e1b1 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs write newline
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 44cdc814-62bf-4033-9e5f-a12af7dc4de3 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7407d28d-99e4-4b14-b14a-af2bcd5a30c8 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1ebfef7f-f252-4e21-aae1-29f8a5981718 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d8ea1117-5b30-456d-8dac-147a32ff1d04 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7ecbb288-1d78-4a7c-89a9-b47f84a32448 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 21cfe2ab-2b31-4a46-97d2-828433491170 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Beyond oracle: Verifier-supervision for instruction hierarchy in reasoning and instruction-tuned llms
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 48029d7e-f467-4632-9309-499a21e736a7 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2ac75d33-6d5b-4bf5-b07b-3fe8b882cb8c · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 45661ab3-f3cb-4ea9-a9c9-efe335a0f535 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Process reward models that think
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f128a48d-731f-46f1-8d57-728379350875 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs How to Correctly Report LLM-as-a-Judge Evaluations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e763ad97-8017-4944-a0c3-90190430f9c1 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ec19ef6e-2772-44da-aae4-b49188f278c2 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8ad72153-e9aa-4dd1-a290-1e3c637781a2 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 768b5491-8865-4ee8-963c-b45f37ef4ee0 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Autoharness: improving llm agents by automatically synthesizing a code harness
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a75a7856-c5f8-409c-b872-eaec0a77c1f8 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3a597ab0-7276-4732-b31c-d285b5cfc3a8 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 91b8ae1e-4a0e-402a-a76c-98d6d5cb03de · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Natural-Language Agent Harnesses
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 89ed779a-d700-45f5-963a-a3e25281d343 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 67002705-f036-44ce-95a4-3a7e5675ba47 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Beyond outcome verification: Verifiable process reward models for structured reasoning.arXiv preprint arXiv:2601.17223
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation bc17e442-b7b8-4bf7-a2e3-e26fa3ebcc59 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Generalizing Verifiable Instruction Following
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 85b937d4-e414-4c45-83ff-e793b734385e · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs OpenAI GPT-5 System Card
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 119e44ff-95f2-4253-a83a-9b506b9a545c · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 145e1f44-a3ec-43c5-a78b-146fd6307c3c · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Trust- judge: Inconsistencies of LLM-as-a-judge and how to alleviate them.arXiv preprint arXiv:2509.21117,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 51f8e7e6-b35a-42e7-a489-d3107768fd20 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1017f8af-d607-4e1b-b968-50192fc8f6da · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7875a516-346b-4f3a-8e20-5c64e9321aca · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 099524b1-63dd-425c-b722-87a3f0b40fdd · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 92666ef6-494a-4a0a-be5b-9eb7e62098bd · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs AgentV-RL: Scaling Reward Modeling with Agentic Verifier
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fd5e532e-81b4-4da3-aca3-7a62b4d7a490 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 00d1a3ea-39d5-4299-b112-6b2d5ca445a7 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f3f8c785-c41e-426c-b7d3-cbd8aa881af9 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Available: https://arxiv.org/abs/2603.11445
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c90841e7-ff2f-4029-b6d1-7a368a59c826 · outbound
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 939fb746-b159-48e3-a0cc-4d79b29405b3 · inbound
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.