Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2307.10928.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:02.317133Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-07T12:53:50.256862Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 1d49c44f-db24-4451-9392-3c230673a8bd · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 279
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c26c8cb1-9deb-44ad-aa7a-70f4d0e034b4 · inbound
SedarEval: Automated Evaluation using Self-Adaptive Rubrics FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50dac8b0-662b-4f05-b4cc-d4f38e9eb180 · inbound
KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d4bd05-9e66-4595-96ca-9c7d02b0a5a3 · inbound
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a6d43c-35b1-476d-aa3a-e21e8df48cbc · inbound
An Empirical Study of Evaluating Long-form Question Answering FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87c415b5-6dfd-479b-bec7-d0c365df8382 · inbound
A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84066faa-04e3-431f-b039-ec8378157115 · inbound
Are Today's LLMs Ready to Explain Well-Being Concepts? FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd119d8-450b-4eb6-b495-24bfb2c8b879 · inbound
UQ: Assessing Language Models on Unsolved Questions FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4f0aa8-f099-4c56-b6bd-72416660167c · inbound
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8a27f8-de3b-4de6-beb5-98b83b610457 · inbound
Evalet: Evaluating Large Language Models through Functional Fragmentation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 50ece269-fd60-43e7-942d-47cf9998569a · inbound
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6b95f879-f33c-465c-8d05-f16086ceed2c · inbound
LETGAMES: An LLM-Powered Gamified Approach to Cognitive Training for Patients with Cognitive Impairment FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 811466b6-b7db-4b46-a5a3-38f7cf118e1a · inbound
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4ac0ccd0-d7f2-45c6-b523-2aa338fee32d · inbound
Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 63206756-f47c-4bf2-b0aa-e35c1ae753ff · inbound
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1a017449-6d68-4f02-ba82-adcdb4578c77 · inbound
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4e4ad4d8-1c69-442c-b7bb-9fe8fdaac2db · inbound
When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 02a760de-746b-47eb-888d-610fbd272bbb · inbound
BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0943bb3e-9479-4079-8ef0-06aa0e6c431d · inbound
Towards Multi-Agent-Simulation-Based Community Note Evaluation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a962ad72-f002-41aa-9c53-f115f979c3c0 · inbound
LLM-as-a-Verifier: A General-Purpose Verification Framework FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 85b10f22-d2c3-4088-9cb6-b95c00ee20b5 · inbound
LLM-as-a-Verifier: A General-Purpose Verification Framework FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c67380c-9093-4233-93c4-26d016e8022f · inbound
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4152ccf-6391-47f9-9ca8-cfd91ffd4de3 · inbound
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6389a7fe-8a12-42a2-92ea-f27dc9c379b7 · inbound
RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.