Pith. sign in

Paper Citation Record · LEDGER

FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2307.10928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.10928 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:02.317133Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:53:50.256862Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1d49c44f-db24-4451-9392-3c230673a8bd · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.837923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:250b62bfd5e1b035f8c195090e0997d13ceee1aa9b25a9b89de344a21fdc3bba

Observation c26c8cb1-9deb-44ad-aa7a-70f4d0e034b4 · inbound

SedarEval: Automated Evaluation using Self-Adaptive Rubrics cites this paper.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.975242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.975242Z digest=sha256:c885b120d6a553db5abbb58828b3db87c821cfdd77d327201ad01ea3961cd8f3

Observation 50dac8b0-662b-4f05-b4cc-d4f38e9eb180 · inbound

KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems cites this paper.

KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T13:08:23.170852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:08:23.170852Z digest=sha256:ee5bbc25d394659f91c0210850a50b1a931ae195a75a8583de2baa853e444b1a

Observation c7d4bd05-9e66-4595-96ca-9c7d02b0a5a3 · inbound

Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation cites this paper.

Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:02.317133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:02.317133Z digest=sha256:d5117b352f9e9e13aa654de4e4c090183bb0bf109eca3555367d6eb462f15be1

Observation 84a6d43c-35b1-476d-aa3a-e21e8df48cbc · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.107650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.107650Z digest=sha256:ca098fed559167e8be06930982235104195b48781858e1d4fc00bd4e68d514f6

Observation 87c415b5-6dfd-479b-bec7-d0c365df8382 · inbound

A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces cites this paper.

A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T05:17:00.443011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:17:00.443011Z digest=sha256:f802552c6aaa350dd1e4ab59d6ce5898dc8a58e996a268056f5b5179e8c7037d

Observation 84066faa-04e3-431f-b039-ec8378157115 · inbound

Are Today's LLMs Ready to Explain Well-Being Concepts? cites this paper.

Are Today's LLMs Ready to Explain Well-Being Concepts? FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T01:02:52.712447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:02:52.712447Z digest=sha256:739d9cdbac192ebcb5143976f8b8411037c674739593ca2b29875e7990b883bc

Observation ddd119d8-450b-4eb6-b495-24bfb2c8b879 · inbound

UQ: Assessing Language Models on Unsolved Questions cites this paper.

UQ: Assessing Language Models on Unsolved Questions FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.722330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.722330Z digest=sha256:afeb1b4eb60600f56a49921ab730eff62e233e9dcb13eac2e6508cc9449d52aa

Observation 3e4f0aa8-f099-4c56-b6bd-72416660167c · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:57.722583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:57.722583Z digest=sha256:b7ff02ec753c127b482844055ad816926330cb158eda2c7d632c4c7f6185547b

Observation ac8a27f8-de3b-4de6-beb5-98b83b610457 · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:39.895609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:c39ec1b424c9a0a48f3eff9d5a03c6558eb5c7523c27060f8bbb4df5579bb27f

Observation 50ece269-fd60-43e7-942d-47cf9998569a · inbound

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework cites this paper.

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:20:17.465911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T15:15:25.731014Z digest=sha256:cf4447b18885ce0b676b7823d7d3393a57d1327913acb48abeb99eaa04a0a986

Observation 6b95f879-f33c-465c-8d05-f16086ceed2c · inbound

LETGAMES: An LLM-Powered Gamified Approach to Cognitive Training for Patients with Cognitive Impairment cites this paper.

LETGAMES: An LLM-Powered Gamified Approach to Cognitive Training for Patients with Cognitive Impairment FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:18.547399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T21:17:51.805937Z digest=sha256:5f91a6f85c9ae48ccbf14ccc51c6f731d85fb328571d01c7f11f3cc027a6f596

Observation 811466b6-b7db-4b46-a5a3-38f7cf118e1a · inbound

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents cites this paper.

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:05.333681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T02:19:57.341402Z digest=sha256:e7758059ae9095b819ef02214a3c83edd8f1d9287a84b64e19b43c89bfd84be8

Observation 4ac0ccd0-d7f2-45c6-b523-2aa338fee32d · inbound

Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives cites this paper.

Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T01:04:50.196228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T01:00:41.543394Z digest=sha256:1572d48cd64537e4090636369e896b2334eac8de1c0f502279c1eaa502ba7007

Observation 63206756-f47c-4bf2-b0aa-e35c1ae753ff · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.202440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:ccc105f21ee8fd696d6484d627b6ef64422ab0166e8ef8867515f7ad7f7807dc

Observation 1a017449-6d68-4f02-ba82-adcdb4578c77 · inbound

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria cites this paper.

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.343551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T00:58:53.626310Z digest=sha256:faf841a312b6b8d9f7059801b8c7e8c3e2f53007decafba73ca50883a4ed5cdb

Observation 4e4ad4d8-1c69-442c-b7bb-9fe8fdaac2db · inbound

When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs cites this paper.

When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:14.257988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T07:35:18.546331Z digest=sha256:cdca1751a5e56fec9a29fdb55c7ff45f480c24df9d25871a85dc5564226d24e2

Observation 02a760de-746b-47eb-888d-610fbd272bbb · inbound

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning cites this paper.

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:36:23.390469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T14:10:10.215779Z digest=sha256:fd7051d4ce744d8d7c0bbbff190e0a5bee05227226bb33cb884f717c61adb2d2

Observation 0943bb3e-9479-4079-8ef0-06aa0e6c431d · inbound

Towards Multi-Agent-Simulation-Based Community Note Evaluation cites this paper.

Towards Multi-Agent-Simulation-Based Community Note Evaluation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:16:53.938204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T04:13:11.400535Z digest=sha256:a1dc4ae5244159e90d5bce0cbae54d7db8b25656ad4c2f6f2018ed6ad7355244

Observation a962ad72-f002-41aa-9c53-f115f979c3c0 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:53:50.259092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:853e393ba357124ae0c4353d6894eb52b4900591e6e070e2757765da3fb87f05

Observation 85b10f22-d2c3-4088-9cb6-b95c00ee20b5 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:f22757ac2791774649d0d4dd577c288d4ebd9225bff5f30fd5f964d60bcd87d9

Observation 2c67380c-9093-4233-93c4-26d016e8022f · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.657589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.657589Z digest=sha256:f4c9e882d66b08f0cc447fa86fb997030301f1515d7218702b432cc0b122ce85

Observation e4152ccf-6391-47f9-9ca8-cfd91ffd4de3 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:28.395480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:28.395480Z digest=sha256:ade4079a21045d188381d57a8e8f830b75c21085546f58f9b1bf17e4b150ee34

Observation 6389a7fe-8a12-42a2-92ea-f27dc9c379b7 · inbound

RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation cites this paper.

RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:38.951249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:38.951249Z digest=sha256:c6f9d3580d1a11c7c1182f72249e04a346cd4a98c6b066a2fcd228b31b703ca8