Pith. sign in

Paper Citation Record · LEDGER

Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2403.16950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.16950 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:28.566067Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a413f4a9-afcb-478d-8dcb-db1cbab5a1b4 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.365470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:a457968cf782336020b06c251657778f8010bd20cbc0b248a56785a8d470f78b

Observation dcf14201-5066-4262-a930-f65456683b2b · inbound

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives cites this paper.

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:44:55.865958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:44:55.865958Z digest=sha256:b96c2bb1a1f8caefccbc0fb7151201dad7abc24604d89a57128e2b85bebaedc3

Observation 44dc0183-0c48-4173-9e4e-b63d19ed9b6d · inbound

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios cites this paper.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.390831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.390831Z digest=sha256:79ef9521f6a1852c990535dd1f801aedb84edc7da8c63df5324427663db221af

Observation 77254223-cb60-4048-bf6d-6f6806b002ac · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.387370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.387370Z digest=sha256:f30f2879167ab8232255bcd83a986d82234ad91a7da42c15adceb77fc22d2301

Observation 08a2d0c1-e85d-4b7e-b30f-8f3c2221b847 · inbound

Optimizing Compound Retrieval Systems cites this paper.

Optimizing Compound Retrieval Systems Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:28.566067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:28.566067Z digest=sha256:0361b80b8490cbf48374bd180802efcb4f8ea2afb43e44c49b409eafb40e6557

Observation aa8c7127-2558-4955-acec-0829f72af897 · inbound

QBD-RankedDataGen: Generating Custom Ranked Datasets for Improving Query-By-Document Search Using LLM-Reranking with Reduced Human Effort cites this paper.

QBD-RankedDataGen: Generating Custom Ranked Datasets for Improving Query-By-Document Search Using LLM-Reranking with Reduced Human Effort Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:52.486624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:52.486624Z digest=sha256:ba7393f64868fbb2aabb229482ed2f9e2b2602df0ad79c00477d19de30a07f1a

Observation 0536fe8d-b247-49ee-b366-65b2947586ca · inbound

Ranked Voting based Self-Consistency of Large Language Models cites this paper.

Ranked Voting based Self-Consistency of Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:26.836007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:08:26.836007Z digest=sha256:de4f1319236e1513200b93013839d6d1c7f4f5a3d4df57907a25ffaf5fce9e87

Observation 045f3ca9-c561-4f3f-a762-ee3b14c03279 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:35.554333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:35.554333Z digest=sha256:b7eaf7b64983a32c9193f6a119e3975e3068d4802a702e379723302f713b0198

Observation daa50822-7eaf-40f1-ba4a-31023c04a764 · inbound

Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases cites this paper.

Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:45.759286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:45.759286Z digest=sha256:9756dc13b494d0ea70fc34298d06d945ecbf8dc3e5dcc4eb00bdf43015000331

Observation 02c6a2a9-05d9-4fdf-b7ce-ef5d54c8ffc8 · inbound

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models cites this paper.

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:07:14.262841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T10:05:22.562737Z digest=sha256:abe6c7fa1105b98aea0c26b4e0f3eae401ffb266b0c30c98858d1406bb3dd906

Observation 45712385-d5db-491a-9802-e76798fa58b0 · inbound

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models cites this paper.

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:01.132316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:01.132316Z digest=sha256:25799ee2c2c4aa6b66d9e7e40086fa5a0aa05bd908559c54bd86b978a26b5ca7

Observation 24071394-692b-4ee8-9aaf-90635da0d43c · inbound

Deep Researcher with Test-Time Diffusion cites this paper.

Deep Researcher with Test-Time Diffusion Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:21.497194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:24:21.497194Z digest=sha256:9c541618a689ed84269533819749f443382dec9e610fca04f0f4647c216534d8

Observation 9b4a0313-598a-45b9-a452-6ac597410690 · inbound

Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization cites this paper.

Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T16:13:49.902821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:13:49.902821Z digest=sha256:191129ddd825756ed306a95f67bb3052a643f7fe4de032f0b9a4d33a8de55e91

Observation 23d6e99e-59bb-4bdc-9aae-80085bd46aba · inbound

Semantic Data Processing with Holistic Data Understanding cites this paper.

Semantic Data Processing with Holistic Data Understanding Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:08:09.534763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T19:07:49.756349Z digest=sha256:1e10c13b0ccf385ec1f0e5081ad53cace72dbd0fe3d9b461d3e23db8603add9e

Observation 333773b1-c3c1-434b-95d2-a90fbc04dafe · inbound

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild cites this paper.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.133595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:dda25d263cfc2e123530bd3a942ddbcfb339d5f905010b34f55dfe9c468a94c7

Observation a1076fd7-f506-448d-837a-fac402bdfa98 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.214388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:c72e3cbc21173e4bfb54a815fae87049d4a4d7f338c8f54954f1c887ab63dbf9

Observation 4ad14bc4-2452-4807-a925-15e9573003cc · inbound

GRASP: Deterministic argument ranking in interaction graphs cites this paper.

GRASP: Deterministic argument ranking in interaction graphs Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:58:14.909207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:57:55.198779Z digest=sha256:f50c8a5503a668ddaab2dca0ad96b0fff95e33d91c740a759fdf8a16c41a6cc7

Observation a917ec0f-6780-4430-8a9c-2f8683a63e5d · inbound

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges cites this paper.

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.723043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T05:58:59.870335Z digest=sha256:fab381da722c6ae0d67059cc91539a9fd6d58a519d6012148ec4720944c7255c

Observation 99177898-7b7c-45c6-9c21-ca4dc5678be1 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:47.116638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T07:49:36.816100Z digest=sha256:650efa6e21fca2e943ac72fa05afb1730ee9c338d0c1d77fa4d4d29bd85d494b

Observation 895e2b5d-14bb-41bc-aa26-1af41d739c06 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:04:39.390666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T10:17:33.176525Z digest=sha256:c2da9055fe9961bf0fa3f1c563ebf614b2b0518422a0ac0310a9a15313c5c7de

Observation 62ed0cd6-59b0-4254-ba16-c3d9dd522f82 · inbound

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank cites this paper.

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-26T03:58:57.117163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T03:56:28.271760Z digest=sha256:7f26e36c484661c1b27eba972208fa5c2cd38f73fbc60a928fdda8671d51a398

Observation 4c09fb6f-9b4e-43ac-973a-10cf1f96a2c3 · inbound

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 cites this paper.

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:23:00.942005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T02:17:26.872854Z digest=sha256:0d4de566750cf21bed4f2331575616b995289231f6443ae7bfee4d598d89f6e6

Observation 323dacc8-4392-4f0d-b5ad-4560fa5a6002 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.163786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.163786Z digest=sha256:8f3c09a3965e7c3983a7a27f443018b692fae7d4e135963d2825dc12d12867c9