Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals

As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.08809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08809 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:11:41.723677Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1155b71-e107-4bbe-aeae-b27a94747926 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.593705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.593705Z digest=sha256:c5d0b9a3cd6b0a4d179d30e13d62f33eb2193b6d2abff2a23c747b90006c6bb9

Observation 780f3e7d-73e0-4609-81b2-53a1f4c6ece5 · outbound

This paper cites Self-teaching prompting for multi-intent learning with limited super- vision.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Self-teaching prompting for multi-intent learning with limited super- vision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.606910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.606910Z digest=sha256:c9da1b6fc27455745406b919a1bc2a6d99315faab0ee102db7c1f7b4620c7ad7

Observation fbd2193e-118d-4d9c-bd11-fc22c1d5cc00 · outbound

This paper cites An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.630126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.630126Z digest=sha256:f7982a09d5a8c8089c6aaee5857cfffd98c69c39058509c6e1dab0bf6e783593

Observation ec452bde-1b91-4f30-835a-199a972eb0e2 · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Model Guided Tree-of-Thought

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.647592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.647592Z digest=sha256:11190ab1e3f3946341bb74ee1ceb09bce0fe31f170c59682a0b43c73b003ba92

Observation 92527c59-d016-460c-89fa-24f1b07f8594 · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Models for Data Annotation and Synthesis: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.653118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.653118Z digest=sha256:7ef7dcce397e585ed3c5ebf1fa123d80e69d90c959cc20616d43e60e1bf2d220

Observation 8b125de1-3aa7-49cb-8222-c159da1a5b8f · outbound

This paper cites CodecLM: Aligning Language Models with Tailored Synthetic Data.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals CodecLM: Aligning Language Models with Tailored Synthetic Data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.660463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.660463Z digest=sha256:4d97ff0ba1828e73c80862797a7a60361c7d0e03afcde6e869e3d17e18f6e117

Observation 1f8d1606-cbe6-497b-9bbe-dbc658520440 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Finetuned Language Models Are Zero-Shot Learners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.668846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.668846Z digest=sha256:0fe3f0e55c904660fd5a556f09b049afd96bb18812700958ad9bda09fc880efb

Observation eaf3db79-79cd-464d-983f-d6e95b491a94 · outbound

This paper cites Unigen: A unified framework for textual dataset generation using large language models.arXiv preprint arXiv:2406.18966,.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Unigen: A unified framework for textual dataset generation using large language models.arXiv preprint arXiv:2406.18966,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.676658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.676658Z digest=sha256:6bcce0c8b88680a77ae7f048b5589b2a9bfaff98f97cd6b9341f09b7e45540c1

Observation 2e045291-eb2a-472f-b5bc-c7a70f28a214 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.683557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.683557Z digest=sha256:f201dbc2d6d3199490301ea8df046feae8573b7f7cab25e4e4c4107684fe302c

Observation 338a3504-808b-4418-8baf-c4d8a489ba21 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.691968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.691968Z digest=sha256:5923f2e27a344ef7bee8d090f5563e5297e9077a51d1077316070f6747ac6d23

Observation 30e50ffc-9452-48bb-ac5d-c14f7fac065b · outbound

This paper cites ZeroGen: Efficient Zero-shot Learning via Dataset Generation.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals ZeroGen: Efficient Zero-shot Learning via Dataset Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.698533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.698533Z digest=sha256:5125dc104c18d46c33df34d4d616170c072ac3591ead7e4101b3db30f569bbee

Observation 07457aea-99d2-4fe8-a7a1-b60ed49e0470 · outbound

This paper cites Clusterllm: Large language models as a guide for text clustering.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Clusterllm: Large language models as a guide for text clustering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.711658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.711658Z digest=sha256:21423145e183b2f922344e2daeaed903c66afb04e2c2d5292cac2b521b7e4cfc

Observation e8ead15a-d86e-4033-a881-23c70738845c · outbound

This paper cites Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.717627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.717627Z digest=sha256:0dbc562eef7d54b65407453c1ec98b8286a1e06242b96f7330240aea6cf51ec7

Observation 7c32b9ed-be89-420c-b367-3de21fab166e · outbound

This paper cites yi denotes the LLM annotation accuracies.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals yi denotes the LLM annotation accuracies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.723677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.723677Z digest=sha256:bd64fc24701a99bbd1e83791c1598637bad183445271a3bc8d3369aa4abd2033

Observation 70bfafd6-52d9-4e8e-a797-b91e80301d73 · outbound

This paper cites Large Language Models Can Self-Improve.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Models Can Self-Improve

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.623904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.623904Z digest=sha256:790bda766d514d0b3926d63d8c02e9d7f425e8dec104adfdb4a8c0f6bbf97ebe

Observation 8ddb5ba5-2b9b-4000-8a18-fbe2b55a60cb · outbound

This paper cites MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.635830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.635830Z digest=sha256:a18f087873df9df7daeeb83a6f5f1f81fe841946e3e6979bb6c6404f2ca56bf8

Observation 7c8b7f60-0c63-4701-888c-dcdcc7e57049 · outbound

This paper cites Efficient Intent Detection with Dual Sentence Encoders.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Efficient Intent Detection with Dual Sentence Encoders

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.601157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.601157Z digest=sha256:01be92cbc85f3491368832cbbeaf165ebc863b385835721a5f2073d6b79f74ab

Observation 4ace878f-ff85-441e-ae45-d2f847d49ad1 · outbound

This paper cites New Intent Discovery with Pre-training and Contrastive Learning.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals New Intent Discovery with Pre-training and Contrastive Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.704280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.704280Z digest=sha256:1a884b9250894908ac4078933f1d8a9805f840997d5c68a4258387cbbd4d01af

Observation 71e14911-340e-4b2c-bdec-063d2428ef9e · outbound

This paper cites TWEAC: Transformer with Extendable QA Agent Classifiers.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals TWEAC: Transformer with Extendable QA Agent Classifiers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.612448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.612448Z digest=sha256:5e14a3562d08aa435fb44a75317f6b21d199241fe10aa3a0164cb3e7c0f80dfe

Observation 8d39323f-e64d-429f-abf4-1048cb2d4f7d · outbound

This paper cites FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.618338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.618338Z digest=sha256:55974af0e135e7778341011e817b5c84436c7fd050fc7deed94d73567866ac2d

Observation b97cea59-8b75-4c4d-b098-1dd48215103d · outbound

This paper cites Graphprompt: Unifying pre-training and downstream tasks for graph neural networks.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Graphprompt: Unifying pre-training and downstream tasks for graph neural networks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.641752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.641752Z digest=sha256:321180e959c9cab31c4af38736cbf000ea86c73603034ad3faceb1131c6532bf

Pith citing papers

No inbound Pith citation observations are available.