Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals

As of 21 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.08809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08809 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:11:41.723677Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1155b71-e107-4bbe-aeae-b27a94747926 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.593705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.593705Z digest=sha256:519764e8c986291e11830e693b693dd1345a0192def3405d244962d96d1209d9

Observation 780f3e7d-73e0-4609-81b2-53a1f4c6ece5 · outbound

This paper cites Self-teaching prompting for multi-intent learning with limited super- vision.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Self-teaching prompting for multi-intent learning with limited super- vision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.606910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.606910Z digest=sha256:db52adee1a172bed6341e4301b68a859711278525816babe1734a0aa24e7f5dd

Observation fbd2193e-118d-4d9c-bd11-fc22c1d5cc00 · outbound

This paper cites An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.630126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.630126Z digest=sha256:06b3d17bf0e7c8124f0f09d8e4d21c7418a38d9d0ee8ed0d24f6ae0911fd8f87

Observation ec452bde-1b91-4f30-835a-199a972eb0e2 · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Model Guided Tree-of-Thought

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.647592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.647592Z digest=sha256:01ab3f40adae87048a1ccaf245929eff7f94b74d70feb3c6f24faa03ee608091

Observation 92527c59-d016-460c-89fa-24f1b07f8594 · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Models for Data Annotation and Synthesis: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.653118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.653118Z digest=sha256:109fcb4a10ed30ba0d03c0626100c413c3f17653092ccfdb9d16f1ccc8392dfa

Observation 8b125de1-3aa7-49cb-8222-c159da1a5b8f · outbound

This paper cites CodecLM: Aligning Language Models with Tailored Synthetic Data.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals CodecLM: Aligning Language Models with Tailored Synthetic Data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.660463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.660463Z digest=sha256:57a0fe6c1c06ea677ab9f695fd62658790f15114389d4eeb6dc5b7c712f915a2

Observation 1f8d1606-cbe6-497b-9bbe-dbc658520440 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Finetuned Language Models Are Zero-Shot Learners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.668846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.668846Z digest=sha256:74541d0ae03c2928de1446202da308918c52fe6883c2438d6e75680a4a40a30b

Observation eaf3db79-79cd-464d-983f-d6e95b491a94 · outbound

This paper cites Unigen: A unified framework for textual dataset generation using large language models.arXiv preprint arXiv:2406.18966,.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Unigen: A unified framework for textual dataset generation using large language models.arXiv preprint arXiv:2406.18966,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.676658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.676658Z digest=sha256:afeb270f5b9fa0bb357cde3797463cb9c76edf00b2eac89a28517361e3c51fc3

Observation 2e045291-eb2a-472f-b5bc-c7a70f28a214 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.683557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.683557Z digest=sha256:0ff8480876ab80e4d2ef85da8f904e9707232ded3c0e2516116f994cceff08da

Observation 338a3504-808b-4418-8baf-c4d8a489ba21 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.691968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.691968Z digest=sha256:7d97fa854521000f15bc260a98aa48bf6d177d78209ebf546ede9e405d5a0cfa

Observation 30e50ffc-9452-48bb-ac5d-c14f7fac065b · outbound

This paper cites ZeroGen: Efficient Zero-shot Learning via Dataset Generation.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals ZeroGen: Efficient Zero-shot Learning via Dataset Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.698533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.698533Z digest=sha256:019d2d16c37dedeb41a9a6d496629bea5b6383cc9e56bfce60201ac49c781fc5

Observation 07457aea-99d2-4fe8-a7a1-b60ed49e0470 · outbound

This paper cites Clusterllm: Large language models as a guide for text clustering.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Clusterllm: Large language models as a guide for text clustering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.711658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.711658Z digest=sha256:c285110144742b07de9bc7f59d5a36b20ae9f8c28c6a2371cc2174f28f7e88b6

Observation e8ead15a-d86e-4033-a881-23c70738845c · outbound

This paper cites Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.717627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.717627Z digest=sha256:d19e23a9104ae95d05dc14bff4dd008d217d5ec50152391a1bcab533a6c1f7ab

Observation 7c32b9ed-be89-420c-b367-3de21fab166e · outbound

This paper cites yi denotes the LLM annotation accuracies.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals yi denotes the LLM annotation accuracies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.723677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.723677Z digest=sha256:2d9f67243cb20d8e09b2531ed1681a20495e6ad8aebab9732b58c9c5f00152d6

Observation 70bfafd6-52d9-4e8e-a797-b91e80301d73 · outbound

This paper cites Large Language Models Can Self-Improve.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Models Can Self-Improve

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.623904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.623904Z digest=sha256:254260ceee1d9ee001b3903d7de5a2fcd61cfb4948bc25d44efaecd0f4561c8f

Observation 8ddb5ba5-2b9b-4000-8a18-fbe2b55a60cb · outbound

This paper cites MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.635830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.635830Z digest=sha256:19c4dbe5997560393b3e5ca62648790454dc9a56819257d1308473d322422976

Observation 7c8b7f60-0c63-4701-888c-dcdcc7e57049 · outbound

This paper cites Efficient Intent Detection with Dual Sentence Encoders.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Efficient Intent Detection with Dual Sentence Encoders

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.601157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.601157Z digest=sha256:a19f0e029032d0d9ee5f57e14a2f6a5c63959bb35f4abb2c23c6d09e8ffaa990

Observation 4ace878f-ff85-441e-ae45-d2f847d49ad1 · outbound

This paper cites New Intent Discovery with Pre-training and Contrastive Learning.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals New Intent Discovery with Pre-training and Contrastive Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.704280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.704280Z digest=sha256:77aee9344a1eaf2aec1341f5b67c6404d20b43b2dcf2c2f71193b11acd5eca4b

Observation 71e14911-340e-4b2c-bdec-063d2428ef9e · outbound

This paper cites TWEAC: Transformer with Extendable QA Agent Classifiers.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals TWEAC: Transformer with Extendable QA Agent Classifiers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.612448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.612448Z digest=sha256:1e8bbcccfafd07bada8b8c552bb3cd285511091c8df41183826e6bf1f41ce460

Observation 8d39323f-e64d-429f-abf4-1048cb2d4f7d · outbound

This paper cites FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.618338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.618338Z digest=sha256:262e5b520c62ca439f934788fc821fccf07e9e4452297fb5ce6723bd870fc108

Observation b97cea59-8b75-4c4d-b098-1dd48215103d · outbound

This paper cites Graphprompt: Unifying pre-training and downstream tasks for graph neural networks.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Graphprompt: Unifying pre-training and downstream tasks for graph neural networks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.641752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.641752Z digest=sha256:80f53e65dc354128e1f5515e2466a2008135d6602bb8d9f911dbd5d932e39ca6

Pith citing papers

No inbound Pith citation observations are available.