Pith. sign in

Paper Citation Record · LEDGER

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

As of 23 July 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2606.26429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.26429 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:14:26.356160Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T16:05:41.788680Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T16:07:20.325594Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ccfddf6-26ec-4c79-88af-b54cabf3e3cc · outbound

This paper cites Forty-first International Conference on Machine Learning , year=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Forty-first International Conference on Machine Learning , year=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:4940ed8dd30970ff65aeaa50a5e5d1aa377975cc6f57aee319670ecc12122b23

Observation d318c180-a4be-4f7f-9432-e0dfe025a42c · outbound

This paper cites Data Contamination: From Memorization to Exploitation.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Data Contamination: From Memorization to Exploitation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:49:58.248277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:c2f9360db5aaa963264f53c910e645d8e9ff6d315a2f3062a398e79d50879fbc

Observation 7006b46c-66a3-4b39-9e7e-348a32b0f6d6 · outbound

This paper cites Nature Communications , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Nature Communications , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:980566f15b80b2faa14b948d2e8b3df862ba09ee676bfbe0fc201a6ba19ea844

Observation 55cfe3bd-196c-4919-abf4-5e08550773f7 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.280215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:89c061e3fbef3bca6686cc3f5371615e0efd5430fcf0e883e749320a56c3cd7b

Observation 1f0a3836-84f9-489f-aa22-3e19cae5979a · outbound

This paper cites the method of paired comparisons , author=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation the method of paired comparisons , author=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:af9d28d6e9eb846dcb5036990cfa92a5ebb1b568a15bc4cba8ad2c946679b298

Observation cddb38ad-1887-48f8-9225-fa08122fdde9 · outbound

This paper cites 2008 , publisher=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation 2008 , publisher=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:41dba140741ff8de7e3c8b25213363cf7630325fd51112c19e914ba163eb1053

Observation 932600f6-e401-4ef2-8dfb-ef2562857349 · outbound

This paper cites Educational Measurement: Issues and Practice , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Educational Measurement: Issues and Practice , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:72b753e752cfd7921ecf6bdb72b69e1c915eeb511d9255ffe0e6f399d170d600

Observation 3834d3aa-5bf6-4e3f-a311-3fb96ed422c8 · outbound

This paper cites Statistical Science , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Statistical Science , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:454c45870aa0c82bc3d6c7830754d695c9efad630e82c070421ec336a68e4daf

Observation c63ec8d4-b2b7-40a0-9dbf-1ec627767a09 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Measuring Massive Multitask Language Understanding

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.259676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:ad9c04388740fd9d57ed5a17d33e57e18f4a817b6232d7a6e2bb49182b85a562

Observation b02c6b0e-583a-4b47-8a62-6978ad7700b7 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.277852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:cfda8d29b50661bf9164ebfa2f5f3cc4b39b7dbf01ed8e97a9e6cac91f517243

Observation ef1bdfa6-9472-4242-91d3-084e3dc71feb · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Transactions of the Association for Computational Linguistics , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:f2d392269584ee1994111bf0442a416be46a100a322177d97dc54f3ba6839f23

Observation f714d820-4112-418c-b82e-e190103319dc · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.261898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:3348e52fec65a13dbb7d9d18f60f0881924dc2fa4a4fdf214ea3f322e149e514

Observation d4a32d6a-1054-463a-9ed8-fc5924b0ddd7 · outbound

This paper cites Proceedings of the 2018 conference on empirical methods in natural language processing , pages=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:ff68fe316350da6593e799a7e096f6ad9f5d700203214026ffd48e5e4c1ebbc6

Observation 3b68f834-cc3c-49d3-af0a-d232eee38211 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.257368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:7db237322614d469d5b3623fe4c692f23b81889bac3ba69a293116109e8e73a4

Observation bd09cc4e-0987-4f74-a636-001460ca0419 · outbound

This paper cites Proceedings of the 2015 conference on empirical methods in natural language processing , pages=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the 2015 conference on empirical methods in natural language processing , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:e6e493a5c173ce4dc2c8bb26fb0b49e5079cf96130b4d90e45fa58a699731973

Observation e481baff-3803-4266-aab0-8190f3234a61 · outbound

This paper cites Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long papers) , pages=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long papers) , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:ac3d0cf97f410053867f0d539748aa71495097f3258eb5037597fbbe8bb5de0e

Observation a1d1f07e-3724-4663-b11f-7098fe330b44 · outbound

This paper cites Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP , pages=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:23effc29b4d37b8a128295298e04d4aee22a032d3ea85da477930055d031f4fd

Observation 180266dd-f50a-4d41-b524-92071757867a · outbound

This paper cites Advances in neural information processing systems , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Advances in neural information processing systems , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:34413c695c3969f16626dea3e8400e3bef4bc8d7137250179d10fb16ddd92811

Observation 7c4a2a6a-60cd-4d59-9ad0-93cba028c5d3 · outbound

This paper cites Holistic Evaluation of Language Models.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Holistic Evaluation of Language Models

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.264058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:40044ae0bea140b036e0fdb072127dc246ec5de1cb1153a40b600ca776f94474

Observation b350b804-cfc1-4bd5-a8ac-6656a013f835 · outbound

This paper cites Handbook of statistics , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Handbook of statistics , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:c112a39f88bd01800a02acc66ef5970c161af2d0b3c89f211b4e82c9b1254cea

Observation df8773dc-c4a7-4cab-b696-46940a66f01d · outbound

This paper cites 2013 , publisher=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation 2013 , publisher=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:8009e638483244cea03eb141c8ea334ef4ce86024eff4c6e7a0a645a2d998a65

Observation 43f7c667-275a-47ec-afa1-f770cfb4edde · outbound

This paper cites Proceedings of the Conference on Empirical Methods in Natural Language Processing.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the Conference on Empirical Methods in Natural Language Processing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:a5d579f9219d60ce81eac82ba33bed2b7161c3fc7f6f23e9c13717a7c8ae4c99

Observation 51e9c071-aa89-4d4d-95c3-e99a40991ec5 · outbound

This paper cites Artificial intelligence , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Artificial intelligence , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:df67c28a1b8bfaad39461b991f4c71d6af701f20e418b2f093e4900fb604307b

Observation ba47fbbd-40b6-4cc2-8203-e084f6ffd983 · outbound

This paper cites arXiv preprint arXiv:2510.04051 , year=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation arXiv preprint arXiv:2510.04051 , year=

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:58.268798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:9e72d339c819395d9ba6b623a32b207676c501c8e9a1239349ce9e0c9436c2d3

Observation 24721fe9-f405-4d74-a960-f6727dac347e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Advances in Neural Information Processing Systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:e03419360684aaa382e5f87007585cfa27f1adf371df2f4644d62e0963c3d8d4

Observation d1daf105-995a-48d0-b289-1f9ded24c6ae · outbound

This paper cites arXiv preprint arXiv:2510.14966 , year=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation arXiv preprint arXiv:2510.14966 , year=

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:58.250580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:c02fb45ff564cddaad5f603fd89e7211b2a31d3543fd6d9d1d7526de746926d5

Observation 7a62f2c5-7bd4-452c-ad66-f6ade2147e8f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Evaluating Large Language Models Trained on Code

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.252821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:f1efba64589483e052bc7ba73c46d0e6e25ea72ec73a28a6d4d44778e9da5278

Observation 2f24fda3-05b5-452b-9491-866dd4d69ca6 · outbound

This paper cites Program Synthesis with Large Language Models.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Program Synthesis with Large Language Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.270929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:6fd844b954e2ee74ddfa3dbdd49a901bdc46bebef5d5979483bd3ca634b93a1c

Observation f25ab560-0a33-4719-82cd-37e1bf4d79d0 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:131b08a2abab22e1390ff5e89d2511edd170284239fc881d3345aa7fbe012265

Observation 279d9781-488e-41bb-bca2-a93605aea6a9 · outbound

This paper cites Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T01:20:49.385427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:0a993781a1ecd6a2af0baf23e531ec8bc8eea8fe5de3433209b4f330305759e7

Observation 905e5a3a-99ef-46d3-9dcc-e65403a8928a · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:72a83167345def567ca7ab6f6a4cb7eccf1728760ad5925dc6cf49bf78b2e393

Observation ff679207-a8b8-400e-b497-ffd21b12a42e · outbound

This paper cites Rethinking llm evaluation: Can we evaluate llms with 200x less data?arXiv preprint arXiv:2510.10457.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Rethinking llm evaluation: Can we evaluate llms with 200x less data?arXiv preprint arXiv:2510.10457

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:58.255096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:6b5ba4ac2a6cb1058c6f9240ab4abfd72247909505ff80badef1002f97becae9

Observation e4dd826f-c775-411b-a82b-f9e4ef6b7489 · outbound

This paper cites Proceedings of Machine Learning Research , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of Machine Learning Research , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:45b755db29ce3d258ae45b734ad0ec5b3e06ef4e0972a33abd889ccc212f6a53

Observation 8416edd7-a460-495a-be28-9de5f9792b2c · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.273229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:4f444c8c129cc0cefada1b9892c1a0831dd5c185c1f9b75ed563478000e589fb

Observation 5c74ef1d-97e9-40b0-aece-4f8ace94720b · outbound

This paper cites International Conference on Learning Representations , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation International Conference on Learning Representations , volume=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:03246c850efcd59ab136b98c02c79c8f0b6de593ee6540ada46c2bfa801ddc09

Observation ddc59372-a33f-425d-916a-af262d8da50b · outbound

This paper cites Proceedings of the 12th international conference on natural language generation , pages=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the 12th international conference on natural language generation , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:8fcdac2b80036d492e38a750cae772b69ef06f59e4ccbdc36312e6035d463292

Observation 7f8c0343-eb32-43fd-a7d6-0fc0cd8a11e9 · outbound

This paper cites Proceedings of the 1st workshop on benchmarking: past, present and future , pages=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the 1st workshop on benchmarking: past, present and future , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:9e675c1c5f25b0f2b2fcbf61cfeae14f7c840325478ac3669506105d1ab01344

Observation 88ae9e67-0ba4-47ee-9ea4-1d3face42b6b · outbound

This paper cites International Conference on Learning Representations , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation International Conference on Learning Representations , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:766e2e176cbd897b8b5b12710ddd0f0620024c1db95c5f2687f5ef787514ac43

Observation 84652388-98e8-4a29-b7a6-1aa76ff34d75 · outbound

This paper cites Humanity's Last Exam.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Humanity's Last Exam

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.245588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:dc2522c942e0b37d3f1642d5888b0e5adc4c8a4ca648d9a3605e86a66b4f6662

Observation c0ca4663-e1c4-49ec-93db-77df631ccf51 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:59e1381eda10d3d0794b696cb0791ff36bc8c922fcf249519e3c94eaeb28cd05

Observation 28794ce1-a3df-42ff-8f00-9d91e372ff60 · outbound

This paper cites Measuring short-form factuality in large language models.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Measuring short-form factuality in large language models

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.266350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:76b773449445c79ac5b94d41f12760903bdd213ad95dfbce6a303e3108e1c6fc

Observation 78756b5c-460a-4e82-bc85-fe2dd192b7bf · outbound

This paper cites Advances in neural information processing systems , volume=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation Advances in neural information processing systems , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:e9d7ccd6b1e5e36ee96208b56b19add3bf7565582773fc87b1b6c4f718d7975a

Observation 0f8facc7-978f-41d5-8d6d-c91e8094eaaf · outbound

This paper cites 2026 , howpublished=.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation 2026 , howpublished=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T01:14:26.356160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:97c28ef373495fe776becf0b30de7ba01e5433cc2851432259075c72d1af95f9

Pith citing papers

Observation 1c17183e-e198-4ae5-b2d0-0daf2d4730e5 · inbound

Bug Report Specification Refinement with Trajectory Guidance for Automated Program Repair cites this paper.

Bug Report Specification Refinement with Trajectory Guidance for Automated Program Repair DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:07:20.327776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-10T16:05:41.788680Z digest=sha256:48d21dd14021c3fb0f2cb747d8ea3dbc57193dd6a707ade0a3f6f00f529f9f84