Pith. sign in

Paper Citation Record · LEDGER

Answer Matching Outperforms Multiple Choice for Language Model Evaluation

As of 7 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 11 inbound Pith citation observations for arXiv:2507.02856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02856 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:08.655335Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:47.447302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:00:01.346988Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2da168eb-6191-4d7f-8310-64499fb7d47c · outbound

This paper cites • Reversible: Expansion is infinitesimally slow, maintaining equilibrium.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation • Reversible: Expansion is infinitesimally slow, maintaining equilibrium

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:10.621361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:07.892717Z digest=sha256:d9ea8d0df1522928387c2e3a6f6b8ee4029ebbea7566caa6e098292eb4c909f3

Observation 498c18bb-1daf-4009-a314-41e5477a80ee · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:10.418925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:07.958785Z digest=sha256:a4052c3368603b86c4ada216c8615901af279f2acb880b93f15e5b49c6cea56b

Observation 478e6010-5de2-44f0-a451-1b74e04be053 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:10.185150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.044400Z digest=sha256:c15dd73fc0dfd1045a54a8b3b9c004857e516545324b25fea19048754125cec6

Observation 5b9ed96b-4f81-40a1-ba5a-cf33897bef0b · outbound

This paper cites Calculate the total change in entropy, when a sample of nitrogen gas of mass 14 g at 298 K and 1.00 bar doubles its volume in an isothermal reversible expansion.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Calculate the total change in entropy, when a sample of nitrogen gas of mass 14 g at 298 K and 1.00 bar doubles its volume in an isothermal reversible expansion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:09.978120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.111058Z digest=sha256:de8847691ee43949f55ccc616c0d1d8337fa13520d8676636464392a13139f78

Observation a63bb3a2-be4b-4674-aba5-836b42a28dc2 · outbound

This paper cites We find that MCQ estimates the highest accuracy followed by LLM-judges.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation We find that MCQ estimates the highest accuracy followed by LLM-judges

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:10.801232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:07.805474Z digest=sha256:935bc409abca524a6041575578cdd471ffb3b70caf40df451d21c639ff62a482

Observation 7992e3db-e97f-4297-a46e-d39ff3bb736c · outbound

This paper cites " " # The response can have more information than the ,→ ground− t r u t h . # I t can be more s p e c i f i c ( f o r example ,.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation " " # The response can have more information than the ,→ ground− t r u t h . # I t can be more s p e c i f i c ( f o r example ,

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T20:23:08.975801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.655335Z digest=sha256:15978c26cf9039ae97d382971e93a9559cd7a964a05b578fa7215e79700c43da

Observation 06cffa44-e03c-4bf0-b34e-c30b8d3f07e2 · outbound

This paper cites • Reversible: Expansion occurs slowly enough to maintain equilibrium.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation • Reversible: Expansion occurs slowly enough to maintain equilibrium

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:09.822721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.199589Z digest=sha256:dfcb493cd51d54dabc18a094979c61c741b4ce5938b6d93db959ef736d50e665

Observation 6d569b6e-02ca-400c-9298-cf15e9353434 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.674329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.289554Z digest=sha256:844fc5e08537deadb726f0ffe9adde13cefe472d8f1365a4abd1252fca8ba3b3

Observation 1bf6c3ef-6f46-4469-b995-a7d238a450d7 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.489007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.356038Z digest=sha256:5891ea967d001c188104e11b8acaa5afd683f5d889feb975d010f0ccb2c79c78

Observation 67b44c01-64bb-4470-99fb-410eb0e0a570 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.334697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.450396Z digest=sha256:3630938fab9a2fdfdbd27336454455d2f91c4216bddefd11424d2023ca3d8713

Observation 7b28d7db-4c2a-41b7-8754-b16ac26cda0a · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.125448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:23:08.546381Z digest=sha256:025238c45e25115929f9d060e0dc695de87834c42eb86a5d9d44eef1685814a7

Observation dc0d4a58-c0d3-4146-bf3f-70b7017a9a22 · outbound

This paper cites cloze procedure.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation cloze procedure

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.761432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.761432Z digest=sha256:89c4267fbb1807e783bebd61199aa0db3bb15fdfc12ebc6cfe3bc380a3f8afc7

Pith citing papers

Observation 7c5363e6-c712-4c37-aa28-cbdd340f86bd · inbound

Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions cites this paper.

Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:11.483898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:11.483898Z digest=sha256:3502a3e91c7b3f9dbeee999b1791aca8e47806ba2e1498186b7898c3ae0ed5ad

Observation 6e6db20a-4026-483b-a1bd-cca37b5d2f2d · inbound

Entropy After </Think> for reasoning model early exiting cites this paper.

Entropy After </Think> for reasoning model early exiting Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.599288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T11:51:58.579048Z digest=sha256:452a3955757b9ceaa72a2d9c23dd11dab35b289ccc64a3c5facf419e8200baaf

Observation 12c11f98-a00a-4ed0-ba99-d3424204496e · inbound

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants cites this paper.

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:12:22.093472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:12:03.909181Z digest=sha256:44da0a0a4bc5f7188ea843f2caf0f060a9641c3e090461964999afed080979d2

Observation 14106616-9b6b-493c-8290-c9e419e0e622 · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:d30d3d0fd5d16768708e241861130ef6ca628fcd97397ff3c768a11cb9e064ef

Observation 6b6c605c-69f9-42ce-8b0b-e98af22fd748 · inbound

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings cites this paper.

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:31.152523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:32:23.001110Z digest=sha256:667107bff33e69605abe734eec81918aa086b96d0fe36c9a50177951ffb72e1a

Observation 9a7f5675-c8a4-4ee0-8a6c-a4089dc377a3 · inbound

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval cites this paper.

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:27:40.323764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:13:21.687950Z digest=sha256:3f36e62f58f333f2113a6fd8177e0b7a327dee2079055bb973ba2f35faaf0ef5

Observation 0104e497-a442-4600-88db-9684d3bc0a6a · inbound

Improving Cross-Format Robustness in Language Models with Multi-Format Training cites this paper.

Improving Cross-Format Robustness in Language Models with Multi-Format Training Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.335157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:13:22.927828Z digest=sha256:84ebbf04635baf7af42d76c68c052661a15d9f7386f570d6c78e6dd904cec327

Observation caacba0b-b620-4009-8718-1170947ae885 · inbound

Storyline Trees: Hierarchical Representations for Long-Form Narratives cites this paper.

Storyline Trees: Hierarchical Representations for Long-Form Narratives Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:33.787019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:10:05.604537Z digest=sha256:7b60ac41a23ac5890c1b4ce837cb96bd2fc632b1276505641976a5824be915e9

Observation 22d4eb0e-8361-4029-849e-ce002bb215c0 · inbound

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice cites this paper.

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:00:01.348676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T23:08:27.995672Z digest=sha256:1f556cbddfc3d91b77c794345568422503dc044ec8320c59fc5b4017c4d19c8b

Observation c967abae-1ab3-4368-b4b6-5466d675666d · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.011737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.011737Z digest=sha256:8acdbf19d4a021c728b4060e60f8552ea65ceb3de94d7a3a599e2d76a06efc5f

Observation d1c55097-9c9a-49dc-bbdf-d047c057a294 · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:47.447302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:47.447302Z digest=sha256:339c859cd70d2ea3ae863ab9963e36a0f89478c7d7ccbf2e96d44b6dca48bb82