Pith. sign in

Paper Citation Record · LEDGER

Can multiple-choice questions really be useful in detecting the abilities of LLMs?

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2403.17752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.17752 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:19.156591Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8d466641-af47-46df-b93e-fc39786e1f3a · inbound

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models cites this paper.

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models Can multiple-choice questions really be useful in detecting the abilities of LLMs?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:18:31.878279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T22:15:52.638622Z digest=sha256:ed331a2eb63cad7ea5d8090a5d5c291bec69c4870e0059b41eb9c296bffaceb9

Observation 3be4421e-4f5b-46b1-9241-d222b0a3d1ff · inbound

Scaling Decentralized Learning with FLock cites this paper.

Scaling Decentralized Learning with FLock Can multiple-choice questions really be useful in detecting the abilities of LLMs?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:19.156591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:19.156591Z digest=sha256:c099167038b73860d3787d3abf1bfb104b52cf4da9b94d246075eea02dcd2b9f

Observation 5629695e-ec7c-4fc7-9ba2-1f997fd2b78b · inbound

Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked? cites this paper.

Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked? Can multiple-choice questions really be useful in detecting the abilities of LLMs?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:57:03.027966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T03:52:56.761396Z digest=sha256:c4521e44d623a91f68bb93bac3f6c0de7ceffcd3d2e65402f57bd0864be285f4

Observation 9d5420fc-1229-469c-84a3-aff64c8ed38e · inbound

MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints cites this paper.

MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints Can multiple-choice questions really be useful in detecting the abilities of LLMs?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:25:38.225587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:25:38.225587Z digest=sha256:a88061f11e53cde2ed0a33ed9eccd3ed0f6e61d33bdf2e3b0e8a53d6b5d566fe

Observation 0243b014-6b7a-4643-b1eb-828a6c052df6 · inbound

Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation cites this paper.

Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation Can multiple-choice questions really be useful in detecting the abilities of LLMs?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T21:33:29.608043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T17:23:50.486306Z digest=sha256:267f18c61d07a005b536756d531460c37e5cd38d2e122532f06b9fdc365dda11

Observation e7a5022c-b761-49bd-bdca-510e3314b8af · inbound

BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models cites this paper.

BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models Can multiple-choice questions really be useful in detecting the abilities of LLMs?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.953803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:01:10.756844Z digest=sha256:47a7f033583249b28193b79518185cc937a28ac36bf60ec93a1bf8f24719adfa