Pith. sign in

Paper Citation Record · LEDGER

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

As of 12 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2608.04670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04670 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:21:45.118956Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa7468e2-a0f4-4914-bcae-ec7edfd4b7dd · outbound

This paper cites Lewkowycz, A.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Lewkowycz, A

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.849252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:44.956027Z digest=sha256:aeda21ad3aa78093f5fdc5c50f9445c5102bacc01fb317e32852c5c9ac516673

Observation d1fe282e-2c47-455a-9284-f52d84fc41a1 · outbound

This paper cites Chang, X.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Chang, X

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.835347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:44.960843Z digest=sha256:786381f76c75490b2b04a8cd38571818777f0491730aada059d9924594127e8a

Observation e731ee4e-e7ee-40ad-a049-d58207ca6228 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.822239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:44.965029Z digest=sha256:8a8462e91b68b0d12e80ff245116912f5d65e87068e40b40b11f9a9c1eb9384e

Observation 3d39048b-ef78-4893-b00a-e0898b2c173c · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.969886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.969886Z digest=sha256:14a1619deba3807b1dab2927d030c078718816ab8a710305aa9b4175fb76fc6a

Observation f2726dda-cc6f-41a1-acf9-d8277c2e3b49 · outbound

This paper cites The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.974685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.974685Z digest=sha256:aa6cbe8a15c275fde77b5f9f7ab569af39d2203c11626aa95945dd44c907d7f7

Observation 32961ddf-e62a-4474-9931-7139f5ae7dc2 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.979311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.979311Z digest=sha256:c34a5a39a6d398385c55e466ea1efcc1958f544efbdaa283b6d1d950641f9e57

Observation d1664dfb-7a8b-43ae-9d8f-25dcb1d79db5 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.808707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:44.984335Z digest=sha256:58a0195b74be1c4b54b39d1bdc6950f6b92b84de42281cd54e8ceaa70f38ccdb

Observation e4a11fcf-645a-4aaa-98c9-6ec210916ee1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.988297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.988297Z digest=sha256:f4722274aed0c51ac0b7055968fa1ca3856050231144fb9676fac8c12063cb61

Observation 409df814-db12-4d26-87d8-298b6d9b4a98 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.992611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.992611Z digest=sha256:7a2d5120116d3e0b54b855a283cae09b332b3d08bbdd01043a31ae5c77a4c9b6

Observation eb55a51e-25bd-48a6-a8a3-2e1a23e3fbd3 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.794336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:44.996759Z digest=sha256:6533d6d634e979c881917304fe1f6bb2290e5579c621715c7c82ca9a36aff388

Observation 44e5fd0c-1344-4e75-a719-9bb3ce192717 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.000585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.000585Z digest=sha256:9ea714ec5c946f3052b5155aafeb018241bb8de1444a097bdbc4d085f42f2b4f

Observation adf15a65-4f2c-4d03-a251-96731052994a · outbound

This paper cites On the Measure of Intelligence.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark On the Measure of Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.005089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.005089Z digest=sha256:b5461c072d202dd34390004dc0f81080d43d8640875cbf364ea83e153751000f

Observation 5841a2f2-84ac-4dcb-b08f-127a485e314a · outbound

This paper cites ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.009340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.009340Z digest=sha256:a8a8fcee2cabce04f6f18fca14c257450d6f758041b5b66f9006174d2a1218f5

Observation 5ca2b51a-5480-4db7-9ac3-05a694cf0493 · outbound

This paper cites Attanasio, P.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Attanasio, P

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.779829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.013470Z digest=sha256:12ba01fa8de3c980f6a32007305ca8cbb392da92328dd745b4eabbbd4fb06ed0

Observation 50242c54-3c37-4d0a-8311-5e6ccd0e23b5 · outbound

This paper cites Evalita-LLM: Benchmarking Large Language Models on Italian.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.017840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.017840Z digest=sha256:ca8f6fa40d74cdee543c03b5b6a9899487c5725cd890f7847cd6dca0f16cac56

Observation 3becf2ce-aee1-4e6f-97dd-7a5fdfa9ea04 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 16

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T19:21:45.166767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.022090Z digest=sha256:86d0aad444ec0f7364a40f4944ff8226a6981fc745e39f0182c49f563af4c86e

Observation 0b1b81e7-71ef-4cc2-8dc1-049d87926a49 · outbound

This paper cites Tedeschi, F.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Tedeschi, F

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.766098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.026171Z digest=sha256:6454c5614c58c30afac84425185c91bd392d6e6c8fd1f2c68fd5625b6d5cbc91

Observation 3fa49c86-3971-4e5c-bc3c-ff461ea6fd2c · outbound

This paper cites Khoshtab, D.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Khoshtab, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.751518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.029963Z digest=sha256:158affff10e0e37807c8e3bb6b34946e090009fddc29c2d5a1e421de75cdfbd0

Observation b2676a61-15f0-4206-a9e0-216573980de9 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.033785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.033785Z digest=sha256:40c2a132730363cab6c96a4d2a88a8b4425c666ca3810e92af9930c228092d2d

Observation 0aae267c-af0e-4576-99bd-5c6f3ded4d1f · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.736703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.038028Z digest=sha256:ecb4fc0b1a21e3577f1160d69731cfa49bc8456653d994454c63e59c31701646

Observation b6c82470-0b85-460f-bb34-6dd550e20627 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.721199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.041895Z digest=sha256:d2da4d4fdbcdd4a55e16b49b31ecb5eb6642ecff9faae7af2e456f86f6a3a76d

Observation 7dfba636-070c-4208-a330-61e7b51bf8c4 · outbound

This paper cites Donthi, M.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Donthi, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.706277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.046086Z digest=sha256:bd623a91175a2ddbb2c43e047eced62f8eb43b7a7a9cd6a80131ddaf01f0e144

Observation d7760f35-722f-4c16-9172-439e74990dff · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.691013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.050259Z digest=sha256:afb0f4b796c30b29f1ca3cbaa1107455194f7bae729c69643028431ec259e2b3

Observation 6daa9457-2d54-433c-bd73-1eb68edea5b5 · outbound

This paper cites Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:21:45.389273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.054075Z digest=sha256:1ef160de439c8bd0c51a2870e2a646b02f975720608b03317e9021dbf1075859

Observation 09ce6256-3f37-4da6-9d04-c4342578370a · outbound

This paper cites Caramagna, I 200 proverbi italiani più belli e famosi (con significato), 2025.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Caramagna, I 200 proverbi italiani più belli e famosi (con significato), 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.674684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.058791Z digest=sha256:a10e76e449db062aec7a6cd7fbccc1e08128de35155a923a48c934402a09b055

Observation 38eded67-ba3d-48fd-939b-b3ca4da3b4a5 · outbound

This paper cites URL: https://openrouter .ai/, accessed: 2025-06-15.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark URL: https://openrouter .ai/, accessed: 2025-06-15

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.658978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.062942Z digest=sha256:41a55e84b111510365de600d475e51459fec7a37a9032bb7bb5c1da44678f79f

Observation e487e10a-0c4b-457c-b815-695b8298c026 · outbound

This paper cites GPT-4o System Card.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark GPT-4o System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.066685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.066685Z digest=sha256:1e71998547a1c302fd981cbb7e17b4972d30b7cf859176d12b080f760a225701

Observation e269fac7-1e18-4536-b981-0c539fa1a946 · outbound

This paper cites DeepSeek-V3 Technical Report.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark DeepSeek-V3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.071212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.071212Z digest=sha256:b7b02c5eae4ec2ea2194b4bb0033c8d0b9032acf54fc7af5dfba1e742d06fbb2

Observation 5bfe8cad-5a95-4ff0-b287-f69a764c903f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.075878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.075878Z digest=sha256:57471db5866f734dc2a67f179e36e7750aed2541442cbe9946cd4a2a3aa07e7d

Observation ecb7b7ea-3c87-443d-beff-9246f08d6f9d · outbound

This paper cites Qwen3 Technical Report.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.080370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.080370Z digest=sha256:b79c3044a6bfe3322da6bcc453a7e170e2e6f0e8da4ded9f476847e8b3dfb2c3

Observation 5a462dfd-05cb-47d9-9ff0-71c05c9a5435 · outbound

This paper cites Gemma 3 Technical Report.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Gemma 3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.084774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.084774Z digest=sha256:a906bf1b9d4c68d4f45d4b7a9fd83a8ca5e19649f871b6eda593e85358a80bde

Observation af5a729b-cfb3-4a72-876e-fd28f7d1accb · outbound

This paper cites Orlando, L.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Orlando, L

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.644039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.089087Z digest=sha256:39db2c82fbfe32ad16bf3cec9f97310155813d7bad826cd49355e67999b95b44

Observation 60171fa0-2dd6-4f60-9313-81f3b1345f32 · outbound

This paper cites URL: https://huggingface .co/ mistralai/Mistral-Small-3.1-24B-Instruct-2503.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark URL: https://huggingface .co/ mistralai/Mistral-Small-3.1-24B-Instruct-2503

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.629134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.092781Z digest=sha256:01f0bfc73e10760b6a67a019a6eb8969d000b7493d499eb76f0eb8328fd07971

Observation d3a5a1f8-39b5-4461-92ce-32202342d426 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.096571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.096571Z digest=sha256:61fe7e63f5b7595b08e1f7e578cdc98e5b91f886d6dd8a02751d67ae7fd42e67

Observation 005ae45e-b53a-4848-bbfa-7a72dad54bc4 · outbound

This paper cites Etxaniz, G.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Etxaniz, G

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.100736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.100736Z digest=sha256:e0a75b48840b84479bed29606019d3731d8e99f9ec5f0fb968804562c35fe106

Observation 69c0ce65-5727-4e35-b690-39617bbbd0ea · outbound

This paper cites Ranaldi, G.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Ranaldi, G

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.614492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T19:21:45.105191Z digest=sha256:ef90911d97bfcfce323a3e7d06b343dadf60503f226686dc74d5d154183e0c98

Observation f768a60d-86bc-491b-a4d5-762fe23e970d · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.109647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.109647Z digest=sha256:2cec6ecf92b02d2e8b1761cc6521c60be0868cc097162100a92d44830fe57f24

Observation 0ebb7780-776f-4606-a588-55645408e891 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Reasoning Models Don't Always Say What They Think

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.114206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.114206Z digest=sha256:788fbc114ebb2c82f1e5085121f204ac0dc8eaa015704943610bfe7cca1059ca

Observation 0d10fc8b-a605-42a7-8afb-100a10f1f11c · outbound

This paper cites "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.118956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.118956Z digest=sha256:f74d089586719b4b950b32d663bde897e82109cb683283cbfab70fd7c7db081d

Pith citing papers

No inbound Pith citation observations are available.