Pith. sign in

Paper Citation Record · LEDGER

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations

As of 22 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2509.07135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07135 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:16:49.312244Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved32
  • parse uncertain8
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e65df800-a62f-4fd7-bcf7-10705eceaee8 · outbound

This paper cites Language Models are Few-Shot Learners.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Language Models are Few-Shot Learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.069536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.069536Z digest=sha256:0a05d19212ab2c69cade1b735c3754a57705350f9c55cf8816a49bfe11a42b56

Observation 0e6010d7-479f-4ed2-b91b-2236d9014be6 · outbound

This paper cites Kasneci, K.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Kasneci, K

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.054572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.073659Z digest=sha256:2177a7597d21ec3c353d9a1b6cda7c12fec1f889b5c4984c15a683ab532ab1db

Observation a2e3a83b-9b98-486e-82e4-58a45ea40cc5 · outbound

This paper cites Baidoo-Anu, L.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Baidoo-Anu, L

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.044766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.077500Z digest=sha256:14127762ca31152b5b6cd0a2b2d6bd8e878adc6f2ac6a4c6800e8b4e98604e1e

Observation 326e9dee-dbb8-4493-b0df-d5d650d6b192 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.082371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.082371Z digest=sha256:7132cde10feb3117a6c7f7249fda4a5ecf81dc006c45132121f24f27862f25a3

Observation 204fa921-5871-42bf-9b11-5d7a813520c1 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.086389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.086389Z digest=sha256:e108a52e850202023894b64f5bc22aede34280b61ec3403c676b348acd289032

Observation 6667e3b0-de83-4fce-af9c-15949911fea6 · outbound

This paper cites Hendrycks, C.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Hendrycks, C

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.033443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.090279Z digest=sha256:944b585af07274d0213d38ddd1b6f53c37acb8057518f16b9bed0b275e645277

Observation c68c3aef-8b7e-4a85-bb0e-13549259717c · outbound

This paper cites Attanasio, P.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Attanasio, P

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.022100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.098386Z digest=sha256:9fd42d6761e8a06c79b7e7f1e8a65b0cbe354bdf4dc45ae350160b6a8941eeb4

Observation 419d4ade-231c-4b2d-a00d-bbd7f31016b2 · outbound

This paper cites Moroni, S.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Moroni, S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.011347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.101564Z digest=sha256:06675e7777b5799a0a5d492108673f8d0d88e71f1ec11b0bc00500a2fc6ebef5

Observation 25e77af3-452b-4438-8f56-6fc147781370 · outbound

This paper cites Attanasio, P.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Attanasio, P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.993573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.104902Z digest=sha256:9889bad08fd959b90ed1c5edab98b8ddecc7ad89b5099ab2e5287bbe3f96f2db

Observation b2f3cdf6-4b39-4373-b71a-8c7e00d947f8 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.108419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.108419Z digest=sha256:85b6dffafd18d49b63d28ba2288158e30309c7bdeb4dc57447f7890688571898

Observation 9001183f-ac9c-4fb2-8b4a-25112fd11c0a · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.112548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.112548Z digest=sha256:3816db6e61e25a6fc7e01fb958cea91096fac81b02b41454ac22b01d61904ba0

Observation 91c3b74b-c09a-46f5-a217-c44e7966071b · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.116436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.116436Z digest=sha256:771aa782bad0a1386fa8370e42eb707935ca251262e9b2ef71c1ba83b6dd5a8e

Observation 7f42af82-5b87-4985-ade6-f85c4f43ce9b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.982195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.120056Z digest=sha256:b4d8966e638575a7e6e20f61fe12eb03b83497ec97747c68d710676b692d7d3a

Observation ace3c07e-32a7-4e8d-b114-26ce30898445 · outbound

This paper cites Nentidis, K.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Nentidis, K

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.972111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.123598Z digest=sha256:920fb1beda87621c0254fdd433d9b99e34b4bb00746c0a8d0465ca574bcec617

Observation 23bc6808-d427-46ac-8787-18ece3a7d9bb · outbound

This paper cites Rinaldi, J.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Rinaldi, J

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.962017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.126443Z digest=sha256:b5ad2f16d6e50397914da6cd6afc9d4fbf324ab9c07cb5e1e94700e7268c5f38

Observation e89ac205-2c16-4b6f-913b-93f4e1898fb8 · outbound

This paper cites Casola, T.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Casola, T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.952148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.129599Z digest=sha256:e67e9e555e150303d6c45ab5cd5a9aa8130aa40bf6649dea04985639abbd02de

Observation a6800399-d681-4e64-b362-6d9055d06838 · outbound

This paper cites Altuna, G.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Altuna, G

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.942246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.133420Z digest=sha256:5c490f670b5229c04cce90833a2d2e7814c50f22512de1c361dc741f5ec4295b

Observation af934513-ee0b-4917-ba29-535199513247 · outbound

This paper cites Puccetti, M.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Puccetti, M

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.931512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.136745Z digest=sha256:2d1c2c98328fd1e1ad56c730ab1617378ede20c27db3ec130e01157cfcc57451

Observation 8a896b5a-2d01-4cb1-b795-1bfb3a134c95 · outbound

This paper cites Calibrate Before Use: Improving Few-Shot Performance of Language Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Calibrate Before Use: Improving Few-Shot Performance of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.139687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.139687Z digest=sha256:e387e7611f83e7ad0178588bf49060209645b34dc7af9a0d765577b92606b051

Observation 69b2e6f1-05b6-48eb-aa13-7ce947d1ea26 · outbound

This paper cites Wei, et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Wei, et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.921769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.143562Z digest=sha256:e47cbedf93494d2c6bdc03324a22ad2ad849e92b400b3c389109a4bd20739815

Observation d4e9987e-4e7b-403d-985c-f52aff34382e · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.152173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.152173Z digest=sha256:72982cedcdacf1ced9f575d6728c7db6c55b9c493af3cda87c2b958104b9d561

Observation 6fbd38dc-624b-45bf-af5b-2ec106e795f7 · outbound

This paper cites Yang, et al., Qwen2.5 technical report,.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Yang, et al., Qwen2.5 technical report,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.912340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.155721Z digest=sha256:7c43aaf47666c2b5a45867e89194c9849bfb832c554ebded44aac34d62375cf6

Observation 54b2dff8-c75a-4ea7-a8b9-d40a4a847185 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.164060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.164060Z digest=sha256:a55b4c36f50c0a0e608b088a32e13ae28566e2a8de34980238f18d7b2327ad4f

Observation a406451c-01f9-401a-a742-9e474aca0834 · outbound

This paper cites Grattafiori, et al., The Llama 3 Herd of Models,.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Grattafiori, et al., The Llama 3 Herd of Models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.903468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.167514Z digest=sha256:9ff33331d0e9e3f659e15b40eabdc1d4cbf18f36a9959562068f50666b20be29

Observation 6020059f-a93d-45d9-a91f-736fc3a6a39d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.174532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.174532Z digest=sha256:d08763be444bee67bbe6eb6262d73b1f8a53a920f4ad17a9e4dc8cb58519d566

Observation e3bfd345-919a-43ef-a7bf-b711dfee23f8 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.894241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.178213Z digest=sha256:34f07793cc53093db8cd118e1d01f7f40583fe9f576da6cf0e588652bfed148a

Observation 70bba81e-5eb4-47ca-b173-59df023a23f9 · outbound

This paper cites Groeneveld, et al., OLMo: Accelerating the sci- ence of language models, in: L.-W.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Groeneveld, et al., OLMo: Accelerating the sci- ence of language models, in: L.-W

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.184981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.184981Z digest=sha256:ce79e8ce089f2e8456fdd9cbd9d3cc1e729b2dcc7ab64d2968f6b87ee5dc664b

Observation edd361d1-4dd9-4873-b94c-0b59175953ab · outbound

This paper cites Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.188509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.188509Z digest=sha256:414051777658f8207deb1ec8b2e056e4bfc4d1248a3d1d4b132e772b4571e430

Observation 1b23dd12-0a49-42dc-802d-2af027aeef81 · outbound

This paper cites Orlando, L.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Orlando, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.882860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.192297Z digest=sha256:31403d9c458255642dcb9bffed66f09cf4760d110389984d42c57a79a3787794

Observation 0b6a3a29-9f68-4616-a4b5-988c756e2d2a · outbound

This paper cites Tutti i bambini amano il gelato.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Tutti i bambini amano il gelato

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.871689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.195766Z digest=sha256:fae8d0f48788aa7f76cf6dc245ad777d8351c65b0b46411c6f1443c55bbd316d

Observation ff3760e2-5407-4bdb-9036-30622c17c823 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.181295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.181295Z digest=sha256:a4e7e19e5e76d5563b15eeb053b182dff1caa6d602acc00323c9066e4e58c0e1

Observation 6e278186-accb-4ab0-aded-4346cd79eeba · outbound

This paper cites All children love ice cream.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations All children love ice cream

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.861537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.199048Z digest=sha256:a3d3ab012bd8963537452ed179dcf1c740073ae045430b8f7466e93c9bae6b94

Observation e807ec9b-b50e-44e1-bb32-b151c542ae4e · outbound

This paper cites Logic Example Domanda:Se e solo se Giulia a luglio non va in vacanza in montagna, va poi in vacanza al mare ad agosto.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Logic Example Domanda:Se e solo se Giulia a luglio non va in vacanza in montagna, va poi in vacanza al mare ad agosto

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.850957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.203529Z digest=sha256:9013da4380405ab2a05282149ad402f70eb5669f033dfd14f909fa8be8e15b28

Observation c801bf5b-9fc0-44a7-98a6-1562e4648bb8 · outbound

This paper cites Carolina ha acquistato molte borse, dunque ha speso molti soldi.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Carolina ha acquistato molte borse, dunque ha speso molti soldi

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.841115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.206629Z digest=sha256:ec5596d22a3efe95f547d543a8f5df7eb6c29bb136aced5f80c935accdc70259

Observation d916dd1e-97df-414f-9e65-696f04cd43ef · outbound

This paper cites Stasera non ha piovuto, dunque è andata in motorino.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Stasera non ha piovuto, dunque è andata in motorino

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.832486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.209773Z digest=sha256:b78ccc03be81d849456713fbc2d2c3194c72cc06f2d92b33be19fb149a215dc4

Observation e193d783-787d-499f-88d0-07bbdd8f1c22 · outbound

This paper cites Ha già man- giato albicocche a pranzo, dunque a cena non mangia le fragole.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Ha già man- giato albicocche a pranzo, dunque a cena non mangia le fragole

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.823889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.212754Z digest=sha256:e2be2c9fa3bc2b6e2911a41820dcf6243157accb6de95c4dc3c2bf3efc99ff77

Observation 02ddee4a-1adf-4f59-bd04-7e046c8d81ac · outbound

This paper cites Clara ha superato gli esami, dunque ha studi- ato molto.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Clara ha superato gli esami, dunque ha studi- ato molto

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.815477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.216582Z digest=sha256:913c7fce75db9b71a2d1dcb2ed1ab4d002317ec77f1f48a8fef115c89204b096

Observation 89998033-3302-4271-87da-d6ce58e0dd1b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.805946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.219515Z digest=sha256:51251ac6b9ee45296bfcf06c44aa9792101ed774ab4508b080c52f64f1954035

Observation 788372c4-1cc8-4420-8ddf-7e92d3bf499e · outbound

This paper cites Carolina bought many bags, therefore she spent a lot of money.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Carolina bought many bags, therefore she spent a lot of money

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.796767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.222516Z digest=sha256:a5c71c2644f317fcca189f674bdd8a1a58d41d54f7b073032b32951b28630238

Observation 79086ea3-cc52-4a2c-9711-0ce53e29b796 · outbound

This paper cites Tonight it did not rain, there- fore she went on her scooter.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Tonight it did not rain, there- fore she went on her scooter

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.787663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.225393Z digest=sha256:e4de5ae4cfc110ba11456cf08d2f4321905b4ec99ad78fec24f8c310655fac1f

Observation c37903f5-366b-4eb0-a422-0eeab25fa248 · outbound

This paper cites She already ate apricots for lunch, therefore she does not eat strawberries for dinner.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations She already ate apricots for lunch, therefore she does not eat strawberries for dinner

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.778524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.228744Z digest=sha256:cb480b6c217df6d24d4ac0f2a3ed8e7ef48e2bb86def9602cec59aabf3ea78f5

Observation 01fc384f-0306-4105-b761-39bc7b6b7202 · outbound

This paper cites Clara passed the exams, therefore she studied hard.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Clara passed the exams, therefore she studied hard

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.768441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.232201Z digest=sha256:76e4ebbd307e0a1f20978dfdcd13aa5df25c71b49ba60cb76f30cdd8d8e87cf3

Observation b955e2fb-3393-4b54-9464-e3e0decb00da · outbound

This paper cites Riccardo does not play ten- nis, therefore he did not play football (Correct Answer: 3) A.3.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Riccardo does not play ten- nis, therefore he did not play football (Correct Answer: 3) A.3

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.760046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.235146Z digest=sha256:5b631f6326795f686dcb68a86462c293572759d033973e89d4355d707e8448cc

Observation dc5abd4a-3dff-498b-8d32-e9f70ca1bce1 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.751081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.238359Z digest=sha256:aa70dcb63824d4ce5649753fc60fb3fa2def782372f1d2b2cd98b4f9c370ae16

Observation d5b6dd9d-a3bc-4621-86ae-0004792b549b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 49

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.741993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.241354Z digest=sha256:437c43d488d6a5dfc559eeae6c272adc32839cebf590cf4856b805b61f43c66f

Observation 06ee3f3e-d069-4cb8-9fb5-17392069896f · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 50

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.732772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.244491Z digest=sha256:117c8ecea5d46d521afebb54fc23284074c1931a403b1ba3439c14b2c36138e3

Observation e3ea7f55-0f8b-4746-b42e-23c6ee37c60e · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.724675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.247704Z digest=sha256:7a05a48afe8dfc3648ab2789421170bae3b8d3190176a1cdca2601e0550dbfdb

Observation b12fa6f6-4463-4830-a9e2-31ee8344aa56 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.715691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.251112Z digest=sha256:f4cdee41f6f6f09e754df5febfb38014adbe4a6f9ff9e0a5059232789c3b5871

Observation 4d6b9822-df21-4664-aa4d-345e56eca906 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 53

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.706488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.254326Z digest=sha256:b0c8f73769621b27d951c407b414c13d1c41fc50b05bbebb48e23752c15dddae

Observation 52b337fb-7527-463e-9865-aeb06616114f · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.698528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.257527Z digest=sha256:e9ae21c2d3edede79e02f9447d61693039509b379d4435a48c1115d25f209120

Observation 0e9fc822-787c-4c40-a52d-759ec30c0132 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 55

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.690350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.260719Z digest=sha256:242c4152503e9c282bcf53fe9251d0e644ede86ba02485c2c532ffe6e6f048a2

Observation 1e68f2cf-7f91-4d6b-a60a-f724970bb7b3 · outbound

This paper cites Chemistry Example Domanda:A quante moli corrispondono 5 mL (d=1,8 g ·cm−3) di un composto avente una massa molare di 450 g·mol−1? Possibili risposte:.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Chemistry Example Domanda:A quante moli corrispondono 5 mL (d=1,8 g ·cm−3) di un composto avente una massa molare di 450 g·mol−1? Possibili risposte:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.681710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.263970Z digest=sha256:2ba099aac3c2340e3ee6e1fbc35b97bc6c3428911908751db95be1daf7248e5c

Observation dfa12a9b-b86c-4518-af17-7203273c4e4b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.672857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.267210Z digest=sha256:e874dfe03f7238b2c2f7de16b3b54fe082a22e75791d316bb5db94e25b0d1c7d

Observation f7dd84eb-8475-456b-97b8-c15416b26e80 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.663843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.271160Z digest=sha256:09ae271f5e7b19d1cb297f172ecb1a641418b2c1e503473384d72f4f78c28ec0

Observation 4465b9ea-ee69-4b40-8315-7483de6bb1ed · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.655370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.274285Z digest=sha256:81b886646e1bfc4ba5b5e55c9bd328cee400a3223642d101415aa39a2c5ef7bc

Observation 4aa63aa9-1018-4f60-8830-53869bab582b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.645513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.278030Z digest=sha256:ce6f9668de8c5a1a7a0c7db33ad67d9c7461414ed4db5ba3108c34482ac7f3d2

Observation d527f1af-176b-460f-a817-8d8c7490ab44 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.636019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.281242Z digest=sha256:fd198fcf9fbdde88dfcae60d89dca7738ad95dba72ebec2871c029b3686265d8

Observation 121ac5b2-ba64-47a5-84bc-0df35dfabf48 · outbound

This paper cites Mathematics Example Domanda:Dati tre segmenti AA’, BB’ e CC’ tali che: AA’ = 2 cm, BB’ = 1,5 * AA’, CC’ = 2,0 * BB’.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Mathematics Example Domanda:Dati tre segmenti AA’, BB’ e CC’ tali che: AA’ = 2 cm, BB’ = 1,5 * AA’, CC’ = 2,0 * BB’

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.626694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.284391Z digest=sha256:1d7c8979890af660822263df4854cc81dba1280fda33041d7a42531e2a5bd002

Observation beb19267-748c-4cab-99ec-03486af1a569 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.616832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.287982Z digest=sha256:aef599c4320176defd9e1b7e555743ca2c51a4c60d5b9d49764458ac062a9eae

Observation 98692b07-e034-4301-b79a-aca303a168cf · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 64

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.607039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.290935Z digest=sha256:78fd649c0bcfbbe222c59101b2a19dd5ab91cb6f3ba01ab09db2211f7b09fdf6

Observation 4d423c89-8307-4191-b0c9-6ecccd559b36 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 65

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.597102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.293642Z digest=sha256:2c84f69e4b15b6b786e109aad61025fd0257ef09f5357ea41470a5f2133c862b

Observation 0e97a9d2-5a33-4ca5-8f7f-f38a82323d18 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 66

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.586934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.296835Z digest=sha256:5749793410bad4227618c7960da73abb243441b31673c94169c13ae1d2d9760f

Observation a02d19f7-f5bd-43fb-8159-821f7c13b79a · outbound

This paper cites Which triangle is possible to construct with these sides? Possible answers:.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Which triangle is possible to construct with these sides? Possible answers:

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.576247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.299876Z digest=sha256:cebc1a0389e57bf59680ecb51dd7eaa4903bd58d1782f933505612bd9f99bf3a

Observation 2acda169-967c-40c5-b247-e08c418936e8 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.563265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.302833Z digest=sha256:5795f77b5b8dceb1801f9db60240679c8e3211df75cfc73e7d9c2c00a6a2c882

Observation 61a9a47e-e0f4-481b-8d34-bdc0bb827362 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 69

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.552433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.306274Z digest=sha256:9b3589192b0612adeff2ad14713b8b6ae21cd54d465b45f513f61669f3e3a698

Observation bd18b8fd-9f9f-4327-beaa-2e163ca0c557 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.542972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.309383Z digest=sha256:2f359a4dd679fc284e6fcf76f4841a42af61b5586958e3af8e68f6ad577e3477

Observation b8559297-e408-42e4-be3c-c5cf45afb821 · outbound

This paper cites Per-Subject Model Performance Table 4 Per-subject accuracy (%) on MedBench-IT for Standard (Std.) and Reasoning (Reas.) prompts.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Per-Subject Model Performance Table 4 Per-subject accuracy (%) on MedBench-IT for Standard (Std.) and Reasoning (Reas.) prompts

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.531635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:16:49.312244Z digest=sha256:a69b864bce091572d63f5e6af1b44beb6c6d7116b4b485aef7d4eccbd039fdf5

Observation 612f47fa-c3d0-476a-80cc-d7b96b97530e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.093791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.093791Z digest=sha256:c56013715c71c1909a67e11df2ef7a574c6010ac7f39d0eceef17f7ba8e1fe74

Observation 36afa985-006f-4174-a06b-ab01ab22e8ac · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.147542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.147542Z digest=sha256:e93dffd4c10c7d351001db1a34420c8e05e4bf8e3b321ac047d2375d25ba183d

Observation b483a8fa-bc0c-4e15-bb33-0fc1cc6dc9c4 · outbound

This paper cites The Llama 3 Herd of Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.170935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.170935Z digest=sha256:6fc1fe7515f05aa98776ebd7a8c87cc61422be2b12e2f0ea4a5e23cc2d4da12a

Observation eb6dd17c-2406-4603-9855-6aa9ff542f6c · outbound

This paper cites Qwen2.5 Technical Report.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Qwen2.5 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.160095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.160095Z digest=sha256:fffbb61d1b4c3215637e8a462ac024dc86c97d85c8e47475bb3c1b36411caeda

Pith citing papers

No inbound Pith citation observations are available.