Pith. sign in

Paper Citation Record · LEDGER

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations

As of 23 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2509.07135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07135 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:16:49.312244Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved32
  • parse uncertain8
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e65df800-a62f-4fd7-bcf7-10705eceaee8 · outbound

This paper cites Language Models are Few-Shot Learners.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Language Models are Few-Shot Learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.069536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.069536Z digest=sha256:0a05d19212ab2c69cade1b735c3754a57705350f9c55cf8816a49bfe11a42b56

Observation 0e6010d7-479f-4ed2-b91b-2236d9014be6 · outbound

This paper cites Kasneci, K.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Kasneci, K

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.054572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.073659Z digest=sha256:741980a04ec949c1279e438ed61dd957c5dce0c2818dc0773999aff44914c8ed

Observation a2e3a83b-9b98-486e-82e4-58a45ea40cc5 · outbound

This paper cites Baidoo-Anu, L.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Baidoo-Anu, L

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.044766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.077500Z digest=sha256:2ac810d4c14c9d4e33cd059dd057a9ce4ee2c0ce1a1224415426f26b36c2cdfc

Observation 326e9dee-dbb8-4493-b0df-d5d650d6b192 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.082371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.082371Z digest=sha256:7132cde10feb3117a6c7f7249fda4a5ecf81dc006c45132121f24f27862f25a3

Observation 204fa921-5871-42bf-9b11-5d7a813520c1 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.086389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.086389Z digest=sha256:e108a52e850202023894b64f5bc22aede34280b61ec3403c676b348acd289032

Observation 6667e3b0-de83-4fce-af9c-15949911fea6 · outbound

This paper cites Hendrycks, C.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Hendrycks, C

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.033443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.090279Z digest=sha256:2ebbca22b3dd63d950fe8b0d47d58753cb542b1f9ddde79dfc61a744dff03c62

Observation c68c3aef-8b7e-4a85-bb0e-13549259717c · outbound

This paper cites Attanasio, P.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Attanasio, P

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.022100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.098386Z digest=sha256:8eaad7b30f3bdde0464f63596485102cf1ae8d350e6c5fc2c9236ce04d63e6e1

Observation 419d4ade-231c-4b2d-a00d-bbd7f31016b2 · outbound

This paper cites Moroni, S.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Moroni, S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:50.011347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.101564Z digest=sha256:d00268e97a82bb50ca7069a2914f16176d5442cfc2341b073e705549f30bde3f

Observation 25e77af3-452b-4438-8f56-6fc147781370 · outbound

This paper cites Attanasio, P.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Attanasio, P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.993573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.104902Z digest=sha256:6103abd013823e354a49f13d7fe249d75b99ad3954f27c23a04197972418f2d4

Observation b2f3cdf6-4b39-4373-b71a-8c7e00d947f8 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.108419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.108419Z digest=sha256:85b6dffafd18d49b63d28ba2288158e30309c7bdeb4dc57447f7890688571898

Observation 9001183f-ac9c-4fb2-8b4a-25112fd11c0a · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.112548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.112548Z digest=sha256:3816db6e61e25a6fc7e01fb958cea91096fac81b02b41454ac22b01d61904ba0

Observation 91c3b74b-c09a-46f5-a217-c44e7966071b · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.116436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.116436Z digest=sha256:771aa782bad0a1386fa8370e42eb707935ca251262e9b2ef71c1ba83b6dd5a8e

Observation 7f42af82-5b87-4985-ade6-f85c4f43ce9b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.982195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.120056Z digest=sha256:1fbe9d0ccea1e59332d567b50ad5dccc6f3eea74ab5c65c185148f8f97bddf58

Observation ace3c07e-32a7-4e8d-b114-26ce30898445 · outbound

This paper cites Nentidis, K.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Nentidis, K

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.972111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.123598Z digest=sha256:cdf28466542c3001727e719a4f021e00a8215ccd9fc6a5e88b46c2cd7c4ce554

Observation 23bc6808-d427-46ac-8787-18ece3a7d9bb · outbound

This paper cites Rinaldi, J.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Rinaldi, J

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.962017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.126443Z digest=sha256:a4a76aa96cc0541cba7e0874c28eae046dc901a17b363c35d0310259d3566ed1

Observation e89ac205-2c16-4b6f-913b-93f4e1898fb8 · outbound

This paper cites Casola, T.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Casola, T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.952148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.129599Z digest=sha256:3813f99850694300f6c8c03273e173ad691ebfebf922f9a0f32f3de4a8dd2ee8

Observation a6800399-d681-4e64-b362-6d9055d06838 · outbound

This paper cites Altuna, G.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Altuna, G

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.942246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.133420Z digest=sha256:85ac8d50a8f1fa7d95bd0ac809353f195c6d2e207ed433d620d86d53944c1660

Observation af934513-ee0b-4917-ba29-535199513247 · outbound

This paper cites Puccetti, M.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Puccetti, M

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.931512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.136745Z digest=sha256:247560064f350bf47b710386bbf382928f520a05e23efe7853dca2113b70589f

Observation 8a896b5a-2d01-4cb1-b795-1bfb3a134c95 · outbound

This paper cites Calibrate Before Use: Improving Few-Shot Performance of Language Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Calibrate Before Use: Improving Few-Shot Performance of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.139687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.139687Z digest=sha256:e387e7611f83e7ad0178588bf49060209645b34dc7af9a0d765577b92606b051

Observation 69b2e6f1-05b6-48eb-aa13-7ce947d1ea26 · outbound

This paper cites Wei, et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Wei, et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.921769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.143562Z digest=sha256:f5a8939fc1938a37f5bcafef13e584ac4061cb5e9d885613651789c0d61d6d0b

Observation d4e9987e-4e7b-403d-985c-f52aff34382e · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.152173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.152173Z digest=sha256:72982cedcdacf1ced9f575d6728c7db6c55b9c493af3cda87c2b958104b9d561

Observation 6fbd38dc-624b-45bf-af5b-2ec106e795f7 · outbound

This paper cites Yang, et al., Qwen2.5 technical report,.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Yang, et al., Qwen2.5 technical report,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.912340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.155721Z digest=sha256:cc43b8841e839ef088d0e8380aa9ed2aef1143c574bc0d63fce119880ae2d098

Observation 54b2dff8-c75a-4ea7-a8b9-d40a4a847185 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.164060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.164060Z digest=sha256:a55b4c36f50c0a0e608b088a32e13ae28566e2a8de34980238f18d7b2327ad4f

Observation a406451c-01f9-401a-a742-9e474aca0834 · outbound

This paper cites Grattafiori, et al., The Llama 3 Herd of Models,.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Grattafiori, et al., The Llama 3 Herd of Models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.903468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.167514Z digest=sha256:5358f9feef2b7ee5bbc3352317280b98ce3440c8df616ed746ca2c8ce16bae47

Observation 6020059f-a93d-45d9-a91f-736fc3a6a39d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.174532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.174532Z digest=sha256:d08763be444bee67bbe6eb6262d73b1f8a53a920f4ad17a9e4dc8cb58519d566

Observation e3bfd345-919a-43ef-a7bf-b711dfee23f8 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.894241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.178213Z digest=sha256:30c782e11c571f9c2fbd229edaa303299b351fcff1dcd7825696043908de83f0

Observation 70bba81e-5eb4-47ca-b173-59df023a23f9 · outbound

This paper cites Groeneveld, et al., OLMo: Accelerating the sci- ence of language models, in: L.-W.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Groeneveld, et al., OLMo: Accelerating the sci- ence of language models, in: L.-W

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.184981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.184981Z digest=sha256:ce79e8ce089f2e8456fdd9cbd9d3cc1e729b2dcc7ab64d2968f6b87ee5dc664b

Observation edd361d1-4dd9-4873-b94c-0b59175953ab · outbound

This paper cites Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.188509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.188509Z digest=sha256:414051777658f8207deb1ec8b2e056e4bfc4d1248a3d1d4b132e772b4571e430

Observation 1b23dd12-0a49-42dc-802d-2af027aeef81 · outbound

This paper cites Orlando, L.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Orlando, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.882860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.192297Z digest=sha256:584878059cd00d15e02b86a7f32f507bac88e6924b9dc1491996a7946f62170a

Observation 0b6a3a29-9f68-4616-a4b5-988c756e2d2a · outbound

This paper cites Tutti i bambini amano il gelato.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Tutti i bambini amano il gelato

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.871689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.195766Z digest=sha256:0ae650bc179e82600470a94054eb875a8a0e053ebd51aef444981723cb582079

Observation ff3760e2-5407-4bdb-9036-30622c17c823 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.181295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.181295Z digest=sha256:a4e7e19e5e76d5563b15eeb053b182dff1caa6d602acc00323c9066e4e58c0e1

Observation 6e278186-accb-4ab0-aded-4346cd79eeba · outbound

This paper cites All children love ice cream.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations All children love ice cream

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.861537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.199048Z digest=sha256:6dfde8e969055804beda8825ee23d5662eadea4d13a4736ce79eaef9f893196c

Observation e807ec9b-b50e-44e1-bb32-b151c542ae4e · outbound

This paper cites Logic Example Domanda:Se e solo se Giulia a luglio non va in vacanza in montagna, va poi in vacanza al mare ad agosto.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Logic Example Domanda:Se e solo se Giulia a luglio non va in vacanza in montagna, va poi in vacanza al mare ad agosto

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.850957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.203529Z digest=sha256:81f2bf5d1b368428b5021eb86785132e637ff18ca646ca08b3a2a8318b35366f

Observation c801bf5b-9fc0-44a7-98a6-1562e4648bb8 · outbound

This paper cites Carolina ha acquistato molte borse, dunque ha speso molti soldi.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Carolina ha acquistato molte borse, dunque ha speso molti soldi

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.841115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.206629Z digest=sha256:ec2a956d8be0f02280f46e8578de119734e68dab681b4bcab476faaaf9aefc1f

Observation d916dd1e-97df-414f-9e65-696f04cd43ef · outbound

This paper cites Stasera non ha piovuto, dunque è andata in motorino.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Stasera non ha piovuto, dunque è andata in motorino

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.832486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.209773Z digest=sha256:bbc0409b8f6cf0f2e7dff806526d1cbaa2955e11bcdbf1f70461f128c80d407b

Observation e193d783-787d-499f-88d0-07bbdd8f1c22 · outbound

This paper cites Ha già man- giato albicocche a pranzo, dunque a cena non mangia le fragole.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Ha già man- giato albicocche a pranzo, dunque a cena non mangia le fragole

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.823889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.212754Z digest=sha256:093b0325df35beec24d599be046571b43eb1d814d787f257f3e17041938d981c

Observation 02ddee4a-1adf-4f59-bd04-7e046c8d81ac · outbound

This paper cites Clara ha superato gli esami, dunque ha studi- ato molto.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Clara ha superato gli esami, dunque ha studi- ato molto

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.815477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.216582Z digest=sha256:0a1cb3cf2ba4d6622c383eed4d344b514c0ce368ed5f5754bea877eb9f3e8c74

Observation 89998033-3302-4271-87da-d6ce58e0dd1b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.805946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.219515Z digest=sha256:af9c2dc50dacef6ee5dccd55f9cde64b56c1388ee790804ec19013f40529fb18

Observation 788372c4-1cc8-4420-8ddf-7e92d3bf499e · outbound

This paper cites Carolina bought many bags, therefore she spent a lot of money.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Carolina bought many bags, therefore she spent a lot of money

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.796767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.222516Z digest=sha256:bd038ba3afb96d7efcd58b820651b871a265e4f03536acde861f740faad66717

Observation 79086ea3-cc52-4a2c-9711-0ce53e29b796 · outbound

This paper cites Tonight it did not rain, there- fore she went on her scooter.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Tonight it did not rain, there- fore she went on her scooter

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.787663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.225393Z digest=sha256:bc1a33ed06fa491d39088a52fadc91283560e27bf65ffa13880c2149d92d64ea

Observation c37903f5-366b-4eb0-a422-0eeab25fa248 · outbound

This paper cites She already ate apricots for lunch, therefore she does not eat strawberries for dinner.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations She already ate apricots for lunch, therefore she does not eat strawberries for dinner

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.778524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.228744Z digest=sha256:cb14f0b7b4550da7b851453a1b4e291e2bb0ff7994d9cb4c7e9e4d512177b829

Observation 01fc384f-0306-4105-b761-39bc7b6b7202 · outbound

This paper cites Clara passed the exams, therefore she studied hard.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Clara passed the exams, therefore she studied hard

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.768441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.232201Z digest=sha256:790499a8c6ccdcce58bf6310065906c30c75df82dc3bfe9af24a0213f2cc87b0

Observation b955e2fb-3393-4b54-9464-e3e0decb00da · outbound

This paper cites Riccardo does not play ten- nis, therefore he did not play football (Correct Answer: 3) A.3.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Riccardo does not play ten- nis, therefore he did not play football (Correct Answer: 3) A.3

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.760046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.235146Z digest=sha256:82a1004d00c5b275fc80fd2e49a3de84f20fdb36b9b2794871a621e47dd61181

Observation dc5abd4a-3dff-498b-8d32-e9f70ca1bce1 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.751081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.238359Z digest=sha256:ce973ecfb1df13a02709eb0f0ece7388f6c40d2db6edbd1f89515105be0818c5

Observation d5b6dd9d-a3bc-4621-86ae-0004792b549b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 49

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.741993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.241354Z digest=sha256:47c7950f7ffac1951561871c6d9dd9a8ca61592b0fef048ebb8ac7522646b26f

Observation 06ee3f3e-d069-4cb8-9fb5-17392069896f · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 50

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.732772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.244491Z digest=sha256:0bdd4ecd06c33f1cf7c71e4af3f2b0a9a8e675cd81b76c4244bc0c442588e202

Observation e3ea7f55-0f8b-4746-b42e-23c6ee37c60e · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.724675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.247704Z digest=sha256:635f2af5db87fdfd7a77a26954961cab522a9815de1e99a7c98ec48b970aac51

Observation b12fa6f6-4463-4830-a9e2-31ee8344aa56 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.715691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.251112Z digest=sha256:9ad0604a80789755436d877784aeb50eb5b2cb52a5c6c4848ba42e1fb1c7eff4

Observation 4d6b9822-df21-4664-aa4d-345e56eca906 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 53

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.706488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.254326Z digest=sha256:efd1cb2a1b6fd9312a8083b1d153b14b559a534e8d3889b9469349e70e799e50

Observation 52b337fb-7527-463e-9865-aeb06616114f · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.698528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.257527Z digest=sha256:3632b06c5027af6514ae37fd09c12f737488b6d9e56aeff685597f28d26bc57d

Observation 0e9fc822-787c-4c40-a52d-759ec30c0132 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 55

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.690350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.260719Z digest=sha256:5399292b31a038f65dc5691f071b0d8760f848dcb4fddbc22da32d51a1332faa

Observation 1e68f2cf-7f91-4d6b-a60a-f724970bb7b3 · outbound

This paper cites Chemistry Example Domanda:A quante moli corrispondono 5 mL (d=1,8 g ·cm−3) di un composto avente una massa molare di 450 g·mol−1? Possibili risposte:.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Chemistry Example Domanda:A quante moli corrispondono 5 mL (d=1,8 g ·cm−3) di un composto avente una massa molare di 450 g·mol−1? Possibili risposte:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.681710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.263970Z digest=sha256:dda1a66402cb8fd9886e927f33cb690bad89d4687e91b1daf392b9af3c7652c5

Observation dfa12a9b-b86c-4518-af17-7203273c4e4b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.672857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.267210Z digest=sha256:4571e37ff70ca7f11b8da41980737d5996c4b86cd818704b85012a0912bceecb

Observation f7dd84eb-8475-456b-97b8-c15416b26e80 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.663843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.271160Z digest=sha256:4e3642c524f99b09089a6a21852e7d62dd2203e3fe52b7e92ec3dcab8d1ca915

Observation 4465b9ea-ee69-4b40-8315-7483de6bb1ed · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.655370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.274285Z digest=sha256:df0118e1304312fa28d0c91dd96b4fe7a7f25b9220769dacd0bc5094b71149e6

Observation 4aa63aa9-1018-4f60-8830-53869bab582b · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.645513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.278030Z digest=sha256:a0ac22b9b49617d0708e12d68f2187405fc2eb73186f7b20a28cd0ab4e587f34

Observation d527f1af-176b-460f-a817-8d8c7490ab44 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.636019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.281242Z digest=sha256:11d8167d578a22a822d1dd27f1c3f5300f5ffcdb99f55a8640fa01cd73fe76c4

Observation 121ac5b2-ba64-47a5-84bc-0df35dfabf48 · outbound

This paper cites Mathematics Example Domanda:Dati tre segmenti AA’, BB’ e CC’ tali che: AA’ = 2 cm, BB’ = 1,5 * AA’, CC’ = 2,0 * BB’.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Mathematics Example Domanda:Dati tre segmenti AA’, BB’ e CC’ tali che: AA’ = 2 cm, BB’ = 1,5 * AA’, CC’ = 2,0 * BB’

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.626694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.284391Z digest=sha256:77e75e5172865dd188691ac745b6cbac7c7b823fd922d4ff4c53dbfe1615e6f0

Observation beb19267-748c-4cab-99ec-03486af1a569 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.616832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.287982Z digest=sha256:f000483e3362360af46187da9bf08370e2f647c03f80155791676dc701762ac3

Observation 98692b07-e034-4301-b79a-aca303a168cf · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 64

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.607039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.290935Z digest=sha256:63f076629670c26776826ee05a9e295a42c5b3f85cd275b0ea4cebc2c28bfb7a

Observation 4d423c89-8307-4191-b0c9-6ecccd559b36 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 65

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.597102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.293642Z digest=sha256:3445fa4ad3fa0c73d7ae585904b47b4318c369d98b8376ddf4170813de3701ce

Observation 0e97a9d2-5a33-4ca5-8f7f-f38a82323d18 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 66

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.586934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.296835Z digest=sha256:8d3ffa16eb39d93142eef16b11e5bedfbfe2f0f25a8cd19e7c3ae350aac86bca

Observation a02d19f7-f5bd-43fb-8159-821f7c13b79a · outbound

This paper cites Which triangle is possible to construct with these sides? Possible answers:.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Which triangle is possible to construct with these sides? Possible answers:

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.576247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.299876Z digest=sha256:c55531fa7fb9b8f452626715d5267b6abf7d48191d04119828406a56d2986ed0

Observation 2acda169-967c-40c5-b247-e08c418936e8 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.563265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.302833Z digest=sha256:473e664f6b1078623884d770d17e2d26e73f14be221ade111edeaf2715ff0782

Observation 61a9a47e-e0f4-481b-8d34-bdc0bb827362 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 69

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T16:16:49.552433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.306274Z digest=sha256:49c87fb27d635043a5b29bf3430b266ffbde70c12c5d8115f6267fe7790d08c1

Observation bd18b8fd-9f9f-4327-beaa-2e163ca0c557 · outbound

This paper cites an unresolved cited work.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:16:49.542972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.309383Z digest=sha256:5630f59de5d00867e924ca8887f813522b02b3197df76c74a5919469e158aba4

Observation b8559297-e408-42e4-be3c-c5cf45afb821 · outbound

This paper cites Per-Subject Model Performance Table 4 Per-subject accuracy (%) on MedBench-IT for Standard (Std.) and Reasoning (Reas.) prompts.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Per-Subject Model Performance Table 4 Per-subject accuracy (%) on MedBench-IT for Standard (Std.) and Reasoning (Reas.) prompts

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:16:49.531635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:16:49.312244Z digest=sha256:ca995fd688faea52a667b4598a8f9836856e026f535281ec7dc9f7c8fa471ec6

Observation 612f47fa-c3d0-476a-80cc-d7b96b97530e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.093791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.093791Z digest=sha256:c56013715c71c1909a67e11df2ef7a574c6010ac7f39d0eceef17f7ba8e1fe74

Observation 36afa985-006f-4174-a06b-ab01ab22e8ac · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.147542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.147542Z digest=sha256:e93dffd4c10c7d351001db1a34420c8e05e4bf8e3b321ac047d2375d25ba183d

Observation b483a8fa-bc0c-4e15-bb33-0fc1cc6dc9c4 · outbound

This paper cites The Llama 3 Herd of Models.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.170935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.170935Z digest=sha256:6fc1fe7515f05aa98776ebd7a8c87cc61422be2b12e2f0ea4a5e23cc2d4da12a

Observation eb6dd17c-2406-4603-9855-6aa9ff542f6c · outbound

This paper cites Qwen2.5 Technical Report.

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations Qwen2.5 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T16:16:49.160095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:16:49.160095Z digest=sha256:fffbb61d1b4c3215637e8a462ac024dc86c97d85c8e47475bb3c1b36411caeda

Pith citing papers

No inbound Pith citation observations are available.