Pith. sign in

Paper Citation Record · LEDGER

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

As of 22 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.15241.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15241 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:47:50.268649Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4ebc507-d62e-44e1-bf94-5f791056258f · outbound

This paper cites Medico 2025: Visual Question Answering for Gas- trointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medico 2025: Visual Question Answering for Gas- trointestinal Imaging,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:45.720994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:45.720994Z digest=sha256:9f66f9e60b875e93e051404f3a6aa84694fb67948d8e1f7e11c10c20cfe53a55

Observation d16a7216-3a40-44fc-9818-4085e9660cdb · outbound

This paper cites Kvasir- VQA-x1:A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Kvasir- VQA-x1:A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy,

Reference 2

Resolution
verified exact
doi, observed 2026-08-01T23:48:20.137832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-01T23:47:45.844349Z digest=sha256:352baa73cf5c3eb996f6b2a4d4eac21efdc1bd5bdf01b06e3d7ad0ca6c5c094c

Observation 5dec4e0d-4588-455c-bb74-84a48943852c · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LoRA: Low-rank adaptation of large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:45.957189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:45.957189Z digest=sha256:287d78c37fd4e850210f93c3863a931acee0209aa66344a8e765d614c8aac2d1

Observation a9dcffa5-4f19-4203-a93e-ec59e81f761d · outbound

This paper cites Qlora: efficient finetuning of quantized llms,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Qlora: efficient finetuning of quantized llms,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.069256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.069256Z digest=sha256:070cc8c24d93c2e69c63919613eefe42e34031d65cdeda84f027876bcdfb9e08

Observation 26945d1b-331b-498f-8161-b968d2a1e030 · outbound

This paper cites Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.167259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.167259Z digest=sha256:602356d29849ba6461549dcd58f42832ca9fa7417c4b845757bb09bab3c850b1

Observation 4f05b3ad-951d-4b70-a0a6-f6accc12d987 · outbound

This paper cites A Survey on Medical Large Language Models: Tech- nology, Application, Trustworthiness, and Future Directions,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA A Survey on Medical Large Language Models: Tech- nology, Application, Trustworthiness, and Future Directions,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.282042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.282042Z digest=sha256:437d72eb9874955e4c9bbda7b2abf90ea3ad2a61987bb8e0a9acf1853272c532

Observation da9bb747-cbe2-4829-9677-7b9f50b3bf53 · outbound

This paper cites VQA-Med: Overview of the medical visual ques- tion answering task at imageclef 2019,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA VQA-Med: Overview of the medical visual ques- tion answering task at imageclef 2019,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.428697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.428697Z digest=sha256:797656a7194addc2c1927422b6b6af9b32b5155f4372f5008b117b08e4bae450

Observation f643b1a8-7b63-49ad-9eb8-d0d57fb0216f · outbound

This paper cites Medical visual question answering at imageclef-vqa med,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medical visual question answering at imageclef-vqa med,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.633914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.633914Z digest=sha256:ece2f1410f7ae69536ad0fbe45a0d10ffd85593c8f70be459ba037e3816a030d

Observation c5a89455-2369-4dff-9f7d-0e0ca65a1abb · outbound

This paper cites Medical visual question answering: A survey,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medical visual question answering: A survey,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.734353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.734353Z digest=sha256:c68bf9f110f480d2683443c10b6207bba8a636947508a9c26f119e04ec8a2dac

Observation 9151dd43-1ebb-4203-a109-1c10c3687e90 · outbound

This paper cites Overview of ImageCLEFmedical 2025– Visual Question Answering and Synthetic Image Generation for Gastrointestinal Tract,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Overview of ImageCLEFmedical 2025– Visual Question Answering and Synthetic Image Generation for Gastrointestinal Tract,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.891922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.891922Z digest=sha256:fcae9c6d81e26626915a00adf537c6bf7cb7802a5142de49d5b8929b0a59156d

Observation 5339cb1b-4d73-48f7-81ab-d46c12c1717d · outbound

This paper cites Kvasir-VQA: A Text-Image Pair GI Tract Dataset,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Kvasir-VQA: A Text-Image Pair GI Tract Dataset,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.102764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.102764Z digest=sha256:bec7e89e9239d47132aa9d344bff6319be1e09b9a60ce6c36e6a4658d753c788

Observation 5cd66580-afca-4f33-b966-b0ed6097ac59 · outbound

This paper cites Exploring Vision-Language Models for Medical VQA on Gastrointestinal Images: A LoRA Fine-Tuning Study,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Exploring Vision-Language Models for Medical VQA on Gastrointestinal Images: A LoRA Fine-Tuning Study,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.257062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.257062Z digest=sha256:210a72f948dc4917d24403976eb7eb129337efd0c55301ac4b23920454ccec92

Observation 1d0c46b8-516a-4dcb-b199-cdff2161440a · outbound

This paper cites LoRA-Enhanced PaliGemma for Efficient Visual Question Answering in Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LoRA-Enhanced PaliGemma for Efficient Visual Question Answering in Gastrointestinal Imaging,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.412718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.412718Z digest=sha256:05d85cedfe9ddae7cd85c4fe1d4942560d802cfe2cd95878cedec798a07e2ff1

Observation 4d614b8a-4db9-48eb-91a6-cee613eb9f89 · outbound

This paper cites Multimodal Explanations: Justifying Decisions and Pointing to the Evidence,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Multimodal Explanations: Justifying Decisions and Pointing to the Evidence,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.576978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.576978Z digest=sha256:4a25b2e0a5e4360630ae8425a47e55b272dff3875c49e0fd0f855ef043ded160

Observation b1360515-aeb9-4a11-9ef2-c5df9b42e6f5 · outbound

This paper cites Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.742968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.742968Z digest=sha256:19ff438bfebdd3f0235949b5d652edc8f2a683e27af3e289829a0e11265eb298

Observation ea716100-f6f1-4e61-a763-5206f7367ae0 · outbound

This paper cites Towards Faithful Model Explanation in NLP: A Survey,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Towards Faithful Model Explanation in NLP: A Survey,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.856628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.856628Z digest=sha256:f631e97c77416cdd4ade83e79325cf8b605c3772a0c5573922ed5d8dd27b5675

Observation bd9a8709-8a16-4e3c-88e5-016dfb7b7596 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.009591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.009591Z digest=sha256:55a3ceca62f43ac1a1f944eccc7ede65353b114f1b65fbfc68a4323d7a11372d

Observation 68ccbab9-1bf5-4d6f-b779-5f07e51735f3 · outbound

This paper cites From Answers to Explanations: Self-Probing Efficiently Fine-Tuned Vision-Language Models for Medical VQA at Medico 2025,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA From Answers to Explanations: Self-Probing Efficiently Fine-Tuned Vision-Language Models for Medical VQA at Medico 2025,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.150605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.150605Z digest=sha256:4f361f17b0c08f5beaae9f4c2162abb5c32bb3fa85cae557302410e15f088529

Observation 681e5945-f427-4ea2-ac2d-4f2a37530fb9 · outbound

This paper cites Curriculum-Guided Fine-Tuning for Multimodal VQA in GI Endoscopy (Team Lama4Vision),.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Curriculum-Guided Fine-Tuning for Multimodal VQA in GI Endoscopy (Team Lama4Vision),

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.294746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.294746Z digest=sha256:1c87d6f61f2b4ce3c254e1bf5f153f553686037eba986b2730fab4841dd041b5

Observation c5ca3c1e-3da0-492c-a78d-8d709f5d43aa · outbound

This paper cites Qwen3 Technical Report.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.399084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.399084Z digest=sha256:b290dc619ef3774910f559313409717f97ca76ccb83ff1c7d6abd345f806723b

Observation 7236eafa-da93-4966-8f40-deb8c5dc4ebe · outbound

This paper cites Medico 2025: Visual Question Answering for Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medico 2025: Visual Question Answering for Gastrointestinal Imaging,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.480029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.480029Z digest=sha256:91d19bcc1927a592d72f7bfc486ee2f4b157e982f1b0fe0a4abfc8447f934557

Observation 493f3572-b12f-474c-9412-4f751380382f · outbound

This paper cites ROUGE: A Package for Automatic Evaluation of Sum- maries,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA ROUGE: A Package for Automatic Evaluation of Sum- maries,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.556305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.556305Z digest=sha256:ea74ce359c034dd600d673a95994359330b6e330c36202f02f3472e6bd06b64f

Observation a82743ab-0539-45c4-9570-acededdcdc94 · outbound

This paper cites METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.632097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.632097Z digest=sha256:1557155e43de4ca2b1d5bff67fb95a4aebfaad07ae028b49972c5d9f7553bd8d

Observation 110fcc60-9559-4100-bdb4-b083f2e2a41b · outbound

This paper cites chrF++: words helping character n-grams,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA chrF++: words helping character n-grams,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.711015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.711015Z digest=sha256:b60142f26af113b847f59031cd3453ca9d8998a218146def7f019f5226219f94

Observation 6abbe02b-a3a5-43f5-82a6-6458c5dd0dff · outbound

This paper cites BLEU: a Method for Automatic Evaluation of Machine Translation,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BLEU: a Method for Automatic Evaluation of Machine Translation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.751684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.751684Z digest=sha256:9032573bdf7fb967b34307eb928e2a8210a5f221c406e49572e3a11357cbd8ab

Observation 24df5ea4-9ff7-43fc-aa81-9d2341311d4b · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BERTScore: Evaluating Text Generation with BERT,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.819943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.819943Z digest=sha256:de78b8f4077c583cc5b9594143f2bd85ded9263c2c854a2d63ca4277c363ccce

Observation f988d8ec-5902-40a0-8837-923347d84bd7 · outbound

This paper cites Evaluate: A library for easily evaluating machine learn- ing models and datasets,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Evaluate: A library for easily evaluating machine learn- ing models and datasets,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.903305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.903305Z digest=sha256:257d357f5c4219c386cb71e478b1422f4b5ed65ffcae9d11691feb47a7de8511

Observation c0d5858b-8470-4f7e-bb7b-c91f24e72121 · outbound

This paper cites Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.998530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.998530Z digest=sha256:becedb071b3429a73582047704afca508be01641d91b31b8f9471b7782172c74

Observation 1bec4554-0cb9-4d1a-a21d-ea1453b008dd · outbound

This paper cites Medico 2025: Visual Question Answering (with Multimodal Explanations) for Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medico 2025: Visual Question Answering (with Multimodal Explanations) for Gastrointestinal Imaging,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.080126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.080126Z digest=sha256:83534ad52cf0eeded91c16ea570c602cf2a180425368e21361679b5178c74338

Observation 3a0891e2-cfa0-4662-abf6-c849e2cb7fd8 · outbound

This paper cites Enhancing Encoder-Decoder Architecture to Visual Question Answering Task for Gastrointestinal Images,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Enhancing Encoder-Decoder Architecture to Visual Question Answering Task for Gastrointestinal Images,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.157256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.157256Z digest=sha256:636ffa729148255184e2e1a43dca22ad3dee522ba11aa1d95794dd0d4dabd0ff

Observation 959a3a71-2560-4da1-9f26-fae5a11d93e7 · outbound

This paper cites BLIP-2-based Visual Question Answering with Multimodal Explanations for Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BLIP-2-based Visual Question Answering with Multimodal Explanations for Gastrointestinal Imaging,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.265736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.265736Z digest=sha256:1b0f3b157cceb66deda9db4133923bf566ff7896fb9ae6525a31ea1d7161220c

Observation dc4b55f5-baa4-464c-9c54-7fe22d247155 · outbound

This paper cites X-VQA for GI Diagnostics: Multimodal Visual Question Answering with Confidence-Aware Explanations,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA X-VQA for GI Diagnostics: Multimodal Visual Question Answering with Confidence-Aware Explanations,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.378765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.378765Z digest=sha256:e15008a423504e3a6bd13f4ecda3baf2eafee47addf2f393580b7bafd78aa0f1

Observation 01ce6139-da6c-4c35-be3f-bb35e21e2f5c · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.460925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.460925Z digest=sha256:b5688ef69d9cc071283f7d3eec35548e35c6669240db63d2b82c5ade0f7ced17

Observation 0c948514-a116-437c-8b8f-0b9341fe9c28 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA PaliGemma: A versatile 3B VLM for transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.538430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.538430Z digest=sha256:da40f420619859a71c4ed4a2a9c091f9f8166909d06c5902368cebdfc87d37cd

Observation 21735fab-982d-4862-af1c-9ea6dc7c9c46 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.649987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.649987Z digest=sha256:7e403477749acfdf118a2004fc4a2fb207b0947bf441d7830d9d2d2f9f5f1411

Observation 4c965876-d298-4e34-88c4-468c00889d1b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.734091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.734091Z digest=sha256:3291188078aceccf65113a0205d60d58506a16b1f2c0f9cce86eece94bc86752

Observation 964f9428-2927-4230-98fd-36e4e4208916 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.798966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.798966Z digest=sha256:1653841dade8631ae6f9625f31af2d187971a7da4a167010b832da53ecbf4c51

Observation e52d2d06-c041-4263-afe5-8600e111f4ea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.903613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.903613Z digest=sha256:15e1c91548285c17a0efa44c7d0842fc05f71eae45a3fddfcd565c552aace32b

Observation edd60e04-6a9a-4b1d-a3df-b01248df2e21 · outbound

This paper cites Image Segmentation Using Text and Image Prompts,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Image Segmentation Using Text and Image Prompts,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.994765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.994765Z digest=sha256:8fc9d5dc0fcbb72096728390e6d8efae127b3a1d84843cd8a7d4b13c058e08b1

Observation 8429f8ad-901e-4982-a43f-f08385b01322 · outbound

This paper cites Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.067973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.067973Z digest=sha256:5e7100ddc9b4e86cb7998eb3492e8f089ec7b5dca6c857cb72f4f5e87cac9f7c

Observation 120f8509-c27b-444f-b5f5-9c9a2c5543c4 · outbound

This paper cites Quantifying attention flow in transformers,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Quantifying attention flow in transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.149119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.149119Z digest=sha256:c50b8203e54ee142137d9a78b9e311486d581562825bdf4896a5d29ed642d8d7

Observation 20d0c7a3-558b-4114-a3c7-48f4f355eed8 · outbound

This paper cites Curriculum Learning: A Survey,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Curriculum Learning: A Survey,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.203408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.203408Z digest=sha256:d9fcc641e9f3987cb284b97bd259516c445c64481122c0356c9523b8d1ef80ed

Observation aaa4a646-4320-4ef4-a2f7-34572a60d652 · outbound

This paper cites LLM-FP4: 4-bit floating-point quantized transformers,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LLM-FP4: 4-bit floating-point quantized transformers,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.268649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.268649Z digest=sha256:ab50c83e20c0842b3820befd4967da52975056b70d68867f3f2cee3cd6f8908e

Pith citing papers

No inbound Pith citation observations are available.