Pith. sign in

Paper Citation Record · LEDGER

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

As of 18 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.15241.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15241 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:47:50.268649Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4ebc507-d62e-44e1-bf94-5f791056258f · outbound

This paper cites Medico 2025: Visual Question Answering for Gas- trointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medico 2025: Visual Question Answering for Gas- trointestinal Imaging,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:45.720994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:45.720994Z digest=sha256:41ed0468ff3363d2c99ece33625192d47164ac6fd9a3bf9167fc6996a09f7051

Observation d16a7216-3a40-44fc-9818-4085e9660cdb · outbound

This paper cites Kvasir- VQA-x1:A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Kvasir- VQA-x1:A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy,

Reference 2

Resolution
verified exact
doi, observed 2026-08-01T23:48:20.137832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-01T23:47:45.844349Z digest=sha256:5cb73d1f1714b5ef60ad84b9549891fd01e2ed60aceaa6a6977632164d800ef7

Observation 5dec4e0d-4588-455c-bb74-84a48943852c · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LoRA: Low-rank adaptation of large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:45.957189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:45.957189Z digest=sha256:cad36225528e066d6226d4654010a06c32306bb10acde70cfecc98e9c7041d87

Observation a9dcffa5-4f19-4203-a93e-ec59e81f761d · outbound

This paper cites Qlora: efficient finetuning of quantized llms,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Qlora: efficient finetuning of quantized llms,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.069256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.069256Z digest=sha256:d7bf822cb25f4e5735e8410b321be4f111ec80217495633f6281be54eb1730bd

Observation 26945d1b-331b-498f-8161-b968d2a1e030 · outbound

This paper cites Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.167259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.167259Z digest=sha256:08f163a6b3375761bb7e41733c34a32bb9d05be17ae91c1909db5ad8609e6f52

Observation 4f05b3ad-951d-4b70-a0a6-f6accc12d987 · outbound

This paper cites A Survey on Medical Large Language Models: Tech- nology, Application, Trustworthiness, and Future Directions,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA A Survey on Medical Large Language Models: Tech- nology, Application, Trustworthiness, and Future Directions,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.282042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.282042Z digest=sha256:d0362278912f85ccb22b693109f974794972561525de20252c9b040e231fa761

Observation da9bb747-cbe2-4829-9677-7b9f50b3bf53 · outbound

This paper cites VQA-Med: Overview of the medical visual ques- tion answering task at imageclef 2019,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA VQA-Med: Overview of the medical visual ques- tion answering task at imageclef 2019,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.428697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.428697Z digest=sha256:31ce5768dfe608bea7ace59bb42379ceb4e46c67eb00bc01d5342253f0c4ddd4

Observation f643b1a8-7b63-49ad-9eb8-d0d57fb0216f · outbound

This paper cites Medical visual question answering at imageclef-vqa med,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medical visual question answering at imageclef-vqa med,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.633914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.633914Z digest=sha256:617432cf1f166266d60f69f873b15cdc52b342cc037ce47f74c90752ff76ce49

Observation c5a89455-2369-4dff-9f7d-0e0ca65a1abb · outbound

This paper cites Medical visual question answering: A survey,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medical visual question answering: A survey,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.734353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.734353Z digest=sha256:464261acd4dc107762431f5e86407e1596349539a65c0c1354c2c963201bee2a

Observation 9151dd43-1ebb-4203-a109-1c10c3687e90 · outbound

This paper cites Overview of ImageCLEFmedical 2025– Visual Question Answering and Synthetic Image Generation for Gastrointestinal Tract,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Overview of ImageCLEFmedical 2025– Visual Question Answering and Synthetic Image Generation for Gastrointestinal Tract,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:46.891922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:46.891922Z digest=sha256:32c5ed1eb54cd0e7c4ce25d8ad83df53a4a5877a0bfc17ae50b0f86bd89a5227

Observation 5339cb1b-4d73-48f7-81ab-d46c12c1717d · outbound

This paper cites Kvasir-VQA: A Text-Image Pair GI Tract Dataset,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Kvasir-VQA: A Text-Image Pair GI Tract Dataset,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.102764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.102764Z digest=sha256:f5caaaf17a592e8e25f082ee31803777f38f1d0275d4747295ae18f35d81315c

Observation 5cd66580-afca-4f33-b966-b0ed6097ac59 · outbound

This paper cites Exploring Vision-Language Models for Medical VQA on Gastrointestinal Images: A LoRA Fine-Tuning Study,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Exploring Vision-Language Models for Medical VQA on Gastrointestinal Images: A LoRA Fine-Tuning Study,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.257062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.257062Z digest=sha256:199922b47216c2c88489e5fa1f1d43ee59fff1e156ae0a7662048a1773d7dfa3

Observation 1d0c46b8-516a-4dcb-b199-cdff2161440a · outbound

This paper cites LoRA-Enhanced PaliGemma for Efficient Visual Question Answering in Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LoRA-Enhanced PaliGemma for Efficient Visual Question Answering in Gastrointestinal Imaging,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.412718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.412718Z digest=sha256:5b217e5fc127f56727dd7dc7b7e828c10ca65ca61cf2514b8d7560b235fe3cb2

Observation 4d614b8a-4db9-48eb-91a6-cee613eb9f89 · outbound

This paper cites Multimodal Explanations: Justifying Decisions and Pointing to the Evidence,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Multimodal Explanations: Justifying Decisions and Pointing to the Evidence,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.576978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.576978Z digest=sha256:7cd06e17e2d93bc38a4aeba40d22ce479c81db431be92c9e81471f031d57818d

Observation b1360515-aeb9-4a11-9ef2-c5df9b42e6f5 · outbound

This paper cites Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.742968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.742968Z digest=sha256:c2f197ce9e518a308b535b552ef23e6787b65c91c954bfcb7c7af33a5ce5ec10

Observation ea716100-f6f1-4e61-a763-5206f7367ae0 · outbound

This paper cites Towards Faithful Model Explanation in NLP: A Survey,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Towards Faithful Model Explanation in NLP: A Survey,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:47.856628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:47.856628Z digest=sha256:0d12acdfcb4dfbf4e22cce4ff5410ddcd9e8fe160f1114d3c55badc3ab919585

Observation bd9a8709-8a16-4e3c-88e5-016dfb7b7596 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.009591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.009591Z digest=sha256:3e6c8c6a588f1adcec008a63065dbf0cc29c2040d3abed72fd6bb7d8ad985c17

Observation 68ccbab9-1bf5-4d6f-b779-5f07e51735f3 · outbound

This paper cites From Answers to Explanations: Self-Probing Efficiently Fine-Tuned Vision-Language Models for Medical VQA at Medico 2025,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA From Answers to Explanations: Self-Probing Efficiently Fine-Tuned Vision-Language Models for Medical VQA at Medico 2025,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.150605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.150605Z digest=sha256:a4ded3fe31199d6422f30c5b819c3d2d2d22816373a07e84131bec98e055b463

Observation 681e5945-f427-4ea2-ac2d-4f2a37530fb9 · outbound

This paper cites Curriculum-Guided Fine-Tuning for Multimodal VQA in GI Endoscopy (Team Lama4Vision),.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Curriculum-Guided Fine-Tuning for Multimodal VQA in GI Endoscopy (Team Lama4Vision),

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.294746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.294746Z digest=sha256:6a8ba8174f0baf1db6e2bbdce69aadcea16f004d074d0fe12a8697fd0353c30e

Observation c5ca3c1e-3da0-492c-a78d-8d709f5d43aa · outbound

This paper cites Qwen3 Technical Report.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.399084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.399084Z digest=sha256:4eeaeef3cee05b32d2256d93d00447a774c850ff741b82661c1717cc6fe75793

Observation 7236eafa-da93-4966-8f40-deb8c5dc4ebe · outbound

This paper cites Medico 2025: Visual Question Answering for Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medico 2025: Visual Question Answering for Gastrointestinal Imaging,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.480029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.480029Z digest=sha256:00afc226dd6c1d2893a379a49be9b53e4100e5fbbdbf47227050cadfca1f91ab

Observation 493f3572-b12f-474c-9412-4f751380382f · outbound

This paper cites ROUGE: A Package for Automatic Evaluation of Sum- maries,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA ROUGE: A Package for Automatic Evaluation of Sum- maries,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.556305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.556305Z digest=sha256:738d2aa865c09a8efccb3e971cbb3ffe7e298c8c1a891ec0dcf97350735b5f6d

Observation a82743ab-0539-45c4-9570-acededdcdc94 · outbound

This paper cites METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.632097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.632097Z digest=sha256:3a2801ea31955f00be134310b01cc878201c87da760699b2aba4d25fa8662443

Observation 110fcc60-9559-4100-bdb4-b083f2e2a41b · outbound

This paper cites chrF++: words helping character n-grams,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA chrF++: words helping character n-grams,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.711015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.711015Z digest=sha256:a76335d74dc58c437605f68ab5ac698d57ee25184da2f2733a5ba4b960771d92

Observation 6abbe02b-a3a5-43f5-82a6-6458c5dd0dff · outbound

This paper cites BLEU: a Method for Automatic Evaluation of Machine Translation,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BLEU: a Method for Automatic Evaluation of Machine Translation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.751684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.751684Z digest=sha256:cf5d032b1b9e593e1f0e3a2443592b5dba4aff5e0a7f86a6a117dbfb555910d4

Observation 24df5ea4-9ff7-43fc-aa81-9d2341311d4b · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BERTScore: Evaluating Text Generation with BERT,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.819943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.819943Z digest=sha256:23e675475cf75870740de4ff8286d7a04c3ee7842bac87e97b956ff2c36f3087

Observation f988d8ec-5902-40a0-8837-923347d84bd7 · outbound

This paper cites Evaluate: A library for easily evaluating machine learn- ing models and datasets,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Evaluate: A library for easily evaluating machine learn- ing models and datasets,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.903305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.903305Z digest=sha256:42ed9812e6459587ff52d12b2c0a98b18258d2a4052b383cdc20b3e20518c744

Observation c0d5858b-8470-4f7e-bb7b-c91f24e72121 · outbound

This paper cites Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:48.998530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:48.998530Z digest=sha256:cdce66fb18706bd98da92955d21327d897bd09f4fb5ea112f09ece8def73fd17

Observation 1bec4554-0cb9-4d1a-a21d-ea1453b008dd · outbound

This paper cites Medico 2025: Visual Question Answering (with Multimodal Explanations) for Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Medico 2025: Visual Question Answering (with Multimodal Explanations) for Gastrointestinal Imaging,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.080126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.080126Z digest=sha256:9a81227d56547efb3262d7d43e6fdbfdf7e296f6443a8f8d34ea49185c7c30f5

Observation 3a0891e2-cfa0-4662-abf6-c849e2cb7fd8 · outbound

This paper cites Enhancing Encoder-Decoder Architecture to Visual Question Answering Task for Gastrointestinal Images,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Enhancing Encoder-Decoder Architecture to Visual Question Answering Task for Gastrointestinal Images,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.157256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.157256Z digest=sha256:206ef218fce39e89aba143aba34195941092a75c5ac51f09d2a9426c17fa3865

Observation 959a3a71-2560-4da1-9f26-fae5a11d93e7 · outbound

This paper cites BLIP-2-based Visual Question Answering with Multimodal Explanations for Gastrointestinal Imaging,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BLIP-2-based Visual Question Answering with Multimodal Explanations for Gastrointestinal Imaging,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.265736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.265736Z digest=sha256:822ddad056609ebde54b0f79b90ebc17f6b94cdd624141f8c6b3dc8ba45e548b

Observation dc4b55f5-baa4-464c-9c54-7fe22d247155 · outbound

This paper cites X-VQA for GI Diagnostics: Multimodal Visual Question Answering with Confidence-Aware Explanations,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA X-VQA for GI Diagnostics: Multimodal Visual Question Answering with Confidence-Aware Explanations,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.378765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.378765Z digest=sha256:6fec6d86d7d47964ba687e370ef78fa460bc47498cd30171aa43f5065713f11e

Observation 01ce6139-da6c-4c35-be3f-bb35e21e2f5c · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.460925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.460925Z digest=sha256:ecbdb087653e455a8a4dc7a3c2440a32ef000e8e8b2f352ec2c569f6e2295d14

Observation 0c948514-a116-437c-8b8f-0b9341fe9c28 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA PaliGemma: A versatile 3B VLM for transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.538430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.538430Z digest=sha256:e559a00450784d73801eed0c8dcdf301205f25824274f8daa8eff10b68b6915e

Observation 21735fab-982d-4862-af1c-9ea6dc7c9c46 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.649987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.649987Z digest=sha256:dca1864aed0f7d1444b323c5962f2a773334894682873d90d86a9d51150698ef

Observation 4c965876-d298-4e34-88c4-468c00889d1b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.734091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.734091Z digest=sha256:cb8aebde56317b893ebddaee290d89be05032aa7edd1144050a48b7f6be79239

Observation 964f9428-2927-4230-98fd-36e4e4208916 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.798966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.798966Z digest=sha256:de2d6fde92e117f325b852a5f1e8d11f88d17d32953d66a4b5278bce474b5d1e

Observation e52d2d06-c041-4263-afe5-8600e111f4ea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.903613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.903613Z digest=sha256:ffca41fe5901b67e5784a70ee2e0c7f8a2663eac69ee9aba6b63d56b528113c5

Observation edd60e04-6a9a-4b1d-a3df-b01248df2e21 · outbound

This paper cites Image Segmentation Using Text and Image Prompts,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Image Segmentation Using Text and Image Prompts,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:49.994765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:49.994765Z digest=sha256:a55e5a5ad9de8eaee705e326b670dd4cafdfb97e06a957847175b32dced780e8

Observation 8429f8ad-901e-4982-a43f-f08385b01322 · outbound

This paper cites Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.067973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.067973Z digest=sha256:b71fb7df4fdef22c542ac858ce3344cedec0c2cf08bd7601ce1a0dab5cae59c2

Observation 120f8509-c27b-444f-b5f5-9c9a2c5543c4 · outbound

This paper cites Quantifying attention flow in transformers,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Quantifying attention flow in transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.149119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.149119Z digest=sha256:5d9e77e42c9b4ade1b3978d1902a5dfd1ab944d9e29f151a50b4337e7daa09b0

Observation 20d0c7a3-558b-4114-a3c7-48f4f355eed8 · outbound

This paper cites Curriculum Learning: A Survey,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Curriculum Learning: A Survey,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.203408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.203408Z digest=sha256:af8020784c6aec5e3bd70dd156f6af91fed729d58dec340a24cae42c629f5edc

Observation aaa4a646-4320-4ef4-a2f7-34572a60d652 · outbound

This paper cites LLM-FP4: 4-bit floating-point quantized transformers,.

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA LLM-FP4: 4-bit floating-point quantized transformers,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T23:47:50.268649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:47:50.268649Z digest=sha256:78f3f0dbd0f9a194a4c99b1b1c49d1a4921e5c0ee4a7f165bcf5b4d70ac116c6

Pith citing papers

No inbound Pith citation observations are available.