Pith. sign in

Paper Citation Record · LEDGER

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model

As of 24 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.11737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11737 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:46.996738Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d963877b-8665-4014-b7de-6062df7c19d2 · outbound

This paper cites Attention is all you need,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Attention is all you need,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.436713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.913572Z digest=sha256:729830c9f85591e63228e74ec8cfb0dfbc11a8c3d9081abaf7630f5613a40ee1

Observation 0e4f3ca5-e108-42cf-a603-29d52d2ffb9f · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Bert: Pre-training of deep bidirectional transform- ers for language understanding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.422275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.918969Z digest=sha256:1867c2c2e4f96c4917c8f00067da010f72a5116d75296f3d3f0495e865b645bc

Observation 7baca484-44c1-40e1-a54f-c52b616ccc9a · outbound

This paper cites Language models are few-shot learners,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Language models are few-shot learners,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.408762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.923578Z digest=sha256:012a218bee681ebab5ed064f32887251a904c79a68daa7b0331e9329eaeceaa0

Observation 17c2c583-b58d-45ec-a7d1-a87ee0e4f8fc · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Learning transferable visual models from natural language supervision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.395250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.928138Z digest=sha256:1a7afc2ba7b53d161df650c15a9edecf9ff53ece633ed0b61f92f47fb81306f8

Observation e159b67f-ca98-4f7f-b435-80f1395be3e8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:46.933247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:46.933247Z digest=sha256:e999586feabb06ad08d51f1f666d85c646cfd96de76083d38cf3073e713f02ea

Observation e591881f-ff0f-463a-afa7-b29e370e63e9 · outbound

This paper cites Efficient Classification of Long Documents Using Transformers.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Efficient Classification of Long Documents Using Transformers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:07:47.200152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.937398Z digest=sha256:d0df73a17cd2eebf5c668d4d84266dbeda16ca638029d08c5727ffb96aa5b6a8

Observation 81c8f2a3-ae87-471c-aa2d-a8fbf9998054 · outbound

This paper cites Is space- time attention all you need for video understanding?,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Is space- time attention all you need for video understanding?,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.381666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.942484Z digest=sha256:b4ed3a75d2de64634645fb0321ebd308e6c1d0c22281535a1fc194a60828305f

Observation 545ee55e-bd53-461e-ad63-cb906e78c36d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Learning transferable visual models from natural language supervision,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.367842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.946532Z digest=sha256:19a0bb1d134d4cd2613b5843704b6c4416646700e9c59457c18fe5a31745dbfc

Observation 0ea8f46c-d1d8-4e08-8af9-671cbe95dda5 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Improved baselines with visual instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:46.950506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:46.950506Z digest=sha256:2986a1dfcea486c836363d030589311a066b5c8107b14574fb5311582ef2c427

Observation 93d29c51-5815-4f2e-b7d4-865c2fa0df3d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.344668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.954483Z digest=sha256:23742ef1693821b89520f0a248c1fed27d104fe0fa5213e93ddc524bdf5693d5

Observation 3df4dd59-1181-4bcc-8c3a-6437010fede4 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.330736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.958473Z digest=sha256:9e88d5d054dd6590b8f49088a82bea9ad36df6761096e4ecd643f575ed18a305

Observation b4d9e39c-50dc-49cb-a313-820649899c32 · outbound

This paper cites LLaV A-UHD: an lmm perceiving any aspect ratio and high- resolution images,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model LLaV A-UHD: an lmm perceiving any aspect ratio and high- resolution images,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.315929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.962379Z digest=sha256:b62edd8f5a0f3ee67d2a5b4e0e5ee81b7f3bee6292eada8d39e259bf8877d401

Observation e9a0e98c-c2ba-41df-b785-eaa1a263a7a5 · outbound

This paper cites Llava- mini: Efficient image and video large multimodal models with one vision token,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Llava- mini: Efficient image and video large multimodal models with one vision token,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.301516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.966642Z digest=sha256:e088baba9031af769f24a6ee7ca00e3a731a5ae556111d1bc30b3359e9d60011

Observation bc28a1a9-3bd2-4d21-9834-9a7ea77ade7c · outbound

This paper cites Logra-med: Long context multi-graph alignment for medical vision-language model,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Logra-med: Long context multi-graph alignment for medical vision-language model,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.287974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.970451Z digest=sha256:3caf0a2ebdb00a563b802d08ca1ae40699fdab96bc9f9478de203ae72bc2338d

Observation e50e9563-fb6f-49b6-9318-57d30a71d3df · outbound

This paper cites Llava needs more knowledge: Retrieval augmented natural language generation with knowledge graph for explaining thoracic pathologies,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Llava needs more knowledge: Retrieval augmented natural language generation with knowledge graph for explaining thoracic pathologies,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.273740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.974400Z digest=sha256:814835d35cbb6fc1484e548f33de98ae9016eb9f57d9a666c0ec3c5d0db48da4

Observation 4beb51ea-38c8-4b44-aad1-1f800646325d · outbound

This paper cites Dense connector for mllms,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Dense connector for mllms,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.253762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.978149Z digest=sha256:dae766b54a96d8f0053b5f9273a024f1cdd9eb6cf53e97309da1d409bc495386

Observation 3fadda6c-97ce-4c99-9b48-683d85f0e9c6 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Llava-prumerge: Adaptive token reduction for efficient large multimodal models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:46.981674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:46.981674Z digest=sha256:1606e1fbee7b37276f5827ec6c8e77a7043be22f8d037ad83d29ffbc93208341

Observation 84607495-f5b1-42ad-bdba-cb6fd4fc72f1 · outbound

This paper cites Visual instruction tuning,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Visual instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:46.985422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:46.985422Z digest=sha256:36fd7040fe8d3438a38ff27dbc1c8e7508bdb737d7b02625142ec7eadea84478

Observation b4acc257-de97-46f6-8e80-700df82cb743 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Sigmoid loss for language image pre-training,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:47.230836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:07:46.989033Z digest=sha256:3b33e5c5443d7bd215ce9630007d14156926e087924a392a03599e868aa16a26

Observation c0bc33f2-2cf0-4777-ba82-2fe02a408677 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:46.992637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:46.992637Z digest=sha256:bf042666e33cbb8ddc40ed5179df0d7705d8b888eb0329de3291fcf17fa5ca43

Observation c6146f65-b236-4f8d-9620-4573c52fdc62 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model Yi: Open Foundation Models by 01.AI

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:46.996738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:46.996738Z digest=sha256:589e35a50fd7c0c9f6012b59b69ad4a59f5b329cd730c23ce74c2277aed6547a

Pith citing papers

No inbound Pith citation observations are available.