Pith. sign in

Paper Citation Record · LEDGER

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

As of 21 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.17608.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17608 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:11:06.855494Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c678bbb-6371-4031-a950-bbd2d6a19665 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.812570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.812570Z digest=sha256:0c064475405459a96266cb7425c3841630097fd58d9180823c06dba5fa77a42b

Observation 46f92de4-27b7-46b0-9d78-5e8e9266305d · outbound

This paper cites Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.816752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.816752Z digest=sha256:547e9bd041f2920e96ecb56bd613759d25f02cb273c01fe658a8296a19ed5de2

Observation fc9aecbf-803f-476e-a325-b104ade7fe17 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs An image is worth 16x16 words: Transformers for image recognition at scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.820646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.820646Z digest=sha256:692bf1e1c9d406feaa81c0811384836f73b2267e9fd94e4a2d4e176cc3422c51

Observation 07927673-49a2-4b5d-b93f-51541d26839f · outbound

This paper cites FeatUp: A Model-Agnostic Framework for Features at Any Resolution.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.824017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.824017Z digest=sha256:dc9c378ee8c7b2d973921ec3e76a218a5d49245bd0426aee7cdaca644c214146

Observation 95fb0188-136b-44d8-bb27-6aba2ac748a8 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.827940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.827940Z digest=sha256:4792f0ec41c09ace19a5e6985f35969c1f1a8972a12de07f0a54fb7d23666bb7

Observation 3c670d6b-7212-4179-8c99-b3ca8edcca4f · outbound

This paper cites Improved baselines with visual instruction tuning, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Improved baselines with visual instruction tuning, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.983729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.831872Z digest=sha256:6abdb22d662aa44d5e2cfa6aea00714a5d12d18c4f2109bfcbd44e8baaa6ec14

Observation 8650848e-a45d-4d0a-813d-b0a74636c121 · outbound

This paper cites Dinov2: Learning robust visual features with- out supervision, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Dinov2: Learning robust visual features with- out supervision, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.835300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.835300Z digest=sha256:5da6e6240ca1b97261ace3c59e545b36b4024523e3bc5100e5f7f33b996c3f43

Observation 8114acb9-f031-43d0-a778-e39bd9f25df8 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.838648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.838648Z digest=sha256:efeceb8227f09523a57bd407af056071343c476ae2bef4c3130b6d8e64de1e79

Observation 05c5e45a-10e5-45e2-98c5-33f914cbf4ce · outbound

This paper cites When do we not need larger vision models?, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs When do we not need larger vision models?, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.960347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.842102Z digest=sha256:6f1a233df9bcb1f0909f2ca244d6493a0cfc0ae08ffe70ada6f6422e5709f91d

Observation 44a41b24-1df3-43a7-ae9f-3416e7d382d8 · outbound

This paper cites Lift: A surprisingly simple lightweight feature transform for dense vit descriptors.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Lift: A surprisingly simple lightweight feature transform for dense vit descriptors

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.948957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.845562Z digest=sha256:1efbb7959efd51bbfb96cca1ae2d318ef1583e3e65f788462df5cd9a550c93c6

Observation 66f3a4a0-2c27-4ec8-a72f-ffad3f0ec65d · outbound

This paper cites Dragonfly: Multi-resolution zoom-in encoding enhances vision-language models, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Dragonfly: Multi-resolution zoom-in encoding enhances vision-language models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.937353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.849093Z digest=sha256:9f3da58ed898cd89651952fbee4f61cf2d3968a2883dfd57aa754bc77d9361f8

Observation 5cf71c3b-621d-49c7-9850-7f71fb4a476a · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Cogvlm: Visual expert for pretrained language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.925819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.852153Z digest=sha256:41feae5375ad6a3e3ad974367ad461f6d403ffd2a24fd57bf638a2166175e9ae

Observation 6b0ab73c-bd1c-4a6a-ac5f-5a46d228ac3a · outbound

This paper cites Sigmoid loss for language image pre-training.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Sigmoid loss for language image pre-training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.855494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.855494Z digest=sha256:586cea7b95aa9f75726b367f0b6db5e19ea15b9e070fddf8125ffaf4f2cb0226

Pith citing papers

No inbound Pith citation observations are available.