Pith. sign in

Paper Citation Record · LEDGER

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.17608.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17608 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:11:06.855494Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c678bbb-6371-4031-a950-bbd2d6a19665 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.812570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.812570Z digest=sha256:26e610cd084737cf607f60f371bdc9c8efe7ddaab72dc3de190d066ce249fcae

Observation 46f92de4-27b7-46b0-9d78-5e8e9266305d · outbound

This paper cites Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.816752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.816752Z digest=sha256:1649019e05a63a7d819c49c45bec02154568e58189b798831b2697648eeb657e

Observation fc9aecbf-803f-476e-a325-b104ade7fe17 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs An image is worth 16x16 words: Transformers for image recognition at scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.820646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.820646Z digest=sha256:0aaf5edb526e69edae3beb0493f1761f38d16041f95af97107291702d279d963

Observation 07927673-49a2-4b5d-b93f-51541d26839f · outbound

This paper cites FeatUp: A Model-Agnostic Framework for Features at Any Resolution.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.824017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.824017Z digest=sha256:c01026bed40bfd53e5b644dd19f910a6463421b6df0461f6e0ea360075ca19df

Observation 95fb0188-136b-44d8-bb27-6aba2ac748a8 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.827940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.827940Z digest=sha256:b7bd191c0b0d9c9e1c5cb18e0d7f9e0911b4a4c06f75b932d54d0169733e68d8

Observation 3c670d6b-7212-4179-8c99-b3ca8edcca4f · outbound

This paper cites Improved baselines with visual instruction tuning, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Improved baselines with visual instruction tuning, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.983729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.831872Z digest=sha256:566d4641362854dc16e51369e1fd1cb79abf63fd5216cecbe73cea3bc06cfe5f

Observation 8650848e-a45d-4d0a-813d-b0a74636c121 · outbound

This paper cites Dinov2: Learning robust visual features with- out supervision, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Dinov2: Learning robust visual features with- out supervision, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.835300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.835300Z digest=sha256:ecc522e81e23d0cb077c35e590c73635c34b2adf961b7cda03b016bb83aee6d2

Observation 8114acb9-f031-43d0-a778-e39bd9f25df8 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.838648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.838648Z digest=sha256:c1127b4d9ac6c49dfc89e8036a02871d94eb44b7bf87fac3f33725081b09fc87

Observation 05c5e45a-10e5-45e2-98c5-33f914cbf4ce · outbound

This paper cites When do we not need larger vision models?, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs When do we not need larger vision models?, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.960347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.842102Z digest=sha256:e3b821478a6756baed310829753b021f275bbf118d0599edb7ce63bbc3dc32cb

Observation 44a41b24-1df3-43a7-ae9f-3416e7d382d8 · outbound

This paper cites Lift: A surprisingly simple lightweight feature transform for dense vit descriptors.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Lift: A surprisingly simple lightweight feature transform for dense vit descriptors

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.948957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.845562Z digest=sha256:f482226ce11ab2538e5fd6aebda9f8870b418cf85c13094a7f03544261152a17

Observation 66f3a4a0-2c27-4ec8-a72f-ffad3f0ec65d · outbound

This paper cites Dragonfly: Multi-resolution zoom-in encoding enhances vision-language models, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Dragonfly: Multi-resolution zoom-in encoding enhances vision-language models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.937353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.849093Z digest=sha256:dcb7b797595867d06d21b0b0c6bf8feaef718e2ed038ae00be096e35084b2900

Observation 5cf71c3b-621d-49c7-9850-7f71fb4a476a · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2024.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Cogvlm: Visual expert for pretrained language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:06.925819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:11:06.852153Z digest=sha256:5b72488eee3be05a58d8aad6b8d542461d6aae0dfe99eb0ba883a3ad762c1af3

Observation 6b0ab73c-bd1c-4a6a-ac5f-5a46d228ac3a · outbound

This paper cites Sigmoid loss for language image pre-training.

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs Sigmoid loss for language image pre-training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:06.855494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:06.855494Z digest=sha256:5eee42509198efb44db1a2360fdc4dd241743a17cb3c5918384a85a8661d4699

Pith citing papers

No inbound Pith citation observations are available.