Pith. sign in

Paper Citation Record · LEDGER

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning

As of 14 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2411.12787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12787 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:37:54.891321Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7f1bee5-c5be-4905-b8ca-a43edb0dd259 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.675226Z digest=sha256:904a13b0b93b2b39456d4e09eecb4447ed9b0ab40bc698d99f0db767d18eda69

Observation a26616b5-c94c-419a-a46a-4fef259519f5 · outbound

This paper cites Layer Normalization.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.681289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.681289Z digest=sha256:d546617f33d60d364913014cd5c2d5c31da905f881991aa7820438a95d83a09a

Observation 2b986d2a-0096-41de-86b8-ffc475aaa73e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.687801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.687801Z digest=sha256:2ef8a88591245df003e81d1ab81b6be08fd4cf1f96e9f944d3b35f3fe460d7f3

Observation 5afec7ec-c18a-4f90-ae8a-94c76f4ef2bc · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.695363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.695363Z digest=sha256:0d32d532a5d5f411cd4616126145e910b051f0c248e5c5ec8bea878937d142c7

Observation 66132ade-bf2d-4a37-b9ba-22066cb3ef22 · outbound

This paper cites LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.701624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.701624Z digest=sha256:78814d002211a08dc842dd76ef72e4a65bcd7571e500368865c4125fdc2e586d

Observation 2f5db9c4-0d15-41d3-9a10-32912c279eef · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.709703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.709703Z digest=sha256:8f3b33b9d467754c8f446f306c8ae7d0506f616b267a0dc87c389e519c504ace

Observation 08921ef2-c25c-45a5-97d2-dde53a6983fd · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.715260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.715260Z digest=sha256:13fce737efa37c5caee429aeb42e1bd6dc942305d8417c6d75c286aa48e26023

Observation 0da724d6-0a78-47c3-bf3b-d571344a3133 · outbound

This paper cites Parameter-efficient fine-tuning of large-scale pre-trained language models.Nature Machine In- telligence, 5(3):220–235, 2023.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Parameter-efficient fine-tuning of large-scale pre-trained language models.Nature Machine In- telligence, 5(3):220–235, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.720473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.720473Z digest=sha256:a73a06f184dca7b977f4adb7bb13040fffd5ce7865198d1fd7a5458809dc2109

Observation a6e169af-b8f8-4e5e-bdd5-e05a896e9cc3 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning MouSi: Poly-Visual-Expert Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.726053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.726053Z digest=sha256:8ede68c9a814519ee0e4fab0bebcf83178ce8de9db356a83b034ee9e76c3dfa6

Observation 791b1a21-74a6-4ede-ba5f-e9c4af2a9656 · outbound

This paper cites Deep sparse rectifier neural networks.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Deep sparse rectifier neural networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.718337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.732437Z digest=sha256:1abc0eb86d130edd9e658ddb001af9ef8367e9bd5c87779a4a51e7417dbb7da6

Observation 95eb8dfe-3c88-4319-944f-85912b3452dd · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.738181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.738181Z digest=sha256:517735cba64562b3f6737c67b3ee67771767d728e48ede209fb075005aad2744

Observation acacf141-d1a7-492d-b5b9-3c6f8cf9c8e3 · outbound

This paper cites Harder tasks need more experts: Dynamic routing in moe models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Harder tasks need more experts: Dynamic routing in moe models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.743855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.743855Z digest=sha256:37d4378c60ee44a1bd186c4d581464a8f4ff8575e4d7c5a0d93c2326b1bc6463

Observation 0f0c4c9c-b5a1-4e67-961f-33c073ec933b · outbound

This paper cites RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.748872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.748872Z digest=sha256:3b56ecb148d57e8f8ee56e741685b186ef4baa900f17cced9877b1fadf2040a1

Observation 2cae4764-886c-4fa6-92aa-ae68254158f4 · outbound

This paper cites Unlocking textual and visual wisdom: Open-vocabulary 3d object detection enhanced by comprehensive guidance from text and image.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Unlocking textual and visual wisdom: Open-vocabulary 3d object detection enhanced by comprehensive guidance from text and image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.699366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.754030Z digest=sha256:2274efd219b2297f94d2563253366db6641aaed303cb02ccb78a52eb709b9810

Observation 7ad139bc-c6db-4689-ac7b-37cdb056f429 · outbound

This paper cites Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.759113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.759113Z digest=sha256:e304b6653989321aa16bcba71b0c05e3d66032455a8e1b9eb7fb692b570fcb31

Observation 8e8fdd22-c8c3-4c82-b9b2-96d3439bc3a4 · outbound

This paper cites Lumen: Unleashing versa- tile vision-centric capabilities of large multimodal models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Lumen: Unleashing versa- tile vision-centric capabilities of large multimodal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.682222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.764630Z digest=sha256:a9a6e4983d02a44ba59eccb261ea199c63fa27453c987bca1dec754b8497414b

Observation 7da838a1-9d43-4d58-99a5-22b63f4f3261 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Imagenet classification with deep convolutional neural net- works

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.770220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.770220Z digest=sha256:5f5f7bb34b37eb1bd1ee6fbc42edd7cb501a842e89e135dda289b5e7717abb3d

Observation c7b873ba-d7bf-4ae4-b205-0edc508b4c29 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.775267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.775267Z digest=sha256:1ba4a157d4cd0919cfb676a6687b96c193eecf7369cd6f857aee20e41c14d111

Observation eccca891-7eb6-4474-bd9e-d6cc548c56e1 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Rouge: A package for automatic evaluation of summaries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.780768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.780768Z digest=sha256:1e933a8fe843934c5a8dca2e369f2bca5ddb75e69c72b2056d00d1d60cbea77c

Observation 95104aa3-f445-495f-93da-ecc27bb27c04 · outbound

This paper cites Visual instruction tuning, 2023.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Visual instruction tuning, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.625708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.786295Z digest=sha256:551835db6683b02cec7f8a9463e9ed1b4c4c94606ed9b6a5d5b719ab213af3f9

Observation 480ba212-1816-4dbd-9758-2dd3d696f467 · outbound

This paper cites Improved baselines with visual instruction tuning.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.606297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.792350Z digest=sha256:276fe147a3f2a990cdef5a9e83e4052ceb83d74869c3159dcede33fca7bbb833

Observation 9823dec1-cb94-42d7-a37d-f41e97f3ed15 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.798888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.798888Z digest=sha256:ca62de34e98a325dc629f5c04768d0ebda024361ef48cfb53c4c8856bc945276

Observation 0f5cf207-ebb8-4ff4-a6c4-54c25a6b83e6 · outbound

This paper cites AdaMoLE: Fine-Tuning Large Language Models with Adaptive Mixture of Low-Rank Adaptation Experts.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning AdaMoLE: Fine-Tuning Large Language Models with Adaptive Mixture of Low-Rank Adaptation Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.805675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.805675Z digest=sha256:ef7b8e07bb6ce84121bab762491e7dce8364b0ddb82bd6389a1d0cc063e9962c

Observation 2bf07059-d448-4a95-9510-58703b1ff33f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.811609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.811609Z digest=sha256:01565feafaaaf781a1b444ce77da8f22e0e2f8b01b41016ae04e19940787e243

Observation d93f0c8e-19ec-4a4a-b20d-2101feb765f4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning DINOv2: Learning Robust Visual Features without Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.817637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.817637Z digest=sha256:6237fac25083724b4be91522930ac59e0b2916da604d833f6207d89e18b738df

Observation 16286e8f-2876-4b89-9ad0-6c6ad8ab5879 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.587874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.823531Z digest=sha256:e3065f871e2b66a4f242eb88ddca9f0af8fc205d7ef3c76c4547b18856972b4b

Observation 4167a38d-3880-40f4-a5ef-164b61f5a131 · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning A Call for Clarity in Reporting BLEU Scores

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.828772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.828772Z digest=sha256:781bc5600411a28cd193248baddd473497fffa6deed382faadd194fef0826a65

Observation 61538d55-5b18-46f5-b59e-c707a540807b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Learning transferable visual models from natural language supervi- sion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.570322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.834854Z digest=sha256:de8d96f23fef8bff63d6d64191075958b8f821ebc00e706a7711b5ec703153e5

Observation 780a05aa-8369-4659-8f5a-6edb99520912 · outbound

This paper cites Scienceqa: A novel resource for question answering on scholarly articles.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Scienceqa: A novel resource for question answering on scholarly articles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.547356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.840309Z digest=sha256:a6893db027f24ec45f782e93aae03ff3ce644687e032fe2e9202ceb97a6b0d4b

Observation ecd51877-14f4-4d05-b97d-ea703d8943f3 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.845427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.845427Z digest=sha256:f0c1b91d073c0af9c7180d2568ca3aa2164397976e0d957b91ecfcb4f06a55af

Observation 031c52a2-023b-4533-841e-c6bbe7f3c87c · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.850908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.850908Z digest=sha256:32cfaabf7f80facb33620e3467003bc773456fbc4963fc5a8b8c759b21a02183

Observation 264b8c67-1619-472e-9c4c-6c98e158f314 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.855723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.855723Z digest=sha256:5868295873bab6dc4fb3a980fc5e4ffa49757eaef4766b3bd2f9bcbce63c6cdc

Observation 1a8cebf6-8ead-4290-9bfd-00fc64a31cf5 · outbound

This paper cites Mixture of LoRA Experts.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Mixture of LoRA Experts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.861350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.861350Z digest=sha256:fc20fbd7b3c49d03214d0f8c9f75c80f257e6b1eb86227bfe7cb9d05dcd904c0

Observation 9663fe63-948c-4bfd-9f35-dc7e1365bfa1 · outbound

This paper cites Vision transformer with deformable attention.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Vision transformer with deformable attention

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.511682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:37:54.867653Z digest=sha256:b96bf909e081266a18bd42cebd33184f0548938a2b7a7ba9248c7b3a2bca0ef2

Observation 6aee29cb-5595-4836-be2c-d468dbc7c14c · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.873349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.873349Z digest=sha256:c1af9b2f81c51cf72ffbbed8d0c4b59914330e0f3c8052a2b55d49e7dfd9d1b5

Observation c651cc23-de9b-469f-b8a4-48eb6bd89114 · outbound

This paper cites FoodLMM: A Versatile Food Assistant using Large Multi-modal Model.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.879662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.879662Z digest=sha256:e1e3ea14645b11aed680672bccee4155122a97a282dbc485772b2faa9c5d16da

Observation bce9786f-9c3d-4cdd-8eeb-1345393caf8c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.885564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.885564Z digest=sha256:ce80e5e6f6776ba00a918ed992cacd82c28dc8fac92e02077bcce8d99b94d35a

Observation b787204d-0eae-4bab-a0a2-258eadc739e8 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.891321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.891321Z digest=sha256:3e6439270f1a19081c5bdd3043282c84b546aa4a1163be238409d086653fb426

Pith citing papers

No inbound Pith citation observations are available.