Pith. sign in

Paper Citation Record · LEDGER

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models

As of 11 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2501.08443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08443 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:00.817362Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be73d4bd-ef57-42ee-874f-f9a60d38f61d · outbound

This paper cites Improved baselines with visual instruction tuning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Improved baselines with visual instruction tuning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.506708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.506708Z digest=sha256:9ac8ff57171a22f9bff1e25a3b1bf474caa7bb19193813961962dd1d0622b796

Observation 8490bfac-f1e6-4a41-aa93-ec56dabbd854 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.139224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.513042Z digest=sha256:ed5b08f88e480f094fa2e08718dadb1c16824ee841c4c8279601fdb3748be5d0

Observation d701c504-7e14-4aa5-8d31-d043187974f5 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.518915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.518915Z digest=sha256:646caf5c99b3b237037e8c63a1392418695ec3b05d86bef041744b23048670c7

Observation 89922604-efeb-46ff-9162-fa71417b4beb · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.104871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.525814Z digest=sha256:394905bff2eeffdfc57d4ae326e7c294ad7aefdb08040f6b0357ee05a8672c93

Observation 8d25a3c1-cc19-4567-a9b9-fb99a4f20347 · outbound

This paper cites Talk2bev: Language-enhanced bird’s- 29 eye view maps for autonomous driving.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Talk2bev: Language-enhanced bird’s- 29 eye view maps for autonomous driving

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.075418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.533535Z digest=sha256:d9485be25f9ebf612513b432dca6a78a48d29ff7b8fb73922dffc28c4cf93132

Observation cf590dee-02c1-45e6-80ce-e2caa7310f03 · outbound

This paper cites What can VLMs do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition , 2024.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models What can VLMs do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition , 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.038121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.540953Z digest=sha256:ff8eb8e38d913eb36fea6a5135e4c1490e474c235b865cf1a09d4e628e98e279

Observation 2fef787f-f3b7-4cc9-b564-136d50c348cf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.547361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.547361Z digest=sha256:6dcc47f130ee773cebeda52aed91d63a1caa822a3e4359b2ae28dc3b834aa3f9

Observation 2a057600-aee0-43cd-aabc-e2ed182e82ea · outbound

This paper cites Learning transferable visual models from natural language supervision.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Learning transferable visual models from natural language supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.555358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.555358Z digest=sha256:2955948d24d8380e15b11c50e8332d3ee924a651cb05684d3ce22c1ceb592187

Observation 74fb2fdd-49a9-4c36-b160-f9baa2d2b086 · outbound

This paper cites Sig- moid loss for language image pre-training.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Sig- moid loss for language image pre-training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.974940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.563882Z digest=sha256:124388b6581afa55bdf13be08bb6d7874d5ba0e72dedf0f29bbd360580b7a01e

Observation e70968f1-08fb-4f95-9d73-b04acc9a92c0 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models A Survey on Hallucination in Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.577136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.577136Z digest=sha256:a863d5376ca8263a192d1df4f6800f51688666fa708c3be8b2795e3c11eb5554

Observation 3c60e1f1-274a-4bfa-914c-0e559bffa61d · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.935040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.582477Z digest=sha256:758c8b04a303449e112a072288031cdcab1e9c0d4cd99cf2ae58bf836abfb69c

Observation 653fdad4-a124-4cb3-b982-d97938d745f8 · outbound

This paper cites Teaching matters: Investigating the role of supervision in vision transformers.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Teaching matters: Investigating the role of supervision in vision transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.899169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.592166Z digest=sha256:c297f7aff3137807e66bd2473d493045f0ade8fe699b7384884b414ebf1fee6c

Observation af3d2fe9-75c3-4303-b35a-1ba8b32c0ed8 · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models What do Vision Transformers Learn? A Visual Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.604755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.604755Z digest=sha256:3ebd2e0a7731b670f673b8f14df5185f0aecdc8e093962527aa3e0d4764cae65

Observation 90f23e13-0e69-404d-b768-e6734529e72e · outbound

This paper cites Dense Connector for MLLMs.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Dense Connector for MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.610965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.610965Z digest=sha256:1e3244d26b5bef1608800912f327e3ae97d309ad4ab1a806d5e6f1e66e049008

Observation 8246b6d9-140e-4581-9051-2b84db952a05 · outbound

This paper cites MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.618591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.618591Z digest=sha256:6a84eb00da3695857262d394b307d1873dd89f1365e0a9a1f8280b5d7805fdc6

Observation 4631b653-1938-4e6c-b03f-3a27ec22ea07 · outbound

This paper cites MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.624001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.624001Z digest=sha256:d740b049b8796aca7ab83a1aca35a6af5dcc3a1c272d97ed2fc2897f9b84436f

Observation 898b8b5a-0402-4615-9ee9-2154ab66256a · outbound

This paper cites MoVA: Adapting Mixture of Vision Experts to Multimodal Context.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.630940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.630940Z digest=sha256:a288863bc46660ab5138185d46938682832759ed6852ca179558e8b7691a9e62

Observation f3580f57-c49a-4d0c-a359-49cfc8d51145 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.637401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.637401Z digest=sha256:90ba011809bfdbdb08a98bd6e8d1789bc54a191f481a9f8ccde3ecb1d45a554b

Observation d7d7f11d-e134-46bf-b1f3-369914cb3555 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models V?: Guided visual search as a core mechanism in multimodal llms

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.873148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.644690Z digest=sha256:24d38a7e792474ce542a7714597d118cf0b3702bfabe61891f035ed96a888568

Observation 416e3d2c-97ea-48a3-b3b1-1155cfbbbe3c · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.652869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.652869Z digest=sha256:9c15d05f1efc83b04368c1529155391fb8923ade2ca5185d367e3ce655d31b28

Observation b02f5a5f-4c2b-4d89-96db-5a0f22b816c8 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.661476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.661476Z digest=sha256:cb417ba88f6f52ccc1c07f42446c24a5a328e7253865b856360f01d5bfa6148f

Observation 7b31f994-3db9-4df9-8637-ec7370a7d065 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.669269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.669269Z digest=sha256:550cd927683bd1eaf810d052510147810795b2b15cbf219a2bb12316562ae449

Observation bdfdb95f-bec3-4c6a-ae64-c86047a500fd · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Gqa: A new dataset for real- world visual reasoning and compositional question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.839319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.675932Z digest=sha256:acd47794fb7fa0e2636e122a57c27b3aa16085c57f65e8dc6d19b19f9ccdba34

Observation 33bbe69a-0068-40b3-8dc3-4d50962adb11 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Seed-bench: Benchmarking multimodal large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.810997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.681358Z digest=sha256:9c89a4af58aa7c49032edd9eb26514ff9efbbf34b4d1b9ffce43c73442ce1cf3

Observation 2ff97242-887e-4944-a538-9141a8ab09a1 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.690938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.690938Z digest=sha256:1f2e8628c9ec8d902416e9e6b96043ccf20c077ea1980199e3cf4149bdbf3bfa

Observation 06cc2734-d903-4553-956d-dbfcdbbf5371 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.695885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.695885Z digest=sha256:9e5409610b0e45d417a39011705d77d6ba6e5f0382b309aace2f9a69284110ea

Observation 29ab807f-1c88-4b7d-a8ea-0832ef65912c · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.703130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.703130Z digest=sha256:fa3ac4f4a47c051d40913b23d5be331508403bdf8419ee4dd3dda153f626152f

Observation 0efc949a-7bae-46ad-8973-be955f65c7d2 · outbound

This paper cites A diagram is worth a dozen images.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models A diagram is worth a dozen images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.733399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.709191Z digest=sha256:a33446bfbfeea8ea63b50aa9f68916f32ef9515daeba38976a6e02e44a0b7341

Observation a46e26cb-fe04-4b59-869d-dffdb9d6281a · outbound

This paper cites Learn to explain: Mul- timodal reasoning via thought chains for science question answering.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Learn to explain: Mul- timodal reasoning via thought chains for science question answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.710494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.716041Z digest=sha256:15eb565c63cf425eff03d92dcb54fd55266e1e8543a3999d09f170e2aa43c697

Observation 1fbbf652-46d6-4fd8-b1d8-d0fa655c23b0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.727677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.727677Z digest=sha256:525ceee05d33f89f79cfacaa2ace6e0bbd069b5a68b6dc3c8694efb1194cfad8

Observation 653df0b8-edd1-4d60-9a2d-b005808c9712 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.735323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.735323Z digest=sha256:071adc66fe6b614e98c38c0c56b882de4a611db151fbc5a873b94ee33221ecb4

Observation ef3bfa21-e413-4e3e-8aea-086b37dfa983 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.742679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.742679Z digest=sha256:f76b91081158649fd5caf5367a22343632e09c4e009176594a2b807518b6f526

Observation 75b86e8a-8257-437f-a287-5cc42410321a · outbound

This paper cites Towards vqa models that can read.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Towards vqa models that can read

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.749770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.749770Z digest=sha256:7410b0e351f1985bc7d9e6a73e402890fc57e7bdcea8753abff13882cfed8a90

Observation e9ec0c7f-c951-41dc-82cd-e9054f08ef5b · outbound

This paper cites Grok-1.5 vision preview.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Grok-1.5 vision preview

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.756337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.756337Z digest=sha256:0701b57902613f695ec11289e6c35eb5d34aacacf7c513ba5decb16b15bcba22

Observation 19656f93-4ef4-48a0-ae14-7711b892958c · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Gaussian Error Linear Units (GELUs)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.764139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.764139Z digest=sha256:a198887a4e8d897b3a0bb043cc977fbea239f85bd1083d6f56dea71c9237e529

Observation bfdaa309-176d-42e2-a61c-75346a002ea2 · outbound

This paper cites Mpnet: Masked 34 and permuted pre-training for language understanding.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mpnet: Masked 34 and permuted pre-training for language understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.623934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.771072Z digest=sha256:d57c16798b88a62ea6b35cc49a94908bc26b2bcd61d46694f424c205f37d11cb

Observation 6965958c-569f-4dcd-ac64-8d22f287c9e9 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.599632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.777993Z digest=sha256:ec1dddac1df84c6556032b2921a9252ac94575102e610cb17eeca6ce8ee316cc

Observation ef5fe6bd-c25c-41eb-b6bb-58a347ec8ae2 · outbound

This paper cites Decoupled Weight Decay Regularization.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.783374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.783374Z digest=sha256:1610eb3c63737bcc3d08ea5a0f4438955fa399f9f13e78bf081d2b21cda3ad3f

Observation 2b03c9c9-6175-447d-b0a2-a922ed74190e · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.574404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.789635Z digest=sha256:0ae38d7f890f2891e57740e2bc68a875da55ae51f7e08a8873807ec045732c2b

Observation 237d7dad-4893-422d-b544-c4588fb984f3 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.795831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.795831Z digest=sha256:9d96daab133f055ad339da03dee842a1f8ddc10febc374189b751dd9b33d3708

Observation d5f9fbc8-6894-4604-883a-f29cc4ee3007 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.553064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.801285Z digest=sha256:7ee76d81ce7f4c070cab11bcd7142e134f89e459cfc9ac35c673f0763ad2d466

Observation 08e6b8ce-f8e3-4346-be79-123858363622 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.805901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.805901Z digest=sha256:228dbedf5b94b97e928981aead1694d8909c793209835cfb605bbff7bdbfde5a

Observation 9e17f659-684c-4700-acfd-ede212d60078 · outbound

This paper cites Obelics: An open web-scale fil- tered dataset of interleaved image-text documents.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Obelics: An open web-scale fil- tered dataset of interleaved image-text documents

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.516382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:00.811591Z digest=sha256:ab930096558b529eafe767d3e7921d2e3726e48c10f3f63ffb7cd07a0e761074

Observation b4dce60f-e2f2-4024-b239-f0c80c3ef830 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.817362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.817362Z digest=sha256:7f1e7cdeabd2787d907ddcccb0c995e5718f3ebe23fb4e1a7c243f4a956e9bb3

Pith citing papers

No inbound Pith citation observations are available.