Pith. sign in

Paper Citation Record · LEDGER

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2501.08443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08443 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:00.817362Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be73d4bd-ef57-42ee-874f-f9a60d38f61d · outbound

This paper cites Improved baselines with visual instruction tuning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Improved baselines with visual instruction tuning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.506708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.506708Z digest=sha256:ee0a274bc4824beff3554e3fff9c2039a0faf77a7c660e486ab7480e93f7138d

Observation 8490bfac-f1e6-4a41-aa93-ec56dabbd854 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.139224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.513042Z digest=sha256:eea61d29620bb966850c07db8a6eebbc3d9cd65cda4a8365f3c6fce091e2f2bc

Observation d701c504-7e14-4aa5-8d31-d043187974f5 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.518915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.518915Z digest=sha256:3b0712e6b4fbf20caa8e7035156b322dba9a50fc5abb6f205adc7dd885920207

Observation 89922604-efeb-46ff-9162-fa71417b4beb · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.104871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.525814Z digest=sha256:abdaa52dd407bb5cd34e7e17703dcf9148aad4059fd0cea77746528a7435f1bf

Observation 8d25a3c1-cc19-4567-a9b9-fb99a4f20347 · outbound

This paper cites Talk2bev: Language-enhanced bird’s- 29 eye view maps for autonomous driving.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Talk2bev: Language-enhanced bird’s- 29 eye view maps for autonomous driving

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.075418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.533535Z digest=sha256:a3b44bca2fa45a0dcc088b75f2bbb41ed6a5a65bad5ec37d599fd38c050c573e

Observation cf590dee-02c1-45e6-80ce-e2caa7310f03 · outbound

This paper cites What can VLMs do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition , 2024.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models What can VLMs do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition , 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:02.038121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.540953Z digest=sha256:de8dcba87b544360d62b7bab376859b5d6bfa441c2172b9fba99f202a38eaa37

Observation 2fef787f-f3b7-4cc9-b564-136d50c348cf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.547361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.547361Z digest=sha256:813e55af62e3efc472d1c0635a0a4d6dfadf1acb14643cbfb12ee74011601d4b

Observation 2a057600-aee0-43cd-aabc-e2ed182e82ea · outbound

This paper cites Learning transferable visual models from natural language supervision.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Learning transferable visual models from natural language supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.555358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.555358Z digest=sha256:652dabc279f39239391dd3428892cc4aa3be0ba48ce76500ce6838643542161d

Observation 74fb2fdd-49a9-4c36-b160-f9baa2d2b086 · outbound

This paper cites Sig- moid loss for language image pre-training.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Sig- moid loss for language image pre-training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.974940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.563882Z digest=sha256:466b20bdd3669a5db268559b958f911ac516d6762a0f2ce08beea2a996bcdbe0

Observation e70968f1-08fb-4f95-9d73-b04acc9a92c0 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models A Survey on Hallucination in Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.577136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.577136Z digest=sha256:9f7ab8f63edefbe39f2861de2bb076756e52bf60a21539ad0d1d9f0bd6b98052

Observation 3c60e1f1-274a-4bfa-914c-0e559bffa61d · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.935040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.582477Z digest=sha256:7bd92c777ca7a6dbe8fa6024d3da98079f8008656576613b3f9da1ce3b98477d

Observation 653fdad4-a124-4cb3-b982-d97938d745f8 · outbound

This paper cites Teaching matters: Investigating the role of supervision in vision transformers.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Teaching matters: Investigating the role of supervision in vision transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.899169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.592166Z digest=sha256:f7a10ff5f0c12cc2af174b565080611faa5492e5ef6d0ff7f16db36fbf624378

Observation af3d2fe9-75c3-4303-b35a-1ba8b32c0ed8 · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models What do Vision Transformers Learn? A Visual Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.604755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.604755Z digest=sha256:42cd2bcd56d001994e5818260c967c7177f2dbc8fa2fac02646f368d260f0b7c

Observation 90f23e13-0e69-404d-b768-e6734529e72e · outbound

This paper cites Dense Connector for MLLMs.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Dense Connector for MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.610965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.610965Z digest=sha256:ae0b4020099ce0f1e673c51507f4cecb47a02404bf671c633c793d67ddd99380

Observation 8246b6d9-140e-4581-9051-2b84db952a05 · outbound

This paper cites MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.618591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.618591Z digest=sha256:9a95898f16c4f0cd3ae5460274e801b6c351c9e8b53e3400cada17589e5e933e

Observation 4631b653-1938-4e6c-b03f-3a27ec22ea07 · outbound

This paper cites MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.624001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.624001Z digest=sha256:1f30dcdafbe8e9304bcfe56c03d1d57e1fdded48024de7f031ef371037d9b9c3

Observation 898b8b5a-0402-4615-9ee9-2154ab66256a · outbound

This paper cites MoVA: Adapting Mixture of Vision Experts to Multimodal Context.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.630940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.630940Z digest=sha256:f9e05893574cdc374457a2d903e96706257fdbb00f5eb949eb249e0935436e0f

Observation f3580f57-c49a-4d0c-a359-49cfc8d51145 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.637401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.637401Z digest=sha256:fe4a1a53a7306a527e95e6f5f185c0b0273e23c231070542f89b66092335b3a3

Observation d7d7f11d-e134-46bf-b1f3-369914cb3555 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models V?: Guided visual search as a core mechanism in multimodal llms

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.873148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.644690Z digest=sha256:5ee817c45262573121b3ee36db1e66fdbfaee26b016d2566b16546415ff2e9c8

Observation 416e3d2c-97ea-48a3-b3b1-1155cfbbbe3c · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.652869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.652869Z digest=sha256:d43df94a5aa2382832d42ea7e8e23ec31bebafd96843b6868bb7ebf7c604b4f4

Observation b02f5a5f-4c2b-4d89-96db-5a0f22b816c8 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.661476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.661476Z digest=sha256:3ad69441857fe9bf5275cb0f688615120778475f4db6082ae3d4b5fc58d9245a

Observation 7b31f994-3db9-4df9-8637-ec7370a7d065 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.669269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.669269Z digest=sha256:8e7aa548ed9fcc7829a815439564ccfb8519b2fb3a5b16d35d901249ef366d3b

Observation bdfdb95f-bec3-4c6a-ae64-c86047a500fd · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Gqa: A new dataset for real- world visual reasoning and compositional question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.839319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.675932Z digest=sha256:6fa8783514c73c4b235d4c60a18141f2ee9b729646bc14e6288ff6788c66518e

Observation 33bbe69a-0068-40b3-8dc3-4d50962adb11 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Seed-bench: Benchmarking multimodal large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.810997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.681358Z digest=sha256:0568fecc2ec9bdf8b14fc80b5ac0bde7c8d020083c1b49aecf1f826cd5907536

Observation 2ff97242-887e-4944-a538-9141a8ab09a1 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.690938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.690938Z digest=sha256:a130b07087883b2913d1fcd72194b00946bad7f891759d26f1f5ed3ec0353b5e

Observation 06cc2734-d903-4553-956d-dbfcdbbf5371 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.695885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.695885Z digest=sha256:c19c513a6a7a44b5907688b745633539f4060cd97a971ecf7846bd672462e641

Observation 29ab807f-1c88-4b7d-a8ea-0832ef65912c · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.703130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.703130Z digest=sha256:68227ee556a000ef6071cf39ff02a6a3b2061b4ac5fd5db8f5a32309d0faf118

Observation 0efc949a-7bae-46ad-8973-be955f65c7d2 · outbound

This paper cites A diagram is worth a dozen images.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models A diagram is worth a dozen images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.733399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.709191Z digest=sha256:d66abcd6be36d49d1b8cec5d03050f6eb9bd70ccd61d28d8c35df78b291f367e

Observation a46e26cb-fe04-4b59-869d-dffdb9d6281a · outbound

This paper cites Learn to explain: Mul- timodal reasoning via thought chains for science question answering.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Learn to explain: Mul- timodal reasoning via thought chains for science question answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.710494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.716041Z digest=sha256:cee07b3464596cb9dd83af46bd5dc44be0574bce8c5684f96de62421e3f990de

Observation 1fbbf652-46d6-4fd8-b1d8-d0fa655c23b0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.727677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.727677Z digest=sha256:2db121d3083ec431666cc08979ddd27f0b968b4a6f5386146317e902f22840a4

Observation 653df0b8-edd1-4d60-9a2d-b005808c9712 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.735323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.735323Z digest=sha256:67a84d1be9d964142a90432a69644e05f73e2398d5fa20220700ea105282bd1b

Observation ef3bfa21-e413-4e3e-8aea-086b37dfa983 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.742679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.742679Z digest=sha256:04af4ef3254bb958200b494e3b409d9d47df715c39f01c969a8415797ab582bb

Observation 75b86e8a-8257-437f-a287-5cc42410321a · outbound

This paper cites Towards vqa models that can read.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Towards vqa models that can read

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.749770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.749770Z digest=sha256:b22255c64be23f06da080caf8b47208c4fda2eb3108d985288bff439811bbf4e

Observation e9ec0c7f-c951-41dc-82cd-e9054f08ef5b · outbound

This paper cites Grok-1.5 vision preview.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Grok-1.5 vision preview

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.756337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.756337Z digest=sha256:4cf4254f0d300016c49091039dd9ae24e8cc7a86a75d7dc6e0745bb08cae3dd9

Observation 19656f93-4ef4-48a0-ae14-7711b892958c · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Gaussian Error Linear Units (GELUs)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.764139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.764139Z digest=sha256:e5f61a24f8c393bc0a0a9d614e4d62562fd3c887b8e23bc6b169767034cc306d

Observation bfdaa309-176d-42e2-a61c-75346a002ea2 · outbound

This paper cites Mpnet: Masked 34 and permuted pre-training for language understanding.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mpnet: Masked 34 and permuted pre-training for language understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.623934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.771072Z digest=sha256:49e20f4983a0d71413088529a7635e44bd074aad25cc19cfe727ec5cc6b2de08

Observation 6965958c-569f-4dcd-ac64-8d22f287c9e9 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.599632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.777993Z digest=sha256:c34a52f71d63191a25c29cb4c1a9324dd8c645c7dbfca955a371e71e45234d43

Observation ef5fe6bd-c25c-41eb-b6bb-58a347ec8ae2 · outbound

This paper cites Decoupled Weight Decay Regularization.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.783374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.783374Z digest=sha256:239b46352b15093d6f56e94d5a61f9895afd76881237c2e7fdc9a79f56d95770

Observation 2b03c9c9-6175-447d-b0a2-a922ed74190e · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.574404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.789635Z digest=sha256:e1a61141e6a408bc1de364b6f77214fb4236dc040cff3d36b972997702348d08

Observation 237d7dad-4893-422d-b544-c4588fb984f3 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.795831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.795831Z digest=sha256:b52a27feac29c5a42fd9fcd4033bd09ecbbe444b44b4b5a3c852de5d693478d1

Observation d5f9fbc8-6894-4604-883a-f29cc4ee3007 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.553064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.801285Z digest=sha256:c305869fc2e1c550c2623746b65dc9f5f13522f080420c262c324dea7e57b702

Observation 08e6b8ce-f8e3-4346-be79-123858363622 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.805901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.805901Z digest=sha256:302b791b12f8b1121ebf770328d5b723ae92fddbdcbcf218a7c19c7734fb8693

Observation 9e17f659-684c-4700-acfd-ede212d60078 · outbound

This paper cites Obelics: An open web-scale fil- tered dataset of interleaved image-text documents.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Obelics: An open web-scale fil- tered dataset of interleaved image-text documents

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:01.516382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T01:04:00.811591Z digest=sha256:acedede7d01084d53d9e58f6473eaeee6ea114e54e47f95005c712be39d3ee20

Observation b4dce60f-e2f2-4024-b239-f0c80c3ef830 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.817362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.817362Z digest=sha256:60f4033e6728a31a2f766301341695fb8d99f66ffafdf16714c01f058afacdf2

Pith citing papers

No inbound Pith citation observations are available.