Pith. sign in

Paper Citation Record · LEDGER

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective

As of 10 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2502.01524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01524 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:40.168560Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 142 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6d437cd-146a-4b77-a2d4-69c028ed7f5a · outbound

This paper cites A survey of vision-language pre-trained models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective A survey of vision-language pre-trained models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.867620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.867620Z digest=sha256:25ae6bd222255f510f32a14437e194354bc3f6b073afd0c98a54ef2a519a0c58

Observation 1e443c2d-445c-4abd-8f04-3800e1afa15d · outbound

This paper cites A survey on multimodal large language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective A survey on multimodal large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.872201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.872201Z digest=sha256:94d2c67569b7d3f8d06158986f235168b1da3c65f692922cfc7625a9fa0feb84

Observation 6daf3c52-c715-42ee-8f23-7c07cdbd0bb7 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.875630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.875630Z digest=sha256:ab7c6b86d1f9d74a8720a86b7038354295ba0f6bac49af73ba1a60ee240c73b5

Observation 164ad24e-cf08-4bff-b421-2bb11d9c1dca · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.879606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.879606Z digest=sha256:85deb4ebce0f082df88041995f0c9893b6e7e14365855d5955c8d40de13b380e

Observation b5c87fe2-625d-4855-a27f-335b72e9376b · outbound

This paper cites CoCa: Contrastive captioners are image-text foundation models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective CoCa: Contrastive captioners are image-text foundation models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.883096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.883096Z digest=sha256:26cf3c9bdb792dfb55f1d69f282f3c7e2dc2bef17c6397d42d810d691a0221c1

Observation f426c724-45ed-4418-adba-dd5e62354e62 · outbound

This paper cites Multimodal few-shot learning with frozen language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Multimodal few-shot learning with frozen language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.886565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.886565Z digest=sha256:c564f68e951603e94a22f43e440ffe40f85b60813cf9e58a7932b339cf2e4594

Observation 46a9367c-c561-4218-95e6-3298602c00f0 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.889977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.889977Z digest=sha256:0ab59f1fa2d6574d418013452f4d9e01b665407d7101e1bb7d907b01733be533

Observation 6ede9358-135c-49ac-a1b7-166d4138f130 · outbound

This paper cites Visual instruction tuning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Visual instruction tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.893311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.893311Z digest=sha256:be0ad92cefadc63b6ffd2e797fa7feb0abd2d6ea68102a3c5b74d0b2a4cd41d6

Observation 22c3ea26-8cb7-4be9-93ee-c0317562134d · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.896772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.896772Z digest=sha256:41dd836ae71ff28ed8fd83d2d052ee54a20dec4020155912c05b7db1312571d7

Observation ef4f0941-7721-4946-9159-2ba47d851006 · outbound

This paper cites Enhance Reasoning Ability of Visual-Language Models via Large Language Models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Enhance Reasoning Ability of Visual-Language Models via Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-09T15:04:40.716519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:39.900226Z digest=sha256:f47fab58af044a46fd374aa0a00e0738684d4b70c20537556ea704dc0b7666cc

Observation c987c973-45c6-4dc3-a7b1-1ae724bd7c58 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.903844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.903844Z digest=sha256:032a7acfb716ab7056ce607d2f36dc4916dd24abac09b4a4904dddd6ebe5c028

Observation 6dedc236-47b4-4c04-85f6-228e1aee63f0 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MMBench: Is Your Multi-modal Model an All-around Player?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.907109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.907109Z digest=sha256:0968963e6ae2e2fe24a35dd632a5fe99afb66b183b5c34f98e19cf08c41e5b6c

Observation 6dcf1de6-7740-40ac-864b-634065d7126e · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.910736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.910736Z digest=sha256:5703981ec7ac641555ec3ab24ec9ce6135bab1a163405d95ffad390085a10dbd

Observation cd67d5d3-1d84-4aa8-acaf-d9ba6fbd9704 · outbound

This paper cites Efficient multimodal large language models: A survey.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Efficient multimodal large language models: A survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.914137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.914137Z digest=sha256:ac89c2004f2e69378e8754aef33c7bd3b538966e9af6ef5d4a5c01bb0289776c

Observation 647b4678-9837-4bb3-947a-1a1e691f536e · outbound

This paper cites Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.917367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.917367Z digest=sha256:2d15f74fe3374a0ac00ee61d3fbe39abf3bcb8ff82299f342e0b6dfeb53f404e

Observation a29d29bf-7d17-4e2f-ac16-41e4e9e7d1b1 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Flamingo: A visual language model for few-shot learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.920901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.920901Z digest=sha256:f01e4e1b35f603b530514623ba8e672cbce72e729ff0be243425c4a7b07b10ce

Observation e0b950ca-9483-42a5-b8e3-d2a945e9fbc9 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective ClipCap: CLIP Prefix for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.924340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.924340Z digest=sha256:57fe7b978c43688955de595b104804c5b15340336d34bb2a85cd878ecfefb99c

Observation b564ca1b-22f3-463b-9742-65a1c236502c · outbound

This paper cites MAPL: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MAPL: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.927544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.927544Z digest=sha256:562bbf3a3428708f70c68ddd87e8f0efc00f1c2a3d3ef6c2fcbd508d340fa45b

Observation 298d19fb-73cc-4deb-920e-49a02a55fae5 · outbound

This paper cites Meta learning to bridge vision and language models for multimodal few-shot learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Meta learning to bridge vision and language models for multimodal few-shot learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.930180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.930180Z digest=sha256:909fb3ca80a2653eb34471084f9af16beb04c97f7c1707402be077bd37ef53b7

Observation e2705466-79f1-4220-8889-99aff244a97a · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Grounding language models to images for multimodal inputs and outputs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.932638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.932638Z digest=sha256:0f4213a7115b2fd0fab7282e45343cab5b512d0f28de07ed81304210d3921e75

Observation 498abddf-ebeb-4656-8571-af2692ff48ad · outbound

This paper cites MiniGPT-4: Enhancing vision- language understanding with advanced large language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MiniGPT-4: Enhancing vision- language understanding with advanced large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.935038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.935038Z digest=sha256:eb928811e022c062db37f46f9748ff48f8d1f0db358e96e2e0706c87a6c7dae4

Observation 7f630ceb-7b5a-4204-923b-fb43ffb6d6e5 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective CogVLM: Visual Expert for Pretrained Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.937725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.937725Z digest=sha256:9f804921829898f72880b81aaa0a3bdae0e018660a2c40b93dbc4473120cf2dd

Observation 0269ea58-282b-475a-998b-a714a920f942 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Zero-shot video question answering via frozen bidirectional language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.940432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.940432Z digest=sha256:dce79ec1f7bbeb07682c48285855bdd791c71f81f02c5812bb0625d0246b7dee

Observation 60e8302d-1e73-4e18-bb87-d65cc6f04b2e · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.943147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.943147Z digest=sha256:611b2e5019de58559a503665b78a8c1381f66069aefc0b55afbd47febe6ef08e

Observation fb9d8667-ce3b-4569-8819-10fae47b825d · outbound

This paper cites Improved baselines with visual instruction tuning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Improved baselines with visual instruction tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.946263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.946263Z digest=sha256:494d6ce4c40f942fb54e7bcaec7d644e086691fa6f3932e44b1e6083d67e4f89

Observation 939da5a9-fefe-4e2e-8126-3e87f6f7a102 · outbound

This paper cites Qwen Technical Report.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Qwen Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.948925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.948925Z digest=sha256:2b69a723c1fc28d579e392b1f508bb36324e8b925c1f6240035775334a1ebd42

Observation 9b7213fb-94c0-42ab-b8d1-8f55e76ea7f0 · outbound

This paper cites mPLUG-Owl2: Revolutionizing multi-modal large language model with modality collaboration.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective mPLUG-Owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.952130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.952130Z digest=sha256:0e4f8390c210bfdab6007d1acac40717ac7398a600d27ab2d659d50961360635

Observation 4b0aa42d-dad9-46f8-a9de-823a1f6ef859 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.955349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.955349Z digest=sha256:651fe700edc490b2d5923c7476f52b3ac6f14a67929feff33e6fd2764d86cea5

Observation eb78acd7-16a6-4749-8518-77d4829a4d38 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective LoRA: Low-rank adaptation of large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.958683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.958683Z digest=sha256:c4ed4ee12387fcd7af9a7721734db9d277072c64183704bfa0c396ffa8a7468b

Observation 872f33cc-f8e6-4364-b6fe-8a3940391320 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.961643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.961643Z digest=sha256:d652f4c6288af6f48c3f02bf086daf37b71a5b07429d112e320aecd4ba5fe30e

Observation 205ed56e-b783-45c7-be21-eeb7ba56ca4e · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.965044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.965044Z digest=sha256:22e8dc42801d65a88dbb013cd3ec2bc9fc2a4d8fc901a052509704664b948f3e

Observation 320076f9-1530-4cd8-8348-30a4b87f2868 · outbound

This paper cites VL-Mamba: Exploring State Space Models for Multimodal Learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective VL-Mamba: Exploring State Space Models for Multimodal Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.968203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.968203Z digest=sha256:9e5bd069ccd6e2394448ed1de8cd3530d910f93a3e38a0b90c96bef827fff91f

Observation 5996fdf1-abec-4886-a158-75ea0c9a31fd · outbound

This paper cites VL-Adapter: Parameter-efficient transfer learning for vision-and- language tasks.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective VL-Adapter: Parameter-efficient transfer learning for vision-and- language tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.971642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.971642Z digest=sha256:1500e17dab5cf8738c427cc4c88aec7e2c478e72121a1fcd4d5282aedf836167

Observation 365c8f04-35c6-4d65-96fd-b0fec66bcd56 · outbound

This paper cites eP-ALM: Efficient perceptual augmentation of language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective eP-ALM: Efficient perceptual augmentation of language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.974572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.974572Z digest=sha256:9a5e8e60b7259ae51da77a95deade5fef58d065e1ba68f38039a64a8e63fe8b7

Observation b2703f53-58c3-4817-884a-bc2848fea02d · outbound

This paper cites Modular and parameter-efficient multimodal fusion with prompting.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Modular and parameter-efficient multimodal fusion with prompting

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.977405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.977405Z digest=sha256:8c7a75f350308aff7a69abde599d59ea8460a1fadb5ef7735e294376d572ab4d

Observation 5b3af116-f719-4f8f-b741-c6ed778b100e · outbound

This paper cites Memory-space visual prompting for efficient vision-language fine-tuning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Memory-space visual prompting for efficient vision-language fine-tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.980372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.980372Z digest=sha256:3cf567f38840c64bcc2de631b6e7342f73ea5ecec52775d91ae12ee4da94e9c0

Observation ce314d64-5d73-4e5c-ae4d-192615ba2de4 · outbound

This paper cites LLaMA- Adapter: Efficient fine-tuning of large language models with zero-initialized attention.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective LLaMA- Adapter: Efficient fine-tuning of large language models with zero-initialized attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.983200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.983200Z digest=sha256:80b31fb94ace16e1331ce157127c65f3d4d811e61ddf78b9ed1e6b62bd14411c

Observation d589512b-9208-4fc7-9e5d-dd8029137ec2 · outbound

This paper cites LST: Ladder side-tuning for parameter and memory efficient transfer learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective LST: Ladder side-tuning for parameter and memory efficient transfer learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.986215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.986215Z digest=sha256:9d643749746fc255526f757dfce38381cf423ddd0412f0d05893c167205bd893

Observation cd8b9769-22f6-4504-bf24-bb786ac0c34b · outbound

This paper cites Querying as Prompt: Parameter-efficient learning for multimodal language model.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Querying as Prompt: Parameter-efficient learning for multimodal language model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.989148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.989148Z digest=sha256:91685e5544ad214db05d6cf23e3ba57c67696670da699abca909bf228eae19ff

Observation 174c3a02-9cea-4510-a5ba-43ec66a4061a · outbound

This paper cites VL-PET: Vision-and-language parameter-efficient tuning via granularity control.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective VL-PET: Vision-and-language parameter-efficient tuning via granularity control

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.992123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.992123Z digest=sha256:b4943f51bde061f8befc5d645bac795bf8b3cf9d40280516174e3f8c7824f695

Observation 3531df63-6c86-4d50-9622-6f928072a227 · outbound

This paper cites Cheap and Quick: Efficient vision-language instruction tuning for large language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Cheap and Quick: Efficient vision-language instruction tuning for large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.995219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.995219Z digest=sha256:a3edc1dfb6ae03cccec1ab59f4226fc25e427df40c40de9172b971bef93f8033

Observation 8a02be6a-ca0a-410d-805b-27517172b903 · outbound

This paper cites https://scholar.google.com/, accessed 3 February 2025.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective https://scholar.google.com/, accessed 3 February 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.998099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.998099Z digest=sha256:b4657ee2be7110c207ae6749048b16c7989a5c66f92a69cedd643872971767cd

Observation d57eb153-3b91-44a0-9dff-b049c42df845 · outbound

This paper cites https://ccf.atom.im/, accessed 3 February 2025.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective https://ccf.atom.im/, accessed 3 February 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.001097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.001097Z digest=sha256:a87ec41ce2ae39817c38be6ec23917d79d352a6a8e812de5b9fc22b894479fef

Observation 3238c389-7e2a-4016-92f6-276a361469c5 · outbound

This paper cites MAGMA – Multimodal augmentation of generative models through adapter-based finetuning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MAGMA – Multimodal augmentation of generative models through adapter-based finetuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.004049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.004049Z digest=sha256:8263a972320392e53ad9125e2e07d3ae2be1d2a30f6bf8bd11f84cab076dba8c

Observation 89db0d2b-9a55-4f5f-bbcb-e66c5c8cda71 · outbound

This paper cites Fusing pre-trained language models with multimodal prompts through reinforcement learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Fusing pre-trained language models with multimodal prompts through reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.006975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.006975Z digest=sha256:9b53546c056fe56aecd676c97839029042cd029d5ce7270a7cd567ad9e5fa547

Observation 98198cd3-20b6-4985-87af-7510cc8920bd · outbound

This paper cites Aligning large multimodal models with factually augmented RLHF.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Aligning large multimodal models with factually augmented RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.009864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.009864Z digest=sha256:dca27c5c45aa5aa2a5bbf010c724a3c7a596cdde20443f6151e5dbdd41a520b3

Observation 13933274-9c49-48b4-a150-00ad65d5e30c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal LLM.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Honeybee: Locality-enhanced projector for multimodal LLM

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.012962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.012962Z digest=sha256:e8044d0f1d005800f79faf2479f1bec769a7b62747daa28d081f970116e771a9

Observation 309f6c3a-8f81-4b20-ad4b-905e202c53b6 · outbound

This paper cites Tuning large multimodal models for videos using reinforcement learning from AI feedback.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Tuning large multimodal models for videos using reinforcement learning from AI feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.016197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.016197Z digest=sha256:8837d363cc85bef38fc0b501e8584a790f0f60cf60708ad7c69702a68cc7dac0

Observation ebffaf42-cce5-4b44-a99f-94a273ea8917 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.019790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.019790Z digest=sha256:82173df430d7842f002d3ed8b47a87fdeab97acc6337f964956c219b1fac6d08

Observation c6e04b1b-5c38-4d0a-b2ac-7feecbfcff68 · outbound

This paper cites Attention is all you need.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Attention is all you need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.023001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.023001Z digest=sha256:cb45ce39d4806b0b857a1c27fb38e260719d1fbd97a506fbe1ec5f2daa7423fe

Observation 7e9f69bd-b610-4224-abd1-26373140a28f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.025902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.025902Z digest=sha256:40b0053555d7db2671ee6a22903f73081ddb9ec7a1d7a7cfb63bd2ee5d9c9830

Observation 8daa8679-7691-43cf-8aa1-b3ce9a17c000 · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with disentangled attention.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective DeBERTa: Decoding-enhanced BERT with disentangled attention

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.029101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.029101Z digest=sha256:de07918d42c06112bbdd5059c6bef87590f9f1ef8f36e9805f0d047d3922e5e0

Observation 9c3a6450-afbb-476d-b6be-8b73bff3c758 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.031596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.031596Z digest=sha256:72c3e13248550ad525080fad2728983e89f3d4a10c2c282fdefb8a6fbc3f7764

Observation 3409ea1b-91f4-4a2a-939c-bf0e6b79dd54 · outbound

This paper cites BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.034055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.034055Z digest=sha256:80b50c2e09c9761ad4d39379f0f9371b8a559f681396f44fc11c8d1b055b1961

Observation 5dec0e91-c574-4940-aa2e-156a8acf3e2f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.036831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.036831Z digest=sha256:4cc99fa76999491a874968462f8961fcb4dfdf610be96fd3297f665c511b16e2

Observation 1e1968f0-f36c-48d8-b779-7a1d39c5b6a9 · outbound

This paper cites VICUNA: An open-source chatbot impressing gpt-4 with 90% chatgpt quality.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective VICUNA: An open-source chatbot impressing gpt-4 with 90% chatgpt quality

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.039400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.039400Z digest=sha256:f83a37d886e6f5bc82a22363311e3e7db0a2c10f812f3fa0ed1ca6657a98710c

Observation f2ec3af0-cfb1-4232-84d4-06ff68551f05 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective OPT: Open Pre-trained Transformer Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.042291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.042291Z digest=sha256:d1d3a91578f39ae5761166e80515d95801b4f8f97f81cac213660726ece29e8e

Observation a7f6d2ac-d41b-4473-8924-c4bbb317bf08 · outbound

This paper cites GPT-3: Its nature, scope, limits, and consequences.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective GPT-3: Its nature, scope, limits, and consequences

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.044988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.044988Z digest=sha256:9de80ff11c956d99d0aacfb4c9e051bf3442c832dbca3f226a7ac5711aaf7aa4

Observation 36bb6093-3351-473d-84be-da66d9b7fe31 · outbound

This paper cites GPT-4 Technical Report.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective GPT-4 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.047541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.047541Z digest=sha256:fb00edf87c7dda3b904d44f8063c5c37feb3e65558827043311c34015642c325

Observation 78122603-8fd6-4c47-91e7-f392efa82edd · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.050595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.050595Z digest=sha256:a03fa11a71dcbabb24a46c569699536121b9b4f8a46e64a30c8b7a0de590ce94

Observation 48c4d7dc-be62-45ef-827e-c53a22c52833 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.053774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.053774Z digest=sha256:14320447b518d32245180855d58eb2c31ee2f45875b269054d5781f3e9b5a591

Observation 22dd2925-29ef-4c1e-9b24-b3a4edf80ffa · outbound

This paper cites Scaling instruction-finetuned language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Scaling instruction-finetuned language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.057046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.057046Z digest=sha256:6a01dc76b608cdfeda20b38e1b95c508eb1510679478306f7af62a14c334db9e

Observation 294a6f7b-3705-4a1e-b48b-650688d5e136 · outbound

This paper cites Language models are few-shot learners.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Language models are few-shot learners

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.059978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.059978Z digest=sha256:92cf3b27ab7de1ce7a728db79018f32131ba7905864257cac8937a0451c5da69

Observation 033162df-9bbd-4cce-8b43-068b24bc71b9 · outbound

This paper cites GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model [software].

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model [software]

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.062934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.062934Z digest=sha256:f04db04f3ff1310bc06a0e8f88945f475fbb3afba583d934f99bf1f60626b3da

Observation eb6aef89-43f0-47fd-8ff1-2ac67ddfd510 · outbound

This paper cites An empirical analysis of compute-optimal large language model training.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective An empirical analysis of compute-optimal large language model training

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.066175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.066175Z digest=sha256:304c90cda51c12777f225da36299213b22d0774f8440d1c671d3963df98244c9

Observation a961051f-5c53-4d48-9d8f-d55d5f0e04e9 · outbound

This paper cites Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs [software].

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs [software]

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.069197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.069197Z digest=sha256:bf2bb582a5d5fd1ed2bc0def501829b4bad29d51242b75417294441a93a83e69

Observation 3f2584fb-627e-4d84-a5d0-3f4f72252b71 · outbound

This paper cites Releasing 3B and 7B RedPajama-INCITE family of models including base, instruction-tuned & chat models [software].

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Releasing 3B and 7B RedPajama-INCITE family of models including base, instruction-tuned & chat models [software]

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.072233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.072233Z digest=sha256:792e464bc005f2dd2e6bc7869d18f785dc8aa8b0d9e91f5959b7baa393284b5c

Observation f252bce0-b004-404c-9a07-ae4224e21174 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.074970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.074970Z digest=sha256:7e1ed1ab30ca37fa719e23208f2088188c8083a3dbc2e6b0a07ee67c90399bd3

Observation b3c53c0a-839b-4ab3-833d-851c92904839 · outbound

This paper cites GLU Variants Improve Transformer.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective GLU Variants Improve Transformer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.078158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.078158Z digest=sha256:d61fafb36a310f758291c4ababe76e8e7056cbca8537d2b07e78d2e9afff23e1

Observation fb8b7ef8-928d-4190-8167-b2b3c3b0c29c · outbound

This paper cites Prefix-Tuning: Optimizing continuous prompts for generation.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Prefix-Tuning: Optimizing continuous prompts for generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.081565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.081565Z digest=sha256:9ab4a0279d29d0b0cde25a4e7e6dc94f606a752f50a611d8dbde24a948fff2de

Observation 87cccc94-30f8-40fd-9fff-9c2c956b221f · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective The power of scale for parameter-efficient prompt tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.084499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.084499Z digest=sha256:bbe358cc8b6abe82e7c256f844fbf4483d6b14aecd234ae87b610bcef0fabef2

Observation ecf87842-a444-41d5-8f8c-aa1d74bf042d · outbound

This paper cites Parameter-efficient transfer learning for NLP.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Parameter-efficient transfer learning for NLP

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.087547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.087547Z digest=sha256:21699f149a655c53a163260478bf4340430caa6113869c8d6edd20a0eae4950c

Observation eb915d4a-a094-4644-acfc-726a530bac7a · outbound

This paper cites QLoRA: Efficient finetuning of quantized llms.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective QLoRA: Efficient finetuning of quantized llms

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.090638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.090638Z digest=sha256:a7f773e510f30e1e2bdd1e10418caf1461f25506fd81cfe53001ee381662f676

Observation cdf36ff4-ae60-4626-bac8-637268b8744d · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.093575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.093575Z digest=sha256:287d6608929b8745fcb8e1a335147636c58f7ca0913174e32b7e45f046e8ac4d

Observation 2f36c6d8-c326-4c23-ac71-f4480cf7a73e · outbound

This paper cites Visual prompt tuning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Visual prompt tuning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.096753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.096753Z digest=sha256:dcfd39a2a92a3756d73b101bdf238e9b9a5be63c882e89410781f3e113e8a6b1

Observation 4455784a-1e4c-46c3-971c-1ae0d508118e · outbound

This paper cites Exploring versatile generative language model via parameter- efficient transfer learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Exploring versatile generative language model via parameter- efficient transfer learning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.099802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.099802Z digest=sha256:42b0486cc9dbdc67a69c590d26c1e75f5dc1d004eacc5b2f83d5761631f32b04

Observation 22772174-463a-40ff-bfcf-78d869e83a38 · outbound

This paper cites Towards a unified view of parameter-efficient transfer learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Towards a unified view of parameter-efficient transfer learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.102650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.102650Z digest=sha256:ac2af526c10b365b0a0147bfceabb3cc3fe0d235640edc0d8cace7d28dd3f660

Observation 3caed64f-01e2-421f-84d0-f561d203eecd · outbound

This paper cites AdapterSoup: Weight averaging to improve generalization of pretrained language models.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective AdapterSoup: Weight averaging to improve generalization of pretrained language models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.105545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.105545Z digest=sha256:03edc66430619a3bb1fe4aeb3bc04be2bfa9f7de8768663bf1cde2cc9ca7b75f

Observation 61074ac1-edc3-4e83-9567-321dfa50f885 · outbound

This paper cites Multi-Task Deep Neural Networks for Natural Language Understanding.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Multi-Task Deep Neural Networks for Natural Language Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.108448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.108448Z digest=sha256:bbdd934c3d911695fc78514525c4dd83e7e057fc23481bea5a94e353a63ba3e3

Observation 145829e4-82f3-4838-b307-254595175fa7 · outbound

This paper cites Multitask learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Multitask learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.111423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.111423Z digest=sha256:37c4ebd44af1a0bc73a1811c5b4679d1619807c94d9bef3abb4bcdfb9ea065b7

Observation 9c6b41fd-7495-4f55-b266-31b88dff6087 · outbound

This paper cites A survey of multi-task learning in natural language processing: Regarding task relatedness and training methods.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective A survey of multi-task learning in natural language processing: Regarding task relatedness and training methods

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.114409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.114409Z digest=sha256:47a7d7c9a216ce3e3e52d268a7f75cce15170f9c0321c85b66eacfa75b688308

Observation f1b100b5-75c3-4d63-89ee-b4a63864b153 · outbound

This paper cites Multi-Task Learning with LLMs for Implicit Sentiment Analysis: Data-level and Task-level Automatic Weight Learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Multi-Task Learning with LLMs for Implicit Sentiment Analysis: Data-level and Task-level Automatic Weight Learning

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-09T15:04:40.419710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.117223Z digest=sha256:9203bbaadcb346e558d479c553d6c149e40c319d65dd05de21bc6bc89bb98691

Observation 4f8032f2-8fd3-4bd0-bbef-a8eb6f4cb9b8 · outbound

This paper cites Dai, and Quoc V Le.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Dai, and Quoc V Le

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.120222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.120222Z digest=sha256:04351f37a0d0c21185aa9e1ddfce4681213226e1e7fa0d5fc45503bc37f77ac0

Observation 8508d555-5ee9-4d9b-a066-21277a4d4e9e · outbound

This paper cites Training language models to follow instructions with human feedback.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Training language models to follow instructions with human feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.122588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.122588Z digest=sha256:606761e7358e0605d89bde52b69731dd0bd2c47625f19fde452f32a398f8c132

Observation 7df50e87-abea-417d-bff0-23feca8f92de · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Constitutional AI: Harmlessness from AI Feedback

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.124995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.124995Z digest=sha256:08a424f18254ffd7ddcca4ff8ef0709f52b7852931d278d2d67e2172c9e0713f

Observation dcdfa431-b162-47d8-9177-e59c42c01592 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.127769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.127769Z digest=sha256:f14390c92778107e790c0d310def937fa744558292273b5de6d01e21a1266bae

Observation cde0bf21-9b97-4abb-b73d-705f0490cd7e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Proximal Policy Optimization Algorithms

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.130188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.130188Z digest=sha256:6a88c92bca4fe80d79ceb698d4b3df5c7df05d00207a527082fd127c568eb23c

Observation f5d59ee6-009d-4382-943c-f41311cb25d2 · outbound

This paper cites RLHF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional human feedback.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective RLHF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional human feedback

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.132714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.132714Z digest=sha256:4c74f197760c372aa9668169fff60c74034c2f0bfadeaa99507ceabe4f7f8ba9

Observation 470ca58c-b3ba-44ca-9021-e1f81ed16205 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Learning transferable visual models from natural language supervision

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.136023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.136023Z digest=sha256:3e89f75b9666b118ba1c4a75761046a78802b3f674603dc496f9d350d2633e37

Observation a5bf2872-84f5-4092-8062-cc6226d6a5a0 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective An image is worth 16x16 words: Transformers for image recognition at scale

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.138949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.138949Z digest=sha256:475350b4795c2bc6d80d44780f9111fe4e3a7213cad4e984526851e7ed0655db

Observation 61784e65-4a2a-4e55-bcb6-77e2f0f611f4 · outbound

This paper cites High-performance large-scale image recognition without normalization.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective High-performance large-scale image recognition without normalization

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.141909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.141909Z digest=sha256:e8d89b59b6987cce6257731f57fc2180674e6a71df4d7e45e974c26418ce8fd6

Observation 6255f9a5-f207-44ff-be21-1b3092b36b96 · outbound

This paper cites Scaling vision transformers.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Scaling vision transformers

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.088602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.144780Z digest=sha256:34aa24123bc0cbe8e4f3b92b9e4e9989682b5a227d62155768b5ef10c7631742

Observation afb0647e-2f7b-4235-8b34-c99094102e1c · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Reproducible scaling laws for contrastive language-image learning

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.080405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.147567Z digest=sha256:725337ecc7938269cd4f2ed380372d34f592c698a77c2dc8c0a5c84a69645d27

Observation 5a6b4a6b-ebb5-4b10-bc51-6d3a821661ef · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.150536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.150536Z digest=sha256:bedf1535c34e6137dd3fe46f471f3b5bf5838fa513603d81dad6a55cf0fa8e90

Observation 00784461-af2d-4d8b-9519-093e8373a2aa · outbound

This paper cites OpenCLIP [software], July 2021.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective OpenCLIP [software], July 2021

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.072389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.153714Z digest=sha256:e4cf771ab60fa40b2faa9b43e49e5aa00a0e0c6b76ded36924cc59ff2c8449db

Observation 003ac2b0-afb8-40a3-a46c-44fc642ec4c3 · outbound

This paper cites Deep residual learning for image recognition.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Deep residual learning for image recognition

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.064314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.156598Z digest=sha256:d4d7eb6159b66e25a622f9654f9745882d216c1a6edbdc08f2637286be5b32d6

Observation 5c142689-3d6e-4ab9-aa0a-f5ea26551c03 · outbound

This paper cites Linearly mapping from image to text space.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Linearly mapping from image to text space

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.056430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.159739Z digest=sha256:c17a87d143b9b59acc3b5e217b863b24eb180ec9c5d274637687e8a2b66a4a15

Observation db588090-a4db-44ec-9d9e-189bc0681193 · outbound

This paper cites Deformable DETR: Deformable transformers for end-to-end object detection.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Deformable DETR: Deformable transformers for end-to-end object detection

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.048237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.162661Z digest=sha256:a27dd4eac2adec7d04db8f30b27994d31631ded406db5d5d07ef73e8ef577178

Observation dcb2dde3-3c5f-4edb-b9d4-759bb782a0a6 · outbound

This paper cites Aggregated residual transformations for deep neural networks.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Aggregated residual transformations for deep neural networks

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.039690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.165626Z digest=sha256:7ae1756dc7c2f9227b1210afb2b4593e9e8ea11a2b567287c43a84f91faf059b

Observation c7148fc6-e865-4067-9acf-6d075db9a4b0 · outbound

This paper cites Conditional positional encodings for vision transformers.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective Conditional positional encodings for vision transformers

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:41.031431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T15:04:40.168560Z digest=sha256:94933b6988eb50e981bf272ee3b836b4234e9f8fa26cc551dff6d26b37d30caa

Pith citing papers

No inbound Pith citation observations are available.