Pith. sign in

Paper Citation Record · LEDGER

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

As of 23 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.26596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26596 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:40:56.954792Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f88dbb6-2687-4aa1-9f9a-997a3ba183fc · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.926574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.926574Z digest=sha256:dccbd00cf515e2fa4e166c00fcd290bcc9384d0d0e51083bf2fd10b11b860ace

Observation 5d9c196b-f4a7-4bbd-9d2f-052b3c22c11b · outbound

This paper cites Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.935423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.935423Z digest=sha256:b3b9a4f9c26f96e9f18409dd78f81db70c5ae3a6707e47b67abf888295420b4e

Observation 5a61cdb0-9ed5-417f-95a6-a18341fcdc43 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution VILA: On Pre-training for Visual Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.938008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.938008Z digest=sha256:ba3070c276c88f4f78fc76b2d3c84f778768997acc9091b516f2d3c54225590f

Observation 830ae66c-781d-41e5-9326-2c2e2ff03646 · outbound

This paper cites Decoupled Weight Decay Regularization.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.940805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.940805Z digest=sha256:2e161a6792539883ee5deb92bdbd7dfc8eb70121f786e7b63368ebacf270e73c

Observation 3100d93d-404d-4521-a8ec-ab1fecd3994b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.943390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.943390Z digest=sha256:9764137dd8149ee2231bfd5119eb085eca6ade5115fd99540dd43429b5f83108

Observation 790da517-9b0a-423b-a228-e98ba5390d4b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.945948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.945948Z digest=sha256:043b3bfa9dd4b9bf0f51d377dd1125854caf2d6d7d9a82d1f8f0fa29ba2eea18

Observation 25be7dd9-6dfe-479e-a1ac-7ed684938b88 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.949309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.949309Z digest=sha256:c97bf0521be7419066f26c817aedfcb7c8b4c845012c11b3462f623df584d189

Observation 00355b0e-bfab-4df1-8d82-d3624c9629dc · outbound

This paper cites FLGo: A Fully Customizable Federated Learning Platform.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution FLGo: A Fully Customizable Federated Learning Platform

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.951963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.951963Z digest=sha256:9b868861e4be29157b1922ec7c3de3724be74c2c3f5495294bb07d84ed1cf550

Observation abbab7a4-eef6-4abe-8d74-a2efe25f6549 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.954792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.954792Z digest=sha256:12adb3b527195008458577b298ca046f715fa22ee2f0d0f2f6243ad09482d209

Observation a2d46f7d-6730-4e90-80dc-490bfbf64241 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.887890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.887890Z digest=sha256:2584fc41a79d1efa07ee081e28ed363b45960d547fa65dbdc685a135994b8893

Observation 6facfb72-b996-4426-bca3-9d5f5b3d3a2b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution LoRA: Low-Rank Adaptation of Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.929777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.929777Z digest=sha256:eaeef4865a7427fe3013db153aca66a970a72f618108d87fee90e2518affe54b

Observation 4e6b62ce-890d-4902-a46d-f722d15d2379 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.858304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.858304Z digest=sha256:234bb801384912fc1607466bef78f35544c79acaddae287018604a8e5d3751ae

Observation 7f387456-358b-4518-b435-c8a1f63056a1 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.871061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.871061Z digest=sha256:8de55c0dc6d70bf40459f2398ba147dd8906c6c8f5f8588d30ffff9b57b9f697

Observation 2c30a087-2d53-4d67-ad03-b22958e7e088 · outbound

This paper cites DReSS: Data-driven Regularized Structured Streamlining for Large Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution DReSS: Data-driven Regularized Structured Streamlining for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.907968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.907968Z digest=sha256:0da84e7539996adc3dddf62e24518f16d8d72283a88d651b6ba4be5d990446ea

Observation b1902258-08e3-4dc7-bbb3-d3a19e07e521 · outbound

This paper cites Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.932940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.932940Z digest=sha256:58209087743cf90c7e321a6a996348a7edbbf0651dbbc28a82d266d312e949b0

Pith citing papers

No inbound Pith citation observations are available.