Pith. sign in

Paper Citation Record · LEDGER

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning

As of 5 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2607.22013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22013 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:08:22.631516Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:08:19.928596Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 572dfe5e-3e16-494d-8acc-a9815007995b · outbound

This paper cites Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:19.928596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:19.928596Z digest=sha256:dc66009b8243d0b4132dd675ca1dd6601464b601235068995e0a11dddb351302

Observation bf720631-acfe-444d-87f2-30a7480b2c1b · outbound

This paper cites an unresolved cited work.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:20.198809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:20.198809Z digest=sha256:33ba78b0e92ea1b23449a4b449d8de67f57d60045705450e3f2ec6698aae9d5c

Observation af4ddf06-7860-4887-9d81-91297cb5a3a7 · outbound

This paper cites Experiments Settings Dataset.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Experiments Settings Dataset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:20.311163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:20.311163Z digest=sha256:79f4969966930bfaf42e8aab7444957a372557c36f29fce818df4eee2d8af918

Observation a3f7c85f-0fd3-48f5-a645-97c253b8a848 · outbound

This paper cites Existing methods often blur tiny cross-modal cues.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Existing methods often blur tiny cross-modal cues

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:20.727904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:20.727904Z digest=sha256:55495add375aff134e7da442548b0621e3ac61d8c38d66d16f3834471ac104a0

Observation 7aafa06f-d25c-4d6d-b9ed-4123f11dfd77 · outbound

This paper cites Hyperparametersαandβwere 0.1 and 0.2.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Hyperparametersαandβwere 0.1 and 0.2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:20.487322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:20.487322Z digest=sha256:1f710f1d9b0b36aec2030d31b96eebc02cfa31be03fcdf6e597ad155d826ac63

Observation 26f5203e-888a-40b7-8469-4ed0cbe47298 · outbound

This paper cites Joint multimodal entity-relation extraction based on edge- enhanced graph alignment network and word-pair rela- tion tagging,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Joint multimodal entity-relation extraction based on edge- enhanced graph alignment network and word-pair rela- tion tagging,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.478404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.478404Z digest=sha256:79003f8efc58d0d33ca790d081a997c75cecfdb1fc2c0463b1d440d3f00505a0

Observation 6cec807d-f14e-4ae5-b327-0d8017c0d173 · outbound

This paper cites 61966038 and 62266051, and the Postgraduate Research and Innovation Foundation of Yunnan University under Grant No.KC-252513133.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning 61966038 and 62266051, and the Postgraduate Research and Innovation Foundation of Yunnan University under Grant No.KC-252513133

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:20.891583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:20.891583Z digest=sha256:7999755e1f17e7e2373724160d5900d17d351cd4726fbfaa2ee8f48a022ccf42

Observation 05c685ae-b72b-4f81-a279-915143eb7f4d · outbound

This paper cites GPT-4o System Card.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.020971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.020971Z digest=sha256:721875e6bf81b07fb9d67d40beb8fd5e53802d6c5108008946cbc5419619c907

Observation 4473e0c0-249e-4700-9d3d-2a9c925875b7 · outbound

This paper cites Visual instruction tuning,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Visual instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.179902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.179902Z digest=sha256:7dc406452af126edb1c2953073eb67d4b9c136a690ed20071b80915b09989f15

Observation ce604d35-73c2-4938-8a38-085c05f2c630 · outbound

This paper cites an unresolved cited work.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:20.063892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:20.063892Z digest=sha256:bd3b31a74dd38c438798ae17c0e6565b0ad80c6f294c3af3a5c9310b548feaa7

Observation 5780e02f-3a99-4197-8e69-19c607a59e3c · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.271702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.271702Z digest=sha256:b9599f4722ae3c50d1bfae4f5351a67feda8815271fe6a86e58d95b85f9dd84b

Observation 2a635ebe-4c11-456f-8467-c5d86fc0f527 · outbound

This paper cites Boosting the power of small multimodal reason- ing models to match larger models with self-consistency training,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Boosting the power of small multimodal reason- ing models to match larger models with self-consistency training,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.355410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.355410Z digest=sha256:b01ac22c4e5407d9d4f9efdb25b99fdb7c8ebd47257996b6b0efd0481dc42a4e

Observation e07f9f1d-e045-476a-bfe1-e1187881bd0c · outbound

This paper cites Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.408899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.408899Z digest=sha256:af818d4dd5b110ca47d8de66801edfbe1866b02d7cee6f8c234c916e54ebedc8

Observation 80c5a806-f999-40e2-880f-2d9a7e8617d5 · outbound

This paper cites Enhancing semantics in multimodal chain of thought via soft negative sampling,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Enhancing semantics in multimodal chain of thought via soft negative sampling,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.567687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.567687Z digest=sha256:243ce2002863d40424ab357a506a47c0247bb91914b81c5b2558b394269408c9

Observation 9604640a-1b92-4aad-8055-c052569babd8 · outbound

This paper cites Enhancing human-like multimodal reasoning: a new challenging dataset and comprehensive frame- work,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Enhancing human-like multimodal reasoning: a new challenging dataset and comprehensive frame- work,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.638785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.638785Z digest=sha256:228b1aed57552ca94c08aa71ba823c7c529580a763081f0ecf8ed3a9cf11a272

Observation 2243f7c4-c996-46b4-9ddc-09930bbadad1 · outbound

This paper cites Learn to explain: Multimodal rea- soning via thought chains for science question answer- ing,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Learn to explain: Multimodal rea- soning via thought chains for science question answer- ing,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.733047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.733047Z digest=sha256:43fc986b4558fcaff3cab4605e27613adaf2a7dc6ebd4f5f0cd296e0558f497f

Observation ea741cca-4a97-44d7-9790-beda0cfcf815 · outbound

This paper cites M 3cot: A novel bench- mark for multi-domain multi-step multi-modal chain-of- thought,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning M 3cot: A novel bench- mark for multi-domain multi-step multi-modal chain-of- thought,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.795830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.795830Z digest=sha256:6153f7dd5e5521b62ab7c5443e5c736392dc1e67bb62695fa44e19cdf298922f

Observation 244eeeca-a03c-4d74-9700-f779a3922c41 · outbound

This paper cites UnifiedQA: Crossing Format Boundaries With a Single QA System.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning UnifiedQA: Crossing Format Boundaries With a Single QA System

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.860620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.860620Z digest=sha256:6abced714d1577a93d355e9c7787691e81f2eccc42d1130d525d1bd766ab36bd

Observation c45c1e88-c51d-40a4-a3ee-fc9f0ac27808 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Chameleon: Plug-and-play compositional reasoning with large language models,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:21.993735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:21.993735Z digest=sha256:53704b75f971845051238176feedadac85f6ab0ebd0709da635dbbff582b1396

Observation bd4a26c1-2b41-499d-88cb-640293d72da2 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:22.141248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:22.141248Z digest=sha256:da0a670ec642db6125f0d1d9f8b8a4308b9010f9994a6781de2cebf8043a5f6c

Observation 3a77c7c5-ada7-437b-91a1-9d8aa1a48739 · outbound

This paper cites IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:22.246100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:22.246100Z digest=sha256:49764ffafac39e31ca26168cedf8e4274a0f0b09aaf9635eee0a73a79383a474

Observation 959dd084-546c-4c63-8354-82e9659ea67d · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:22.347900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:22.347900Z digest=sha256:697ed75e9b055801d33f5666696c2796d6c2c9d578c52de5c11dd7a2718af8bd

Observation d06f4b06-904b-40ba-835e-e0e2ea92bff7 · outbound

This paper cites Cheap and quick: Ef- ficient vision-language instruction tuning for large lan- guage models,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Cheap and quick: Ef- ficient vision-language instruction tuning for large lan- guage models,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:22.497615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:22.497615Z digest=sha256:ebac2dc9e11e052160174f00f0d64ba86967fdc77eb13167f0764b1c58d49600

Observation 23698837-a64a-4597-b497-e49a11a5adb9 · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language mod- els,.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language mod- els,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:22.631516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:22.631516Z digest=sha256:05c88017d130b4d415a9d7ad3e9f502b8b208fad378ef79425516409f7945648

Pith citing papers

Observation 572dfe5e-3e16-494d-8acc-a9815007995b · inbound

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning cites this paper.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:19.928596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:19.928596Z digest=sha256:dc66009b8243d0b4132dd675ca1dd6601464b601235068995e0a11dddb351302