Pith. sign in

Paper Citation Record · LEDGER

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.22074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22074 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:04:59.948955Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eed6f0b8-80c6-4573-b526-a83907fc9470 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.854687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.854687Z digest=sha256:5856a59b0af340d1233ba6964192dda8edee3e75bc8d674d8ed42e9b4d5c58cc

Observation a7e6307e-cf6c-4465-a6c8-7cd9e00a4038 · outbound

This paper cites Visual in-context learning for large vision-language models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual in-context learning for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.859067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.859067Z digest=sha256:284f90362457bf53a91f738f9d1ba142baefd297c67d26ce0ae4ba54cd9efe7e

Observation daf19fc4-d5db-48f6-9db2-a38e434b1910 · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.862849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.862849Z digest=sha256:0996e4d79161f62bc52d6c1fb8c872fbc3ecebe8eaceb99eef08ceb4389c7faa

Observation e6de482d-ebe9-4236-9aca-1f6b4e645aba · outbound

This paper cites ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.867055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.867055Z digest=sha256:89c9fecddb94e942c0100a1fb8704643bfe588025e9dc8dc187eea59a49d6fcc

Observation 84520e5a-e62f-4d17-8039-a4c51d0e088a · outbound

This paper cites Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.870622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.870622Z digest=sha256:24aad6b65454a2344c77b4f3351ff27cd5233979b56c5ef916ba467e6b2f251c

Observation 76e8ea03-cdaf-457b-8e0d-efae0b1fc29e · outbound

This paper cites What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.343521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.874506Z digest=sha256:9da48f3c11429d28085b2a687453b5b1d5a22f6ddb6fd430c3e36a69c9e17e42

Observation aece26ed-97b7-41bb-a991-f1d456d6b06c · outbound

This paper cites A review of multi-modal large language and vision models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A review of multi-modal large language and vision models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.331761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.878977Z digest=sha256:ec70f877e51f1be0d22204f909f959317c00a4af68ece8ebab4c07743c64255e

Observation 84fc1e5c-0a1e-4554-90ce-54a2f1285286 · outbound

This paper cites Visual instruction tuning towards general-purpose multimodal model: A survey,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual instruction tuning towards general-purpose multimodal model: A survey,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.319186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.883186Z digest=sha256:baef3385700587a7c7c61ffa5c56de082a5b6db9b360f696e5dd7e378094943f

Observation 8ed50742-4e36-44c9-b86f-ad6ce0f8ac30 · outbound

This paper cites Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.306541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.887095Z digest=sha256:19bc08ebd0769740ddf44a430359c653de8e8f946ec91e12270f74ba16ace28a

Observation 4520a800-44e8-4f19-8a52-ca216c3f6b64 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Sharegpt4v: Improving large multi-modal models with better captions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.294680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.890472Z digest=sha256:0c04e2f4e92d01f3fbea5a66c0a7dd2489e40abb25856d30683e8f2535ff67e0

Observation 1103b626-32c0-4f45-913b-80a277aa2a3c · outbound

This paper cites Multi-modal large language models are effective vision learners,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Multi-modal large language models are effective vision learners,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.894097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.894097Z digest=sha256:75a85aadab0217bb5e227f7463572dda827fa03991d6ea476f8a202d57d4e913

Observation c8b58121-f346-4249-8102-92961db53fe3 · outbound

This paper cites Modeling event-pair relations in external knowledge graphs for script reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Modeling event-pair relations in external knowledge graphs for script reasoning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.897499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.897499Z digest=sha256:03d46a45fb7dc86f7cecc34dc63a3c19f9440581a19dc56c43944e4c4636cbf6

Observation 06954d07-b834-4b41-a2bc-79ff76cfe250 · outbound

This paper cites Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.900566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.900566Z digest=sha256:049d1c848ea92b0173dd3ef7b90c4bbd5417928a467dfab74344ac107f44ad42

Observation f9d596e8-1976-4f7d-b1a8-6ad16e1508e8 · outbound

This paper cites Eventbert: A pre- trained model for event correlation reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Eventbert: A pre- trained model for event correlation reasoning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.903623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.903623Z digest=sha256:29323c507b602c5394d84d7fce496806a605f1e63b26f33618300ad04f9039bd

Observation c9e8263d-7004-46bd-b587-ee7ad934f3ec · outbound

This paper cites UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.249670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.906812Z digest=sha256:83f03525a07098be72b88b1d9631461b5a71c656ddd1659bde6726b042280eae

Observation aff49c02-5ff6-4a94-981a-7280134854c4 · outbound

This paper cites Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.910308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.910308Z digest=sha256:2a45cdce2279753d12f7dc0fd7ae870a88625b8f97c6da197caf4c687c4472ce

Observation fb191429-c155-4b43-a552-7a70ab5fff99 · outbound

This paper cites Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.238479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.913550Z digest=sha256:b20ffc3fbb36e52075dc3ab64c58a6cf55d2b7a4b0d68cd318e4079defdd045b

Observation 99b3907a-3287-4daa-9672-4f4f4f0bcb4f · outbound

This paper cites Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.213765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.920890Z digest=sha256:611ed42b4c42e614578d829de99a9cbdc22d1cfa14cd95c8449311a51f5800ed

Observation a92279bc-8044-4e9f-8ccd-e3675463f366 · outbound

This paper cites Adareasoner: Adaptive reasoning enables more flexible thinking,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Adareasoner: Adaptive reasoning enables more flexible thinking,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.199911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.924161Z digest=sha256:de3a1c23ec28ec7c06133a07b5e0378d612f100ce6b1f25a909772f87e050361

Observation ce37a8b1-cc62-422b-8209-403a76c8601c · outbound

This paper cites Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.187066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.927819Z digest=sha256:47e58145fc396e53fea7a779c7f5c23612b39a44608e79a159f4f7aa1cf9a000

Observation d3d7508e-0496-4981-9938-8fcff68440a5 · outbound

This paper cites Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.176327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.931589Z digest=sha256:cc44ea805015192b10e7919bf8dd9c1fa625cd5402f9f240d10caa2cfdcc439a

Observation a7b3811f-b8fb-4bc5-911d-e094928adcc5 · outbound

This paper cites Agent-r: Training language model agents to reflect via iterative self-training,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Agent-r: Training language model agents to reflect via iterative self-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.165443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.935334Z digest=sha256:c3d67092bb687e5914e2616add53e3da9f29d6e5d0c093a28b380af34701e472

Observation 2febc3bb-9713-491c-8298-99b41b4f4f2a · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.938273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.938273Z digest=sha256:6bb80ffbc96cbd079c486ba01d91b2824b9c5a77a391a73d9598f3e9d3bb840a

Observation 5392fb20-1382-4980-9e64-c6eacd653503 · outbound

This paper cites A theoretical under- standing of self-correction through in-context alignment,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A theoretical under- standing of self-correction through in-context alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.153345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.941550Z digest=sha256:22a251214fa6d4e22db89314b0867d96beb76bb11ae5e30bb09f0a766d9b69ce

Observation 45c89ec6-7aed-4a28-8da7-873179191f71 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.945292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.945292Z digest=sha256:4cc83bce0ab10d4e818f95577cc55f73c4d4eed3e3ae03f9a123f5f27bb43850

Observation bcd54209-dbad-47dc-ba69-c710818bff93 · outbound

This paper cites Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.142180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.948955Z digest=sha256:f2c2ef692e51d4cb8e0ef81f6e6fd71f6a965520ccb7cf19c086066f12b04c61

Observation 54312aed-0e1b-4ac1-a559-fe6981de1f74 · outbound

This paper cites 5014–5035.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs 5014–5035

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.225719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:04:59.917050Z digest=sha256:91c2eecd4e406bbcb42727b275224327992079eedaf30d3d6260b4db4ed76cfd

Pith citing papers

No inbound Pith citation observations are available.