Pith. sign in

Paper Citation Record · LEDGER

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

As of 22 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.22074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22074 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:04:59.948955Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eed6f0b8-80c6-4573-b526-a83907fc9470 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.854687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.854687Z digest=sha256:7fbc9f9d961129a26b11323fd1cce72f2d0c1289f7bb73f224c83533840989c2

Observation a7e6307e-cf6c-4465-a6c8-7cd9e00a4038 · outbound

This paper cites Visual in-context learning for large vision-language models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual in-context learning for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.859067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.859067Z digest=sha256:86f044efb1b0bf4c9cba9f5c1760c25f003dbd35537c25bda709fa007360fbef

Observation daf19fc4-d5db-48f6-9db2-a38e434b1910 · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.862849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.862849Z digest=sha256:032c68d7cbbe21f8205cce8e36b9b6120003f1d3baa516e364bbb65e24186646

Observation e6de482d-ebe9-4236-9aca-1f6b4e645aba · outbound

This paper cites ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.867055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.867055Z digest=sha256:0227282483177ca31adb8acfd45212df7db11f0ad3d97283539bf84ad89cf41f

Observation 84520e5a-e62f-4d17-8039-a4c51d0e088a · outbound

This paper cites Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.870622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.870622Z digest=sha256:17256968beb7de40c0043f7ccd7b8f41b5acd8d7e0210a3b7c86ee3f585c6571

Observation 76e8ea03-cdaf-457b-8e0d-efae0b1fc29e · outbound

This paper cites What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.343521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.874506Z digest=sha256:81745c378b75a388cfc9d99d0b8099dd23a6d1db3a08df76600f81b61196a709

Observation aece26ed-97b7-41bb-a991-f1d456d6b06c · outbound

This paper cites A review of multi-modal large language and vision models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A review of multi-modal large language and vision models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.331761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.878977Z digest=sha256:b5db18a5bd903ee7b0621ae4d45e3f7a69df539596f80ddbe8003fe26fb4bed8

Observation 84fc1e5c-0a1e-4554-90ce-54a2f1285286 · outbound

This paper cites Visual instruction tuning towards general-purpose multimodal model: A survey,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual instruction tuning towards general-purpose multimodal model: A survey,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.319186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.883186Z digest=sha256:ebe2618e9d68a933ee0cdf7523254f6c2a1c333eb68096fbab16b602a00218ad

Observation 8ed50742-4e36-44c9-b86f-ad6ce0f8ac30 · outbound

This paper cites Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.306541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.887095Z digest=sha256:abbd7c253fda3311c69d32ab1cb05a6addbf12b176565bdcd5ae7979aa639484

Observation 4520a800-44e8-4f19-8a52-ca216c3f6b64 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Sharegpt4v: Improving large multi-modal models with better captions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.294680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.890472Z digest=sha256:ac9ce177a2c1506f0ff704644f252310d6d49c88a4cabf2e2740c5d4967e1355

Observation 1103b626-32c0-4f45-913b-80a277aa2a3c · outbound

This paper cites Multi-modal large language models are effective vision learners,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Multi-modal large language models are effective vision learners,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.894097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.894097Z digest=sha256:be3fe4394d678cc9f50f2ec795d1b6d93cdbdd7711ca0f8f199f356164d31cc7

Observation c8b58121-f346-4249-8102-92961db53fe3 · outbound

This paper cites Modeling event-pair relations in external knowledge graphs for script reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Modeling event-pair relations in external knowledge graphs for script reasoning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.897499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.897499Z digest=sha256:3ffaf9bf903b567795bf78bd537ec8360355405ec2b5811f1fb01119cdaea51d

Observation 06954d07-b834-4b41-a2bc-79ff76cfe250 · outbound

This paper cites Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.900566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.900566Z digest=sha256:30d2d48614a84c590d3fb9d9ac4a40cf3b0bd376fd9b13293bcab96fd1ce2b71

Observation f9d596e8-1976-4f7d-b1a8-6ad16e1508e8 · outbound

This paper cites Eventbert: A pre- trained model for event correlation reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Eventbert: A pre- trained model for event correlation reasoning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.903623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.903623Z digest=sha256:1a81c54dd3be5e50a1d73143dd5baed4439850364527f5778809811230ed94c3

Observation c9e8263d-7004-46bd-b587-ee7ad934f3ec · outbound

This paper cites UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.249670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.906812Z digest=sha256:f6519f407a384a364085e1fcc4a0b3f23fe7308c7922dd1ccf28ff26c88a26a1

Observation aff49c02-5ff6-4a94-981a-7280134854c4 · outbound

This paper cites Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.910308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.910308Z digest=sha256:0f7b934b4cdfc4f71be90dbc82ece175fdd0991fa425cbe2e48536554a9fc3cf

Observation fb191429-c155-4b43-a552-7a70ab5fff99 · outbound

This paper cites Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.238479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.913550Z digest=sha256:a72a010e23cd2f1f4b458ce5a3e2c00d2440722001420315f207cca0f06c60b0

Observation 99b3907a-3287-4daa-9672-4f4f4f0bcb4f · outbound

This paper cites Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.213765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.920890Z digest=sha256:b947ec20c33054697f7f9359688e257bf4f86696b5505599a629dea47796bac8

Observation a92279bc-8044-4e9f-8ccd-e3675463f366 · outbound

This paper cites Adareasoner: Adaptive reasoning enables more flexible thinking,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Adareasoner: Adaptive reasoning enables more flexible thinking,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.199911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.924161Z digest=sha256:9963ec0965b46b637dbd5a2d75e9ffe876ca536ff692ed2fd2a064f567451c00

Observation ce37a8b1-cc62-422b-8209-403a76c8601c · outbound

This paper cites Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.187066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.927819Z digest=sha256:e0715d841d0335d524007026f627907ba4fa5853ffd93d33beea34f85b847ec3

Observation d3d7508e-0496-4981-9938-8fcff68440a5 · outbound

This paper cites Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.176327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.931589Z digest=sha256:00aa8d2d2bc6d4e415b27201e0fc5f0e6e91c4ca72c141acfe8b4a774f10ee3f

Observation a7b3811f-b8fb-4bc5-911d-e094928adcc5 · outbound

This paper cites Agent-r: Training language model agents to reflect via iterative self-training,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Agent-r: Training language model agents to reflect via iterative self-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.165443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.935334Z digest=sha256:b854f982f1307833afc5371b38f56a1cd19a66d05a9115018ccfa4b79d8966fe

Observation 2febc3bb-9713-491c-8298-99b41b4f4f2a · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.938273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.938273Z digest=sha256:197a8e40ef8b0073975fe4d6ce741bc206fb930c19deb953d27c5bc20a4294ff

Observation 5392fb20-1382-4980-9e64-c6eacd653503 · outbound

This paper cites A theoretical under- standing of self-correction through in-context alignment,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A theoretical under- standing of self-correction through in-context alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.153345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.941550Z digest=sha256:82a8d87b0630a524f886d3c22b63745e8135004b6e22ff1013641e17e766f673

Observation 45c89ec6-7aed-4a28-8da7-873179191f71 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.945292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.945292Z digest=sha256:3d57ca88adf6ff0edbe8189a47848a0313294c68c7bdcbac6b48234f2f820f88

Observation bcd54209-dbad-47dc-ba69-c710818bff93 · outbound

This paper cites Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.142180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.948955Z digest=sha256:4d432fbdb39a7506cee4f3651714f6fb1fbd14efee2cb886ea762f37ddb06b13

Observation 54312aed-0e1b-4ac1-a559-fe6981de1f74 · outbound

This paper cites 5014–5035.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs 5014–5035

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.225719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:04:59.917050Z digest=sha256:950e66959c81abc4521c66e5fed004605659836a4cb97576c2564a378797553b

Pith citing papers

No inbound Pith citation observations are available.