Pith. sign in

Paper Citation Record · LEDGER

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2506.07227.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07227 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:48.054205Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2de8716a-b311-4ff6-96fe-0cdd660016d0 · outbound

This paper cites URL https://api.semanticscholar.org/CorpusID:276612236.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning URL https://api.semanticscholar.org/CorpusID:276612236

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:59.507971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:38.608937Z digest=sha256:7f2ef70d78ee52d3f331775f1b49260a20fde78de7ca42d65ea97135eec53eb0

Observation bca26de9-f16f-4a11-a98c-8a064ec72ebc · outbound

This paper cites GPT-4 Technical Report.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.711206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.711206Z digest=sha256:bb4d388963eecb4478e0bdadaf5c2e1966da7ecf72d0ac4469efc37a4ed21524

Observation d6ad0563-42e0-4b3e-a964-2444405eca28 · outbound

This paper cites Qwen2.5-VL Technical Report.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.884823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.884823Z digest=sha256:b0e6153df452b79cb66428a8f38295efaaf4cdc97eb391c40f5829bd9ab23d90

Observation 2a3bf250-33e3-4e51-84b1-e5994c2a80f7 · outbound

This paper cites Hallucination of multimodal large language models: A survey .CoRR, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Hallucination of multimodal large language models: A survey .CoRR, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:59.171317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.041996Z digest=sha256:4d0a8c8a3ca58ce909d56671c1b6825bc9a006e4ce88520c5a539cd916d99640

Observation 423e27ac-0282-4dda-9410-db59986b3274 · outbound

This paper cites Flux.1 fill [dev].

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flux.1 fill [dev]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.861034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.254218Z digest=sha256:07b5167e027d2fa17cfdbcb30b983a4e426f672b765b942bce11efe9e5ee68a2

Observation 5ecf6e1c-d64d-4ff5-ad5e-c9aecd056bb3 · outbound

This paper cites Ledits++: Limitless image editing using text-to-image models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Ledits++: Limitless image editing using text-to-image models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.484131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.375131Z digest=sha256:b998b40f9b497ba7142357395f3d36714c48942173c2aecf98566129a791064e

Observation 662d1733-67bc-49c9-991b-986ec3bad760 · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Instructpix2pix: Learning to follow image editing instructions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.126442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.449044Z digest=sha256:2cea16dfa7b8d8016359d0aea61e9ff4473696b99ceba8176b277bf535d8b0ab

Observation c0ed1ce1-9362-47b4-9348-0c32bc4c05d5 · outbound

This paper cites The revolution of multimodal large language models: a survey .arXiv preprint arXiv:2402.12451, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning The revolution of multimodal large language models: a survey .arXiv preprint arXiv:2402.12451, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.543224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.543224Z digest=sha256:5f2c6afd3503098a2cf8b6b59af53bcdf0e220fa1afb5785d31f5b721c6a3d18

Observation 06591c43-d4d1-42e2-bec5-d356c2e0dd67 · outbound

This paper cites Vlmimic: Vision language models are visual imitation learner for fine-grained actions.Advances in Neural Information Processing Systems, 37:77860–77887, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Vlmimic: Vision language models are visual imitation learner for fine-grained actions.Advances in Neural Information Processing Systems, 37:77860–77887, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.810907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.598014Z digest=sha256:0a9b582833ff48add1cba80f6c2888709ba45246a47ccc13b52d288df7db73d3

Observation e6a69b11-b329-4a2a-83ce-f3fd8ebd5ac0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.736678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.736678Z digest=sha256:080c75e1e7e8e70582f7be815b2019710d70d30d26662d54a92d9b8df45781d4

Observation c46b15e7-cf14-4543-9487-ea9126ad7257 · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Anydoor: Zero-shot object-level image customization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.415831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.854427Z digest=sha256:c59b1d4f0933ff65b6b770567994f9775ce1bc16e1490ba2672be745d3f9125c

Observation 30ddf113-d876-428b-b26e-898c2a7e035f · outbound

This paper cites UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.028036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.028036Z digest=sha256:555f8e80b182b93897de29ecdd835a3378d33aa96d334a935a16f74be55efe7b

Observation 9c6eee60-6280-4765-813f-2a5265c515b8 · outbound

This paper cites Diffusion self-guidance for controllable image generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Diffusion self-guidance for controllable image generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.204373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.204373Z digest=sha256:59e67ec3df337fe0d8960b40e1402a3fa8a8659235fbc584b5458b97b62cfa80

Observation ceaae5c7-6498-4941-8dce-05f664adf7a8 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.351303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.351303Z digest=sha256:674bc7085814a0e937909a1714784abf4d554fdcbdab3a32d612415f0db7d73c

Observation 2dd79992-184b-4994-89f0-007c48796424 · outbound

This paper cites TLDR: Token-Level Detective Reward Model for Large Vision Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning TLDR: Token-Level Detective Reward Model for Large Vision Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:43:48.602722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:40.500892Z digest=sha256:fddfa83cdace4ed554e5a417439ff3670239a1d8aa7adddeb920336b78f55285

Observation 76f01343-d522-4ef2-82d9-fee5a2b553c0 · outbound

This paper cites Towards a benchmark of multimodal large language models for industrial engineering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Towards a benchmark of multimodal large language models for industrial engineering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.107153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:40.648100Z digest=sha256:a897152eaac64f2f52aca58e255b414d50ffd696ce3fc636fb9afd47742fd8e2

Observation 76a6500e-94ff-48d3-be82-8f5016ad95ea · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.719981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:40.743202Z digest=sha256:9a69d03317b859040f7e60100d577650316804646e1d3f66ef0b793190b0a31e

Observation a92043ad-40b2-42c7-8553-5c2f0f6a01c3 · outbound

This paper cites Gemini 2.5 flash.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Gemini 2.5 flash

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.397687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:40.943773Z digest=sha256:bb59cbb51461ca4a5a022fb20291a47b91737387d42572ef646919904410165e

Observation dee98651-ea66-4bc2-986c-1b7b5327c82f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.065207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:41.114382Z digest=sha256:a851da1cdfde4b5ddf607be2fc4ccd1d8afc44989d994582d3289db30adb7106

Observation e6543063-4d93-461e-a0b3-51ece3c23a4e · outbound

This paper cites The Llama 3 Herd of Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.230598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.230598Z digest=sha256:284bb21f1312c06a0c347c5d9d44b383c2b05885e130a68c469338ed36c7fecc

Observation d1a55e4f-cf1d-44ad-b42f-285b3a768821 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:55.619250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:41.386418Z digest=sha256:8a108de0bb94e5acf1b99c43bc9cb75cef1268793f7d29b3cfb319c51a457f36

Observation 6ac500f9-93a6-4298-a548-a3f564b985ed · outbound

This paper cites Diffusion Model-Based Image Editing: A Survey.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Diffusion Model-Based Image Editing: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.528747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.528747Z digest=sha256:4225d7d23d798cc13a6ccc056e9a98169123d939faca13eccf4c452874ee3146

Observation 047225bb-e58e-49e4-ad4b-298873d44d7f · outbound

This paper cites Smartedit: Exploring complex instruction-based image editing with multimodal large language models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Smartedit: Exploring complex instruction-based image editing with multimodal large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:55.263948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:41.681296Z digest=sha256:377fb352f91eeecfba924beb783a5acd13c1c136fdcbd690fee406227934424c

Observation d8922cbf-9b68-4df8-81f1-1ee09e87f00a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.928422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:41.824325Z digest=sha256:bf9c910d9b93586db913db99380852316f4962a161ed0871be4aa5a8ce8de132

Observation 5b9c44cb-39a4-488a-9d8b-3b679c278984 · outbound

This paper cites GPT-4o System Card.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.006009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.006009Z digest=sha256:5f3511f3da24f2ca395882e0525e37f8858fa726d2dd9492a2157f510323f3b2

Observation d0a575a1-2eee-43b7-9f09-029c7d9a1366 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning A style-based generator architecture for generative adversarial networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.551988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:42.126473Z digest=sha256:b98d9ea6d42b04c7016fd8655029e27068eaf6585347798646c2d0a7eeb1214a

Observation 555ec626-1b0d-4dc4-915a-bc6c28fc7849 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.310898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.310898Z digest=sha256:afe1892e8d46cca4b2744970908f9d0110f77493c92b5622afc0c1f2da453a7a

Observation 2e22816e-3bbe-49f3-a114-f95cd95da233 · outbound

This paper cites Evaluating and improving compositional text-to-visual generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating and improving compositional text-to-visual generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.203404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:42.448452Z digest=sha256:29ab391b88e27b6b7803136c833899990e8d4d04f211443523deb565f7a9d307

Observation d0a127ad-0c9a-430a-b980-f740fee99f64 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.594952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.594952Z digest=sha256:ddf1871e5f1f5a630e26b088f1e67e9d665d0d47a9c848419da18904b447d297

Observation e5c16a2e-1828-48e6-a089-e465466d583e · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating text-to-visual generation with image-to-text generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.827596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:42.736228Z digest=sha256:38b8cc01f5deb7e7080d20a7cd26527b17838b37f0b5f9557851fdc9694ff8b6

Observation 01af5ce7-54ed-4e9c-807b-54fd8acf05e5 · outbound

This paper cites Flow Matching for Generative Modeling.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.883509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.883509Z digest=sha256:f456c917d31546ae7d4224a5ad63c47a9d566474665467c153220522ce877c8b

Observation 27be5de5-b3ff-4872-a905-992872b5b4d5 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Improved baselines with visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.066529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.066529Z digest=sha256:5bf62ef077a67e5f66d4be8d9d99b39d78dd911aa1c575c05dbbf5ec64e54b9a

Observation 848a6b9b-f94d-49a4-889d-173ff694dbe8 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Step1X-Edit: A Practical Framework for General Image Editing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.205553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.205553Z digest=sha256:7be5c04d7bd1e94e0db56e867da5af39dec0f3f55be82330f8a11196b57545a8

Observation af6e4322-9d20-45ca-b28b-218f0f494c27 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.306278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.306278Z digest=sha256:7abd9022c8b286f5b1b7127a0350ab4dfdf8a5b1f7428ad75b6b955a7d1ecdb2

Observation fefda47f-026b-4584-8273-f95ab4f3f7a0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.466271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:43.488660Z digest=sha256:775f8f0faeddf108cd9ed94585ad6d30dacbf30e6128e15cc7356134c4921824

Observation cd9cbec6-efba-4ef9-8ee6-195b3fe737c3 · outbound

This paper cites Customizable image synthesis with multiple subjects.Advances in neural information processing systems, 36:57500–57519, 2023.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Customizable image synthesis with multiple subjects.Advances in neural information processing systems, 36:57500–57519, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.106964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:43.656077Z digest=sha256:aa2a3e0d0b7feccea850898a39f6e1e7db3c418403638f39f06508acaa6cb45a

Observation ad8ec765-f698-4aa8-8ebc-9c83b96732ad · outbound

This paper cites MagicQuill: An Intelligent Interactive Image Editing System.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MagicQuill: An Intelligent Interactive Image Editing System

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.759971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.759971Z digest=sha256:9d4b2e2587fb11f946341d1888abeb4dcf6aaeeb4d781bb2128abebd23ab2c6a

Observation d76e4595-fe58-4982-b974-0dcafeb317c7 · outbound

This paper cites Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.858271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.858271Z digest=sha256:9b7e97996c1bee85cd49e45972bafc825e7804562bed87e5680fa30dd07e155c

Observation a609cc02-e81c-44b7-99a2-3fd8f5996f94 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.759445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:44.006676Z digest=sha256:9353eb5600e20b1e1c726c9a419177c51a060cedca289b337c34640f200c6509

Observation 4d5993ce-9291-4d6a-abb7-e2b9c4a38b10 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.132494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.132494Z digest=sha256:5f9f6d87f358e692ce2ba6c9401876c6128d1a216c9ebf71226ab9b964b8fc41

Observation 568f2b0c-7928-450d-8e16-dcbddcbffc46 · outbound

This paper cites DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.281953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.281953Z digest=sha256:68fa03f2253bf473f1c6c82f0eae7d7ed5b17e81819bec1ee6749ced6d6f537a

Observation 2fc74a18-437b-4103-9df5-adaac60d1a1c · outbound

This paper cites Docci: Descriptions of connected and contrasting images.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Docci: Descriptions of connected and contrasting images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.419574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:44.426315Z digest=sha256:1df88b306808d7e9eaee3b527f1491af3ce6e46481777df3639ae84f9faf78c5

Observation 01206368-d747-4e59-9a64-aa0084941692 · outbound

This paper cites Drag your gan: Interactive point-based manipulation on the generative image manifold.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Drag your gan: Interactive point-based manipulation on the generative image manifold

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.139918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:44.593681Z digest=sha256:aef665c7155578169d666a325a1aecb0823a14f13569df9ee6f5b6522cea327b

Observation 7d1111be-3583-4d87-91b0-2f76161e47d2 · outbound

This paper cites Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.803144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.803144Z digest=sha256:6c6a56ffe3d59db11f56407b01fdb4e24af9b4a85ad29fbd3ce5d1421e5fce04

Observation adc2cbee-e611-4ec5-87e2-4bb6e832c22b · outbound

This paper cites Synthesize diagnose and optimize: Towards fine-grained vision-language understanding.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Synthesize diagnose and optimize: Towards fine-grained vision-language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.859033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:44.947642Z digest=sha256:0587ef23c6f30c664769104d26703d83b11e4d4292bd7d2ea0ef3d6f04ac40f1

Observation 8550cc06-d2f2-421a-b074-ece6b8b6db09 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.523948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:45.089647Z digest=sha256:90589c43104cecdee42861d2e83db00dc1744f6df1655501f76d6ab5060e9e3c

Observation 2d389b6c-0f67-4f42-af7d-3ab502b2dc24 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.197125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:45.285969Z digest=sha256:6de0cad24dd836027443db646ab99ae4a42ed88875df3305454b91c39e5e34bf

Observation d7148344-4a87-4cec-9daa-1b2ce1559118 · outbound

This paper cites Towards vqa models that can read.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Towards vqa models that can read

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.871131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:45.403062Z digest=sha256:83748f26ce814b4a8d04dbfa62f0d0444eecbf139820935952cb37881d3c85e7

Observation 96e3fc7b-ebf8-4fd2-b77a-5842f29604d0 · outbound

This paper cites Denoising Diffusion Implicit Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Denoising Diffusion Implicit Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.549277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.549277Z digest=sha256:1df91df65e34d7ec12bec6cee181e9b84669b8da61be2ab1360a2b5a540c6882

Observation 044c6697-aa23-4e2f-943f-064832cfdae0 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Emu: Generative Pretraining in Multimodality

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.643297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.643297Z digest=sha256:92c552dceb53ac3675cf7a439071c9c956a58232b48352b395b7f27e970098a8

Observation af8ef4b5-8b9f-4203-bcd4-7f186fdc9c5e · outbound

This paper cites Generative multimodal models are in-context learners.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Generative multimodal models are in-context learners

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.613669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:45.803662Z digest=sha256:4c57f0ea74ce164326fe30d1159ae8bf7a2b3632205d1829dcb9b9c720fb1fd9

Observation c5c494f0-50f1-4823-b2b6-bad93ebce212 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.997544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.997544Z digest=sha256:2dd87cbea037a70cce1c60738379ae833bb7e41cf63e67beac4091f84669ce37

Observation a0777960-5a47-4dc7-89ae-a04082045c34 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.258784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:46.166326Z digest=sha256:6a60d3b8788cb0e247ac1efcce139418972223f5d3cbd6afbabe0792b16960d2

Observation 00dbf96f-a28d-4884-b1fd-07ac49ee689a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:46.340293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:46.340293Z digest=sha256:194d54524c924066e54e535a0992a72f8aece12eb2f6412459caccf3c7c38901

Observation 70ba53b7-68aa-4836-99a8-2ce4865f6875 · outbound

This paper cites Image inpainting with external-internal learning and monochromic bottleneck.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Image inpainting with external-internal learning and monochromic bottleneck

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.954470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:46.490587Z digest=sha256:524fc76c639deadb79bb9aac7a2c9807d23b9749068486a1e791a61d664c9e6b

Observation 6e62a470-9110-4471-9613-977118ae06a6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Emu3: Next-Token Prediction is All You Need

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:46.605395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:46.605395Z digest=sha256:90e93d9d5e541ba69aeb7078528b8e76a6696ea0bbd7228b68deb5e756fe9810

Observation 8655a29c-b26d-4aa5-8575-aa726a637913 · outbound

This paper cites Uni-paint: A unified framework for multimodal image inpainting with pre- trained diffusion model.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Uni-paint: A unified framework for multimodal image inpainting with pre- trained diffusion model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.620544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:46.748089Z digest=sha256:94d5958d9efc3968d01d771c5efc44013590d6c56cab0597ab6047fc5ecad607

Observation 0faa7296-ebbd-4f28-8003-dc62ddc19aa5 · outbound

This paper cites Imagebrush: Learning visual in-context instructions for exemplar-based image manipulation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Imagebrush: Learning visual in-context instructions for exemplar-based image manipulation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.316942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:46.871389Z digest=sha256:c519e1d903f72ebdddd4d23834246ac18d2a7db49e58a96532278f2f56c867e1

Observation 89789c16-fd0a-4782-a737-d1b9b93d86ca · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.046094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.046094Z digest=sha256:45a35437029ae5fea0a4caf7c40accc1735a758e40bd01e72abc6c0b53b0b567

Observation f38df3e5-9889-4c8b-a62a-44c5a6f99045 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.201274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.201274Z digest=sha256:aaf8f00c0c654e6b59c80a58b512015e102de7de674beaa146a2e84a0d7ebb74

Observation bcbbe6df-eef7-4337-be5b-744b37dbf1b9 · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.364932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.364932Z digest=sha256:cac4b5b3b29ba350448d35183f5d5bdfec3db8adba5ea56e1ce335bbac569d78

Observation 1b62c9c2-7faa-47de-b423-b5ce211a1ec8 · outbound

This paper cites WalkVLM:Aid Visually Impaired People Walking by Vision Language Model.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning WalkVLM:Aid Visually Impaired People Walking by Vision Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.468600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.468600Z digest=sha256:81e11a7bfc1f9729045cbe840fed6581c1413569cf700aac64e84f180f051dc2

Observation 71db210a-92bf-4f4d-8565-0b4df8778aea · outbound

This paper cites Mmvp: A multimodal mocap dataset with vision and pressure sensors.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mmvp: A multimodal mocap dataset with vision and pressure sensors

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:48.999053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:47.607411Z digest=sha256:4507c9b7d397487fae022132921b2cb669ec9dd1db8375c78d1a076e34ab286f

Observation 91ac0d10-9a62-47b6-8b17-530ea3d2b2dc · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.780118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.780118Z digest=sha256:1dc4c7f1d4227e6c9d9cf0da9038c1a53b2b266f3365709be80b9eae0c141943

Observation c9513c2c-afe2-4653-ae9d-c04995f7b64d · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Magicbrush: A manually annotated dataset for instruction-guided image editing

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.935836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.935836Z digest=sha256:cd6f88a2224969410a98c162d05bfc50c36bd87c99da5f0c682208de238329e6

Observation 47a0b8e6-6264-4127-9748-95e7cd8a76be · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:48.054205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:48.054205Z digest=sha256:2be8757a7721848fcaec736b62ce2ec16b5f1249c724aebe6086a6f0587233cc

Pith citing papers

No inbound Pith citation observations are available.