Pith. sign in

Paper Citation Record · LEDGER

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2506.07227.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07227 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:48.054205Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2de8716a-b311-4ff6-96fe-0cdd660016d0 · outbound

This paper cites URL https://api.semanticscholar.org/CorpusID:276612236.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning URL https://api.semanticscholar.org/CorpusID:276612236

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:59.507971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:38.608937Z digest=sha256:7f353bfc69ccb3b2251f318e3000512d6f02570ea31543bf6d72226779af2ea6

Observation bca26de9-f16f-4a11-a98c-8a064ec72ebc · outbound

This paper cites GPT-4 Technical Report.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.711206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.711206Z digest=sha256:f212987f5cad8f1a28fa800b9ca6212a4ce1f22d58cf790dab92efae9b79b67c

Observation d6ad0563-42e0-4b3e-a964-2444405eca28 · outbound

This paper cites Qwen2.5-VL Technical Report.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.884823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.884823Z digest=sha256:e3515dcfdf19124b1f3fdfec985c625d923c5207dcf9008c5a91e0e3e92732a5

Observation 2a3bf250-33e3-4e51-84b1-e5994c2a80f7 · outbound

This paper cites Hallucination of multimodal large language models: A survey .CoRR, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Hallucination of multimodal large language models: A survey .CoRR, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:59.171317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:39.041996Z digest=sha256:03f93e78c679f7c5b06db4dfa46084700b4d2da4ec878cbde5aa2c90e835ce22

Observation 423e27ac-0282-4dda-9410-db59986b3274 · outbound

This paper cites Flux.1 fill [dev].

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flux.1 fill [dev]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.861034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:39.254218Z digest=sha256:3e48f1f0e2eea923d20795b86681707ab095c5cc99ace58713fb0728d989ce9a

Observation 5ecf6e1c-d64d-4ff5-ad5e-c9aecd056bb3 · outbound

This paper cites Ledits++: Limitless image editing using text-to-image models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Ledits++: Limitless image editing using text-to-image models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.484131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:39.375131Z digest=sha256:30ab69eb459690737ed8394ddfdd09f9ebf7e844151a3ca103c66203ecc19f1e

Observation 662d1733-67bc-49c9-991b-986ec3bad760 · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Instructpix2pix: Learning to follow image editing instructions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.126442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:39.449044Z digest=sha256:65d60c634b66a62361426dc78c8a44a14554150bf96d858f5eef74484eec1f69

Observation c0ed1ce1-9362-47b4-9348-0c32bc4c05d5 · outbound

This paper cites The revolution of multimodal large language models: a survey .arXiv preprint arXiv:2402.12451, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning The revolution of multimodal large language models: a survey .arXiv preprint arXiv:2402.12451, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.543224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.543224Z digest=sha256:893720ce6c0037cbacc8b734f8533ad588a2389acc4c052de7be54c920aef862

Observation 06591c43-d4d1-42e2-bec5-d356c2e0dd67 · outbound

This paper cites Vlmimic: Vision language models are visual imitation learner for fine-grained actions.Advances in Neural Information Processing Systems, 37:77860–77887, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Vlmimic: Vision language models are visual imitation learner for fine-grained actions.Advances in Neural Information Processing Systems, 37:77860–77887, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.810907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:39.598014Z digest=sha256:f1cb482b2e897c5eb889b5dfea90d94816596ba55c866fec09fe72919b6cf610

Observation e6a69b11-b329-4a2a-83ce-f3fd8ebd5ac0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.736678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.736678Z digest=sha256:6fa1feca0315d5a8c662024ecf6148266a361900407f77ebeca7f4ede00b10c0

Observation c46b15e7-cf14-4543-9487-ea9126ad7257 · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Anydoor: Zero-shot object-level image customization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.415831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:39.854427Z digest=sha256:094cdfe926683092155f77b501aab1919d1ef9b62df989043f825c9af66ad555

Observation 30ddf113-d876-428b-b26e-898c2a7e035f · outbound

This paper cites UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.028036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.028036Z digest=sha256:a79887815982e6eaa0fe862770fc377714355b1b094f3ca23598934e99b34e4c

Observation 9c6eee60-6280-4765-813f-2a5265c515b8 · outbound

This paper cites Diffusion self-guidance for controllable image generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Diffusion self-guidance for controllable image generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.204373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.204373Z digest=sha256:1df942da46f114b870abddbcec603d495942f51de7c57ad3af7cbf3638bc5c24

Observation ceaae5c7-6498-4941-8dce-05f664adf7a8 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.351303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.351303Z digest=sha256:66f37bd6c87677b3effd682940fa0b24168e370f2cd24494adcd98431d457502

Observation 2dd79992-184b-4994-89f0-007c48796424 · outbound

This paper cites TLDR: Token-Level Detective Reward Model for Large Vision Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning TLDR: Token-Level Detective Reward Model for Large Vision Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:43:48.602722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:40.500892Z digest=sha256:89cc2a4f40c1d8408eb1bd1cb42bd3c30c3a3c353d92eee5e7173f7805db849f

Observation 76f01343-d522-4ef2-82d9-fee5a2b553c0 · outbound

This paper cites Towards a benchmark of multimodal large language models for industrial engineering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Towards a benchmark of multimodal large language models for industrial engineering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.107153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:40.648100Z digest=sha256:81ae87225c980f3df8fd3e4b8bfaf2db75524e4c89c6d7818b3ee9ed51b98c7a

Observation 76a6500e-94ff-48d3-be82-8f5016ad95ea · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.719981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:40.743202Z digest=sha256:defb850ee9ee72c9bb1a8fd46c0e353d1135532d7b6ee200098ebb89d05ba1f7

Observation a92043ad-40b2-42c7-8553-5c2f0f6a01c3 · outbound

This paper cites Gemini 2.5 flash.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Gemini 2.5 flash

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.397687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:40.943773Z digest=sha256:a945871fd742613df003a28e9fb07aa916d86cbafac69938846a65ebe00db561

Observation dee98651-ea66-4bc2-986c-1b7b5327c82f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.065207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:41.114382Z digest=sha256:8b47e55628c57c017566bb65fa2788abada08fefed1b188e9624a925c52213e5

Observation e6543063-4d93-461e-a0b3-51ece3c23a4e · outbound

This paper cites The Llama 3 Herd of Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.230598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.230598Z digest=sha256:e1b130d1699432da34c2b44f8aeab0c85be0f270ca3651e14f50f8a84b126ae6

Observation d1a55e4f-cf1d-44ad-b42f-285b3a768821 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:55.619250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:41.386418Z digest=sha256:f0aadac222fd90267f2ae4947affb952d30fc25f5c637ad5adaa167c1f6c47eb

Observation 6ac500f9-93a6-4298-a548-a3f564b985ed · outbound

This paper cites Diffusion Model-Based Image Editing: A Survey.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Diffusion Model-Based Image Editing: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.528747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.528747Z digest=sha256:277184a356deb6faac84684d4e7819ad8b104193a8575619fae6171e603079ee

Observation 047225bb-e58e-49e4-ad4b-298873d44d7f · outbound

This paper cites Smartedit: Exploring complex instruction-based image editing with multimodal large language models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Smartedit: Exploring complex instruction-based image editing with multimodal large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:55.263948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:41.681296Z digest=sha256:4a38069fdb4c15c5dceca80125d67336f4b1871d873e93713e9ab4a2018be25c

Observation d8922cbf-9b68-4df8-81f1-1ee09e87f00a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.928422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:41.824325Z digest=sha256:b9631c8264ba64a47b3f47824fe224c54ac451346c30e8cd81ed085f02ce3198

Observation 5b9c44cb-39a4-488a-9d8b-3b679c278984 · outbound

This paper cites GPT-4o System Card.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.006009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.006009Z digest=sha256:d037434b6a129d72081ccc9b69ddab1e1cfb37250739cd3d6e39afeee94caf9d

Observation d0a575a1-2eee-43b7-9f09-029c7d9a1366 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning A style-based generator architecture for generative adversarial networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.551988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:42.126473Z digest=sha256:16a76678cc5cb29e7cce10e1afa7b87ad39c46e9a9892ccf968db696b4fc063b

Observation 555ec626-1b0d-4dc4-915a-bc6c28fc7849 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.310898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.310898Z digest=sha256:d311161e9c3bab6f4c34d94879850e0f7d282ca9dbb8c8c48d05009ef6fa8859

Observation 2e22816e-3bbe-49f3-a114-f95cd95da233 · outbound

This paper cites Evaluating and improving compositional text-to-visual generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating and improving compositional text-to-visual generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.203404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:42.448452Z digest=sha256:23f446019ab9b1bca768292a1b18dc1830d7806b28803d534dda88d8e7c62ab8

Observation d0a127ad-0c9a-430a-b980-f740fee99f64 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.594952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.594952Z digest=sha256:337871f0b22d04d806d36ef140d83b15b1f3d624f8598d264b081f5337dfd75d

Observation e5c16a2e-1828-48e6-a089-e465466d583e · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating text-to-visual generation with image-to-text generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.827596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:42.736228Z digest=sha256:7a91d4f6bff1686d17df5e86e820bb81edbb67e338b93351361c5d56def21ea3

Observation 01af5ce7-54ed-4e9c-807b-54fd8acf05e5 · outbound

This paper cites Flow Matching for Generative Modeling.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.883509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.883509Z digest=sha256:f26ddf27950bda44512ec745f9471267b9fe32160404aad6227cef311dd4c2f5

Observation 27be5de5-b3ff-4872-a905-992872b5b4d5 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Improved baselines with visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.066529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.066529Z digest=sha256:3db8baa64d1d1b2f8bfb06a1338c633da61fbc3f34add2a5a146c66c1c4d7249

Observation 848a6b9b-f94d-49a4-889d-173ff694dbe8 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Step1X-Edit: A Practical Framework for General Image Editing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.205553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.205553Z digest=sha256:fd0e3a319485047fd9ad63801e4f22f80690ea348d67edea748b294af2c64dff

Observation af6e4322-9d20-45ca-b28b-218f0f494c27 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.306278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.306278Z digest=sha256:2d547f1f56401a9782993b5988f015336df43c88e17e84ff9f0d935b192571ef

Observation fefda47f-026b-4584-8273-f95ab4f3f7a0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.466271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:43.488660Z digest=sha256:4cafb148750771d1cc4227ad9e6344d98672b88e6be4d13c8b257a2a5cb63cf8

Observation cd9cbec6-efba-4ef9-8ee6-195b3fe737c3 · outbound

This paper cites Customizable image synthesis with multiple subjects.Advances in neural information processing systems, 36:57500–57519, 2023.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Customizable image synthesis with multiple subjects.Advances in neural information processing systems, 36:57500–57519, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.106964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:43.656077Z digest=sha256:43e15eabaa97288019afba60d16f304c4db58335e85a3bc0ff486b55ea1b78bc

Observation ad8ec765-f698-4aa8-8ebc-9c83b96732ad · outbound

This paper cites MagicQuill: An Intelligent Interactive Image Editing System.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MagicQuill: An Intelligent Interactive Image Editing System

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.759971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.759971Z digest=sha256:1ca8beb602d249a1c15a5de75415a69ba0844ddfb8b5709a216c6af9578c5962

Observation d76e4595-fe58-4982-b974-0dcafeb317c7 · outbound

This paper cites Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.858271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.858271Z digest=sha256:dc15e87cc1c5edeafaa436d096f89d0b24858fc2b6dd65d955bf7ecfc98deba8

Observation a609cc02-e81c-44b7-99a2-3fd8f5996f94 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.759445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:44.006676Z digest=sha256:a9bd80aa1267e1e0c9f438bf4a28e50d47039ba17ac5ab46a117bfac50bbb9bc

Observation 4d5993ce-9291-4d6a-abb7-e2b9c4a38b10 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.132494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.132494Z digest=sha256:69396782c262dcf55dfd1f8a52bc68ddcd33fd51e115105c0d4a105294e03695

Observation 568f2b0c-7928-450d-8e16-dcbddcbffc46 · outbound

This paper cites DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.281953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.281953Z digest=sha256:7228293b06d1a6c47a13460f7ac7c8fd584405b2dcbde8e742cdc31c440a6614

Observation 2fc74a18-437b-4103-9df5-adaac60d1a1c · outbound

This paper cites Docci: Descriptions of connected and contrasting images.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Docci: Descriptions of connected and contrasting images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.419574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:44.426315Z digest=sha256:de104baa8993134fab405ef1a00bbd9b3f1a393899945d4ef048988c8ba5a356

Observation 01206368-d747-4e59-9a64-aa0084941692 · outbound

This paper cites Drag your gan: Interactive point-based manipulation on the generative image manifold.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Drag your gan: Interactive point-based manipulation on the generative image manifold

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.139918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:44.593681Z digest=sha256:bf609725deac31f4d15b50a7e589a2b52f15438b39e8ebc5862f194743a9bb87

Observation 7d1111be-3583-4d87-91b0-2f76161e47d2 · outbound

This paper cites Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.803144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.803144Z digest=sha256:843f0e8b5f055951b2248c95952e3d6e460d73af649f06d52faecc7659196d26

Observation adc2cbee-e611-4ec5-87e2-4bb6e832c22b · outbound

This paper cites Synthesize diagnose and optimize: Towards fine-grained vision-language understanding.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Synthesize diagnose and optimize: Towards fine-grained vision-language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.859033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:44.947642Z digest=sha256:95c3dd0125388bd9eb5d47ff2334600566768cf3a42a1618640c6a1620a64c64

Observation 8550cc06-d2f2-421a-b074-ece6b8b6db09 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.523948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:45.089647Z digest=sha256:93e4ad7762abca5cd518e45b356d1bc6de9537c9091896a78430b9064a07d676

Observation 2d389b6c-0f67-4f42-af7d-3ab502b2dc24 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.197125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:45.285969Z digest=sha256:80aa38f62227c8dde643462d730792af720b6d883c2a6b3e0d103b607786c66d

Observation d7148344-4a87-4cec-9daa-1b2ce1559118 · outbound

This paper cites Towards vqa models that can read.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Towards vqa models that can read

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.871131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:45.403062Z digest=sha256:6df0f266537529eaba47207d3dcd9cc33b9560ad3f3be8e31ad8b7767ead817c

Observation 96e3fc7b-ebf8-4fd2-b77a-5842f29604d0 · outbound

This paper cites Denoising Diffusion Implicit Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Denoising Diffusion Implicit Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.549277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.549277Z digest=sha256:394dabccb18eb9b68d3b29f60ed72da1b04472de11c8d05c009361fcae4954e7

Observation 044c6697-aa23-4e2f-943f-064832cfdae0 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Emu: Generative Pretraining in Multimodality

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.643297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.643297Z digest=sha256:5d27c90a6d93c46e19ab5a3adb58d9cbfc180b7dcb83e04180264c5dae770f8a

Observation af8ef4b5-8b9f-4203-bcd4-7f186fdc9c5e · outbound

This paper cites Generative multimodal models are in-context learners.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Generative multimodal models are in-context learners

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.613669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:45.803662Z digest=sha256:101da1e704e3fe019b99cc0e2298d46816697c8be5acda175f0f9fc0f866f91f

Observation c5c494f0-50f1-4823-b2b6-bad93ebce212 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.997544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.997544Z digest=sha256:6a5b99fe17663aa71410ea5729a5ea2094c7b64e5ac0309dfdb0fa0bd80c0678

Observation a0777960-5a47-4dc7-89ae-a04082045c34 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.258784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:46.166326Z digest=sha256:cd33c43684a3744f3d0b29aec881a236a856044e0873d857cc52089f2df526f1

Observation 00dbf96f-a28d-4884-b1fd-07ac49ee689a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:46.340293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:46.340293Z digest=sha256:8aadda537d71daaf94a41c2d5c315bc5c5293908df83a28f6c23e3376da375d4

Observation 70ba53b7-68aa-4836-99a8-2ce4865f6875 · outbound

This paper cites Image inpainting with external-internal learning and monochromic bottleneck.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Image inpainting with external-internal learning and monochromic bottleneck

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.954470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:46.490587Z digest=sha256:c539a118f38d3116b0e2756960966299ef405e52f9f03a079411859d1d626d3c

Observation 6e62a470-9110-4471-9613-977118ae06a6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Emu3: Next-Token Prediction is All You Need

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:46.605395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:46.605395Z digest=sha256:85a8f1edd264ac976964d86ab885c14f63d3c591b36b37b85444d2284018e66b

Observation 8655a29c-b26d-4aa5-8575-aa726a637913 · outbound

This paper cites Uni-paint: A unified framework for multimodal image inpainting with pre- trained diffusion model.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Uni-paint: A unified framework for multimodal image inpainting with pre- trained diffusion model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.620544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:46.748089Z digest=sha256:77e1d8208a04c365d7e05a1e2519f31f7f8f95474ba83bdd3b4cd421138f45c4

Observation 0faa7296-ebbd-4f28-8003-dc62ddc19aa5 · outbound

This paper cites Imagebrush: Learning visual in-context instructions for exemplar-based image manipulation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Imagebrush: Learning visual in-context instructions for exemplar-based image manipulation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.316942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:46.871389Z digest=sha256:45dd31deec50885de5c797658b581f952eefed6209bcdd44c8ed67eb6a7534a2

Observation 89789c16-fd0a-4782-a737-d1b9b93d86ca · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.046094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.046094Z digest=sha256:f6c18f292fd742eb37a0adfe0f6efa5269d9200e5df119c35a325750b070cbc7

Observation f38df3e5-9889-4c8b-a62a-44c5a6f99045 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.201274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.201274Z digest=sha256:27f4005bfa324f69f49f1404f2070b820035ac875cb43dcec4e0babf070de5ba

Observation bcbbe6df-eef7-4337-be5b-744b37dbf1b9 · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.364932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.364932Z digest=sha256:efa851e9a696f254acd0305e2eb8978202e7098df11c8bcfb46b75b8b2d7d9a0

Observation 1b62c9c2-7faa-47de-b423-b5ce211a1ec8 · outbound

This paper cites WalkVLM:Aid Visually Impaired People Walking by Vision Language Model.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning WalkVLM:Aid Visually Impaired People Walking by Vision Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.468600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.468600Z digest=sha256:6043919e8f822307798b8dd44fd0ff0087f9baa87b3c5df742c52d991171c0ad

Observation 71db210a-92bf-4f4d-8565-0b4df8778aea · outbound

This paper cites Mmvp: A multimodal mocap dataset with vision and pressure sensors.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mmvp: A multimodal mocap dataset with vision and pressure sensors

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:48.999053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:47.607411Z digest=sha256:28ce474d4266779d2a0ccb5a929d3f312c44a41181974664de3c361c4aa03f71

Observation 91ac0d10-9a62-47b6-8b17-530ea3d2b2dc · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.780118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.780118Z digest=sha256:62f6be3b2bb0421977bd2e0fe06242b852bafb38ec42fd69152832eb7699056c

Observation c9513c2c-afe2-4653-ae9d-c04995f7b64d · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Magicbrush: A manually annotated dataset for instruction-guided image editing

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.935836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.935836Z digest=sha256:2e134cfcdc5f1e85e73c9282e205a3ad4b3984b23a27ad9ff13f263ca099332b

Observation 47a0b8e6-6264-4127-9748-95e7cd8a76be · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:48.054205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:48.054205Z digest=sha256:8b5dbe16946265e82fe16e21b62ddb52233a12e67105a08b37d8c568f3b6ce86

Pith citing papers

No inbound Pith citation observations are available.