Pith. sign in

Paper Citation Record · LEDGER

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions

As of 18 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2412.08169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08169 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:12:07.179154Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:28.031451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:34:30.784913Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f8cdedd-622e-4853-b6a9-ed83f7b357fb · outbound

This paper cites Breaking common sense: Whoops! a vision-and-language bench- mark of synthetic and compositional images.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Breaking common sense: Whoops! a vision-and-language bench- mark of synthetic and compositional images

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:12:07.763634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T18:12:07.112140Z digest=sha256:50dc4273b0801493aca8f546810620a9657afa5b22a3634ebc6bcb910ad437e5

Observation f9d9a30f-0b7f-494e-96f5-762f3ad6c00e · outbound

This paper cites doi: https://doi.org/10.1016/ j.patter.2023.100695.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions doi: https://doi.org/10.1016/ j.patter.2023.100695

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-11T18:12:07.135148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:12:07.135148Z digest=sha256:6177ad207fc8c7e3326edadd14721ef9a7eb66a91a4a64f170e9d3c998721dbd

Observation 93fad4b1-89f4-4dc2-a7ad-a93ce25f29ad · outbound

This paper cites URL https://aclanthology.org/2024.acl-long.573.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions URL https://aclanthology.org/2024.acl-long.573

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:12:07.708370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T18:12:07.167874Z digest=sha256:eed0ac74f25914229accd727a6cb39fc9465e707b429d5f8f3136cd3b630b922

Observation 309376bf-6da2-4696-a456-22a6732e593c · outbound

This paper cites URL https://aclanthology.org/2024.cmcl-1.2.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions URL https://aclanthology.org/2024.cmcl-1.2

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:12:07.691364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T18:12:07.173573Z digest=sha256:cb7c217fd58c240678a8954b7076fb9d78d418dc822b2c9cfe64ab5bbba47781

Observation 6a097fee-0b66-4e80-8326-43f799065baa · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Adding conditional control to text-to-image diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:12:07.179154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:12:07.179154Z digest=sha256:3d484514ab0266a6d9501f799e59d68f3d757706f538ef98d8feab7a9dde15d0

Observation 758aa00d-56d5-45a0-90e4-1a325c2481b8 · outbound

This paper cites URL https://doi.org/10.1179/isr.1984.9.1.47.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions URL https://doi.org/10.1179/isr.1984.9.1.47

Reference 1984

Resolution
unresolved
no resolver link, observed 2026-08-11T18:12:07.124765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:12:07.124765Z digest=sha256:b9a5f19d12e70cbf36b2ca69457955238b583d1ae750da6ece68fa2b885ec662

Observation cef3b270-d3df-48af-bbda-ac29d8af5352 · outbound

This paper cites Dongxu Li, Junnan Li, Hung Le, Guangsen Wang, Silvio Savarese, and Steven C.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Dongxu Li, Junnan Li, Hung Le, Guangsen Wang, Silvio Savarese, and Steven C

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-11T18:12:07.162821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:12:07.162821Z digest=sha256:a4fbbd5c371f266a72cb09f6018686040edd32eb59f4def82b10f7e151840e50

Observation e74f4de3-9f15-4b24-8de6-4a4c54377f0d · outbound

This paper cites Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou

Reference 2019

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:12:07.538149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T18:12:07.152081Z digest=sha256:2099fb977b7e3a8e75d7cf59830b6fc21c60fbd94f980f066fd94933cf3a1aaf

Observation 8afd3b1d-c2fc-48b6-b10a-9097fafb0045 · outbound

This paper cites Edward J.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Edward J

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:12:07.725176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T18:12:07.157575Z digest=sha256:e7492cfabd4028f031cd20f258284ec0c7c5ae2137a49928455b3a9467717b27

Observation bad0fa4e-ec4c-4cbe-9796-a663b80d2476 · outbound

This paper cites Alexander Gomez-Villa, Adrian Martín, Javier Vazquez-Corral, and Marcelo Bertalmío.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Alexander Gomez-Villa, Adrian Martín, Javier Vazquez-Corral, and Marcelo Bertalmío

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:12:07.743561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T18:12:07.147209Z digest=sha256:86e58d2e4ec0ed3e0a1be92c1176bd50ae7f3efa567af7f0a1d71c7773a47880

Observation a5ff66d0-9839-4b01-9305-0a7c681de59a · outbound

This paper cites Diffusion Illusions: Hiding Images in Plain Sight.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Diffusion Illusions: Hiding Images in Plain Sight

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T18:12:07.119101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:12:07.119101Z digest=sha256:1b5b15b2b9ef3e6860aec41116dc4250e85eacfd9995efb9dbf2438661ef967b

Observation b93caf8e-cf7a-43cf-958a-f21aeebd2d03 · outbound

This paper cites Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models.

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T18:12:07.140617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:12:07.140617Z digest=sha256:d2a637e0306bf804bd0184d278ef545eac07d625d3477150e508da6b48fc8d31

Pith citing papers

Observation 38a83140-adc5-433f-9cbe-d2b3b3269727 · inbound

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions cites this paper.

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:30.891394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:34:28.031451Z digest=sha256:4f7f38eefbcff4fbdda6dc8a75abb84927dee19e8eeede6e63df85b43881017b