Pith. sign in

Paper Citation Record · LEDGER

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2607.23052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23052 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:51:55.185610Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58a3f014-74e5-45b1-8ce7-25ac52ce3384 · outbound

This paper cites H., Kim, Y., and Ghassemi, M.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models H., Kim, Y., and Ghassemi, M

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.040699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.040699Z digest=sha256:f5b6f5946da9f176b2f27d3f381b5a2a020d7c59bd9a35f31be73535f2bf30e7

Observation 13533fab-9929-48d4-8652-9622f5cb2882 · outbound

This paper cites Conceptual 12 M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Conceptual 12 M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.047256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.047256Z digest=sha256:584a4b8838b725fa0bc718bb880a2ac329af8c506b273da45c5c3fe63cf2e443

Observation aa921af7-d115-45ae-8bea-18dc7d70a05b · outbound

This paper cites K., Winn, J., and Zisserman, A.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models K., Winn, J., and Zisserman, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.052134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.052134Z digest=sha256:d795d772436942bb13bf7d7f1819fa4d7bee349ea81dc8e08d8b8ffa429c7477

Observation 24551c83-1460-4f98-aeba-40aacfa03330 · outbound

This paper cites and Kembhavi, A.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models and Kembhavi, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.056791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.056791Z digest=sha256:d9327c8f8212389f1c51748833966e8a9759a21a5382754eea4b7f0a6e5fcaee

Observation 6b3108a0-c8d8-4af3-8ee9-ba1db47fe5ce · outbound

This paper cites SugarCrepe : Fixing hackable benchmarks for vision-language compositionality.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models SugarCrepe : Fixing hackable benchmarks for vision-language compositionality

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.061416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.061416Z digest=sha256:f1b26494e1cccf7a53b5338c843127af794ad45ca5ded72f231ab84691a41eda

Observation 7f541cde-4f27-4976-8d43-b2f191a3c8af · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.065958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.065958Z digest=sha256:fd0c98439970296347c24199370e976cb7777e7530574bef74e352b19759f0f6

Observation caabac4d-b5dc-4fe0-a7e6-3ef28c0466ab · outbound

This paper cites ComCLIP : Training-free compositional image and text matching.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models ComCLIP : Training-free compositional image and text matching

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.071221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.071221Z digest=sha256:ae99b237ef0886d34b586ff5de0d5bba1bf82191e78967880470ece00149259a

Observation 7c3c3956-8da3-43ee-ace1-f8484f7e78ee · outbound

This paper cites The hard positive truth about vision-language compositionality.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models The hard positive truth about vision-language compositionality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.075416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.075416Z digest=sha256:891cd31e9227a0e13a2cb5361b387b262af6cf790218e5ad0dd08d6b204709ac

Observation cae2886b-b8dc-4e43-88a5-80d1d012052a · outbound

This paper cites Is CLIP ideal? No.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Is CLIP ideal? No

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.079762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.079762Z digest=sha256:6978ad1c9a6d1ec2aef3e0ebba33554acb55caa8ad894bb955701b971873b22e

Observation 2a007bba-9c7a-4924-9416-19d03327575a · outbound

This paper cites an unresolved cited work.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.084523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.084523Z digest=sha256:393406a149d0e837d4b05d42c10b815a16eb65695e6229717dd6418199be3699

Observation 74fe2927-dffc-435c-a646-a8a339fd6222 · outbound

This paper cites Does CLIP bind concepts? probing compositionality in large image models.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Does CLIP bind concepts? probing compositionality in large image models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.088705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.088705Z digest=sha256:d76a3c9a1f58d34254a66c40178f72862f7396217c2fd8c67fbddf00d1a44a53

Observation 6ae2b34e-a7e2-43b7-83b3-c45d0cc2ff4d · outbound

This paper cites an unresolved cited work.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.093182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.093182Z digest=sha256:6aa433e31543a1f6ff39b247b28a3d2bea1a2bbebb4ee569ccb285e3a44095d0

Observation d12192ed-7958-4fb0-8741-fdf4ed35a1ff · outbound

This paper cites O., Gandhi, M., Gao, I., and Krishna, R.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models O., Gandhi, M., Gao, I., and Krishna, R

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.097450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.097450Z digest=sha256:97408f515145b7883dbd1de16fcbe45c69dd676fa59f35e90c4331f0b84a12ec

Observation 1090c92c-04dd-4290-ac91-10cb4884c3b4 · outbound

This paper cites Simple open-vocabulary object detection.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Simple open-vocabulary object detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.101937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.101937Z digest=sha256:2e22e581ffad5b90364ebb4c4753cd591084591b53741b6a890a7c1467238a4a

Observation 4c1da4b2-69ae-424d-b48e-15ffc98e7c42 · outbound

This paper cites GPT-4 Technical Report.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.106219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.106219Z digest=sha256:e44e33326cef09fe36d5e5e17669cdbd98414db6cb841e372d301db8b8a99247

Observation 09aec2fb-9d41-4208-84ee-822341bc6ff6 · outbound

This paper cites Know `` No '' better: A data-driven approach for enhancing negation awareness in CLIP.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Know `` No '' better: A data-driven approach for enhancing negation awareness in CLIP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.111104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.111104Z digest=sha256:1009819342972c2dbdb19507c44b852a19365c7b5d62aaca137e4e16c9fef060

Observation 2b9ca59a-fa36-472f-ba1c-06d66830cb88 · outbound

This paper cites D., and Hein, M.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models D., and Hein, M

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.115548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.115548Z digest=sha256:58e7a356d40eea38beb561ea55f5ea13d7d49378010689348f0fe7b3298ad586

Observation fd4964f9-92eb-4ed3-9f96-2fc34e05430e · outbound

This paper cites How and where does CLIP process negation? In Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR), 2024.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models How and where does CLIP process negation? In Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR), 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.119841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.119841Z digest=sha256:ebe8bdcdfc0a81f7e73197a1a3c7e43a1f5ce1187ffb2b2f0866b08a7761c1e8

Observation 95082a60-a57d-4155-a591-c13bd1f6ff7d · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.124066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.124066Z digest=sha256:d2bbc8df94a625a3c0f4ae8d405297bf9b611f32e9480c5dc8c739e1ed4ca14d

Observation 46766c74-500f-48df-bcc8-3207d26a422e · outbound

This paper cites Collecting image annotations using A mazon ' s M echanical T urk.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Collecting image annotations using A mazon ' s M echanical T urk

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.128319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.128319Z digest=sha256:36621a910547010f81b103abf24520d0bd7974df2baa0132a8250071d9735a86

Observation fa702f0c-dd86-4801-843f-45e1e28413ec · outbound

This paper cites T., Argus, M., Fischer, V., and Brox, T.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models T., Argus, M., Fischer, V., and Brox, T

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.132836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.132836Z digest=sha256:8443e23fc7338f8f23b1af41809ee72b82e2771252f9cd9959ddd640387b72b6

Observation 9d4f5c1b-8fab-4d58-817c-c5001745e9c2 · outbound

This paper cites Learning the power of `` No '': Foundation models with negations.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Learning the power of `` No '': Foundation models with negations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.137512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.137512Z digest=sha256:8c7458a69e355f7b00ad34fa8920b26077d553eb688d953ec622639a9f54aa30

Observation 86b8ee81-e9e4-48b9-8ace-1dafd5305f5f · outbound

This paper cites ViperGPT : Visual inference via Python execution for reasoning.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models ViperGPT : Visual inference via Python execution for reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.142033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.142033Z digest=sha256:cb5edf36d70c81243fcc1dd17f8419f643bfdf21afcc326a13ad3a464fb212e2

Observation 7872f528-de3f-4204-9fbc-128d58e54e64 · outbound

This paper cites Winoground : Probing vision and language models for visio-linguistic compositionality.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Winoground : Probing vision and language models for visio-linguistic compositionality

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.146312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.146312Z digest=sha256:dec0abf3fe3e89085abadeb2f00b6490933c4e65028a7efb4c3b480d37dab84b

Observation 53d5e963-cea0-4615-b065-5da6c47c9663 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.150635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.150635Z digest=sha256:6d03d2a6a8949eaba05ac99c033e2625373d98efa2ce10ba2c70ce7a71faa4fb

Observation a79a2f97-e8b3-4ba6-8924-74995bc2423d · outbound

This paper cites A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.155521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.155521Z digest=sha256:7e79a677fcd5abee402327af4405a3ae8b5acfd3a492c02b14ccb89f09d7a5ea

Observation d70ffbe1-fcd4-4dca-9c33-d94f8b4560e2 · outbound

This paper cites Y., Lee, M.-L., et al.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Y., Lee, M.-L., et al

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.160075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.160075Z digest=sha256:c770b433a62fec233ab9fb3bc23e8aaaf09e0e1e11e484d09e82022f51fe9935

Observation 36eed525-17d3-4ff9-a9d6-4ae99f7e1fd5 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.164387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.164387Z digest=sha256:234a177ebb92be809c27aea9cb971a244959e54f068d1aa7926e4a8c40668396

Observation d9dbb6c2-5cd6-4cfd-8377-75206d1f2dfb · outbound

This paper cites LiT : Zero-shot transfer with locked-image text tuning.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models LiT : Zero-shot transfer with locked-image text tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.168632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.168632Z digest=sha256:b5024918dc901d069c6f5412dbe3828227436de69d5e102e2749b3f6a449ee39

Observation 28425009-383c-4a7c-8af8-eefeed24dcb8 · outbound

This paper cites Sigmoid loss for language image pre-training.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Sigmoid loss for language image pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.172818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.172818Z digest=sha256:7af8d5bb42563869041b5e320d2ab07f29a5b8a824f4b541eb5964a3acadc33f

Observation 79305521-3cee-41e5-8875-166d05a4e18a · outbound

This paper cites NegVQA : Can vision language models understand negation? In Findings of the Association for Computational Linguistics: ACL 2025, 2025.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models NegVQA : Can vision language models understand negation? In Findings of the Association for Computational Linguistics: ACL 2025, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.177250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.177250Z digest=sha256:d6d5df55f57d7a258ae43588871849300135f4dbe954199fc793b075abd41e99

Observation 277d1105-0be4-4aac-862f-7ad1ef78776f · outbound

This paper cites VL-CheckList : Evaluating pre-trained vision-language models with objects, attributes and relations.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models VL-CheckList : Evaluating pre-trained vision-language models with objects, attributes and relations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.181466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.181466Z digest=sha256:20b7db10017ebfeeb65bdf3b342ae982104370215f1e2aeac8ea3997c1421971

Observation 23fbf513-596f-416a-a010-1e59a8d823ce · outbound

This paper cites Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.185610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.185610Z digest=sha256:28b3e8414b6306bc007db7119aee58096c09827e096aeb08c9e6fcfacfb66977

Pith citing papers

No inbound Pith citation observations are available.