Pith. sign in

Paper Citation Record · LEDGER

When and why vision-language models behave like bags-of-words, and what to do about it?

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 69 inbound Pith citation observations for arXiv:2210.01936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.01936 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 69 of 69 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:52.486750Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T19:36:29.281490Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1c5deeee-8208-4051-8fa1-6835cb49f8d4 · inbound

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps cites this paper.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.627473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.627473Z digest=sha256:9c733f54112a57e23486be49fc011e4518f43df25575a032f47aaee43fab8cec

Observation 96a26a43-d54f-4170-b4aa-97439c27ece7 · inbound

Enhancing CLIP Conceptual Embedding through Knowledge Distillation cites this paper.

Enhancing CLIP Conceptual Embedding through Knowledge Distillation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:24:06.780353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:24:06.780353Z digest=sha256:f60dd6d2d4d535a5e6e57f335cc29c8f8ba9beca19547c54b70f20d35cb6b951

Observation 50cdc964-c232-4af0-8d39-bb20326e98bb · inbound

VladVA: Discriminative Fine-tuning of LVLMs cites this paper.

VladVA: Discriminative Fine-tuning of LVLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:26.523972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:26.523972Z digest=sha256:a0fe394798211f4902ee85d12f361bef1185e6ded8618920d480ec515517c715

Observation 877fa6fe-ab78-42a7-8366-1b8646e0b093 · inbound

RelationField: Relate Anything in Radiance Fields cites this paper.

RelationField: Relate Anything in Radiance Fields When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:59:14.424665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:59:14.424665Z digest=sha256:cfa4617d29e6654af2a318e2c4200293a96741009794119afc6ad96007f03098

Observation 1c27b363-2f48-49de-8803-0715862d5e01 · inbound

ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding cites this paper.

ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:38.758354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:32:38.758354Z digest=sha256:53d4b7fc2fce4422d976d2acea78721d86453f46cd3f01df3b9c74dd16f584bf

Observation 29cd9db1-eda2-4a78-ad7d-ccab212b63f4 · inbound

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency cites this paper.

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:26:00.653529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:26:00.653529Z digest=sha256:3ac3cbc8bef1d8613051ab0f1435a74de25fc52345c02347874925694f6a4a2b

Observation 1a69c6cc-4c96-4ed1-b5c7-fb92fb010c16 · inbound

FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis cites this paper.

FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:39:25.093138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:39:25.093138Z digest=sha256:cc80310ca6f06f6cf873080c2e10ab71cfb20f448ca83471f2cea4fc564d26bd

Observation a402747d-6017-4b4e-a935-9af29a8907bf · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:24:27.539998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:f4f28a62fd982cb2a49ec97ac8b8bae61573d0c317aee1a12cdfe6fe68d0dd1f

Observation 4323f224-8df8-4547-8e42-746b7bad621f · inbound

Decoupled Global-Local Alignment for Improving Compositional Understanding cites this paper.

Decoupled Global-Local Alignment for Improving Compositional Understanding When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:52.486750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:52.486750Z digest=sha256:2755cf62f74300fab70009c61b3e601d389513be60c043d8221d62b335a5d002

Observation 44242c63-bf64-44cd-9389-36287455d129 · inbound

Multi-Modal Language Models as Text-to-Image Model Evaluators cites this paper.

Multi-Modal Language Models as Text-to-Image Model Evaluators When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:11.456920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:11.456920Z digest=sha256:b988177c26f8703c5bdbdeecb2271faf876198c5f4e088e6e8f1ce639cb3d724

Observation d722abd3-f017-4e8d-a970-bc9b23a064c6 · inbound

Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models cites this paper.

Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:28.262510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:55:28.262510Z digest=sha256:0b209613b2fc386d30666b35a49afe509011561a5e94a004eb01f3ce6716bf7f

Observation dff164eb-6c6a-4acf-8cb4-bee651993ea6 · inbound

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models cites this paper.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:42.791864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:42.791864Z digest=sha256:23513ae86e332a07882d9aad32b4159a8569be8a82082764b347f9063431b637

Observation 3b905e8a-b3f0-4abc-a0ef-56ac6cc61eab · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:00.457612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:00.457612Z digest=sha256:5c04253d9390eb70f202dcfaeb46b6f76b909907836702a620cc0a3e9d650718

Observation 5978c2a7-5ae2-4d2b-a6dc-1796a0f67198 · inbound

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis cites this paper.

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:12.731024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:12.731024Z digest=sha256:f3b09f41417e5c250af31f40746076f36f862994cc0cc87d91c824cbb2ae8824

Observation 2650a0af-46a6-4c1d-a009-fa9c1ab10f53 · inbound

IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth cites this paper.

IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:56.159838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:56.159838Z digest=sha256:a40f246e457725d15169da9859222a1eef7ff230ee1f5a13cf951c808909414a

Observation 21d5b03d-47ce-423f-916d-6834d1354101 · inbound

From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models cites this paper.

From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:53.099753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:53.099753Z digest=sha256:b31788b13be235900ff34e90c934223b5b4ca94a941abf8db67688e417438500

Observation ceca9211-8036-4760-80a7-1391c7ff2020 · inbound

CIVET: Systematic Evaluation of Understanding in VLMs cites this paper.

CIVET: Systematic Evaluation of Understanding in VLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.339325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:36.339325Z digest=sha256:86efe5d01a9ed05b34f5816dd438ac26eb09aa7e30ad96c9089994b56554a28c

Observation 3ed8fddf-ba6d-48cb-8c12-f3952ea698e1 · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.695389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.695389Z digest=sha256:5e452e7ea9d70d4d2e6d5b5de1a1419c8faca13c552f3a28b864ee2389ff45f8

Observation 48d636dd-8c85-446d-acd7-e7f440c730b8 · inbound

ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation cites this paper.

ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:45.958345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:45.958345Z digest=sha256:e7f05c59591c9fdc9ac81720c363f4a3e2a5c561abaa20d88274ca7c9f7691fe

Observation e72a3b76-2f3c-414e-b73d-298ca1be321a · inbound

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models cites this paper.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.147656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.147656Z digest=sha256:a1d67b34c8a5ac5ea4a72f46fe9f4d85dbd9d1a2c1fead06793eaf965cfb6613

Observation 720d82eb-b1da-4161-9cf6-1fba16e76b26 · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.449939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.449939Z digest=sha256:c75a001b41032f09a0dcc8a377b31038956aa01cc615e999cf74c328fe26f077

Observation d715b9ac-d1c6-4571-9849-5ace6a338478 · inbound

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities cites this paper.

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:01.402381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:36:01.402381Z digest=sha256:7a83a5de57ff71dddced02748f771fa665e6e67f36480100c2284ad18bd24d35

Observation 444a16d3-f9e4-41d5-9686-898887bac9e4 · inbound

Negation-Aware Test-Time Adaptation for Vision-Language Models cites this paper.

Negation-Aware Test-Time Adaptation for Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:58.775271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:58.775271Z digest=sha256:fc4fc93891eec82d9ef51e5a742de665ecb04ca41947b579057de97615406b1b

Observation 0e3e5cde-ade4-4b49-9635-629d88a904b2 · inbound

Trade-offs in Image Generation: How Do Different Dimensions Interact? cites this paper.

Trade-offs in Image Generation: How Do Different Dimensions Interact? When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:01.848706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:01.848706Z digest=sha256:c2506ab4cccd3c0d735cd53b6ab75e2980b009d7cad3f5797bba9d5fd0e54070

Observation 5b0e5a8a-4ba5-4085-a5dc-479711bd9875 · inbound

MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding cites this paper.

MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:40:01.070516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:40:01.070516Z digest=sha256:b6445984295fa6c70da81b55e06f49c5e4fa7e631e9e0c2cb1f54b148e6ead51

Observation b0787d3e-1fb8-4ba3-8582-6cbabf57d47e · inbound

Native Hierarchical and Compositional Representations with Subspace Embeddings cites this paper.

Native Hierarchical and Compositional Representations with Subspace Embeddings When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:39.596159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:45:39.596159Z digest=sha256:3d26363eaa4ce9d02364390326348f20b62b70e03ed8af3f4c1fb166ad9a3bf3

Observation a406c379-ad15-4816-8977-5b9deee8d4c0 · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:10.958883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:10.958883Z digest=sha256:d181b5bf5b9a07eb562b60312d3d947763230297a9204afb146678c1550d2f93

Observation fca59d4c-535f-4bc2-94db-56f283c6cb16 · inbound

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation cites this paper.

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.449457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:38:58.449457Z digest=sha256:43a8f8faa79a5139b5d0de5fa52f1a0ee4086590b14c92d400ef6778b9c0ece3

Observation f8da70f2-04a0-4eb4-ad7a-0cb71a05e33c · inbound

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs cites this paper.

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:41:30.190913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T14:41:20.403259Z digest=sha256:cecf7201bb06d7859fb538e6cafac1585cb617979a92c4d445a46a931005f61f

Observation 50be0de3-0ac9-46a6-a88a-d9493641a75e · inbound

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval cites this paper.

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:21:21.345749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T12:18:06.259724Z digest=sha256:2bd2b18dfb69341025ce8781ec6bb84ae5b1d84f0a8d8b653ad6e1e19f303c76

Observation a026177b-d17f-4e53-a456-1f2c21502a7a · inbound

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models cites this paper.

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:17.224410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:17.224410Z digest=sha256:ba67ef87fb8e074f6d5579ba84da817dad5dbb641d22d0b37650b54e2c8d20f6

Observation bac79f28-4979-4612-8ae4-10343a704e77 · inbound

Contrastive vision-language learning with paraphrasing and negation cites this paper.

Contrastive vision-language learning with paraphrasing and negation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:38.983791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:38.983791Z digest=sha256:8a79d6a2a35f26264bc4259d6d1b3c71fa0494de93cbe27c2cb471c52895e8c8

Observation 5df9f551-6620-434d-a92a-ae991542f506 · inbound

SPHINX: A Synthetic Environment for Visual Perception and Reasoning cites this paper.

SPHINX: A Synthetic Environment for Visual Perception and Reasoning When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:21:30.812869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T04:19:26.808804Z digest=sha256:0902d371a5081a12b321a409c7cf097b8ab7ec973ece4fdcbf133feacc1238aa

Observation 60f5526d-83ee-4184-8ab0-574212d155b7 · inbound

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models cites this paper.

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:51.598480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:51.598480Z digest=sha256:3ef8b2bd3cf3446922c5f93235c74d9aa53b91f86aeb4cb655a2d3f7243ad9dd

Observation 22650ba6-8ce3-4633-843e-02d983482f9a · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:21:18.885153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:8913d1d5e1bd811b26bdccd132d2d422a596ddd06b28077f959c9ab87f49bc62

Observation fd44e1be-d98f-4032-8e77-aff850b1e7f6 · inbound

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation cites this paper.

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:03.873100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:03.873100Z digest=sha256:d5f035e0ee141545969e9ec63d61bef47cff09bac681f4581598b6928aa57ac7

Observation ba776f33-078c-43d5-a1df-22b946cefd29 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.781116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.781116Z digest=sha256:ad0e6d0c790e65c3640e87aaed510fe60cd4ad4a350f2c9bebaad0ff7c7503da

Observation 3f5f8a49-aa49-4479-a17f-bc0fdcfb7987 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:9c45c611849bacdee57b3d296292b02b8aeb7450e4b810856191082f950d1901

Observation 24ebf0a7-2571-49e4-beb6-ea2d10b44608 · inbound

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs cites this paper.

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 35

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T09:19:53.771783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T09:18:34.853946Z digest=sha256:4ef063ee6918e64c7414726e9840ea6c3b2fb342c94a01a84344745a1e42ab77

Observation a984be64-eca3-412b-9df6-7fb35f0d700b · inbound

Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning cites this paper.

Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:48:11.172994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T19:48:03.278822Z digest=sha256:fc3d58cb775f59c536c01ac45f8080209260fa131ac45eaf97a3211acebaa212

Observation 11950544-6bb9-47fc-b7f1-1b33fdc3dfb5 · inbound

NSFL: A Post-Training Neuro-Symbolic Fuzzy Logic Framework for Boolean Operators in Neural Embeddings cites this paper.

NSFL: A Post-Training Neuro-Symbolic Fuzzy Logic Framework for Boolean Operators in Neural Embeddings When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:04.121070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:52:36.777259Z digest=sha256:516f77bb51d067665a186dfa0a8ca8ea922af375e5a80e0db5cb8cc526d4dbd6

Observation a31ff7df-3579-42df-8726-b996d065e135 · inbound

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding cites this paper.

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:31:03.898475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:26:55.369840Z digest=sha256:845681137951812e8cefea2995e6fcd9ce0e5fe826676427a9341bbcbe806d0c

Observation 8ab51118-45e0-44e5-985c-f3a8e5f74b6e · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:36:04.871204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:46db017198b20b99d1f5973507dd6c56691ab8621af296ad49d6c1dae2823ce1

Observation cc21d4ef-fd35-4d23-8679-164dbccdc380 · inbound

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce cites this paper.

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.741519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T00:55:56.146885Z digest=sha256:4288305804ae2a1f062c7ea2c79a02b255e772a9b2baff1d8daaa45aa085c3f9

Observation 5fd61976-d684-4fb7-a2b1-815f2c3b182c · inbound

DCR: Counterfactual Attractor Guidance for Rare Compositional Generation cites this paper.

DCR: Counterfactual Attractor Guidance for Rare Compositional Generation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:01:11.745909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T13:06:56.964670Z digest=sha256:3c1250deeb8ebbc3da7c318ff0e4394db00248767966c0dc79380c3ecae58734

Observation 8ac142af-b99a-4ded-bb73-7c7862bdc22d · inbound

Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs cites this paper.

Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:27:29.700340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T07:26:02.947081Z digest=sha256:8bcb9e0141fa069fc984b1c3b2d9eb482b32093822bf9f387f4d4becb0b682ff

Observation 29fa1dec-d6fb-4b8e-a0f3-2a0850ee0988 · inbound

A Composite Activation Function for Learning Stable Binary Representations cites this paper.

A Composite Activation Function for Learning Stable Binary Representations When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:07.934212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:03:42.456988Z digest=sha256:c4940ae03da43c591fa97e8274b199f13d7dce7f3dcd0daecc7f421bf0c6b0ef

Observation e9375fac-c45b-4910-824a-7ea38570c089 · inbound

Letting the neural code speak: Automated characterization of monkey visual neurons through human language cites this paper.

Letting the neural code speak: Automated characterization of monkey visual neurons through human language When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:12:07.069071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T02:12:00.335761Z digest=sha256:91b0d6542cf204dc176f7ddf171bd2b88a35c43757bf3a481380777d2e61cf62

Observation deb95c45-a754-48f7-aa21-abff4ff4c1fb · inbound

Letting the neural code speak: Automated characterization of monkey visual neurons through human language cites this paper.

Letting the neural code speak: Automated characterization of monkey visual neurons through human language When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:02.977300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T21:17:53.615009Z digest=sha256:c7ddb6e4b67c773a7830abd21a5d20a01c0c602a0dd6df0f945841b9a013fc77

Observation 74997e67-f9ae-4f3a-9c1a-35e44398aed0 · inbound

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning cites this paper.

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.639657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:32:00.361991Z digest=sha256:bcc6a09e7705228cabe63a4f9fff7e44c7fbe1aab22f3b9bd360cdb1ba76cf7f

Observation 73d649b2-e73c-4c8d-961c-d21814af32b2 · inbound

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency cites this paper.

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:27.778099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:35:08.286755Z digest=sha256:a230de9f31338583f5aa8df4211be71364777dc0ca03af452b054612c5e6f6cb

Observation 69e4e4ef-8b55-4479-bf43-c039682694d7 · inbound

When Vision Speaks for Sound cites this paper.

When Vision Speaks for Sound When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.772599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T22:12:52.160596Z digest=sha256:9788c010c5e6f16ede6e7fae366cb919e17a3f4d154a3a58b66a44a65899310a

Observation 3beef4a7-9412-4812-8e8b-6b6c89f5b183 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:13:16.266930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:49e78385ef4ce1d2cd704eaec6f6b86b73455864f348753c89329d8217e3daa8

Observation 8d80a690-8f81-4c46-b93c-095273bc7ce0 · inbound

Advancing Creative Physical Intelligence in Large Multimodal Models cites this paper.

Advancing Creative Physical Intelligence in Large Multimodal Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:34:02.910384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T21:09:37.228058Z digest=sha256:9989cb81490361f46400960d5bc51774f8cd8550d29cc43dbdf926ebbdf45611

Observation 97ce575d-8ba8-4720-8fef-d05a526d1cd5 · inbound

A Systematic Study of Behavioral Cloning for Scientific Data Annotation cites this paper.

A Systematic Study of Behavioral Cloning for Scientific Data Annotation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:38.982894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T16:23:08.402194Z digest=sha256:9c1345db72f0845cd7ca2c68e17baddf2e43c049e346a475e0906fd4dc313a24

Observation 068f2855-f599-4563-8fce-1f71c2db2db8 · inbound

Compositionality Emerges in a Narrow Depth-Connectivity Regime: Architecture Constraints and Solution Manifolds cites this paper.

Compositionality Emerges in a Narrow Depth-Connectivity Regime: Architecture Constraints and Solution Manifolds When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:29.816466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T18:22:56.676469Z digest=sha256:2d6d72ae4a7ce5602ec216f0940dd342561674bce9a4fa6f37087833dd8974da

Observation 0ec36d21-fe43-4b22-b26f-4a82aba2d2a5 · inbound

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs cites this paper.

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:29.295857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T18:08:37.978569Z digest=sha256:2c1d97c22bd272b93bb9bddd2b97e13fa1956e1c0d151832390a613c011165e9

Observation 099a458a-b108-4c3e-a69f-0695d16a60e3 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.738424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:a466d818767d5cc0de285caf39ba3020ba10df0b1aa82f7e8f53c1283cc62a26

Observation 855a6172-026e-477f-b46d-b12025a6dd25 · inbound

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors cites this paper.

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T05:54:18.335565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T05:53:26.394245Z digest=sha256:6b3d18bb44c46b8fa2109326ea0f587ca7a7d7bb6af2d388bd0d9538786238fb

Observation 3ee4f559-ea08-47b3-9dc0-b4a6325d463b · inbound

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP cites this paper.

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T19:36:29.283063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T19:26:49.805499Z digest=sha256:c0ff0c93f5f45a144e11d89bc574d65c0712454d4362d1dbb9d2d1e030da6ed4

Observation 6a0b2fbd-b3b5-42cb-b2ee-62af29f5c5bd · inbound

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP cites this paper.

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T15:51:57.429693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:51:57.429693Z digest=sha256:18f7d448b1167b945d34ac10fe34000db80b967461a6952c84196830a138f799

Observation 9105c657-07af-48d0-9cae-0870952e771e · inbound

CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification cites this paper.

CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:13:02.469296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:13:02.469296Z digest=sha256:fe7a925853f1ff2ab07760cc1614512344790c6d66913a3f96a921a1058c51f6

Observation 4164caff-3f99-47c4-8d2d-6a1baa7a8a81 · inbound

Trajectory-aware Cross-view Geo-localization with Sequential Observations cites this paper.

Trajectory-aware Cross-view Geo-localization with Sequential Observations When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:57.326327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:57.326327Z digest=sha256:e55234566176f7d6fcf4414a5e2a2be3ff368935a47e531ee473dcc25f02c936

Observation 27680bac-a6b1-4bf0-8faf-6e63a090ef6f · inbound

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models cites this paper.

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-30T23:29:03.410996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T23:29:03.410996Z digest=sha256:ad48f87672aea6f3a361ba0020975761b20277524f1d32917ea6c0db21ac6be9

Observation 40d68c93-ccc3-4738-8008-65f4bbac032b · inbound

Foveated Probes Recover Localized Binding Information in Vision Foundation Models cites this paper.

Foveated Probes Recover Localized Binding Information in Vision Foundation Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:24:47.569466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:24:47.569466Z digest=sha256:dd625592b587376de158ff1889ad7777e0d5cf57ee90271386461bb8f42aed5d

Observation 0a9b5cbe-0297-4743-80ae-cc5f2ef9f23f · inbound

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning cites this paper.

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:10:13.240029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:10:13.240029Z digest=sha256:37cc6047e5eaab583966f5b28842ddeb7a5fe6fb5972cade08b7948423633948

Observation 8d2b224c-6208-4a48-b279-761cbc5e347b · inbound

SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation cites this paper.

SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:29:19.221100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:29:19.221100Z digest=sha256:3c73d3fe2f3bc4a0cbea81351d91927fbdda68934f9b7154d3f8d3bb3a3b1011

Observation a2e2df52-ce0a-4a6e-95b3-26c4c3e16a94 · inbound

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination cites this paper.

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.565379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:14:42.565379Z digest=sha256:afab6809d264a383c3b3037575e9d13bfa5c1b8b323c45993f5ccd084a4aecc9

Observation f7e661eb-b70a-42f9-85e1-b675b0a4ffa6 · inbound

Rethinking Text-Based Image Retrieval in Specific Domain cites this paper.

Rethinking Text-Based Image Retrieval in Specific Domain When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:23:02.412149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:23:02.412149Z digest=sha256:cb9a21b3cd58a7fd68c291b3c4ecc6fcaace38378cf5d1b6f0a9427acad94a39