Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 69 inbound Pith citation observations for arXiv:2210.01936.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:52.486750Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T19:36:29.281490Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 1c5deeee-8208-4051-8fa1-6835cb49f8d4 · inbound
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a26a43-d54f-4170-b4aa-97439c27ece7 · inbound
Enhancing CLIP Conceptual Embedding through Knowledge Distillation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50cdc964-c232-4af0-8d39-bb20326e98bb · inbound
VladVA: Discriminative Fine-tuning of LVLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877fa6fe-ab78-42a7-8366-1b8646e0b093 · inbound
RelationField: Relate Anything in Radiance Fields When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c27b363-2f48-49de-8803-0715862d5e01 · inbound
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29cd9db1-eda2-4a78-ad7d-ccab212b63f4 · inbound
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a69c6cc-4c96-4ed1-b5c7-fb92fb010c16 · inbound
FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a402747d-6017-4b4e-a935-9af29a8907bf · inbound
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4323f224-8df8-4547-8e42-746b7bad621f · inbound
Decoupled Global-Local Alignment for Improving Compositional Understanding When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44242c63-bf64-44cd-9389-36287455d129 · inbound
Multi-Modal Language Models as Text-to-Image Model Evaluators When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d722abd3-f017-4e8d-a970-bc9b23a064c6 · inbound
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff164eb-6c6a-4acf-8cb4-bee651993ea6 · inbound
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b905e8a-b3f0-4abc-a0ef-56ac6cc61eab · inbound
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5978c2a7-5ae2-4d2b-a6dc-1796a0f67198 · inbound
Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2650a0af-46a6-4c1d-a009-fa9c1ab10f53 · inbound
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d5b03d-47ce-423f-916d-6834d1354101 · inbound
From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceca9211-8036-4760-80a7-1391c7ff2020 · inbound
CIVET: Systematic Evaluation of Understanding in VLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed8fddf-ba6d-48cb-8c12-f3952ea698e1 · inbound
On the rankability of visual embeddings When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d636dd-8c85-446d-acd7-e7f440c730b8 · inbound
ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72a3b76-2f3c-414e-b73d-298ca1be321a · inbound
Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720d82eb-b1da-4161-9cf6-1fba16e76b26 · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d715b9ac-d1c6-4571-9849-5ace6a338478 · inbound
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444a16d3-f9e4-41d5-9686-898887bac9e4 · inbound
Negation-Aware Test-Time Adaptation for Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3e5cde-ade4-4b49-9635-629d88a904b2 · inbound
Trade-offs in Image Generation: How Do Different Dimensions Interact? When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b0e5a8a-4ba5-4085-a5dc-479711bd9875 · inbound
MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0787d3e-1fb8-4ba3-8582-6cbabf57d47e · inbound
Native Hierarchical and Compositional Representations with Subspace Embeddings When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a406c379-ad15-4816-8977-5b9deee8d4c0 · inbound
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca59d4c-535f-4bc2-94db-56f283c6cb16 · inbound
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8da70f2-04a0-4eb4-ad7a-0cb71a05e33c · inbound
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 50be0de3-0ac9-46a6-a88a-d9493641a75e · inbound
GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a026177b-d17f-4e53-a456-1f2c21502a7a · inbound
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac79f28-4979-4612-8ae4-10343a704e77 · inbound
Contrastive vision-language learning with paraphrasing and negation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df9f551-6620-434d-a92a-ae991542f506 · inbound
SPHINX: A Synthetic Environment for Visual Perception and Reasoning When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 60f5526d-83ee-4184-8ab0-574212d155b7 · inbound
Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22650ba6-8ce3-4633-843e-02d983482f9a · inbound
Adapting MLLMs for Nuanced Video Retrieval When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fd44e1be-d98f-4032-8e77-aff850b1e7f6 · inbound
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba776f33-078c-43d5-a1df-22b946cefd29 · inbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5f8a49-aa49-4479-a17f-bc0fdcfb7987 · inbound
Vision Language Models Cannot Reason About Physical Transformation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ebf0a7-2571-49e4-beb6-ea2d10b44608 · inbound
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a984be64-eca3-412b-9df6-7fb35f0d700b · inbound
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 11950544-6bb9-47fc-b7f1-1b33fdc3dfb5 · inbound
NSFL: A Post-Training Neuro-Symbolic Fuzzy Logic Framework for Boolean Operators in Neural Embeddings When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a31ff7df-3579-42df-8726-b996d065e135 · inbound
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8ab51118-45e0-44e5-985c-f3a8e5f74b6e · inbound
Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cc21d4ef-fd35-4d23-8679-164dbccdc380 · inbound
AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5fd61976-d684-4fb7-a2b1-815f2c3b182c · inbound
DCR: Counterfactual Attractor Guidance for Rare Compositional Generation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8ac142af-b99a-4ded-bb73-7c7862bdc22d · inbound
Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 29fa1dec-d6fb-4b8e-a0f3-2a0850ee0988 · inbound
A Composite Activation Function for Learning Stable Binary Representations When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e9375fac-c45b-4910-824a-7ea38570c089 · inbound
Letting the neural code speak: Automated characterization of monkey visual neurons through human language When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation deb95c45-a754-48f7-aa21-abff4ff4c1fb · inbound
Letting the neural code speak: Automated characterization of monkey visual neurons through human language When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 74997e67-f9ae-4f3a-9c1a-35e44398aed0 · inbound
SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 73d649b2-e73c-4c8d-961c-d21814af32b2 · inbound
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 69e4e4ef-8b55-4479-bf43-c039682694d7 · inbound
When Vision Speaks for Sound When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3beef4a7-9412-4812-8e8b-6b6c89f5b183 · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8d80a690-8f81-4c46-b93c-095273bc7ce0 · inbound
Advancing Creative Physical Intelligence in Large Multimodal Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 97ce575d-8ba8-4720-8fef-d05a526d1cd5 · inbound
A Systematic Study of Behavioral Cloning for Scientific Data Annotation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 238
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 068f2855-f599-4563-8fce-1f71c2db2db8 · inbound
Compositionality Emerges in a Narrow Depth-Connectivity Regime: Architecture Constraints and Solution Manifolds When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0ec36d21-fe43-4b22-b26f-4a82aba2d2a5 · inbound
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 099a458a-b108-4c3e-a69f-0695d16a60e3 · inbound
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 855a6172-026e-477f-b46d-b12025a6dd25 · inbound
Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3ee4f559-ea08-47b3-9dc0-b4a6325d463b · inbound
Sparse Attention for Dense Open-Vocabulary Prediction in CLIP When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6a0b2fbd-b3b5-42cb-b2ee-62af29f5c5bd · inbound
Sparse Attention for Dense Open-Vocabulary Prediction in CLIP When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9105c657-07af-48d0-9cae-0870952e771e · inbound
CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4164caff-3f99-47c4-8d2d-6a1baa7a8a81 · inbound
Trajectory-aware Cross-view Geo-localization with Sequential Observations When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27680bac-a6b1-4bf0-8faf-6e63a090ef6f · inbound
TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d68c93-ccc3-4738-8008-65f4bbac032b · inbound
Foveated Probes Recover Localized Binding Information in Vision Foundation Models When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a9b5cbe-0297-4743-80ae-cc5f2ef9f23f · inbound
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2b224c-6208-4a48-b279-761cbc5e347b · inbound
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e2df52-ce0a-4a6e-95b3-26c4c3e16a94 · inbound
A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e661eb-b70a-42f9-85e1-b675b0a4ffa6 · inbound
Rethinking Text-Based Image Retrieval in Specific Domain When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.