Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:17:18.706692Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 2 inbound Pith citation observations for arXiv:2412.08111.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:17:18.706692Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T07:03:50.311891Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:28:31.463894Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6019bf36-f00a-4b30-b782-93305f5e8777 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Is bert blind? exploring the effect of vision-and-language pre- training on visual language understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4eed0d19-faf3-46b5-a20d-4a21a4c5c149 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Probing for constituency structure in neural language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e646a91a-fbc2-4911-90c8-ca0bc6699a50 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Multilingual nonce dependency treebanks: Under- standing how language models represent and process syn- tactic structure
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d6b22f7-4b94-446c-9886-aac1077dfafb · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models On the difference of bert-style and clip-style text encoders
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8cba4eb-a3ec-47a6-9c69-b14f7948943d · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Chi, John Hewitt, and Christopher D
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 71378638-50fb-4638-b466-0162390bdcf4 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Manning, Joakim Nivre, and Daniel Zeman
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9fe3ef78-d8e8-4c99-b5bf-75dd78c927e1 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models BERT: pre-training of deep bidirectional trans- formers for language understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bf7b950d-4060-4efe-b4d9-8340c5c986f0 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models BERT: Pre-training of deep bidirectional trans- formers for language understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e3cf3b0a-fd45-4b6b-b48e-34ac5c0be9fb · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caca6ae7-7e20-4daf-bdb1-d0e7cc91d130 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Colorless green recurrent net- works dream hierarchically
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3a31ba6c-75b1-4622-be6e-684d8b3f0611 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Sensi- tivity of generative vlms to semantically and lexically altered prompts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c1405538-ff00-4ef2-aaea-7a105d0dff9b · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 677c0db1-90c0-40b7-a157-a2acc708d23e · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models A structural probe for finding syntax in word representations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 56c5bccf-9f3c-4f60-9fb9-caa52f971d4c · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models What does bert learn about the structure of language? In ACL 2019- 57th Annual Meeting of the Association for Computational Linguistics, 2019
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4db364f5-c572-4b5b-a255-f4d78eb9ea8d · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ccb044e4-0f1e-4e8a-89da-f3ab2c09700e · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f1ed2467-88fd-4bf2-b118-fad102901963 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Which sentence embeddings and which layers encode syntac- tic structure? Cognitive Science, 2020
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 133d3228-15a9-48e8-927a-4534ac325bad · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Schr¨odinger’s tree—On syntax and neural language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 48a3b413-c8e1-40f7-b350-a08af8dcc826 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6a1956c3-93f8-4a8d-ab4b-590c99d5d7dd · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8f9fac10-ee92-45e2-bc7a-639fec5a99f6 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Do Vision-Language Models Understand Compound Nouns?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a91ce7-b2fd-40f6-a32b-550ad0f48fca · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models How is bert surprised? layerwise detection of lin- guistic anomalies
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad616650-589a-44ae-be2e-6d33b2fd00d3 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Align before fuse: Vision and language representation learning with momentum distillation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 46398e1e-8f8f-450e-b40c-8097809491b9 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models BLIP- 2: bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d313713e-8e90-4089-b5e6-c751fabe1aab · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Open-vocabulary semantic segmentation with mask-adapted CLIP
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7e1cdd25-9d3f-4f16-931f-fa74357cd5fa · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Syntactic structure from deep learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5f946700-7bed-4cc2-9abe-e2606de341c4 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Roberta: A robustly optimized bert pretraining approach, 2019
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 117266b7-204b-4788-b677-0da0499104f7 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Ivanova, Idan A
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4fe35e11-429c-43e9-b133-d07860532979 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5408055f-c43a-46e5-b383-c558767329d5 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Probing for labeled dependency trees
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2b125f22-2375-42f9-843a-3a7792f93625 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Learning transferable visual models from natural language supervi- sion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c40b0a3c-c1d8-4ece-9279-d9c1d5ea36dd · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3787f038-ae33-4800-b201-95417e4d5cd4 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models COLA: A Benchmark for Compositional Text-to-image Retrieval
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9618cb7-6cdb-472f-bfbf-d27a4a45bf9f · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Sentence-bert: Sentence embeddings using siamese bert-networks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e87015-ed3a-44f9-9a33-8957bf81fb91 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Pho- torealistic text-to-image diffusion models with deep language understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb973c5d-be57-4494-b1b1-76da5e961fa2 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Laion-5b: An open large-scale dataset for training next gen- eration image-text models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7822352c-70f4-43b9-ac57-0043fb4a0abb · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models A gold standard dependency corpus for En- glish
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 92b7924f-07bb-4102-95ee-707b1b530efe · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Flava: A foundational language and vision alignment model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dc01ae02-374e-49f8-a8e2-52702ec72184 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Masked language mod- eling and the distributional hypothesis: Order word matters pre-training for little
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eccbb23f-cb63-47c5-a8d5-41c8517b142b · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Winoground: Probing vision and language models for visio- linguistic compositionality
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a26be74c-d5c4-4b1c-9571-742184488dbb · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Diffusion lens: Interpreting text encoders in text-to-image pipelines
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 09e6efa3-3876-4da0-bd53-ee661e9a3ec4 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f2cc2e4c-c08a-4a81-b9c3-2ab518d250a7 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d64dbb61-b9fa-43b3-8413-3c62234ef341 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers, 2020
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bd83f3a9-6cc6-4b8f-9a41-596d762ec943 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a418555b-8f7d-4cb4-8f58-b38d546dc182 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3fc2a1-b756-4885-88cc-545aeed48329 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evalu- ated the ’ViT-B/32’ variant of CLIP – ViT base model trained with a image patch size of 32 – publicly avail- able at the following HuggingFace Link
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 90e9cde7-665f-4a2d-8531-a047f4c0dd79 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the FLA V A pre-trained Model available at the following HuggingFace Link • Unimodal Language Models (ULMs)
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1cb343cf-0ceb-40d7-8adb-ee2c083ca917 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the RoBERTa-base model available at the following Hug- gingFace Link
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8ca19960-d6ad-4cb7-aaa7-d97cca0f17c7 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the RoBERTa-large model available at the following Hug- gingFace Link
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1afe41bf-d967-471c-9c1c-fe4e47ae9c23 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the MiniLM model avail- able at the following HuggingFace Link • Sentence Language Models (SLMs) [34]
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e610d0ee-a44f-49cc-a46f-c99e0f57f13f · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We eval- uated the sentence-MiniLM model available at the fol- lowing HuggingFace Link
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 774188bf-b95c-4093-8799-c95eec0dfa3a · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba3fe6db-30a5-4673-b72d-45470004cf23 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Here the images are provided as input with a patch size of 32
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fdd8f2ef-8711-48ac-ab49-3bcee58c7206 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models This model is also trained us- ing the WebImageText dataset [31] consisting of 400M image-text pairs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f59ff747-1040-4bc8-836d-fa883f722f00 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the LAION-CLIP-ViT-B/32 model publicly available at the following HuggingFace Link
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5768344e-c8bd-4ae8-b21e-f62ac05dd190 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models The pre- training process utilized 5 billion image-text pairs from the LAION dataset
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8274126c-057b-468e-844d-0591854cc171 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models The pre- training process utilized 5 billion image-text pairs from the LAION dataset
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 59378c4d-841c-4fba-9c5e-75da82a50f9c · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 323ccb49-4f2c-4eee-a333-82ac94525e1a · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca105c4b-8f26-4afb-ae78-2d71f4eab987 · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Implementation For instructions to run an example probing experiment, please A.1
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 02fb1b29-5f8a-4f94-a7c8-2f0362d632fd · outbound
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dbc94a0-1715-4f50-b44f-1103a5f822db · inbound
Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4e8c99a8-b049-46bc-91c0-30a2ee5c68f3 · inbound
Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.