Pith. sign in

Paper Citation Record · LEDGER

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 2 inbound Pith citation observations for arXiv:2412.08111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08111 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:17:18.706692Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:03:50.311891Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:31.463894Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy46
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6019bf36-f00a-4b30-b782-93305f5e8777 · outbound

This paper cites Is bert blind? exploring the effect of vision-and-language pre- training on visual language understanding.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Is bert blind? exploring the effect of vision-and-language pre- training on visual language understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.434314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.479900Z digest=sha256:d73a76614a7435dd1cb101c310f7e56723446f9f8624c891bdee70bb8f717464

Observation 4eed0d19-faf3-46b5-a20d-4a21a4c5c149 · outbound

This paper cites Probing for constituency structure in neural language models.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Probing for constituency structure in neural language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.422276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.484593Z digest=sha256:6373ae62e42c615be321ebb7065ec52eca4430c5b72a4c1ef248c63b1032640b

Observation e646a91a-fbc2-4911-90c8-ca0bc6699a50 · outbound

This paper cites Multilingual nonce dependency treebanks: Under- standing how language models represent and process syn- tactic structure.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Multilingual nonce dependency treebanks: Under- standing how language models represent and process syn- tactic structure

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.411796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.488458Z digest=sha256:5193914df31281fa64e6026b7c5f8dec8a7d31498a3ec135848681fcd984dc90

Observation 3d6b22f7-4b94-446c-9886-aac1077dfafb · outbound

This paper cites On the difference of bert-style and clip-style text encoders.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models On the difference of bert-style and clip-style text encoders

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.398997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.491912Z digest=sha256:e433016458460f211a1ce90707d0300163025ea66fecb23c1ecba573ebac43d7

Observation e8cba4eb-a3ec-47a6-9c69-b14f7948943d · outbound

This paper cites Chi, John Hewitt, and Christopher D.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Chi, John Hewitt, and Christopher D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.388244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.495524Z digest=sha256:bd117fc657182575bc2b8a037b33115c119401919c6549bd4dcc69627f329ee7

Observation 71378638-50fb-4638-b466-0162390bdcf4 · outbound

This paper cites Manning, Joakim Nivre, and Daniel Zeman.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Manning, Joakim Nivre, and Daniel Zeman

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.371160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.503732Z digest=sha256:5cf6c6161cffbed37c5c5c4cf6d33510f93d453cc3e09f93ab3e286a9208b592

Observation 9fe3ef78-d8e8-4c99-b5bf-75dd78c927e1 · outbound

This paper cites BERT: pre-training of deep bidirectional trans- formers for language understanding.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models BERT: pre-training of deep bidirectional trans- formers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.361151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.507619Z digest=sha256:411c4739d546ff0d168aebc68acd5a4b022f541c8949320a2e37ae7fc7a542b0

Observation bf7b950d-4060-4efe-b4d9-8340c5c986f0 · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.350894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.511367Z digest=sha256:4561ecf3c19d876b3d23c60852139c96a038d6e965401ae93209614f88ff8bdf

Observation e3cf3b0a-fd45-4b6b-b48e-34ac5c0be9fb · outbound

This paper cites SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.515449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.515449Z digest=sha256:700a7b8b6408fe898e5616ca525459be983a2cf953991d4ac05c979af78f3bae

Observation caca6ae7-7e20-4daf-bdb1-d0e7cc91d130 · outbound

This paper cites Colorless green recurrent net- works dream hierarchically.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Colorless green recurrent net- works dream hierarchically

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.341328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.520505Z digest=sha256:c2585896835c85fb7aded15966226a100e077ab49fbdccb1fa26d72559e4889a

Observation 3a31ba6c-75b1-4622-be6e-684d8b3f0611 · outbound

This paper cites Sensi- tivity of generative vlms to semantically and lexically altered prompts.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Sensi- tivity of generative vlms to semantically and lexically altered prompts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.329383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.524801Z digest=sha256:9ba4dd4a8d621f69577e96dfcd8f63bf0581e1d2930812d0bdd95ee5032bba06

Observation c1405538-ff00-4ef2-aaea-7a105d0dff9b · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:17:19.319189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.527968Z digest=sha256:b6b2738991132893af03b07e7afb20600cd5157b4be13d8dd0119a96384f0f24

Observation 677c0db1-90c0-40b7-a157-a2acc708d23e · outbound

This paper cites A structural probe for finding syntax in word representations.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models A structural probe for finding syntax in word representations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.308438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.531413Z digest=sha256:c1df6105018f8c4fad109a3e1aeb14cd12b7e80effb7934690af59eee9ddd277

Observation 56c5bccf-9f3c-4f60-9fb9-caa52f971d4c · outbound

This paper cites What does bert learn about the structure of language? In ACL 2019- 57th Annual Meeting of the Association for Computational Linguistics, 2019.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models What does bert learn about the structure of language? In ACL 2019- 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.297044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.535029Z digest=sha256:f9f187b2d93bb5ad1cb74e8767bf46ac4a0dde0becad965bce388b4cbfdb0e90

Observation 4db364f5-c572-4b5b-a255-f4d78eb9ea8d · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.286419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.539492Z digest=sha256:c86413cd2c58083f3e50b9357d6b80abb162a76a6464402fdb371df9598998bc

Observation ccb044e4-0f1e-4e8a-89da-f3ab2c09700e · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:17:19.275226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.543043Z digest=sha256:344c658c6b1bc6bd6dd4671afcf36d7077b73d6ccec38dc6f785448bf4305fbc

Observation f1ed2467-88fd-4bf2-b118-fad102901963 · outbound

This paper cites Which sentence embeddings and which layers encode syntac- tic structure? Cognitive Science, 2020.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Which sentence embeddings and which layers encode syntac- tic structure? Cognitive Science, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.265510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.546500Z digest=sha256:92d12712b7ca802f1f196df2081bf4dc12e6c164fb1b3d2e71e05e0b9730cd28

Observation 133d3228-15a9-48e8-927a-4534ac325bad · outbound

This paper cites Schr¨odinger’s tree—On syntax and neural language models.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Schr¨odinger’s tree—On syntax and neural language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.255855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.550133Z digest=sha256:6b13f6ce6ecd06ee612a8d59a2a186242a1436f4314c1161b9315aeff8fb8d16

Observation 48a3b413-c8e1-40f7-b350-a08af8dcc826 · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:17:19.245882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.553669Z digest=sha256:2e199d8a1ed22db32bbc9b659a21651e7973ef20f7708bd0564288c64036df0c

Observation 6a1956c3-93f8-4a8d-ab4b-590c99d5d7dd · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:17:19.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.557348Z digest=sha256:ee9f02096f1acb3160ddc54d28a40aacd339fc5d06c66da13f3d61bc69897cab

Observation 8f9fac10-ee92-45e2-bc7a-639fec5a99f6 · outbound

This paper cites Do Vision-Language Models Understand Compound Nouns?.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Do Vision-Language Models Understand Compound Nouns?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.560850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.560850Z digest=sha256:45bf22b4f222fe0be3b7351618e240c3f125ca162300f3e8f363fb131778d5c2

Observation 25a91ce7-b2fd-40f6-a32b-550ad0f48fca · outbound

This paper cites How is bert surprised? layerwise detection of lin- guistic anomalies.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models How is bert surprised? layerwise detection of lin- guistic anomalies

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.223254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.564664Z digest=sha256:77cf4b07e1c6b2ddf6a8d1e9b504867b848f470d385aa8ea4d60e19577bf7bb7

Observation ad616650-589a-44ae-be2e-6d33b2fd00d3 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Align before fuse: Vision and language representation learning with momentum distillation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.210718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.567757Z digest=sha256:1399985ff946ee8ff96b6823b56ce2eb9b75b0aa5555410d8030b0384eb95b33

Observation 46398e1e-8f8f-450e-b40c-8097809491b9 · outbound

This paper cites BLIP- 2: bootstrapping language-image pre-training with frozen image encoders and large language models.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models BLIP- 2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.198846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.571182Z digest=sha256:1cf9a95fe1ffd1d3523f27e57e4d0cb0fccdc0ba2e91e780bb6ce9dfa733eae4

Observation d313713e-8e90-4089-b5e6-c751fabe1aab · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted CLIP.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Open-vocabulary semantic segmentation with mask-adapted CLIP

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.187249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.574332Z digest=sha256:00cf66417d580b5bb11017b552fa4f570cfbebe7525b4a47562bfc42a36d4069

Observation 7e1cdd25-9d3f-4f16-931f-fa74357cd5fa · outbound

This paper cites Syntactic structure from deep learning.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Syntactic structure from deep learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.177707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.577550Z digest=sha256:574afad1397fb15eac66fb7f5179de661b482023f69dd98363e3780d98964d25

Observation 5f946700-7bed-4cc2-9abe-e2606de341c4 · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach, 2019.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Roberta: A robustly optimized bert pretraining approach, 2019

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.166960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.580868Z digest=sha256:54c58bf08486c5793e9c8730f7c569228ec084ea8c0be9a23916a79cec4cd7ca

Observation 117266b7-204b-4788-b677-0da0499104f7 · outbound

This paper cites Ivanova, Idan A.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Ivanova, Idan A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.155563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.583978Z digest=sha256:3b16d5fd3c72519ac760b31d6226b0bacd4d6002edf651543121c94de704c574

Observation 4fe35e11-429c-43e9-b133-d07860532979 · outbound

This paper cites Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.145111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.587732Z digest=sha256:90e834da0e43f0939f8b6eb7abeb351790c525cc522377a44721a639b0fd22ce

Observation 5408055f-c43a-46e5-b383-c558767329d5 · outbound

This paper cites Probing for labeled dependency trees.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Probing for labeled dependency trees

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.134062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.591332Z digest=sha256:04a0b57d899561ba3daf6361c9d8ebb16c6e0c5a0bbba4699e3b43609fdbf985

Observation 2b125f22-2375-42f9-843a-3a7792f93625 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.123192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.594879Z digest=sha256:5a0e81092574a7026ef7c763d3009e492c675468a203823c2e7723c75c63b52d

Observation c40b0a3c-c1d8-4ece-9279-d9c1d5ea36dd · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.598938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.598938Z digest=sha256:675833199d79065e58ddc6ae2f00f26757b781e75078dc044b5c983ca0bbb723

Observation 3787f038-ae33-4800-b201-95417e4d5cd4 · outbound

This paper cites COLA: A Benchmark for Compositional Text-to-image Retrieval.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models COLA: A Benchmark for Compositional Text-to-image Retrieval

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.604010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.604010Z digest=sha256:fe04a51082afb9f5ed60dadb5a18ea6a56a1060746a1e671ab65841975d2e34c

Observation a9618cb7-6cdb-472f-bfbf-d27a4a45bf9f · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.609169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.609169Z digest=sha256:849d9792cc55ad038e95031e22d056df5943f3da0fbf471878f49ee052ca9d0d

Observation 41e87015-ed3a-44f9-9a33-8957bf81fb91 · outbound

This paper cites Pho- torealistic text-to-image diffusion models with deep language understanding.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Pho- torealistic text-to-image diffusion models with deep language understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.103820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.613343Z digest=sha256:edb3d6cea0609161143568724b2a8bcfdda3bc5b3db4381b1ea4cda9587f5ecd

Observation fb973c5d-be57-4494-b1b1-76da5e961fa2 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.089775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.616962Z digest=sha256:d41e6babbdcaf1576e578f8c4157e2e1520c3fb62577bb525739d928be240cd5

Observation 7822352c-70f4-43b9-ac57-0043fb4a0abb · outbound

This paper cites A gold standard dependency corpus for En- glish.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models A gold standard dependency corpus for En- glish

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.075282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.620442Z digest=sha256:03e4a74a48c675b26edb284ce57d18fd42674d848331e1cd9ae5cb58a02c847e

Observation 92b7924f-07bb-4102-95ee-707b1b530efe · outbound

This paper cites Flava: A foundational language and vision alignment model.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Flava: A foundational language and vision alignment model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.062988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.623857Z digest=sha256:8958ea45a4cc09b863522efd25c1653d43615cf9a83a54c6df196ec0189d89b9

Observation dc01ae02-374e-49f8-a8e2-52702ec72184 · outbound

This paper cites Masked language mod- eling and the distributional hypothesis: Order word matters pre-training for little.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Masked language mod- eling and the distributional hypothesis: Order word matters pre-training for little

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.049766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.628174Z digest=sha256:fc13b39eed943894125f4884f1bbeba20e065a08a374ca717222e5188be30a01

Observation eccbb23f-cb63-47c5-a8d5-41c8517b142b · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.035755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.632017Z digest=sha256:20e3b1854a670f58b2980286ad087158031238d04668ba22c3da31afc77550d0

Observation a26be74c-d5c4-4b1c-9571-742184488dbb · outbound

This paper cites Diffusion lens: Interpreting text encoders in text-to-image pipelines.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Diffusion lens: Interpreting text encoders in text-to-image pipelines

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:19.022176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.635341Z digest=sha256:decd2cb397c577565dc677da3ff1d2689b2491ef13e4340f66ab179f0e670c3d

Observation 09e6efa3-3876-4da0-bd53-ee661e9a3ec4 · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:17:18.999522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.643883Z digest=sha256:762482f1c17e2af4bfef8e9cd5616d65257ed40ae94b78c4aa68523ad24db657

Observation f2cc2e4c-c08a-4a81-b9c3-2ab518d250a7 · outbound

This paper cites Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:17:18.752866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.647520Z digest=sha256:53c5c8aea7de3411cadd89317ba69cf826468c641eed45194736ac5a9281b2d6

Observation d64dbb61-b9fa-43b3-8413-3c62234ef341 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers, 2020.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers, 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.987321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.651685Z digest=sha256:511a32c0fa1d06963f202d9d4bbec4e71d3b80af4b05c22f423b7818072ef955

Observation bd83f3a9-6cc6-4b8f-9a41-596d762ec943 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.975073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.654595Z digest=sha256:71f661ba1848a29fcf859c3308e050c5f512541c653635f2768a11c6a32b6609

Observation a418555b-8f7d-4cb4-8f58-b38d546dc182 · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.657659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.657659Z digest=sha256:b683748a9a1767d6087d6f863d34bc5c8ae1890e40d7e0e21cf5d8c6bd05dd55

Observation 5c3fc2a1-b756-4885-88cc-545aeed48329 · outbound

This paper cites We evalu- ated the ’ViT-B/32’ variant of CLIP – ViT base model trained with a image patch size of 32 – publicly avail- able at the following HuggingFace Link.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evalu- ated the ’ViT-B/32’ variant of CLIP – ViT base model trained with a image patch size of 32 – publicly avail- able at the following HuggingFace Link

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.952195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.664228Z digest=sha256:e45856f641f29990bf2cac7c27b7a16da8cfe0b1615998c61b83f7832c79559c

Observation 90e9cde7-665f-4a2d-8531-a047f4c0dd79 · outbound

This paper cites We evaluated the FLA V A pre-trained Model available at the following HuggingFace Link • Unimodal Language Models (ULMs).

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the FLA V A pre-trained Model available at the following HuggingFace Link • Unimodal Language Models (ULMs)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.940988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.667682Z digest=sha256:a5165633f8a2cec71ecc525086dfbdcf2ba6ed85d6d259995a92e5aee7c0bfe7

Observation 1cb343cf-0ceb-40d7-8adb-ee2c083ca917 · outbound

This paper cites We evaluated the RoBERTa-base model available at the following Hug- gingFace Link.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the RoBERTa-base model available at the following Hug- gingFace Link

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.930226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.671158Z digest=sha256:d08ba39a76fb769472d1e09533ed2aadda7c3e0ad485f73758b89f884cbc006b

Observation 8ca19960-d6ad-4cb7-aaa7-d97cca0f17c7 · outbound

This paper cites We evaluated the RoBERTa-large model available at the following Hug- gingFace Link.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the RoBERTa-large model available at the following Hug- gingFace Link

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.918174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.674810Z digest=sha256:f4299db87ef13343cca610e0e7fc9cdac77d63030061f20362a495a621ac62f4

Observation 1afe41bf-d967-471c-9c1c-fe4e47ae9c23 · outbound

This paper cites We evaluated the MiniLM model avail- able at the following HuggingFace Link • Sentence Language Models (SLMs) [34].

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the MiniLM model avail- able at the following HuggingFace Link • Sentence Language Models (SLMs) [34]

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.905951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.678733Z digest=sha256:9c3f69508c2c2e027a4c95dee336f7399b9aea205cec64fad07ccb0e7b0658a5

Observation e610d0ee-a44f-49cc-a46f-c99e0f57f13f · outbound

This paper cites We eval- uated the sentence-MiniLM model available at the fol- lowing HuggingFace Link.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We eval- uated the sentence-MiniLM model available at the fol- lowing HuggingFace Link

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.890664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.682486Z digest=sha256:04fa8480e1ac5429ed371e24d179022c7f2dcdb8b9734d9e59803a285890b46b

Observation 774188bf-b95c-4093-8799-c95eec0dfa3a · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:17:18.878568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.685782Z digest=sha256:3624e85753d4c824dc9698038af3ec666d3f896edfa5b7a2433a003a2d1c7db2

Observation ba3fe6db-30a5-4673-b72d-45470004cf23 · outbound

This paper cites Here the images are provided as input with a patch size of 32.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Here the images are provided as input with a patch size of 32

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.865292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.689005Z digest=sha256:0393ef85eaf798674e596cc4e9064b5acbbb3b061c5f920ccc976b92c2a7732e

Observation fdd8f2ef-8711-48ac-ab49-3bcee58c7206 · outbound

This paper cites This model is also trained us- ing the WebImageText dataset [31] consisting of 400M image-text pairs.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models This model is also trained us- ing the WebImageText dataset [31] consisting of 400M image-text pairs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.853085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.692477Z digest=sha256:9542165a1ca0a59d3a938bc1405b38548ccf9008b479673d8fd77f10ebe77548

Observation f59ff747-1040-4bc8-836d-fa883f722f00 · outbound

This paper cites We evaluated the LAION-CLIP-ViT-B/32 model publicly available at the following HuggingFace Link.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models We evaluated the LAION-CLIP-ViT-B/32 model publicly available at the following HuggingFace Link

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.841819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.696111Z digest=sha256:2e2c6af478a01584d392b1cfbd2e7b4113b2e38f0bdf3c8de5e5e058aadea89e

Observation 5768344e-c8bd-4ae8-b21e-f62ac05dd190 · outbound

This paper cites The pre- training process utilized 5 billion image-text pairs from the LAION dataset.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models The pre- training process utilized 5 billion image-text pairs from the LAION dataset

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.830896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.699387Z digest=sha256:9134f320e63b35b42e30c66049a1367abc9bf415c4af040f1501258735d6a18b

Observation 8274126c-057b-468e-844d-0591854cc171 · outbound

This paper cites The pre- training process utilized 5 billion image-text pairs from the LAION dataset.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models The pre- training process utilized 5 billion image-text pairs from the LAION dataset

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.818325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.702763Z digest=sha256:fa532fc03ec85e1154a23139f0b32994da4bf8ba0c1fc7d02fab1d2850d303a8

Observation 59378c4d-841c-4fba-9c5e-75da82a50f9c · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T18:17:18.805286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.706692Z digest=sha256:034f9118802c28d6c2285baef86ec6cc578e216cd6a1dec87d25e45ee94682f9

Observation 323ccb49-4f2c-4eee-a333-82ac94525e1a · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.499800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.499800Z digest=sha256:c07dfd68f22ec2feb6604e88be7dca528836dbe407909c77de950486c4b3dedc

Observation ca105c4b-8f26-4afb-ae78-2d71f4eab987 · outbound

This paper cites Implementation For instructions to run an example probing experiment, please A.1.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Implementation For instructions to run an example probing experiment, please A.1

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:17:18.962354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:17:18.661287Z digest=sha256:90c86f5f989889947fc4b5c027408d765bcd643f36f160ca38a1b0f9b23f81cd

Observation 02fb1b29-5f8a-4f94-a7c8-2f0362d632fd · outbound

This paper cites an unresolved cited work.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.639779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.639779Z digest=sha256:774a40bd66a0e343d36e01aac7600bf853d41ca427ad918d3920210e715af706

Pith citing papers

Observation 7dbc94a0-1715-4f50-b44f-1103a5f822db · inbound

Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding cites this paper.

Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:48:27.693701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T23:48:14.217344Z digest=sha256:fb4b296fd408a9ce8bde33bdb4eb6dd96ea1a6ea10d02af2fcf7232be7385d88

Observation 4e8c99a8-b049-46bc-91c0-30a2ee5c68f3 · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.465164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:e212eb76a4ff0e0acccf7b0e236e3d94440042b123e0103348b09862256793ce