Pith. sign in

Paper Citation Record · LEDGER

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.08000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08000 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:48.742052Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b196439a-1861-4339-b27d-57c23a0700be · outbound

This paper cites Learning transferable visual models from natural language supervision.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.403863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:44.404148Z digest=sha256:aca4cf7b935bcf8eec8faf1df75286b2658ceb58a301489fdb013febb9a2a1eb

Observation c88cb4d2-17f0-4076-a45d-90484707a1c3 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Reproducible scaling laws for contrastive language-image learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.526529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.526529Z digest=sha256:13466052e9c25ed7ce75c07de8bef02348362e064e7ad116c3731e1385d496cb

Observation c011eae3-7836-4a3b-b2ea-870a6c0caea6 · outbound

This paper cites GPT-4 Technical Report.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.629556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.629556Z digest=sha256:c24bd5395d44fa04bda854330bd8e5d224a7dffbbec78c2479edd65b5a79dc7e

Observation 0a4e9da3-a4ec-46ae-bc55-b4275ba7210d · outbound

This paper cites Visualinstructiontuning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Visualinstructiontuning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.252715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:44.743931Z digest=sha256:9459da7a395c5c95eb7ffadc7df9929fe03820d04bb59757ae47b222a9b9429b

Observation c41c3b77-5c8b-437a-9899-718f0553e9b6 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.845309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.845309Z digest=sha256:465752a24cd4909a2ae83a0e80824f17e6d7a2d317cd012549ab97118d770248

Observation 40584af3-8ba4-42ac-aa75-446cec0fdab3 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.933692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.933692Z digest=sha256:d694e8d73ecef0c8c746a46818545491efb43f31ebb7650e67b4b3c46ff15090

Observation 83d88bc5-375e-4231-84ed-706503d94dec · outbound

This paper cites Measuring Robustness to Natural Distribution Shifts in Image Classification.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Measuring Robustness to Natural Distribution Shifts in Image Classification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.071477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.071477Z digest=sha256:6494c9cead5f260e5a52db77a36b6c99d838104a1c8c719a98b2926f9c7d8bd1

Observation cdeac91e-f2f1-46a2-b25d-d012f3d9d14f · outbound

This paper cites The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.197907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.197907Z digest=sha256:7541c55b600b378a69610279ba9df88eb5560ca773804c6fc4cc402b71d74527

Observation 9b5d6ad1-921b-4dcc-8427-8a7f921c63ca · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.174976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.324151Z digest=sha256:1e5e9edb840b9e22cd54494f1fefd87b477aed43e1cc24baa3753e6a8b23efc8

Observation 1e89ab05-e4e0-417c-a567-7af021e4aec6 · outbound

This paper cites Data determines distributional robustness in contrastive language image pre-training (clip).

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Data determines distributional robustness in contrastive language image pre-training (clip)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.066083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.443919Z digest=sha256:99131d877ef4bc9f8660abb9c53b7ffba32242a874b790ba060801944c32fd77

Observation 6e24f533-29e8-4405-80ef-4b3ca13b1423 · outbound

This paper cites The neglected tails in vision-language models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The neglected tails in vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.583592Z digest=sha256:965e5a813caee6d9bba72cc7d49e41b5e513cf1c8123eb40e4fdaff8939cec53

Observation 98585498-f97c-4134-b086-41f08b14aaf2 · outbound

This paper cites zero-shot.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models zero-shot

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.840210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.659673Z digest=sha256:5a522e51a8523785b6dfdf0b6a3fbd79579c01caef99c8f7abd180d8fda25a97

Observation 4ee03997-7c78-4265-b26f-1a83e7a71c79 · outbound

This paper cites Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.745787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.745787Z digest=sha256:e4c5f34c409c6d97ae2e4f6e36a2bfe705cae95c5e93f717cfbc7d36d5418ca5

Observation fd56ea5e-25a5-4d91-9f4c-6d6187645e1b · outbound

This paper cites Deciphering the role of representation disentanglement: Investigating compositional generalization in clip models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Deciphering the role of representation disentanglement: Investigating compositional generalization in clip models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.676249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.842927Z digest=sha256:98a68c260d7923c7a4f97e1d03aa9119c6fb4744ff9857748c9373a31407f05a

Observation 25838902-9322-43f2-ba82-ce2c655a5bce · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.923104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.923104Z digest=sha256:e9ad4621d8637def5e8d95df5127ed52decfc24b5e6530251f8976761bf8eec8

Observation 6f1add8e-df0b-4439-932e-52027ab2d4e3 · outbound

This paper cites @ crepe: Can vision-language foundation models reason compositionally?2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10910–10921, 2022.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models @ crepe: Can vision-language foundation models reason compositionally?2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10910–10921, 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.096827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.096827Z digest=sha256:8b83e3d6c8abc2c7c0ed36372dffd11b14b25d3f5743780fc8982e78f39b7b1c

Observation 207f01f4-20ed-485a-81b0-a50c81a42ec9 · outbound

This paper cites Sugar- crepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Sugar- crepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.545135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.233303Z digest=sha256:db47814f01eaa0e027b04aaa09dec2f4fa825ae2ae02102ac93c6fb4ef8ca7e5

Observation 60d29984-4d5e-4197-be41-e06da54d3ebe · outbound

This paper cites A Sober Look at the Robustness of CLIPs to Spurious Features.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models A Sober Look at the Robustness of CLIPs to Spurious Features

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.319612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.319612Z digest=sha256:04661d6db5962b5cc6b387a34a90a120b426411d27628d1d4fcff6bfcec73168

Observation ba669502-0b95-4bcf-994b-d60523b9e356 · outbound

This paper cites Word association norms, mutual information, and lexicography.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Word association norms, mutual information, and lexicography

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.407432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.419843Z digest=sha256:c8df9619116a7c2b643371aeb07e1ea754cd0fc925e8a7e01529e0759dc34d28

Observation 28885a78-4344-42a2-8dc0-c171b6de1729 · outbound

This paper cites Improved baselines with visual instruction tuning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Improved baselines with visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.490093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.490093Z digest=sha256:e00310b6e48d8f3eb64398c005698b8921b72c499f1afb3810d1507b706a8248

Observation f16bf04a-f7e2-4076-9a3b-7985e3121b76 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.563950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.563950Z digest=sha256:f81e51b7b0870de8396c88fc6a9a7ae6737af6b12278788b8b385804d0fb718d

Observation b72d4b65-5d59-420b-a6f5-e07168d0e385 · outbound

This paper cites Two multivariate generalizations of pointwise mutual information.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Two multivariate generalizations of pointwise mutual information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.247942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.701812Z digest=sha256:fa4de9a21e782df25b5e4051e958d040ab38f3b1f4632d384273a018ba428822

Observation 1dd04bef-0152-489f-bd1f-a6b49fee3bc1 · outbound

This paper cites Lawrence Zitnick, Devi Parikh, and Dhruv Batra.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Lawrence Zitnick, Devi Parikh, and Dhruv Batra

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.128937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.817205Z digest=sha256:f08f0cad7b1c4f6e8c237452beda38946ea3906354bbc62b3a6b99fa049e3a48

Observation 2121567b-7e69-49a4-bde3-1243b2b32d97 · outbound

This paper cites The Llama 3 Herd of Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.919345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.919345Z digest=sha256:8a71a61ac164560e65c582776532cac50a487c0a952a532861493e3e58f4e7f1

Observation 2ca3ebd8-02d4-4d88-aed2-02bb93b9094d · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Flux.https://github.com/black-forest-labs/flux, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.040675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.040675Z digest=sha256:5f688f914d21b56a5384a499348c1424c1abdf2d0612b82eadb44e0b81684365

Observation 29bd6e86-cc0b-483e-b38c-3a47cc355b84 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.160787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.160787Z digest=sha256:0ff8e3b00b97552b095bddf68cfe30ea136be3d35a8c95957077700ae1e9a188

Observation 250aebb7-906b-4c67-9cf9-ee4314c75f4f · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.262926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.262926Z digest=sha256:dbbe51d787a95f828bed83090065fd97fa093ccd2a94d18e66d7f3ffa102add8

Observation 339cc28c-f88e-461b-a376-9e1d25dc31cc · outbound

This paper cites Towards vqa models that can read.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Towards vqa models that can read

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.999667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.314389Z digest=sha256:6fe8110b428c277318c0db5569144aada50c46f8ac8a7b47b5ea11ab01e9c6dd

Observation 09724e1e-11fb-4555-88d7-7bb5ee6c3fcb · outbound

This paper cites Recognition in terra incognita.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Recognition in terra incognita

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.800409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.418010Z digest=sha256:6088023cf1338a2c181a1516ed9d43ab7ad4d84147d5ea38f6d76c767c2b1a1a

Observation 266d094f-4272-4fc4-ad49-9c6785ed3322 · outbound

This paper cites Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study.PLoS medicine, 15(11):e1002683, 2018.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study.PLoS medicine, 15(11):e1002683, 2018

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.443100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.520665Z digest=sha256:080f2ace3d4a0abd6097e617784f646bbc8a6f9e025b1f9f997db713fa0e26b7

Observation 47a9c4c8-3021-4671-9691-255ae3e0b563 · outbound

This paper cites Hashimoto, and Percy Liang.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Hashimoto, and Percy Liang

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.603989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.603989Z digest=sha256:a261099778fa7e148228ee122a93a3ed6eb71a2053440b18b6cdaa7d6c193985

Observation 3eb9b301-b9c7-4258-9f7d-2ec1efb4b97e · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2020.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Shortcut learning in deep neural networks.Nature Machine Intelligence, 2020

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.185552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.674980Z digest=sha256:42c32bbf57cecc0000aea12d1a1872b47dcf992f3cc27309342363f51940f96e

Observation cd90cc6b-1161-4104-862c-38303358836a · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.863554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.791137Z digest=sha256:4f0a703bb330b2287b7a57f161651c8fadab9203025cd5d43f6f4eb991f12682

Observation a353761a-8675-4775-ba49-8e40afb2a124 · outbound

This paper cites COVR: A test-bed for Visually Grounded Compositional Generalization with real images.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models COVR: A test-bed for Visually Grounded Compositional Generalization with real images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.888275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.888275Z digest=sha256:47c7c7400a70820bbaffe32d52c0719147b27b7b5aeaecce7faaa22cc07d5979

Observation cddf270b-ab97-4fbf-939b-36699f877d03 · outbound

This paper cites Does CLIP Bind Concepts? Probing Compositionality in Large Image Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.029297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.029297Z digest=sha256:a8f84688e14eda6420921b1e0357284dbf246ee9a8a15aec7fcf87112b547a3f

Observation e72a3b76-2f3c-414e-b73d-298ca1be321a · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.147656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.147656Z digest=sha256:504cada6a5e7db84abf169e06cb08f94db69b9457bdd24140eae41979fa263f5

Observation 9cc3661c-57d7-4361-ab1d-fe160a633f48 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.223641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.223641Z digest=sha256:2fc9914c69228fc5d10c8bb58afd3d9e4f1ae186ebae442723f31efeef23652e

Observation c98760e6-266c-46ff-82b8-da61dc20f370 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.297322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.297322Z digest=sha256:5a87d9781fb7ae387606c2e71f74db3ee03d361bdb8afdf7f595d873681e65fe

Observation 8b821a8e-93e2-45bc-a01f-1e1abb7fc48b · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.326309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.326309Z digest=sha256:3ff982275bc4b86bc1640a9624985619d26ad645b1ae6891697c306641b1560a

Observation 08c3f604-a780-46fe-b6b7-cb5b067c7cd4 · outbound

This paper cites MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.390311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.390311Z digest=sha256:e0714ead88736e03759728d22100b035312c1203b4849dcad8c5a4dcac9c1342

Observation 353242e8-6191-4d6f-b25c-f2d98f403f79 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.461452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.461452Z digest=sha256:12c245b0ecdcc453ed6fe3c856cd72a02882b39b27d456c64fcedddbe22b2627

Observation 63fcb153-04db-4f4e-a61f-25e58fd4085f · outbound

This paper cites Segment anything.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Segment anything

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.683465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.536131Z digest=sha256:816ce2899f11ad39ae42c4a1bdc4d873697d992e1e3cdfd00ba4ba37b2b11013

Observation 0293cf60-53c7-4324-bdb7-91f2a2f60ea7 · outbound

This paper cites Robust fine-tuning of zero-shot models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Robust fine-tuning of zero-shot models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.612236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.612236Z digest=sha256:cd060c0a16d4691cb4470d986fe40908c41d6042846bf434231344288378473d

Observation 9adf4580-b315-4ca6-9b5c-3bc902cd207a · outbound

This paper cites Openclip, July 2021.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Openclip, July 2021

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:49.533161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.687684Z digest=sha256:e523e100e042e90318bbaf4da8c94ae63bc7663cca39653a7fbb03f6e0bf3705

Observation 2026a014-5c8c-4b95-bae7-1d3c8799e045 · outbound

This paper cites visualizable.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models visualizable

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.387084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.742052Z digest=sha256:7cb30754d1b7ef050c6ca017ad91b4a559cb951ea9ff3046d5a52728b98b09ba

Pith citing papers

No inbound Pith citation observations are available.