Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:18:57.661358Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2412.15632.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:18:57.661358Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 37bb6e02-4aac-4a64-8cbe-564d83c5e0ac · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Crepe: Can vision-language foundation models reason compositionally?,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1788f531-0ce9-4c44-b93a-a7d6520b7a00 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space When and why vision-language models behave like bags-of-words, and what to do about it?,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation db3469b1-70c0-4030-8bb6-2b39f234beea · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Sugarcrepe: Fixing hackable benchmarks for vision- language compositionality,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 53ac6195-82bb-4ec6-9cb1-d71815a5b94f · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Winoground: Probing vision and language models for visio-linguistic compositionality,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 25796259-1cb0-470d-82b7-a449dd88ce4d · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space COLA: A Benchmark for Compositional Text-to-image Retrieval
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05fe36df-6be9-46e5-877f-6dad7a267371 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31006a19-18c7-4ed1-af96-16552750bdb6 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Structure-clip: Towards scene graph knowledge to enhance multi-modal structured representations,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b613d16e-7adb-4c98-bbc9-a6196b92e7a5 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b8d73d-bf08-4953-994e-1bf9232eaef0 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Learning transferable visual models from natural language supervision,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f0ce826-e019-4799-bd27-4f8c3efcf2ac · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space An image is worth one word: Personalizing text-to-image generation using textual inversion,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 60773a53-94d1-439e-9cca-7e6f2965f182 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec736375-a958-41d4-a5f3-3504adbf91c4 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Iterated learning improves compositionality in large vision- language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6367ebdf-2a54-4e1a-a946-85f400089516 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Teaching structured vision & language concepts to vision & language models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8004868c-9524-4656-8781-6ecf5af4e3be · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7586e22e-d738-467a-a41a-bb27528d1909 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space What you see is what you read? improving text-image alignment evaluation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dc631507-500a-4777-9dc2-5c53167b68bc · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Multi-concept customization of text-to-image diffusion,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 78b8c5c1-e5ff-4d9d-8569-71505f8d4367 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space “this is my unicorn, fluffy
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f62df5e6-da21-48d9-af5b-d67c7f0d1c39 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Zero-shot composed image retrieval with textual inversion,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2a8a207b-7ac2-4684-a505-9f6f1846edcd · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a435dafd-da81-42b1-a77e-a6c09b3169af · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Pic2word: Mapping pictures to words for zero-shot composed image retrieval,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6bd0fe77-c636-4ffd-97b7-492e42619c17 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 073fccc2-f276-4bf8-b538-fbd48eb9c2c8 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Distilling the Knowledge in a Neural Network
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ddab9d-f98c-4e18-94d1-b48c4df647a0 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5519a82-dcd4-434c-923e-40758c30a318 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Dinov2: Learning robust visual features without supervision,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b141dfb6-7d6c-49e7-93f6-e2033813fe5b · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Clip-kd: An empirical study of clip model distillation,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 65dfa721-ce19-4a1f-90ee-9557804774f8 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Laion-5b: An open large- scale dataset for training next generation image-text models,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 288107f5-27ff-4dbd-9dd6-d61e965cf183 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space ImageNet Large Scale Visual Recognition Challenge,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c710def6-d9ed-48ce-9a1f-4e1e00f6905f · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Language models are few-shot learners,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa58cf2b-edd1-4682-90eb-c81fe48fae97 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Im- proved baselines with visual instruction tuning,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05fd301d-9e33-4f55-8e85-c8c628e07f38 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space The Llama 3 Herd of Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a827271-5192-4cea-9f8b-a310bbdeefa2 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Visual genome: Connecting language and vision using crowdsourced dense image anno- tations,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7f940723-8aba-41c2-b6a7-49928ded111f · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7371fb4-37cf-4302-9f3a-e2616b6e1c59 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space The batch size is set to 256, and loss weights λgpt are swept over {0.5, 0.75, 1} to determine the best model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 251db144-0fd7-45b0-9ab5-32e91c7175d8 · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space A little girl sitting on top of a bed next to a lamp
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f027e09c-400e-4ddb-9587-9f81e14c00ef · outbound
A New Method to Capturing Compositional Knowledge in Linguistic Space Training the textual inversion network Θ takes 18 hours in total on a single A6000 GPU
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.