Pith. sign in

Paper Citation Record · LEDGER

A New Method to Capturing Compositional Knowledge in Linguistic Space

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2412.15632.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15632 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:18:57.661358Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37bb6e02-4aac-4a64-8cbe-564d83c5e0ac · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally?,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Crepe: Can vision-language foundation models reason compositionally?,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:58.062691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.535883Z digest=sha256:12706c518e74947cd2880bd9ca3e1c79497edacb31bdd1a77bd97c11723ae9e5

Observation 1788f531-0ce9-4c44-b93a-a7d6520b7a00 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?,.

A New Method to Capturing Compositional Knowledge in Linguistic Space When and why vision-language models behave like bags-of-words, and what to do about it?,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:58.052441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.541141Z digest=sha256:7db90c098f4bff6fd7ba294246ddf8ad68275595bd772fbea1d2f311705f6748

Observation db3469b1-70c0-4030-8bb6-2b39f234beea · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision- language compositionality,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Sugarcrepe: Fixing hackable benchmarks for vision- language compositionality,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:58.041156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.545117Z digest=sha256:3fe717d8abc9d29a1f4231734c775029345539d7a17fbf70c5d09c223bdc3832

Observation 53ac6195-82bb-4ec6-9cb1-d71815a5b94f · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Winoground: Probing vision and language models for visio-linguistic compositionality,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:58.030314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.549466Z digest=sha256:904177d429cb943950fa5b260b3ca22552d1e4c857863d9ddc536757fd1f92de

Observation 25796259-1cb0-470d-82b7-a449dd88ce4d · outbound

This paper cites COLA: A Benchmark for Compositional Text-to-image Retrieval.

A New Method to Capturing Compositional Knowledge in Linguistic Space COLA: A Benchmark for Compositional Text-to-image Retrieval

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T11:18:57.790881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.554111Z digest=sha256:713e19a0d19a5d3a084b3286a17abdaaf25e123d5d98056d4ed58d327263ae62

Observation 05fe36df-6be9-46e5-877f-6dad7a267371 · outbound

This paper cites TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.

A New Method to Capturing Compositional Knowledge in Linguistic Space TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.558663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.558663Z digest=sha256:17436cd19e24de6c7e2471c020fc2a03da1dd1d0b39814f9c746368163e97133

Observation 31006a19-18c7-4ed1-af96-16552750bdb6 · outbound

This paper cites Structure-clip: Towards scene graph knowledge to enhance multi-modal structured representations,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Structure-clip: Towards scene graph knowledge to enhance multi-modal structured representations,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:58.019850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.563573Z digest=sha256:43ad58394bf8c5661bfba6acc87f693772d3a3e080264d432328a7e54a514447

Observation b613d16e-7adb-4c98-bbc9-a6196b92e7a5 · outbound

This paper cites Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs.

A New Method to Capturing Compositional Knowledge in Linguistic Space Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.567117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.567117Z digest=sha256:d480d59c0f142dad9b7eb62d441b8eca479df86be12dc435019f104b2d65224b

Observation e2b8d73d-bf08-4953-994e-1bf9232eaef0 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.570839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.570839Z digest=sha256:dd21080f652391dcecef0f233288ec753dc0947cf9e1a008127878b589cda883

Observation 8f0ce826-e019-4799-bd27-4f8c3efcf2ac · outbound

This paper cites An image is worth one word: Personalizing text-to-image generation using textual inversion,.

A New Method to Capturing Compositional Knowledge in Linguistic Space An image is worth one word: Personalizing text-to-image generation using textual inversion,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:58.002876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.573849Z digest=sha256:9476c14cf741a1b9f11a9564a362067d8c0fd7623272a7eb9aba37ef2d9d8806

Observation 60773a53-94d1-439e-9cca-7e6f2965f182 · outbound

This paper cites Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations.

A New Method to Capturing Compositional Knowledge in Linguistic Space Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.576776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.576776Z digest=sha256:d8b442b036657f131c6eeeb8eac4e5e8d998caf2792ecf628246acbcaebb22d3

Observation ec736375-a958-41d4-a5f3-3504adbf91c4 · outbound

This paper cites Iterated learning improves compositionality in large vision- language models,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Iterated learning improves compositionality in large vision- language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.992874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.580037Z digest=sha256:f9688d2bb780ab0c709030cce742928c392bf35853ae6b22b05f8a2f51f6b7b4

Observation 6367ebdf-2a54-4e1a-a946-85f400089516 · outbound

This paper cites Teaching structured vision & language concepts to vision & language models,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Teaching structured vision & language concepts to vision & language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.983542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.582717Z digest=sha256:bdccc68fd6fe4cde4d46a5d54f324d3bbd9d5a1889865f3c7f8046ba73e13e29

Observation 8004868c-9524-4656-8781-6ecf5af4e3be · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

A New Method to Capturing Compositional Knowledge in Linguistic Space CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.585495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.585495Z digest=sha256:b846535b84856dc076cf8445752651fa5d58e4520cdafed414ad3c9c90e6285a

Observation 7586e22e-d738-467a-a41a-bb27528d1909 · outbound

This paper cites What you see is what you read? improving text-image alignment evaluation,.

A New Method to Capturing Compositional Knowledge in Linguistic Space What you see is what you read? improving text-image alignment evaluation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.974398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.588638Z digest=sha256:b77841893187035a00fcd24a724bc086e966a1f733d7ccba9f7b32708ad876d3

Observation dc631507-500a-4777-9dc2-5c53167b68bc · outbound

This paper cites Multi-concept customization of text-to-image diffusion,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Multi-concept customization of text-to-image diffusion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.965507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.591574Z digest=sha256:34c06ee3b985aeaaf463548d2ca67da6b3b76a37642ba5442b1b2bb067be9632

Observation 78b8c5c1-e5ff-4d9d-8569-71505f8d4367 · outbound

This paper cites “this is my unicorn, fluffy.

A New Method to Capturing Compositional Knowledge in Linguistic Space “this is my unicorn, fluffy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.956312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.594637Z digest=sha256:73ed196415cadd253df5108f0e21b69607ba3e8337f2e65ec8fb1b86e8acf5ce

Observation f62df5e6-da21-48d9-af5b-d67c7f0d1c39 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Zero-shot composed image retrieval with textual inversion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.947517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.597686Z digest=sha256:3a9c8568ba27e68874851fd81a86d241e066f7ef903249f10d548bc0efb46cd2

Observation 2a8a207b-7ac2-4684-a505-9f6f1846edcd · outbound

This paper cites iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval.

A New Method to Capturing Compositional Knowledge in Linguistic Space iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.601309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.601309Z digest=sha256:0e6a3c5fda4838b9685e646bf9d198565b6b1bc60559cc70e3cf8419099197ad

Observation a435dafd-da81-42b1-a77e-a6c09b3169af · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Pic2word: Mapping pictures to words for zero-shot composed image retrieval,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.936721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.605099Z digest=sha256:d6906093eed15492d54449946465d4cf74f2ed1ee07f36f60cee342338b3a1e4

Observation 6bd0fe77-c636-4ffd-97b7-492e42619c17 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.925960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.608072Z digest=sha256:9f3fdd18ec261f04c663763c262723fa7b852f9ce4cb6ec92a86d9d6efbda087

Observation 073fccc2-f276-4bf8-b538-fbd48eb9c2c8 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

A New Method to Capturing Compositional Knowledge in Linguistic Space Distilling the Knowledge in a Neural Network

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.610992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.610992Z digest=sha256:5f3e0e6a805142b7b545ffd8e76c5b9a6db0bd091c811992b34e16872d7f594e

Observation 22ddab9d-f98c-4e18-94d1-b48c4df647a0 · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

A New Method to Capturing Compositional Knowledge in Linguistic Space Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.614664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.614664Z digest=sha256:148823199250bc387dae19c00d13a923da0d046657a49400caacbd7dc527c15b

Observation d5519a82-dcd4-434c-923e-40758c30a318 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Dinov2: Learning robust visual features without supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.915094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.618411Z digest=sha256:f62ff36400b89af2c3f8d41987d8726d9c32b724b5dac138b629a7dbd1579016

Observation b141dfb6-7d6c-49e7-93f6-e2033813fe5b · outbound

This paper cites Clip-kd: An empirical study of clip model distillation,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Clip-kd: An empirical study of clip model distillation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.903039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.621870Z digest=sha256:ef68944a76a3ef2633839499d3a61bbcea492d1a481eac5b380fb3c7fbd834ba

Observation 65dfa721-ce19-4a1f-90ee-9557804774f8 · outbound

This paper cites Laion-5b: An open large- scale dataset for training next generation image-text models,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Laion-5b: An open large- scale dataset for training next generation image-text models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.892792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.625365Z digest=sha256:eef6dd3fc76b9c726cdb5a3edb3bc0a0e404a25359a252a6af4ac12da0624a5e

Observation 288107f5-27ff-4dbd-9dd6-d61e965cf183 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge,.

A New Method to Capturing Compositional Knowledge in Linguistic Space ImageNet Large Scale Visual Recognition Challenge,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.882735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.628781Z digest=sha256:c7d4a7cde39200761b309dced6de1bb7777a5d0c0cdb9c6ecabad795700a5e2d

Observation c710def6-d9ed-48ce-9a1f-4e1e00f6905f · outbound

This paper cites Language models are few-shot learners,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Language models are few-shot learners,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.872207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.633436Z digest=sha256:e67821664a57fdf194ac3af73e9c78c4057792fdb5238985b4c07f8810bf57b6

Observation fa58cf2b-edd1-4682-90eb-c81fe48fae97 · outbound

This paper cites Im- proved baselines with visual instruction tuning,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Im- proved baselines with visual instruction tuning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.861448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.638079Z digest=sha256:8d206b9d900b2143ebaa36381f81a6fceb83bd948ca4f1693d1178b71d83f6a8

Observation 05fd301d-9e33-4f55-8e85-c8c628e07f38 · outbound

This paper cites The Llama 3 Herd of Models.

A New Method to Capturing Compositional Knowledge in Linguistic Space The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:57.641777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:57.641777Z digest=sha256:d4cd35a6c3e46c61ec5603827a72dbecc099e5301cc20e30d9b0b3f19da0e8f7

Observation 1a827271-5192-4cea-9f8b-a310bbdeefa2 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image anno- tations,.

A New Method to Capturing Compositional Knowledge in Linguistic Space Visual genome: Connecting language and vision using crowdsourced dense image anno- tations,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.849579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.645697Z digest=sha256:3ec9f7e69723c4c84b26dd09b69c959413e31320fc08122add8c288fe2122e5c

Observation 7f940723-8aba-41c2-b6a7-49928ded111f · outbound

This paper cites an unresolved cited work.

A New Method to Capturing Compositional Knowledge in Linguistic Space Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:57.838069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.650102Z digest=sha256:c2f378670c33773b2907b41b53bff36988bc95bd0e4aae138d2ade57d23ad800

Observation c7371fb4-37cf-4302-9f3a-e2616b6e1c59 · outbound

This paper cites The batch size is set to 256, and loss weights λgpt are swept over {0.5, 0.75, 1} to determine the best model.

A New Method to Capturing Compositional Knowledge in Linguistic Space The batch size is set to 256, and loss weights λgpt are swept over {0.5, 0.75, 1} to determine the best model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.826661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.654083Z digest=sha256:259f310b1dbffa4331674bc85cbc08c55a63a0d20ef17b3781691169470e1e7b

Observation 251db144-0fd7-45b0-9ab5-32e91c7175d8 · outbound

This paper cites A little girl sitting on top of a bed next to a lamp.

A New Method to Capturing Compositional Knowledge in Linguistic Space A little girl sitting on top of a bed next to a lamp

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T11:18:57.802219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.661358Z digest=sha256:b4a6858d986beac6089080e062483eec77a6ac4802d930991383495b0d6003c4

Observation f027e09c-400e-4ddb-9587-9f81e14c00ef · outbound

This paper cites Training the textual inversion network Θ takes 18 hours in total on a single A6000 GPU.

A New Method to Capturing Compositional Knowledge in Linguistic Space Training the textual inversion network Θ takes 18 hours in total on a single A6000 GPU

Reference 150

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:18:57.813978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T11:18:57.657771Z digest=sha256:3bc43c9cf5d82484b5ec2f1a7f8f8b6956f28877fb77520829dbec97068ab795

Pith citing papers

No inbound Pith citation observations are available.