Pith. sign in

Paper Citation Record · LEDGER

CF-VLM:CounterFactual Vision-Language Fine-tuning

As of 17 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 4 inbound Pith citation observations for arXiv:2506.17267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17267 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:07.332466Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:51:29.406876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.501635Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact7
  • verified fuzzy33
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db04e06f-f11b-45e7-a857-2b6f14e0b7e2 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

CF-VLM:CounterFactual Vision-Language Fine-tuning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.943940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.943940Z digest=sha256:55c4b0e214c6912db47500086e68a15c664832fd3d3d5ec9627e4d04dc4bd538

Observation dfbd2932-fb48-49b2-b13d-92c274c9d958 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning A Survey of Vision-Language Pre-Trained Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.949603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.949603Z digest=sha256:e87507e0451a747663bf6fb0cb476886c65427bb36886b804ccfd8f3f9c8c5c5

Observation 37109915-1310-47d6-9543-ff27ed9b2992 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Flamingo: a visual language model for few-shot learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.954808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.954808Z digest=sha256:2a6063f693ed70fa0bb820c183d247f9632a6cb47ecde79e5b2d92f2c6181366

Observation fcc8c1a0-c985-4ed2-b869-d631b3815a97 · outbound

This paper cites Causal Inference with Large Language Model: A Survey.

CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Inference with Large Language Model: A Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.959963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.959963Z digest=sha256:b1af25ab0a3901f3a569beb9282d25576bae91d7a397201b778811d130854d22

Observation 60fea556-4575-45ff-9b62-2d850b2a616b · outbound

This paper cites CELLO: Causal Evaluation of Large Vision-Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning CELLO: Causal Evaluation of Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.966031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.966031Z digest=sha256:0ab5d6400a0a99b0a7bab4ea0d2926ab7170639f475be51446e1ad8b0d5e9f12

Observation 99555b6c-afd1-4d6d-8953-d6324ebaf0e1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

CF-VLM:CounterFactual Vision-Language Fine-tuning Learning Transferable Visual Models From Natural Language Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.971026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.971026Z digest=sha256:dc4c90ea5c1a1d39f812a3b4a97a757f5460306dba4f16ad5d67a6c68c22a8d2

Observation 8ff0d1b4-07be-48bd-ae98-24a0b4f88ee3 · outbound

This paper cites Measuring progress in fine-grained vision-and-language understanding,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Measuring progress in fine-grained vision-and-language understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:09.057731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:06.977108Z digest=sha256:3e451143ff7d339c00fe1427ccf59da96df78906f21eb6a301e216f34940da18

Observation 8300da25-7d91-4df9-abed-eaece9b8617e · outbound

This paper cites Synthesize, diagnose, and optimize: Towards fine-grained vision-language understanding,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Synthesize, diagnose, and optimize: Towards fine-grained vision-language understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.985378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:06.981847Z digest=sha256:cfe45c3c9969254c5c0cab4881777eb431779198b245c9f61ce8905971f46d92

Observation e3947761-44ac-4821-adb9-1a20a4b3d0b9 · outbound

This paper cites Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.966532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:06.986025Z digest=sha256:cd17e22facd79906935f4e9d2245616b927757c30eebd86c67a45a8283c69932

Observation 7d2999e9-fa44-4ff1-b40a-137410a2ded7 · outbound

This paper cites Benchmarking zero-shot recognition with vision-language models: Challenges on granularity and specificity,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Benchmarking zero-shot recognition with vision-language models: Challenges on granularity and specificity,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.947634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:06.990755Z digest=sha256:84d0afef131a245288dbf1e1fe4fd5b0b0ab1c2c578673f6089727221554da5d

Observation ab451ce1-26b1-4cb7-81d3-32137ef45aaa · outbound

This paper cites Vilta: Enhancing vision-language pre-training through textual augmentation,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vilta: Enhancing vision-language pre-training through textual augmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.931918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:06.995224Z digest=sha256:0797c8cacfe64afd629081b8fc0c7e4036964190c84767aa7a2fcd950c7a3db3

Observation b637aff9-66da-4a03-bcb8-bc90540662f9 · outbound

This paper cites Towards vision-language mechanistic interpretability: A causal tracing tool for blip,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Towards vision-language mechanistic interpretability: A causal tracing tool for blip,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.914022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:06.999305Z digest=sha256:a48a8d929fa147d5dd828fab71c8c74c1afb9310fe1b392aae8aa1b2f5629c8a

Observation a1d558dd-dfec-4602-8c5c-0d506ba8ad8f · outbound

This paper cites What matters when building vision-language models?.

CF-VLM:CounterFactual Vision-Language Fine-tuning What matters when building vision-language models?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.895947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.003354Z digest=sha256:48d626c885c19f851ce631244aee976b586ee7a75e9cb0168ebf1665f478b141

Observation b9f42f80-8a62-4c16-b903-fd65508ef49a · outbound

This paper cites Fine-grained alignment for cross-modal recipe retrieval,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Fine-grained alignment for cross-modal recipe retrieval,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.878291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.013604Z digest=sha256:96b0a2984fa46c84b1e25b7c7b0ff4ca4d20c2fc16cc59d730d9f8d582da3eb6

Observation 0c8a2ede-f477-4e9f-bc3e-ac43428afc04 · outbound

This paper cites Localized triplet loss for fine-grained fashion image retrieval,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Localized triplet loss for fine-grained fashion image retrieval,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.858955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.018556Z digest=sha256:0ff1be82c3d0b212576b7c7e98ad749c741830f350dc6f5b88fbbc2fd9c87602

Observation 110a4311-512f-4704-9fe0-b9aaca3441f3 · outbound

This paper cites TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.

CF-VLM:CounterFactual Vision-Language Fine-tuning TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:08.073739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.023061Z digest=sha256:f51953e80adff75e73352ce78f2e53d0512fba35336e00903c7de2cbb745b980

Observation 3c99b13f-56ef-4af6-9ab3-ca42fbd4ccb9 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Facenet: A unified embedding for face recognition and clustering,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.028168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.028168Z digest=sha256:cf0d4cf18f9a7aa0fe6bc10cbc1325bf09b72477fab4b33fb1e8db7da8e9a7e0

Observation 8bbcaeb3-aba4-445b-8ba4-599906698382 · outbound

This paper cites CPL: Counterfactual prompt learning for vision and language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning CPL: Counterfactual prompt learning for vision and language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.825807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.037577Z digest=sha256:7043df635e960d59c4656e31ae918fa42d63bb8ff7df82f885cfe35e33d0bf40

Observation f2cdb927-37fb-41da-80bd-bbd97f3162c2 · outbound

This paper cites Attention Is All You Need.

CF-VLM:CounterFactual Vision-Language Fine-tuning Attention Is All You Need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.046597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.046597Z digest=sha256:a31fad28f66cba8c95bd421601b7fd13b6468579c36f8fff516fdad129c7fc9e

Observation b666d97b-07a6-4e8f-afd5-494e7f79f3db · outbound

This paper cites A survey on evaluation of large language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A survey on evaluation of large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.051054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.051054Z digest=sha256:71c7ac0f119c0e9e7e163188a7ee5d5d7bc3980fd8a2ed77bc6947d118b10250

Observation ba152358-f5c9-440e-9f5e-efe7a008b601 · outbound

This paper cites A survey of visual transformers,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A survey of visual transformers,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.808874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.055484Z digest=sha256:57ecac80d5c8f74adedd7c38b9226ad466a21097d108d152b068b0bf7b9b1c1a

Observation 846e6f77-4295-47f2-ba90-0ee581c751d9 · outbound

This paper cites A simple framework for contrastive learning of visual representations,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A simple framework for contrastive learning of visual representations,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.059548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.059548Z digest=sha256:598c3bc109749c922c9c71d5098e15022f2bca16226c30d950875ae39a35ff2e

Observation e4f02ba1-a538-42dd-8408-84b1e0268399 · outbound

This paper cites Understanding contrastive representation learning through alignment and uniformity on the hypersphere,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Understanding contrastive representation learning through alignment and uniformity on the hypersphere,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.781994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.064955Z digest=sha256:7c958566e35299abb9630760043c91fc16442362f533db92acf40349941831e2

Observation 02c019f0-7f4e-4b7e-b6b5-d13924477716 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

CF-VLM:CounterFactual Vision-Language Fine-tuning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.069602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.069602Z digest=sha256:d29ff9266518019102335176e9021d37140f27f39bb9eacac95016b647815df7

Observation 3b9e9e6a-efdb-4d44-a706-1028d67ab742 · outbound

This paper cites Cogs: A compositional generalization challenge based on semantic interpretation,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Cogs: A compositional generalization challenge based on semantic interpretation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.764968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.074611Z digest=sha256:c7e7413b471ea2f8c93788c7a2d96b4a1fdfdff4010d528a14e1d271d6b12658

Observation 43a53351-d77c-4251-b4fd-cea102c36b13 · outbound

This paper cites Learning what makes a difference from counterfactual examples and gradient supervision,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Learning what makes a difference from counterfactual examples and gradient supervision,

Reference 28

Resolution
verified exact
doi, observed 2026-08-07T05:01:07.374921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.085658Z digest=sha256:cce91e8eb21853c4a9f41ae823687c389289830c2ca9b6484f85605e964d45a2

Observation 08b38393-f8f0-44a3-b187-ce2bd65a3274 · outbound

This paper cites Teaching clip to count to ten,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Teaching clip to count to ten,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.748372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.090433Z digest=sha256:480f8a09fadf257fa74a449518b701572f0eb7d0bf48317a203366d12efffdb3

Observation e06c7165-c2c9-4efa-b367-550c11e41768 · outbound

This paper cites DISCO: Distilling Counterfactuals with Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning DISCO: Distilling Counterfactuals with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.094854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.094854Z digest=sha256:7fbbcc52ecfae3b6dadde1631df821881294da1adc3602d0906e0fbce161a104

Observation 2bad5841-e5b9-4482-a9f0-64aaf70d9a71 · outbound

This paper cites Counterfactually measuring and eliminating social bias in vision-language pre-training models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactually measuring and eliminating social bias in vision-language pre-training models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.100956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.100956Z digest=sha256:6fc598af78a2b0bdea2f96c6ab9fffbb81c2c6c3abd905c485aa856ae6c5624d

Observation e09ab765-4ce8-4e63-9431-0f3250b93549 · outbound

This paper cites Counterfactual attention learning for fine-grained visual categorization and re-identification,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual attention learning for fine-grained visual categorization and re-identification,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.842305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.105617Z digest=sha256:ddb7670aadbe5801500c5698e0219405bb23886f3b74cb6065f21ab16bfb96ac

Observation af420799-f337-4ec0-b5f4-19b03960bb1c · outbound

This paper cites Counterfactual samples synthesizing and training for robust visual question answering,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual samples synthesizing and training for robust visual question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.730534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.110453Z digest=sha256:4a223590337a3e37154a8dc5b2f75b1b6d2c59c161e85e04bd7b03cf72c62ea0

Observation 5733a09a-3318-489a-98d1-376fb7498080 · outbound

This paper cites Qwen2.5 Technical Report.

CF-VLM:CounterFactual Vision-Language Fine-tuning Qwen2.5 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.114672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.114672Z digest=sha256:f99dfd683de32f5b6d96329743885b1c2d083166c2f42412a511a3a76aac12e8

Observation f4593c00-6d8a-4f6d-aa1d-a6a6c4f83ef0 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.711237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.119470Z digest=sha256:504be4dd91101cdceb09afebee71e27fe047cb6b5f271d49f666378d864e5600

Observation fda6020a-f7a7-4d80-8212-4c0c609522ea · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.692774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.123742Z digest=sha256:3d0cb7d1287f6f21091483a03633aca828fc7dcc159ab7844da54a4dfc933e41

Observation 060bdcea-5d8c-4cec-9572-47b5edeed4be · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

CF-VLM:CounterFactual Vision-Language Fine-tuning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.128034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.128034Z digest=sha256:a0f073138784848fa3daa6902c92dc931cef27bf933c7defeb3f79db6ce6bedf

Observation efa34ba6-5906-48d1-9a18-de51b7693da2 · outbound

This paper cites ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs.

CF-VLM:CounterFactual Vision-Language Fine-tuning ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.132424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.132424Z digest=sha256:9e95b4a161dbb7d05df9a9ef1683cdd6c2c6aacd6bd3bebb429afcbaa8984458

Observation 866c9125-f995-4dfe-bd87-feb0efc052a2 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

CF-VLM:CounterFactual Vision-Language Fine-tuning When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.674987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.136591Z digest=sha256:afcfbd0ca255da596364d954daa9ec8c1a44feb81e5d58130772e837e75fecbd

Observation e0bd20ed-6479-47b0-bb65-0026b8d5f9d2 · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

CF-VLM:CounterFactual Vision-Language Fine-tuning VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.140504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.140504Z digest=sha256:b44aa733b927dc7c6ca2bc8f92f847460a1ffb6d286ccc2ca0e931781f5db024

Observation f12b6b44-3b06-4d04-a7a4-ca0e8df4fa74 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge,.

CF-VLM:CounterFactual Vision-Language Fine-tuning ImageNet Large Scale Visual Recognition Challenge,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.657970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.145231Z digest=sha256:b364066ed24a4543e5db63c647f7aefe513f404cb572ea093b97574e25647378

Observation 799bc5f2-8288-4427-b732-aea50db657ef · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,.

CF-VLM:CounterFactual Vision-Language Fine-tuning From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.640639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.149485Z digest=sha256:95e78f601e0b91d943ff56f97efe701247a2a523c30a8a84668126f195e885ee

Observation ed03be52-36d3-4416-990c-4cafb667f01d · outbound

This paper cites Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding.

CF-VLM:CounterFactual Vision-Language Fine-tuning Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.741382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.153634Z digest=sha256:d7c18bb26582e8bc4d2dc604f41a3311f6150f67c534d6406fd045df7973fa6d

Observation c6d5e468-171b-47c6-a514-c88c7e26d27c · outbound

This paper cites Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-modal Structured Representations.

CF-VLM:CounterFactual Vision-Language Fine-tuning Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-modal Structured Representations

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.697666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.163971Z digest=sha256:d020847ff7ecfdefc90b9ace5a0eff384b57588bb7624b5bf89c1b6984e2b0e2

Observation 036de79c-a746-4d6b-8d94-91e7d50d6c8f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

CF-VLM:CounterFactual Vision-Language Fine-tuning Improved Baselines with Visual Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.168343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.168343Z digest=sha256:64eb9a7b277af895b8c2b84294e9f6c2bd141cbf7fbcf53fa2a715ddf7a05bd1

Observation 556d253d-ee8b-4367-bd2b-bd573994ec36 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

CF-VLM:CounterFactual Vision-Language Fine-tuning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.173286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.173286Z digest=sha256:7837d9136819431305d8de0ce1a6561e072281235561da9a44269926f63e3428

Observation a3d19d3a-9461-4759-90de-978b5b273bc3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.178097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.178097Z digest=sha256:5df425b941f2e61e8dded95b5bf2870d0a107969d1b54ef34139f82868e7322f

Observation 3838e51a-9024-44e9-8b6c-4840d6b2b301 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

CF-VLM:CounterFactual Vision-Language Fine-tuning SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.182593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.182593Z digest=sha256:9e56b1d180ffdc7af56714cacdb36b5f383599e1c989fe252d9d78586f6eacb9

Observation 4d6eedb6-ee1e-4ec7-a636-f31252676b0a · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.187178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.187178Z digest=sha256:d2a16eec10a6766a0aff0b47ac89f357dff2f51d4b032e2e94cb0a0e4f5672c1

Observation 8903ba8e-e6da-43b9-b714-2c24b5ee9f14 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.191460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.191460Z digest=sha256:6a7771fabdbc28bab45d62b60a889287602ffb51120e8e4fc3f22fd0856e8aef

Observation ce3d1c0b-f348-4bdc-a973-23c90907feb9 · outbound

This paper cites Vision-and-Language Pretrained Models: A Survey.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vision-and-Language Pretrained Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.195795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.195795Z digest=sha256:fdfde512f2788d8c98e9e3b6a39f6ffedd60e2df369b485c79e8c07fb0526345

Observation 6f46b262-c841-40b7-afca-bda312bc53a6 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Exploring the frontier of vision-language models: A survey of current methodologies and future directions,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.200219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.200219Z digest=sha256:d18b6a50407bd10bb5731acf46484afb60cd9ca0a21416c1865540a663b1272e

Observation 04362e3a-ca1b-4aa9-b3c2-5162f4df9c23 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vision-language models for vision tasks: A survey,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.204431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.204431Z digest=sha256:be65e27d6aa29dcc8e469fcfeebc230cd025f1c4cffa6553aca45b3feedb5615

Observation 1064166d-50a6-44ac-a19e-7dead8da82db · outbound

This paper cites Counterfactual vision and language learning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual vision and language learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.613419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.208880Z digest=sha256:ff26267b8e1196c8fef644da6c4604ca56b5d79731e0958fc67d9388d5e0c267

Observation 5256965e-acca-49f2-83fb-b036e5e749f4 · outbound

This paper cites Counterfactual reasoning for multi-label image classification via patching-based training,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual reasoning for multi-label image classification via patching-based training,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.597922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.213460Z digest=sha256:f50607cc1c48901f51ba327cfa6757c0c88fb11d7d5a861078b0f019ee349081

Observation 75ad03d2-c6ca-4845-a852-39b0a661e028 · outbound

This paper cites Causal Graphical Models for Vision-Language Compositional Understanding.

CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Graphical Models for Vision-Language Compositional Understanding

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.719377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.218215Z digest=sha256:ef5c33d05714c472c3ae6080ca1f9b3eeb937106ed0823a7b1c11ed8a9da57d7

Observation f4e29da0-3367-43e2-b371-50cbf99390a1 · outbound

This paper cites CPL: Counterfactual Prompt Learning for Vision and Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning CPL: Counterfactual Prompt Learning for Vision and Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.433090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.222289Z digest=sha256:2046942f5e3f42093cd2516532bb82e5c11714b44306606900c2a0775e3a3c5d

Observation 82b5f82b-87b4-4f61-8a8e-38e472eac53b · outbound

This paper cites Counterfactual Visual Explanations.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual Visual Explanations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.227257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.227257Z digest=sha256:fdef23fd4521ee8d64565abf54b499e5751c32721351dde33fd06b94be95f715

Observation 8e4108df-7b85-443d-af3d-3ff07f794393 · outbound

This paper cites Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identification.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identification

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.410653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.231457Z digest=sha256:012a1bf8d5e7e31fbf5e61c956bf7b27a734717e1e99d8d7fbd93c612b30ac6a

Observation af0ade69-265d-4dfe-87f4-5bac436fca09 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.580872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.236954Z digest=sha256:0c20e30443116df70a3a15c5ded5f9ee66e711391b584342f8dfc27c8c9bba1b

Observation 38178384-683d-4c49-b0d8-2d9434e28870 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.565173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.241166Z digest=sha256:9f5b35d7b7ba32fcfb6c474cc7c61878c8f0bf4ed42f3160fe81335540c12ce7

Observation b4b2c7b5-42f4-4330-97b6-da6084a79ff5 · outbound

This paper cites Rules: • Modifyonly one thing(either one attribute or one causal link).

CF-VLM:CounterFactual Vision-Language Fine-tuning Rules: • Modifyonly one thing(either one attribute or one causal link)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.547774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.245312Z digest=sha256:e07b5915a3d20f5edc52393c15ca8dd8424333e935c7d4c7ff8978a34084d88b

Observation a125a548-442b-48f1-8e97-5a7935069a43 · outbound

This paper cites Example Input: A young woman holding a racket hit the ball, and the ball flew outward.

CF-VLM:CounterFactual Vision-Language Fine-tuning Example Input: A young woman holding a racket hit the ball, and the ball flew outward

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.528325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.250450Z digest=sha256:a013047b7f43d806be082a621fb0724fcae7af14abe60e4cf7fb35ac5c53304a

Observation 98563f15-079f-4b01-877a-75ae0aa6db83 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.512290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.255371Z digest=sha256:becd042ebd51123f248a30cdb2253771a6bdd3c949099eaf57c1e04921e98d34

Observation 87f61e64-b74d-41b0-90d6-958147e1b9ef · outbound

This paper cites Output (causal):.

CF-VLM:CounterFactual Vision-Language Fine-tuning Output (causal):

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.495564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.260715Z digest=sha256:205234a72052d3e3390aa0d987d480cd5042e1a16eeeacfe3a6f9a9192da5144

Observation b3fa6ee5-d2a5-4657-a4b6-db4745d88fe1 · outbound

This paper cites dirt” with “paved roads.

CF-VLM:CounterFactual Vision-Language Fine-tuning dirt” with “paved roads

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.474047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.265588Z digest=sha256:f8dffcc1d63dc131dafd0d2c5bdda28345cd162b707c7c8e9d9441b8a3d00fcf

Observation 37f1a45c-cfb8-476f-877b-b90457564489 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.446664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.269999Z digest=sha256:3f1cb44b023440411b90683eb494aab5747c23f99e2d8465fd5a673e1cb86338

Observation b49d51f6-6c25-4ba7-827c-78703bf04ab2 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.428518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.273979Z digest=sha256:58a08fc092a8c0657674e94ab855038912b41f39f569d4258fcdcca997155f68

Observation efcd3bf2-24ca-48b3-95c4-1d7b15b217ad · outbound

This paper cites causal decision points.

CF-VLM:CounterFactual Vision-Language Fine-tuning causal decision points

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.412119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.277833Z digest=sha256:ef9e712e5e163471d8d61be85ac7d3481032a0f3898dc1b53501f95b090b3640

Observation 1a1cde11-e036-45e2-b861-38daf8f12ede · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.393260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.282997Z digest=sha256:12da207282a4db1a4c7ff1c528bced40e47020670adea4e3894a378ced92e730

Observation 5504e3cd-2e3b-4cfa-854b-3e4a660c462e · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.375706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.287322Z digest=sha256:3e5eda453207096155885ff714405eb325ce9e9c55dc0708c88cc05a652d9e1e

Observation 72662f1d-9f16-4588-bd8f-3ce329cae31a · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.357028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.292067Z digest=sha256:d5788759e4d2b1a3bc224c076e6ab4ca3c038230d6a74d47bd02e90486278299

Observation b4b55706-9edf-4850-8635-25e55d99d1eb · outbound

This paper cites Two people on motorcycles riding them.

CF-VLM:CounterFactual Vision-Language Fine-tuning Two people on motorcycles riding them

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.338390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.297273Z digest=sha256:d30f28ee1846dbd9b4491c0bd88c8a6a082df50ec8112d36a82af06736ccaee5

Observation ef86dff8-a5a1-4677-9571-c66da1bcdad5 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.316820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.302045Z digest=sha256:0f219aeb5588c61813f5491fea9d2659bf2fbcaff1f7b37d08194147b8dae26c

Observation c2267686-b567-446f-825c-7242bcfec410 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.298808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.306355Z digest=sha256:929b93da0bde55db47ffc03b29bc4e3b9905ecccbdfa534ea800e1ce9e341db8

Observation dc45ff25-b0fd-42b3-9c54-1e74ba3d9c7a · outbound

This paper cites causality.

CF-VLM:CounterFactual Vision-Language Fine-tuning causality

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.281488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.310308Z digest=sha256:06e4d12a55bec0168529979d1a23ceba75556616a282d469f72c9ca8d71dee65

Observation 5e3b4834-cbbe-463d-ad1c-215fa2526338 · outbound

This paper cites parallel realities.

CF-VLM:CounterFactual Vision-Language Fine-tuning parallel realities

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.260724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.314460Z digest=sha256:05aa2b7cc48d5d0122d934ebc07e619fdd574ffe17324b4b7845183f23e93b8a

Observation 443bbec6-8f01-47e4-a298-1b68d0e77f84 · outbound

This paper cites These are employed in Lcsd to help the model learn semantic scene boundaries.

CF-VLM:CounterFactual Vision-Language Fine-tuning These are employed in Lcsd to help the model learn semantic scene boundaries

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.243881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.319372Z digest=sha256:3288a41932dfffd4600a30daa1049d75cce4887ac9e4f3c749c22fbce5831ea7

Observation 7aa85fba-ece3-4611-85b4-63191fb81046 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.224933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.323874Z digest=sha256:4e890d679cb18fc371452ebb01f31f099c755aa1ae8e9cf5b3dc43f4673be9e2

Observation 527bc56b-0dc5-4d77-bf56-110aaa7a9aa8 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.210112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.327832Z digest=sha256:c0e59fd00561367878f899129cd5cbf94e8d3522f20e19081c0b932c3748088d

Observation 39db48cf-7edf-4131-b359-6611f6b84ee1 · outbound

This paper cites kicking a ball.

CF-VLM:CounterFactual Vision-Language Fine-tuning kicking a ball

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.194421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:01:07.332466Z digest=sha256:b119dff69e7d658dd0bb6c9f96548c94f6789e4c7d091037fd3e416161e36040

Observation d6f96fae-5544-4683-a0ec-c5087db8e7fc · outbound

This paper cites COGS: A Compositional Generalization Challenge Based on Semantic Interpretation.

CF-VLM:CounterFactual Vision-Language Fine-tuning COGS: A Compositional Generalization Challenge Based on Semantic Interpretation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.080997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.080997Z digest=sha256:907c064cbcf9333a378cd9bfa50458e392105421a5c328604ded2e6c15f09cfc

Observation f17ba9ae-74f7-42d7-bef8-edc49c67d258 · outbound

This paper cites What matters when building vision-language models?.

CF-VLM:CounterFactual Vision-Language Fine-tuning What matters when building vision-language models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.007888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.007888Z digest=sha256:dfef6245a690f6109acc68b6aa86cbc9f145a201fcb3424d0168c8658f2ff2fe

Pith citing papers

Observation ed3a27c0-7fbc-410e-b5e1-f1305cb5945a · inbound

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration cites this paper.

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:51:29.406876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:51:29.406876Z digest=sha256:c3388e173eea16a92a50e1d3f998bb2868411415219483b59a35345863c092da

Observation f99d24fa-3e34-4deb-bb1e-d630323c6d50 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:31.005765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:31.005765Z digest=sha256:6cce514f02148b2b4bc58ef57fd43b04e6d4418ddba8abd22bb8b42801af3ad5

Observation e43294dc-c047-41b0-97e7-665272a05f21 · inbound

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning cites this paper.

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.503204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T09:09:09.716984Z digest=sha256:5851ab7db65810b9d81f1415326218182043fd3c651d204ee6226b366984bf22

Observation 03766e84-061f-41f4-9881-13016b986e34 · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.148082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.148082Z digest=sha256:0447649c6b66ebbe4980d7e445bcab544c342b903ae3210cb2ee28bc65ca6015