Pith. sign in

Paper Citation Record · LEDGER

A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2307.12980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.12980 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:03:01.075876Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

63
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9c5416f6-065f-4477-9113-6cde064a86f5 · inbound

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment cites this paper.

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:03:01.075876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:03:01.075876Z digest=sha256:5530ad576cda4a80d99424630354bc7280822b2e01b4e789956be4e5d27be3ce

Observation 13bba754-af8e-4ba5-8e0f-0d6a91c26c46 · inbound

System Test Case Design from Requirements Specifications: Insights and Challenges of Using ChatGPT cites this paper.

System Test Case Design from Requirements Specifications: Insights and Challenges of Using ChatGPT A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:14:54.870259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:14:54.870259Z digest=sha256:398c7e12d482a7f751dc561b018d25c936c52f041ce22779cd87522806e0a44c

Observation 10e43c3f-40ac-4895-b0c6-0a51d9118e07 · inbound

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds cites this paper.

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:53:37.074234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:53:37.074234Z digest=sha256:9f428f1aca86fba7fb8e056119dc7ceba9339dd928e885b6af61b3165e2df8d7

Observation f3a5d9af-a1de-4234-a681-400af1d768c3 · inbound

Generative AI Literacy: Twelve Defining Competencies cites this paper.

Generative AI Literacy: Twelve Defining Competencies A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T05:53:47.900026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:53:47.900026Z digest=sha256:aa1f966a967d0d5342eef6beeb16e730abdbc2ebfe79c9bda91a723171cff387

Observation 556452ff-d018-4e74-8828-a394b84a5658 · inbound

ReNeg: Learning Negative Embedding with Reward Guidance cites this paper.

ReNeg: Learning Negative Embedding with Reward Guidance A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:13:43.534804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:13:43.534804Z digest=sha256:12c2fdb78d08947bfec9c3f584cf314e59dddfd1fe4fe59e968c563d47bd1dbb

Observation dd9e3983-8336-4470-8675-323b929cae20 · inbound

Multimodal-to-Text Prompt Engineering in Large Language Models Using Feature Embeddings for GNSS Interference Characterization cites this paper.

Multimodal-to-Text Prompt Engineering in Large Language Models Using Feature Embeddings for GNSS Interference Characterization A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:30.336120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:30.336120Z digest=sha256:f059a5fa673c65783a856db57cde528d04db8a6c91371369af9d9219069c4b91

Observation 4379868f-7f41-4e18-aae1-962ed81c0083 · inbound

LLM-Agents Driven Automated Simulation Testing and Analysis of small Uncrewed Aerial Systems cites this paper.

LLM-Agents Driven Automated Simulation Testing and Analysis of small Uncrewed Aerial Systems A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:50:28.300216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:50:28.300216Z digest=sha256:9af706c2178e9c83090bd9613b4e025b80e92e28566b52f5c4a05cab3ffa1d26

Observation 25da61a2-2cd4-4644-b654-fbae490bd101 · inbound

One Head Eight Arms: Block Matrix based Low Rank Adaptation for CLIP-based Few-Shot Learning cites this paper.

One Head Eight Arms: Block Matrix based Low Rank Adaptation for CLIP-based Few-Shot Learning A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T11:16:55.756205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T11:16:55.756205Z digest=sha256:267c2afdf31c2c57213da933482fe1d9fbdc8a3767857e2bfb08b0183cfd3312

Observation a7f0e0f4-ef5e-4514-a31a-a8c44d895255 · inbound

CoDe: Blockwise Control for Denoising Diffusion Models cites this paper.

CoDe: Blockwise Control for Denoising Diffusion Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T17:10:17.260982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:10:17.260982Z digest=sha256:3533c57622976632bb85b2385ba28bc41ef1df8ca5ae26896c9b18525ea0d7b2

Observation bbb966cf-fbbb-4db4-b73e-930bc7e284ea · inbound

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models cites this paper.

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:31:23.388282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:31:23.388282Z digest=sha256:0fdcf3419a9ebc618849a9863485cb6150050ccccf419cbac8d7d672fdbb0642

Observation f8af49c9-438f-4f61-8df0-50cd2d34fbae · inbound

Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI cites this paper.

Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 203

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:51.198362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:51.198362Z digest=sha256:0014d334549429d34a8d25987a8407dfaf94f7b82f1c226c728bd35f7053db72

Observation 2fd5c4a3-f0f8-4472-b0af-28b870ad8721 · inbound

A Survey of Personalized Federated Foundation Models for Privacy-Preserving Recommendation cites this paper.

A Survey of Personalized Federated Foundation Models for Privacy-Preserving Recommendation A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:32:16.526928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T09:28:32.185398Z digest=sha256:131e0f7017b5a540bdce0166d802377a5624047cc20dcd5ea3b6a9304fa918c8

Observation 482ae711-fa47-4911-ab40-5dd7c88d34a5 · inbound

Visual Textualization for Image Prompted Object Detection cites this paper.

Visual Textualization for Image Prompted Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.064646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.064646Z digest=sha256:16b07c9609db4752f8b772319d32760bb05d600040e6ee1ee433cb66fcf0a4e1

Observation 26032027-4fc6-4c4b-95fc-8895bc238252 · inbound

A Survey of AIOps in the Era of Large Language Models cites this paper.

A Survey of AIOps in the Era of Large Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:36.620888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:36.620888Z digest=sha256:bc555f7f2151d2396a2ecc1d18f2700dd477b78a71d8c253f3fe48ea6818d5e7

Observation fe50a836-005b-4f99-9b63-5f8df3e74e80 · inbound

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding cites this paper.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.391857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.391857Z digest=sha256:a587d10468f7623b28b8cce3fca4cb25d78891c74850def6bfc753321dff7188

Observation 56635091-ef64-4704-98cf-59e530b62c46 · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.010676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.010676Z digest=sha256:298ecd07bcd5114bc866448f3a2a6b86710aa8d4399382526e2e816c5d36e003

Observation a0d3e291-6774-4c4b-b33c-5f3ddc703e08 · inbound

Adapting Vision-Language Models Without Labels: A Comprehensive Survey cites this paper.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.585528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.585528Z digest=sha256:a005bbe95a852437cdd17494034fed1a5513374888ceb73150b6008020673365

Observation 29bec3bd-5ffb-488e-a6ab-8b0bec84386a · inbound

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models cites this paper.

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:06:18.622208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:06:18.622208Z digest=sha256:452c8e177293192707e2282159b9cc73ff8c9d00c3a5de2495b822ed4fb086a0

Observation dbdf9168-c7a5-47e5-81e3-ca9122ddf634 · inbound

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds cites this paper.

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.719205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.719205Z digest=sha256:ea9b5170ea22498b3c58c6828a8af20669a0f14a23564983f6c1aac129659504

Observation 66fad044-bbab-42a8-bb45-6fae5efdc267 · inbound

AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models cites this paper.

AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:24:38.922212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:24:38.922212Z digest=sha256:feafa0364d5cda9ed373fbb1e34868c13df1fde643709e608826833e7e100e04

Observation 76c436a8-ec59-4ec4-9f7c-4f762a8b5f5f · inbound

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? cites this paper.

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:18:32.283981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T21:15:20.717705Z digest=sha256:e0bba86d219ec83ce6e6cc5b067300749879c84c35929627d86f63c08a8e2d66

Observation 8fafcd6a-d497-4509-98ef-931aa2a2a01b · inbound

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection cites this paper.

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:57.219371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:07:17.174815Z digest=sha256:4b4cf5a33fa7184ea3c8474cd385e3f0b4f336dfaffd400212dd2bcfbaf3c13b

Observation cf9950a2-3e3e-40d7-b85a-7b2b6288e9c7 · inbound

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages cites this paper.

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:38:21.886516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T14:33:36.100966Z digest=sha256:12e16e5f1a84cb9e95d4f2a08ad8552a51d7e06602683233f8ff16463a46759f

Observation 0f3cf654-188c-4da7-8cb4-4d8c14fd32b8 · inbound

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media cites this paper.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:08:20.409821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T14:05:14.737146Z digest=sha256:1db6f860aab8e1789ef264cbe50c33e159ee0b168a1f92ab03a4d78ab13624b9

Observation da5629af-2778-4f7e-9875-4d730b62fbeb · inbound

EPIG: Emotion-Based Prompting for Personalised Image Generation cites this paper.

EPIG: Emotion-Based Prompting for Personalised Image Generation A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.999944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T06:54:05.515501Z digest=sha256:32f7e5cd0e2a5306c99ccccc10896ff9f37ae7bfa559bbda968d2fd6b69e458b

Observation 29510b11-2cee-45ee-927b-eff60a3f63ee · inbound

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models cites this paper.

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:05:42.782901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T03:02:12.747334Z digest=sha256:33b82e9c770e135d54f8f50575304394f2344befe341b5cbd1ad5d7dd89d3fb7

Observation b7cda6a1-eada-4c41-8416-e152951f0bc6 · inbound

Visual prompt engineering for video models cites this paper.

Visual prompt engineering for video models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:13:12.198478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:13:12.198478Z digest=sha256:81be4a430375369f487a61c4ca823b1b029e59a6823dd3ca586ce6ad6d52ebf1