Pith. sign in

Paper Citation Record · LEDGER

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2108.10904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2108.10904 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:14:03.546300Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:10:05.334285Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c79ce03d-e652-4e1a-9612-2a20441e7ebf · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:38:09.573674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:ae75a95ed58d2de0d8734403f3949a7e1eef0873672904e26e9cc18c463d4654

Observation d3e8c198-2f4b-45a5-9024-463d8f5a0c97 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.756707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:2100e8ebc88b652afad43215b1af21f0e21e95a9eb76adafb8a63598a7ab73c2

Observation 16095bf5-2c75-47be-9efb-9967ab6dc02c · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.338088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:5557f68fcf6bf981f1f24e2e397c6a45888e6282e0d77a2e8e416f31f694c750

Observation 1e62c38a-c5ba-4473-8fc7-b04d20a8b76c · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:53:08.385289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:8020104b19614584b046e226088e67cdc51227392e6d01d5c33d72a7ff9bc6b9

Observation 1611710a-17ef-4a83-a512-ce11cd81fa89 · inbound

A Generalist Agent cites this paper.

A Generalist Agent SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:24:49.969674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:24:49.833638Z digest=sha256:0df4ca46971cde9d65655d93e5b721bd5d5a05318e2647b20673e7cfdf6f0ba8

Observation eed7dac8-d443-494c-8abe-df58544933c5 · inbound

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation cites this paper.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.252855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:5463df3c0b044ad0bba360f675bac2a8b2e7e7ac7e8ccad5dee35bc45853b9b6

Observation 150303d5-51e5-4690-8ec8-b4083d854709 · inbound

Inner Monologue: Embodied Reasoning through Planning with Language Models cites this paper.

Inner Monologue: Embodied Reasoning through Planning with Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:10:45.889197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T20:10:43.912935Z digest=sha256:cafbca2e7193264483e797a8f3de521f001cb3da2ad2b989e963f9ba078b6482

Observation c47735b6-4b4a-4bc9-a0c3-540ba7b778eb · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:29:06.091876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:df9895eae6c4ee60a3d06e58f2d82e406430f9e0a84106f9b17a1daaf991b537

Observation 46221846-5e44-4ae2-81cf-afb669f6962e · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:07:42.982914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:cf4a97dacd52c8c435f1a1d11a406a05ea71801aa4e0bce6f66a7a1e474d1f57

Observation 140417e2-36d2-4f70-aac2-476191209903 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:56:42.300910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:3641efc263aa6ec46fb29242dc70c2833c0884342b6d668ecc3bcd73ccab680e

Observation 03da7090-bb13-4508-8d66-f508abdd6740 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 290

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:25:59.818454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:9d91b12ec43228f8fda7f51d91ae53037690c817d4c5d44f290e3b84bea1f6dc

Observation c3724561-48ee-48d5-9eca-06119d69a7ea · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.208526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:b7566407196e36e95f52abc2b6e7797344d18e34012588a764309ad26261a34e

Observation 8354bd25-586f-4f95-91f4-16e04aef353f · inbound

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding cites this paper.

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T20:14:03.546300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:14:03.546300Z digest=sha256:e1c2bdfc4fbd3c5b4d2ef0c7d6835433b37ba1b8220e4923de881c48f990664a

Observation 16d8b0d0-1aa6-4358-8cfd-db03e911ef78 · inbound

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval cites this paper.

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:40.666585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:40.666585Z digest=sha256:d889a62547f70799872737842a5552f8afb77985f9b6dba463c0f0348da45e69

Observation 77bd1637-6379-4ef3-8d25-928d3b36ff97 · inbound

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model cites this paper.

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:41.477507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:41.477507Z digest=sha256:6143ba0d4b19658487cfc8f9832d95baf334b150964b6750fff6ab8db57d470b

Observation b0e924db-1927-410b-88be-ec97b3d2c838 · inbound

FREE: Fast and Robust Vision Language Models with Early Exits cites this paper.

FREE: Fast and Robust Vision Language Models with Early Exits SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:45.667153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:53:45.667153Z digest=sha256:676fa3170455c8e3c430f45f713e6ba3781cefd3f438c0e04b019c2e7a51548c

Observation 49ce8a1a-37bd-4601-aaeb-0f09e6ec81a7 · inbound

CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization cites this paper.

CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:28.993874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:28.993874Z digest=sha256:8476fc1d8434b433cf244d656bb2274623dc8d4a8988933d53b57fc85dc1b2c8

Observation 32f16b24-5e42-4983-ad77-b0bbfb6583fc · inbound

SensorLM: Learning the Language of Wearable Sensors cites this paper.

SensorLM: Learning the Language of Wearable Sensors SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:06.395146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:06.395146Z digest=sha256:95c63d1d35192a8b0792da6b3e2aed1b80f0127e9e449f3c641a5c7dad59958b

Observation b92bfa76-840a-4600-bc97-171cba9da0f1 · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:02.444346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:02.444346Z digest=sha256:a3951a3c979ab953078bbc2e97d5076951147b5f35e7f6fe4d7121e2ae9717dc

Observation 940e16ad-a44d-49b6-aea4-fc124200a517 · inbound

Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data cites this paper.

Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:51.039188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:51.039188Z digest=sha256:fd8f6e67cb55f0e14f8dd2429c9d0e832ef91fccf9f0462f33f9f38946d29a2e

Observation 354e4939-3157-42f6-b4cd-52aaf15c6c4d · inbound

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges cites this paper.

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:03.339629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:03.339629Z digest=sha256:ceb48c9d383ac4d9d038495f0a361e97fc9f1d4b5c5a2c00704373e51c9e1c28

Observation 8c065b3a-771b-4689-aae5-c083f0391f89 · inbound

From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach cites this paper.

From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:45.349914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:45.349914Z digest=sha256:534d474c11729b8057fd3af0d8a1f283564d272651f6cc6fc5eaafa06ba79c5f

Observation 7feecfc5-30d9-4655-bd45-b597bfe9ae9d · inbound

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models cites this paper.

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:54:10.057658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:54:10.057658Z digest=sha256:5a898b2cdec5142a7a59a5a4c21e95796ac00a01910e5201708a20a141a06e18

Observation 9b31139e-1020-4c3a-ba56-e8306c567ce0 · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.077148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.077148Z digest=sha256:672ee8b399617fd2d74481a52e9e8f15fcf00ef41d27dd9bde5cc952c9f25538

Observation 2a718101-06aa-46d8-a604-39b2520da10e · inbound

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey cites this paper.

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 214

Resolution
unresolved
no resolver link, observed 2026-08-06T11:20:07.002898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:20:07.002898Z digest=sha256:8f3553835b5b8962496e6193cc67fa3c7846c0fcfcf3475879f221614df7292f

Observation 57360d68-7f3d-4b8d-bc83-b0b18d956ae8 · inbound

Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models cites this paper.

Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:58.076344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:58.076344Z digest=sha256:dfe6ee14ef9f1986252ad1ae188eca18bfaad292ee3f51b7731a608e05209af8

Observation c5700a91-3ab5-44d5-b93b-1c8fa37b2bb9 · inbound

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline cites this paper.

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:38:37.414051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:38:37.414051Z digest=sha256:d5672421e2e054fed76e1359efbe7a0926803b60f65ed5ebcd97a3285d7a8822

Observation a81d2586-752a-418f-a8ad-eab72abd079d · inbound

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework cites this paper.

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:45:36.961944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:45:29.400398Z digest=sha256:9ebacdb1aa2198bb535b1857c7b9d8da63baa10141bc8114de34fbfe4cb4fdd8

Observation 5e42c1d0-96fb-429a-bbdd-c8fbb123c2e5 · inbound

Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method cites this paper.

Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T17:14:57.572731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:14:57.572731Z digest=sha256:600e786d557fa13eff1f4f114230a25c8b5b8d139e95fcc47643d2a6707d0f7b

Observation 08bb4adf-603c-420a-b56e-f78030cdd1a4 · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.008483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:15fffc4f39e4c8f398e0731e27b8d041cf0770d56fe89997279fa3579ff94e42

Observation 7830cbe1-7e0c-47c9-b9f5-715b7dc3aaad · inbound

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images cites this paper.

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.372269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:33:02.946839Z digest=sha256:78ee10e2cabbf928597751ffe91f9b92a8d304436e8445ba97990c667a5e4e1b

Observation 6d9a4388-4867-484a-a86d-89a600935572 · inbound

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation cites this paper.

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:41:25.994420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T10:00:25.846913Z digest=sha256:5c6b68cda6ad9bc5c437968104e63002adcda6a71b40adb6ede126d64d957fb7

Observation 02ec7dbd-701c-485e-bbcb-d116d23e87f8 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:11.367780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:48fc8cbe3975711e6b7aa75d5e219f7bcf76218ae5c23e26c5bd9ec831656366

Observation f7df252e-0aa0-48c8-8a74-640ef2e21b82 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:28.761944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:40e19c4bfab621416f1020e9147e1cbe5b93f096197fbfd902fdd7595cd10d35

Observation b8422f70-2647-45b4-9bfb-232a5327bb6f · inbound

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments cites this paper.

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:44:57.697014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:40:33.082748Z digest=sha256:19ac45fb259d33977e33390a153f817428ac08810a6adb3d3c474f31906e5fa9

Observation 3f559bf9-22c8-41cd-a996-5ea9fef6de58 · inbound

ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation cites this paper.

ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:03.089296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T09:45:35.383450Z digest=sha256:a6ed8d707a254e83c835bce1a3516adb5f1e7031a513b0c7eebb8ee42156bf40

Observation c548dcb9-23c9-4aeb-95e7-9b654439f2dc · inbound

WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning cites this paper.

WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:58.489179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:53:11.223341Z digest=sha256:b4dfe6feac355f7677c226e1528a1ae9664605527bd3a77902b1dc7f6b602928

Observation a2d6d632-ba4e-4dca-a174-d876a3603909 · inbound

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition cites this paper.

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:10:05.335820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:38:32.529040Z digest=sha256:694c81aece2be593fe9360488103157d89b739c82af66ba1cb3c0f96844c5f08

Observation 8fef56a8-4df4-4064-bd39-324d808df81d · inbound

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models cites this paper.

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.507790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T07:07:00.265141Z digest=sha256:5cd47f9a5f39bb8391120c4e3ddaaef7a1de9a8bfb12a3147c7b4573d18a1371

Observation 87bd443d-8580-4a0d-8b66-6712e854a642 · inbound

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models cites this paper.

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.354415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T20:46:16.298898Z digest=sha256:cf781c242a431996c942b74303a117a35888a58ddb2cd1154f1643ee0b053f4e