Pith. sign in

Paper Citation Record · LEDGER

Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2403.07750.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.07750 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:53:04.939884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:24:55.402733Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6ba864a9-bf1d-4ce1-b278-b2937d6ac593 · inbound

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey cites this paper.

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-10T04:36:37.535244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:36:37.535244Z digest=sha256:da49f9a96c469cfa86c1f2e084dd79a5637376906cbb47058af6701393f947d0

Observation 9f099f6a-0515-4613-a439-f37cb91ae3c0 · inbound

Vision-Language Model Dialog Games for Self-Improvement cites this paper.

Vision-Language Model Dialog Games for Self-Improvement Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:21:38.283921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:21:38.283921Z digest=sha256:59f2cdbff7af0dce83e24bfcecc1735eb7c2c0c81d576348149975f018673f4f

Observation 800043fb-8e38-4df4-baa3-8aec5c4a1828 · inbound

LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance cites this paper.

LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:04.939884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:53:04.939884Z digest=sha256:b4d4a56a3d5626436ca2a6e31fc29e2e97a55faad4569037de74d27f6216ff61

Observation d002c5df-03fe-4763-a28f-23eee2621d6b · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:55.438877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:53.424779Z digest=sha256:3625da2ac220f92cf5f5058a9dfbb18554d3aa81e3ad06f44df4ec3d2de2d337

Observation fe936764-5bae-4536-aa89-35828de598f0 · inbound

Mining Contextualized Visual Associations from Images for Creativity Understanding cites this paper.

Mining Contextualized Visual Associations from Images for Creativity Understanding Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:09:27.077293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:09:27.077293Z digest=sha256:7f5158babd4a80acacdd4e22f80c322958df7d8c39591bc73576510dbbc75640