Pith. sign in

Paper Citation Record · LEDGER

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.04292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04292 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:19:53.886239Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.252122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e78c84d3-b0fa-4346-abd8-c07bf685350a · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:21.010148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:ebc1f73a329dfb8f4ce7373262f181ce4a98374f1523641fa844b6593eb0829b

Observation 527ce582-fe2c-4f10-88d6-e09feb5f7953 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.859394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:2aae43761501f3cef55482cc5a9ea7cb17d5bd9bca9146147c2feb6695c031ff

Observation 62fa00f9-d82b-4ced-a41e-43fc4c60dcbe · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:44.622251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:f47dbb9514c92e7d5040a103f97a9b7dc26ae42bf988f33b77ca2048a926d308

Observation 79da0965-9f65-40f9-9e30-40f02bcb645f · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.218371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.218371Z digest=sha256:0c38ae815ac62f13f621d7692edbb6479fb477701ea68794f16d46acd10f46cc

Observation 06be4929-c3df-439c-8cd8-4a42f4069f70 · inbound

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG cites this paper.

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.170740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T07:19:44.125479Z digest=sha256:420d7d5522d7892ccf0b886770b1cda2609d300fb5781d87f258dd7a275d2014

Observation 0ce5236c-ea7f-4cb5-88f4-1e587a8e86e5 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.423530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:e8e12365344595eca7217ef8da692812fe332dbd78f94ce2d0ecb46fd4f26e19

Observation c764cc1d-0431-4631-ac7a-56970e5dd1b3 · inbound

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception cites this paper.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.255293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:84e127d918016340382e416a49fb19c5c656d936eea447b740701974ba5ce5aa

Observation 39f0b923-903b-4568-9ae5-62839bfda0a8 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.822604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.822604Z digest=sha256:a9e75885e659da506d0f23776c0633f7e43a1a32b6ff0dfb5746939e9bf36fdb

Observation cac9af94-b7cd-47cc-baff-5f9517ab36f1 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:53.886239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:53.886239Z digest=sha256:e0df2bc41343205eb38c90ddd6b4c115a3abc077143bd96c76a79a51fbade9eb