Pith. sign in

Paper Citation Record · LEDGER

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.04292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04292 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:59:02.069866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.252122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e78c84d3-b0fa-4346-abd8-c07bf685350a · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:21.010148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:454bf3b22014085435202df7450c8d8dd355984e4d97bfea52c73603d71a4b54

Observation a245d924-bf0d-4859-acda-1ba019377bd3 · inbound

LLMs are Also Effective Embedding Models: An In-depth Overview cites this paper.

LLMs are Also Effective Embedding Models: An In-depth Overview VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-11T13:59:02.069866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:59:02.069866Z digest=sha256:7d70ef48afb8deaa4a2bd48bb4e319c95d23533d794a3d2f90d8035c5285524b

Observation 5224de59-74ab-49fe-bb42-b9ca6e9f4159 · inbound

O1 Embedder: Let Retrievers Think Before Action cites this paper.

O1 Embedder: Let Retrievers Think Before Action VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:12.449077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:12.449077Z digest=sha256:cf91d39dd9db1511ee9bc8c10b757e66fdabdeededb90408c4478a340c33eda8

Observation 527ce582-fe2c-4f10-88d6-e09feb5f7953 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.859394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:b8ab082d129a409efe5a62174c361c843c1064df9b2ccf32c9d2f4f4382c1f4d

Observation 62fa00f9-d82b-4ced-a41e-43fc4c60dcbe · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:44.622251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:731e384fc78d569ad3694539c4e1498e02a137fc5a595c3ac524e7d0b9f8dce5

Observation 79da0965-9f65-40f9-9e30-40f02bcb645f · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.218371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.218371Z digest=sha256:672d4dc7c07f56d60288d1c2e1777f3e76b92e7116281b5d86e309323698931f

Observation 06be4929-c3df-439c-8cd8-4a42f4069f70 · inbound

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG cites this paper.

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.170740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T07:19:44.125479Z digest=sha256:2609fbe79c4f7d587b28cad9031c6bf836156ed3e760166925fb95b3e7bb2aff

Observation 0ce5236c-ea7f-4cb5-88f4-1e587a8e86e5 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.423530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:862ff038560a71950711ce346fb065b0d8cf9d6f3dbb76cd6b8e4d8522388ae5

Observation c764cc1d-0431-4631-ac7a-56970e5dd1b3 · inbound

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception cites this paper.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.255293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:332f844d508753e76cd3e55d7482b030820cbeab36dced790e754b04a7a51bae

Observation 39f0b923-903b-4568-9ae5-62839bfda0a8 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.822604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.822604Z digest=sha256:fd222cfc242f3a7e42aed30357d9b6491e473f5681e0d623e8a3cc70f55863e5

Observation cac9af94-b7cd-47cc-baff-5f9517ab36f1 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:53.886239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:53.886239Z digest=sha256:ded3f5b3defd7b64a46ce1a157d5431d352976d8843775e2416aa9ae5d5e5718