Pith. sign in

Paper Citation Record · LEDGER

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2505.10917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10917 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:41:00.805264Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T18:45:00.174957Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3e4f6cf0-67ff-405f-88ce-3095010eaaa4 · inbound

Fine-Grained Zero-Shot Object Detection cites this paper.

Fine-Grained Zero-Shot Object Detection VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:41:00.805264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:41:00.805264Z digest=sha256:01106cf2b2584fdc8d8493808bbee645ba75862e80f6a702b72765b35d681228

Observation 67c046c7-93fe-4e25-8faf-bfd9bf4d1412 · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:52.569023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:52.569023Z digest=sha256:c747b2ed8c0945b0c9cef90b84390f46cd3f8fafc7f2b3feb058a3862dffe0ea

Observation af063bff-0d23-43a7-bdbf-c9a6ab11253d · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.446584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T11:38:01.947806Z digest=sha256:9178e9ea1f64450bcc6c5899c0e413adb4628733ecaee67f0cf65bb88cc4c956

Observation c0ab3302-a89a-4ce4-8e50-5d72fee52e11 · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.177236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:42:46.422756Z digest=sha256:edd47d6f56416612b8e81927b1a0ddce5f781f54d6c74c575c98d3d9352505af

Observation daa69f7a-4758-43b6-8c68-3db61f49861e · inbound

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition cites this paper.

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:42:26.088290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T17:40:30.038337Z digest=sha256:1fd3ac8e0175039e7700b7213dbe6710bbbd18850921049d2e2cd44c81390f60

Observation badd34c8-4d75-4c39-b6c2-9bc6af3ccad1 · inbound

MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein cites this paper.

MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.712771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T07:28:52.030720Z digest=sha256:0ffc40d80ca6e813945bf32fcc7d5638a2bb992d2976f2470c1e8bb3054d91b3