Pith. sign in

Paper Citation Record · LEDGER

X-VILA: Cross-Modality Alignment for Large Language Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2405.19335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19335 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:24:46.461321Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T03:51:25.455093Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a786d3c0-69ef-421c-8c9f-00faf4e84fca · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos X-VILA: Cross-Modality Alignment for Large Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.458475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:90c9a701c26ac5bb361b98f74003a90598536474c403957910fc5bc5818e2f43

Observation eee9d9b1-d53a-44b3-aa83-7c96bf292564 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation X-VILA: Cross-Modality Alignment for Large Language Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:03:33.635353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:44c2a232a4d4a2c5886d3e9cd4c9b67cd532ed87e4673b4dfb76d37522e9cb97

Observation e1bfa684-4684-409b-bf2e-c390f353b7fb · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths X-VILA: Cross-Modality Alignment for Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.461321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.461321Z digest=sha256:92e114eb09d155a48e5e8b0544179d0a150472da8ef9427051696beb835149b0

Observation 33892029-21d1-4f12-b64c-54f9cd3d0a5e · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models X-VILA: Cross-Modality Alignment for Large Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.600967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.600967Z digest=sha256:37f65130a81e42543c09fa681c24923a9585f1c3c0e069e63d8267ed988c05b2

Observation 4806fd98-19bc-45ea-9e7f-f0b3208da644 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation X-VILA: Cross-Modality Alignment for Large Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.881448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.881448Z digest=sha256:bdb206884216c9127aa91f16db9bfb82956e18b99cafe317b86a99b69421bcbb

Observation abff7c5e-1703-4c7e-8ceb-c800c58692f3 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models X-VILA: Cross-Modality Alignment for Large Language Model

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.775698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:2c21211c7e852614e3e82b1ab5381fcfe857a70da4bed829b841a246426f2a4b

Observation 6c494ba1-dcda-4fb4-94f3-cb94560ffa83 · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language X-VILA: Cross-Modality Alignment for Large Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:16.066143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:16.066143Z digest=sha256:659b34df0c1c4150be607e8565854dfe6f1f292ebfd695a20f53b9271a69cabc

Observation 3ba61ad1-768a-456e-a447-82f18c302af3 · inbound

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards cites this paper.

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards X-VILA: Cross-Modality Alignment for Large Language Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:27.201831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:27.201831Z digest=sha256:df42d77206a698b7142e6c4edba42808980b391b46699645a108734abeb7e9e1

Observation aea9a6a0-b6fd-4e56-af63-735f14128493 · inbound

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents cites this paper.

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents X-VILA: Cross-Modality Alignment for Large Language Model

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:19.017616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:19.017616Z digest=sha256:c66830b4f10f541e0b083abea03306bcff51b91e8dbaa2cbfe0a2a88c7219a78

Observation 3a650ec1-b016-4adc-a5a3-61ba922d8f23 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations X-VILA: Cross-Modality Alignment for Large Language Model

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.541496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:7e0afd40d11b9fa32fd118f54bb700886200526ad6989bf477eda47d17b67683