Pith. sign in

Paper Citation Record · LEDGER

X-VILA: Cross-Modality Alignment for Large Language Model

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2405.19335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19335 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:19.867312Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T03:51:25.455093Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a786d3c0-69ef-421c-8c9f-00faf4e84fca · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos X-VILA: Cross-Modality Alignment for Large Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.458475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:100be9621df5a14ca2a73fcfb601c0fe7d5577c0f01d8fbfbef6bcaa7d25885e

Observation eee9d9b1-d53a-44b3-aa83-7c96bf292564 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation X-VILA: Cross-Modality Alignment for Large Language Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:03:33.635353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:23970891db80b3f1fc5db3f50a6c8b58c345f7fe62e417586e4847aa7e2d7de0

Observation 8d7544fe-b2c4-43d9-a0a5-1cc01b9d76ee · inbound

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge cites this paper.

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge X-VILA: Cross-Modality Alignment for Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:19.867312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:19.867312Z digest=sha256:d51fb898fda8e0c0cc0cdceee36aef5236ec6ccd6640686c812f0dfed830af7a

Observation f2facc8b-ee40-44e6-8657-adfca44fe9e3 · inbound

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance cites this paper.

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance X-VILA: Cross-Modality Alignment for Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T19:29:52.758612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:29:52.758612Z digest=sha256:2e4847fe8e1277692bec03ae7b13b108f40e4b960d9fd48e33806beed7380c22

Observation 4bc95560-121a-4c73-8631-2b8e258f4c55 · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks X-VILA: Cross-Modality Alignment for Large Language Model

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.678270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.678270Z digest=sha256:b542efdfef30b7f4d40b12eae00e2562f2ce492611865afb723a7a725166aa19

Observation e1bfa684-4684-409b-bf2e-c390f353b7fb · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths X-VILA: Cross-Modality Alignment for Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.461321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.461321Z digest=sha256:e27f8a00bd9d1764a610e78446c3ab3fde686ac18a308c321cfc3eb92b9a66cd

Observation 33892029-21d1-4f12-b64c-54f9cd3d0a5e · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models X-VILA: Cross-Modality Alignment for Large Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.600967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.600967Z digest=sha256:6f3bbd57a058a7a44d0078b1312a8149ddf9eb39aabf994e85b7148492923b7b

Observation 4806fd98-19bc-45ea-9e7f-f0b3208da644 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation X-VILA: Cross-Modality Alignment for Large Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.881448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.881448Z digest=sha256:c8499a96c5dd5445d772ec8ba92939db0e76eda067e1203d16edf99d6ae900a0

Observation abff7c5e-1703-4c7e-8ceb-c800c58692f3 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models X-VILA: Cross-Modality Alignment for Large Language Model

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.775698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5dbbe8bbd7076a2cf2a3bf65438fec2f51f2af2e2574a4c9e9ab8b9a50a9ae70

Observation 6c494ba1-dcda-4fb4-94f3-cb94560ffa83 · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language X-VILA: Cross-Modality Alignment for Large Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:16.066143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:16.066143Z digest=sha256:844096ac6b346f75a6ceaba3901f22325fd462ae41453ee1f7896cb16b59562e

Observation 3ba61ad1-768a-456e-a447-82f18c302af3 · inbound

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards cites this paper.

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards X-VILA: Cross-Modality Alignment for Large Language Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:27.201831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:27.201831Z digest=sha256:10bf69de97de1d2c75134810c883d36bfb93e08ef529ede7c93433b8d122f6c4

Observation aea9a6a0-b6fd-4e56-af63-735f14128493 · inbound

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents cites this paper.

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents X-VILA: Cross-Modality Alignment for Large Language Model

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:19.017616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:19.017616Z digest=sha256:741cf8a019f2c1472125df0632b3f8d664c8127afca0819f8862b08596587a04

Observation 3a650ec1-b016-4adc-a5a3-61ba922d8f23 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations X-VILA: Cross-Modality Alignment for Large Language Model

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.541496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:49b0b9a118f5f0157f786652fee29fbbe5337d04e372abf7b7a12026dd72e860