Pith. sign in

Paper Citation Record · LEDGER

FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2503.19907.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19907 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:51:45.899451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:08:42.921278Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9b217b4b-27a4-4327-aa33-a7eb093fda5e · inbound

FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers cites this paper.

FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:45.899451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:45.899451Z digest=sha256:7fe20909e5b5481224ec5c286b439feea01086802c4ebc1d1cb345d0083977a7

Observation ac9b5318-3da5-461f-acb7-9f442663a3b6 · inbound

UNIC: Unified In-Context Video Editing cites this paper.

UNIC: Unified In-Context Video Editing FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:43.089813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:43.089813Z digest=sha256:e66f566d6c2627b142eabe0c418d6678869507d067f3094aa4636b0f5aacbfc5

Observation 6d30fc64-2077-4d5d-b3d9-fb1fae6e4570 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:09.305048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:09.305048Z digest=sha256:a6c3c31aa4fb58e4edd29529ee81bfc1f7c9d0b06d39fe1de6c1db3de3d4c560

Observation 43531f67-cbdb-47f5-8b93-7649a5243a81 · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.042362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.042362Z digest=sha256:159167f34d3c1318a56e23755168c67196004e735a44f2f60bee59d4141a1c5b

Observation 326e8b82-529c-4d2c-ae8e-d229d8f0a640 · inbound

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning cites this paper.

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:45:58.733166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:45:58.733166Z digest=sha256:d9288b106b6536fb7c910a57fe0242055f53834c96251483914274ca6670b273

Observation af9e40fa-91d3-4906-8109-06af2f550f28 · inbound

VideoCoF: Unified Video Editing with Temporal Reasoner cites this paper.

VideoCoF: Unified Video Editing with Temporal Reasoner FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:08:43.140676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T00:08:09.479706Z digest=sha256:51255f91a7f91c947bdb8a29545f2d4619b5f0792570bbfdc46f330671d57484

Observation 719cb1d5-e969-4130-ba51-548165db9fb1 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:53.090750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:53.090750Z digest=sha256:505359feebdbfa94ec5087b779795809d92c6b46522e9087ffe469b571ed0609

Observation 6db1d248-700f-4790-a13a-fa9dcf342f7a · inbound

Lighting-grounded Video Generation with Renderer-based Agent Reasoning cites this paper.

Lighting-grounded Video Generation with Renderer-based Agent Reasoning FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:01.722281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:34:39.848940Z digest=sha256:86ff237e6aa0b06bb797a68a8bec01c389d590f76ff6fca7c558cf207d1b89b2

Observation a2f2b7b3-2aab-4796-87f4-1b828f9644d8 · inbound

How Far Are Video Models from True Multimodal Reasoning? cites this paper.

How Far Are Video Models from True Multimodal Reasoning? FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.097620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:6b4d2afa5c2ac32f19e43baf69710f7ed30f3e506c4cfb62136bd929f6d7b880

Observation e88d3aa3-fd8c-4a6f-a198-cd832ffda9f9 · inbound

SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation cites this paper.

SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:30.376329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:57:38.215348Z digest=sha256:ec7d2d36c63186ec7e4a26c9206a33691bec629164bb253f7f1a6c21aaa6d3ba

Observation 22a8db31-2592-4651-8354-a56b870636dd · inbound

MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation cites this paper.

MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:55:10.531147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:53:13.152677Z digest=sha256:9b055cd8cc688061d8ecb19de5cd1eaccb6d09ce3c7d7436806f3a5df5c3f0d8

Observation 57b33d10-cbb2-4eef-9c9e-1334436e31ad · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:48:15.025425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:8ee159ddc87ae8fcb5da6537a7e3f850d61abece5b151f0f078a79a9bd71c671

Observation 3c03f3d8-85e4-441d-9848-c4352dbde82c · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.710354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b2143ea321ad3b4348541fd4164e56773f1b9c10e3e85728201657b312af0cc5

Observation 7b0e3336-454e-454b-9bb0-99acb4d42399 · inbound

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework cites this paper.

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:35:21.904497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T04:30:20.593882Z digest=sha256:12da02a10cfac50bcef21765bbbd9aece09fc97c29fe5ce19ddf2a0978ecabc7

Observation e70ec321-8b18-4fef-8f63-6490a52bd727 · inbound

AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization cites this paper.

AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.363519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:25:18.152134Z digest=sha256:9f118d247929c3a359ade96f1d6d14ea936092fb025a77a7eb2c6b83b90da1f1

Observation 13bc3071-6a48-45aa-b973-465c8921367f · inbound

ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning cites this paper.

ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:42.922951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T16:59:40.925535Z digest=sha256:4626821abc021796e26e32070c06c1371326ae220af5ca4be87646c8c8fa90e2