Pith. sign in

Paper Citation Record · LEDGER

Gen4U: Unifying Video Generation and Understanding via Diffusion

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.06856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06856 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T00:16:43.190961Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact9
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4226e5dd-79a6-49e8-90fb-2d5b018c1d64 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.626831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:93f349540f046dd97bf8e0bdf74eea2eccb3c59a6665bbe9114207ca3ccb2bb6

Observation 05d12c6b-149e-43d2-8139-2d3df8eb2573 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma: A versatile 3B VLM for transfer

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.650395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:e146457e66bb18ca583470fc7b160c3f24975bd0a4a8c0e4667b7437a2cbe256

Observation afb13ef3-a1bd-48c4-9abc-a824187c2672 · outbound

This paper cites Walk in the cloud: Learning curves for point clouds shape analysis, pp.

Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:26:39.079490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:eea24c6fc09f4521f9cd791a59f4123e2afd67f7df780dfa5192564033091013

Observation a3458577-a4a7-4f4c-a660-a63981166db6 · outbound

This paper cites Scaling 4D Representations.

Gen4U: Unifying Video Generation and Understanding via Diffusion Scaling 4D Representations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.630147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:6f16ed76f3cce9bd529e9a76951a97f58213b9201ad3e79c70a684d086cd3087

Observation 9620d0e2-5efe-4e86-8c78-32f82fb9e2fa · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Gen4U: Unifying Video Generation and Understanding via Diffusion Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.647632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:a42ff1a7d1738c1d0e8eaca1df43fe68fafb2dcd6fced67103ed3be722026054

Observation fe557c13-b1d6-4c7e-8957-92ac7a91ba7f · outbound

This paper cites Whatever next? Predictive brains, situated agents, and the future of cognitive science , volume =.

Gen4U: Unifying Video Generation and Understanding via Diffusion Whatever next? Predictive brains, situated agents, and the future of cognitive science , volume =

Reference 6

Resolution
metadata mismatch
doi, observed 2026-07-10T00:26:39.075687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:9f1fbc21e4fc204c97f065bdef6af99ffb3fdd05543bbc48a7ac55b230c3ef8d

Observation 21174190-d36d-49e8-842e-c936c346c4fc · outbound

This paper cites Accessed: 2026-04-17.

Gen4U: Unifying Video Generation and Understanding via Diffusion Accessed: 2026-04-17

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.255962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:9b163443ae6fc5b37f7f93afe18228f6012f2f773ed293a9aae9bb6c1631a464

Observation ae41bede-a22c-4637-8305-daca36e05b63 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Gen4U: Unifying Video Generation and Understanding via Diffusion Gemma 2: Improving Open Language Models at a Practical Size

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.641686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:0b1b1b7e03b51b8d25513210cf04a3604470d1d268273c2313a45b7be6699d9f

Observation 476340bb-1b15-4b50-8511-7afbd7835261 · outbound

This paper cites doi: 10.1038/s41586-025-08744-2.

Gen4U: Unifying Video Generation and Understanding via Diffusion doi: 10.1038/s41586-025-08744-2

Reference 9

Resolution
verified exact
doi, observed 2026-07-10T00:26:39.073411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:bef02c7663dfccaf95cc4eb8cf9658dc3cb35a8a82db1619df848b733f535a70

Observation 9485af90-32a1-488e-815d-c0a10451c24f · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Gen4U: Unifying Video Generation and Understanding via Diffusion Adam: A Method for Stochastic Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.644945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:61ce2856bb823645b067799843ac3bb368d310f46062536ab7159abba3c881e9

Observation 54b15358-3c4a-47c0-b59b-da5c8b931c46 · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.658479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:e311b0d04e73f8d9a1bf00afd3a21fae27adc670ba1ad05aa8353046764d34eb

Observation 1b609cff-0a19-4121-9a35-6161312c2e27 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14386–14397,.

Gen4U: Unifying Video Generation and Understanding via Diffusion A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14386–14397,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.254281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:e4aaf8d043c07f8fea30afbb718eb5add88f67eabda9886ebb536e50cd9e2a2b

Observation c9737ae7-86b5-4b4c-bbf7-8795e08d310c · outbound

This paper cites Walk in the cloud: Learning curves for point clouds shape analysis, pp.

Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:26:39.075356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:10478ee83cac79bf720034f4911d00d306608759e9c395620bdc668bd94d4953

Observation 2247903f-0719-48cc-8eaf-f2e9a1063933 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.653175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:bce2f500c7d6b09504f114261131da948b213d8745f4c2cdd495ea970f4e8dd8

Observation 9e8f8999-e3ee-43a3-a8ea-35e4c9ddc3c3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Gen4U: Unifying Video Generation and Understanding via Diffusion Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.655810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:8548576f7f6ab5bde2db014bd386237399efce468a783b8061a10d5613357de0

Observation 911d4e0d-d8a7-461d-8af8-46884e9f7242 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Gen4U: Unifying Video Generation and Understanding via Diffusion SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.636319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:7ceec7ec81bd2d52100d4d0877ea39456bac1bec2fb32acd8b4b50c1fa6faa95

Observation 12928339-f65f-4e6a-8c5c-3913658eeaa3 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Gen4U: Unifying Video Generation and Understanding via Diffusion Wan: Open and Advanced Large-Scale Video Generative Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.633409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:f8d0ecff6ed5ba1f70ece763f8e42dacf220e452ae660020305ab853b2c73a05

Observation 7846d0c7-704a-4678-bac8-79df2639d526 · outbound

This paper cites URLhttps://doi.org/10.1007/978-3-031-73013-9_23.

Gen4U: Unifying Video Generation and Understanding via Diffusion URLhttps://doi.org/10.1007/978-3-031-73013-9_23

Reference 18

Resolution
verified exact
doi, observed 2026-07-10T00:26:39.072377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:2698eb7b415f6f955933e65aa9bcb7128fe245c33357bbd2ef8b28180da3ccb4

Observation aa4ea3b1-c690-484e-a86f-6a14c1f70746 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Gen4U: Unifying Video Generation and Understanding via Diffusion Video models are zero-shot learners and reasoners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.639035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:7fae6e1ae4281e7331248a4edf86d21579d158a2ee8b3d71a3134e3a0ffb6bb9

Observation 2362a645-e376-416c-bbee-df51a3d874a0 · outbound

This paper cites 14 A Mutualk-NN alignment metric We describe the Mutual k-Nearest Neighbours (MkNN) alignment metric used throughout the paper, following [Huh et al., 2024, Zhu et al., 2026].

Gen4U: Unifying Video Generation and Understanding via Diffusion 14 A Mutualk-NN alignment metric We describe the Mutual k-Nearest Neighbours (MkNN) alignment metric used throughout the paper, following [Huh et al., 2024, Zhu et al., 2026]

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.250928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:161472578740b84469b6ad6807c2f3c5747f58f5013ea4ea499ce0add9ec6aba

Observation bf9563af-9efb-409e-87a2-e19edb9fec02 · outbound

This paper cites Best Single block.

Gen4U: Unifying Video Generation and Understanding via Diffusion Best Single block

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.252441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:e674fdee636f9e464285286b8e9cd36697d9e00a135541927f7d5080851894e5

Observation 249aafaa-771b-43ac-bd02-f4bc5a525261 · outbound

This paper cites The linear decoder presents frames that are sharper, but temporally misaligned with the ground truth, as compared to the attention head.

Gen4U: Unifying Video Generation and Understanding via Diffusion The linear decoder presents frames that are sharper, but temporally misaligned with the ground truth, as compared to the attention head

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.249176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:d13248a6181a7ea086295fbb007c7b0d1aaccc94a4a8584cb5beb0ccc5dfd3e1

Observation f55dd9ea-e6cd-4d07-a874-21a46d25f479 · outbound

This paper cites It plateaus after 10k steps so that the LLM is still updated very slowly to avoid catastrophic forgetting.

Gen4U: Unifying Video Generation and Understanding via Diffusion It plateaus after 10k steps so that the LLM is still updated very slowly to avoid catastrophic forgetting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.247348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:36bc7cbac61135d04eef28593abbb26eebec345e5ec0f79031d110d9c30b4b07

Pith citing papers

No inbound Pith citation observations are available.