Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:49:08.954450Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2412.11621.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:49:08.954450Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f3704557-7b7c-4f31-8cd6-39bd177236fd · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 335ed295-1e65-46e1-88cc-596fc51371fc · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53646f4-ed9d-48e2-9d1c-bd94b34e96ca · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting W.; Fidler, S.; and Kreis, K
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fc4ce99e-b9b6-439a-bbf5-690cbf334087 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 69e13dbf-1d91-4e02-8842-ebf95ba1bfcf · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f0f099a0-f06a-4b87-accf-3b84683ab146 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Prompting Large Language Models With the Socratic Method
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a090152c-e966-4ce8-8258-ab4665fec9f4 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting M.; and Cardie, C
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f6d5fc53-e486-42f4-8cf6-11c77dc60bc4 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Wildes, R
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 42092e8e-5e8b-40ae-a81d-9e27fd1ec5f3 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 95726092-4085-4a18-8bf5-f65d9fe42131 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32aa3d6-7551-4f38-b0ec-450175a8d828 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bbb01c25-d681-4ccf-b080-b6caf2b07682 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting F.; Song, C.; Chen, J.; Gao, D.; Lei, W.; Xu, Q.; Lim, J.; and Shou, M
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 149f6b9a-9f9f-42e0-b150-3b4cc7aab241 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Mistral 7B
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04983b0e-a886-42a0-8d35-e1aeba6aec35 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 74f526f3-895a-4219-895c-7e73c1c056af · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 033db58e-43ba-42b5-9e20-57cec483aad4 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting B.; and Serre, T
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c790b25b-f0cb-4854-9dfb-80459f26cf61 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 545c1ac0-e76b-4f31-9eb4-25ccb4cc5f66 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting LLM-grounded Video Diffusion Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00aa1965-8e9b-4a68-bbe8-4f771952e744 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24affb3b-56d8-43e6-a6a7-830cf003d20b · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Q.; and Lei, S
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7b76ba1a-ecfa-401e-89e3-b9798d9a14b0 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting E.; Eckstein, M
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 62fe9c8b-9799-4a95-8852-8ef2bc278224 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Multimodal Procedural Planning via Dual Text-Image Prompting
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 376e6aa1-4f9b-4510-aa98-6c91b535c724 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ea943d56-b61a-41b4-951f-c5ab3e59e352 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4ec25088-2703-44a6-bcba-f3cfc7e71df8 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting GPT-4 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ae40cd-cc7f-4ce5-b8a7-c7dbc941ceab · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8f02b2-111a-42ec-bdf7-73e276cba1c2 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80bbef9c-ece4-43e5-9016-23f1f6c5d06e · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 14ecd0dd-a6f4-472e-961b-460215eba19a · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 23b1f492-326d-4dd8-b1b0-c5a864c906f6 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 598ba009-f5cd-4530-8272-9e2c58842eb2 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting H.; Sadler, B
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c6ef65e-f41b-4872-84ed-b9c6a83fd178 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9124f636-ed3c-4f6e-ba19-9b26190776b6 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dd4fa7a6-4250-42e6-ac1d-289d73622c29 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc121910-51f0-486a-a7f3-67ca2a509bc0 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 263817a5-8948-426c-bbda-94e44b92593e · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting ModelScope Text-to-Video Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2deac228-7cb7-4aa9-9c5f-88a36854967d · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Z.; Ge, Y.; Wang, X.; Lei, S
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3a5dc640-8b66-4a9b-a0ea-dc86b1d8a70f · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting H.; Miech, A.; Pont - Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39972daa-76e7-4027-bbe8-784006803196 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 784b6a82-36b1-4ffe-9c22-1619ecd18db8 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c17ebd86-dcb9-49a2-8308-8f94be3ee6bc · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Wildes, R
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9820123e-4482-43f6-bbaf-d247a33215a4 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting MagicVideo: Efficient Video Generation With Latent Diffusion Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10658eb0-4a2c-472f-9925-e6ed54ea9942 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6dcd6dfa-435f-49f7-aabc-98b375f1bcb8 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1e4c5177-b20d-49d4-92e9-a7b8df3ffaf0 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad2dbf87-0579-46ce-bd2d-1804976dbdf5 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Fouhey, D
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 63549908-fe39-45f7-ae68-14993bd8d06a · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting , " * write output.state after.block = add.period write newline
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0daf325-8131-4ca7-8b22-a331bb7f1e31 · outbound
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting write newline
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.