Pith. sign in

Paper Citation Record · LEDGER

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2412.11621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11621 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:49:08.954450Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3704557-7b7c-4f31-8cd6-39bd177236fd · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.560441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.771367Z digest=sha256:55139d142c31ca38b2dc49d712c0721441a4d64040aac9a76e43a85147a3420f

Observation 335ed295-1e65-46e1-88cc-596fc51371fc · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.775952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.775952Z digest=sha256:ef263287290740b4d30ad6c5e41fe077edec1a0125e8b576f51f648a8e1289a7

Observation d53646f4-ed9d-48e2-9d1c-bd94b34e96ca · outbound

This paper cites W.; Fidler, S.; and Kreis, K.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting W.; Fidler, S.; and Kreis, K

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.536762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.779774Z digest=sha256:19ce6b7b1bb8fbbec9c78174e8010b8b1c678bf1203249c1fd66b5f20359fc11

Observation fc4ce99e-b9b6-439a-bbf5-690cbf334087 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.524899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.783638Z digest=sha256:06fd656fdc9fcbd570c7b184f03cd7540057daf42ad7adc98cbbcd39141afcae

Observation 69e13dbf-1d91-4e02-8842-ebf95ba1bfcf · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.512793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.787366Z digest=sha256:577132bb54171a5330fee5c384282adae83f7aab9e76330f165c480cf45ad81c

Observation f0f099a0-f06a-4b87-accf-3b84683ab146 · outbound

This paper cites Prompting Large Language Models With the Socratic Method.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Prompting Large Language Models With the Socratic Method

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.791812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.791812Z digest=sha256:53ce42acf9d77c49c4f6d97adfa0eb2c78ef5b358d3916eb875144badb14c714

Observation a090152c-e966-4ce8-8258-ab4665fec9f4 · outbound

This paper cites M.; and Cardie, C.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting M.; and Cardie, C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.501437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.796874Z digest=sha256:bd21933398626155fb0aa08939fd2f7a0c6152c6b5308d212799b5b74f927757

Observation f6d5fc53-e486-42f4-8cf6-11c77dc60bc4 · outbound

This paper cites G.; Wildes, R.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Wildes, R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.490949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.800722Z digest=sha256:e0dc385dec3d443b1eded74464775e5b02f73cdcfbaf11cace4b81b3ac60ebe4

Observation 42092e8e-5e8b-40ae-a81d-9e27fd1ec5f3 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.479735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.804970Z digest=sha256:7a52ab1f99fa726a272f54f9e6b820b282f7de837913ada144f1aa622cee57b5

Observation 95726092-4085-4a18-8bf5-f65d9fe42131 · outbound

This paper cites Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.808773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.808773Z digest=sha256:940894b2e4f523d4c92b74f0924cebfb453249dbe740d940d4b9547ea8ee7b7f

Observation d32aa3d6-7551-4f38-b0ec-450175a8d828 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.467848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.813282Z digest=sha256:2691270a32237643172f1e6a538439f2fc8311ed2f4778d0f06c5cbaede1df6f

Observation bbb01c25-d681-4ccf-b080-b6caf2b07682 · outbound

This paper cites F.; Song, C.; Chen, J.; Gao, D.; Lei, W.; Xu, Q.; Lim, J.; and Shou, M.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting F.; Song, C.; Chen, J.; Gao, D.; Lei, W.; Xu, Q.; Lim, J.; and Shou, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.454996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.817561Z digest=sha256:c32b22dc2f50a2d7ce1bb4f44999ca296fe0d1860dfd844ddcd4e1c7280205c0

Observation 149f6b9a-9f9f-42e0-b150-3b4cc7aab241 · outbound

This paper cites Mistral 7B.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.821527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.821527Z digest=sha256:98f1243620bb643e898a3404467415a907b89e92059c2a24d5f8c7252636af6e

Observation 04983b0e-a886-42a0-8d35-e1aeba6aec35 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.439353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.825759Z digest=sha256:0fdfe4f1dcfad11b4f47969ed5ff0aeb1b1cd9907ec0652edd288369f52f1709

Observation 74f526f3-895a-4219-895c-7e73c1c056af · outbound

This paper cites S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.424290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.829227Z digest=sha256:3fe6b54e0ea9a976ffd5c40608090c3121cb3a3ff701822f475fa6a659a5d76e

Observation 033db58e-43ba-42b5-9e20-57cec483aad4 · outbound

This paper cites B.; and Serre, T.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting B.; and Serre, T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.412834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.832764Z digest=sha256:bb14cd2a49dfcf4a260d204fe982a641ce25f25f4ed8da6d9f0d9229a1eb514e

Observation c790b25b-f0cb-4854-9dfb-80459f26cf61 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.400947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.836447Z digest=sha256:56b5e0ead1811bf371aad79124bcdde3167b8bff970a08755bee45e0d621b7a4

Observation 545c1ac0-e76b-4f31-9eb4-25ccb4cc5f66 · outbound

This paper cites LLM-grounded Video Diffusion Models.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting LLM-grounded Video Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.840008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.840008Z digest=sha256:91b3f42bc04d15288aa64054a62d7acabc0415616cf83acd26f77d47a085c002

Observation 00aa1965-8e9b-4a68-bbe8-4f771952e744 · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.843675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.843675Z digest=sha256:9265112d82d1a242999849318165d157c03caf095c6b1f28508726566ba95cbb

Observation 24affb3b-56d8-43e6-a6a7-830cf003d20b · outbound

This paper cites Q.; and Lei, S.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Q.; and Lei, S

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.388594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.847944Z digest=sha256:89528ee86ef3035a9c2c1156c7758da7db305fd664e1e5d6f368c0aed398b539

Observation 7b76ba1a-ecfa-401e-89e3-b9798d9a14b0 · outbound

This paper cites E.; Eckstein, M.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting E.; Eckstein, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.374221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.851942Z digest=sha256:b50cc40915c3905e70f2fbdc191f1146d1f112c98c859cfe84cd719abe4d13a2

Observation 62fe9c8b-9799-4a95-8852-8ef2bc278224 · outbound

This paper cites Multimodal Procedural Planning via Dual Text-Image Prompting.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Multimodal Procedural Planning via Dual Text-Image Prompting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.856049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.856049Z digest=sha256:892b00c19c10e26fdfa4cb1ff9498beb1d85e9c83e0174e94ef5bfc33eab6e28

Observation 376e6aa1-4f9b-4510-aa98-6c91b535c724 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.359713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.860635Z digest=sha256:b0503921b8673bfa942ce618b37a04f2102ec1731adac8f3a3db526f58f84c37

Observation ea943d56-b61a-41b4-951f-c5ab3e59e352 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.346541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.864619Z digest=sha256:c5be580d25a495b5efa1b33fec33cf9e378167dbffd1b28f687b09a45b16f45c

Observation 4ec25088-2703-44a6-bcba-f3cfc7e71df8 · outbound

This paper cites GPT-4 Technical Report.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting GPT-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.868882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.868882Z digest=sha256:77d1508e6eb4bc59188adebc55fd1833b6a7a2819e38e37aa0c26b302493d9be

Observation 94ae40cd-cc7f-4ce5-b8a7-c7dbc941ceab · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.872878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.872878Z digest=sha256:81cdb9ab1288658c5f464fc7166b401df949739a9ec526ad2a9d433781fb9996

Observation 7c8f02b2-111a-42ec-bdf7-73e276cba1c2 · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.328255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.877444Z digest=sha256:6c8ba17a4b160b72179368f962edd885117c947108de5ee40dce1a00e678bb09

Observation 80bbef9c-ece4-43e5-9016-23f1f6c5d06e · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.317989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.881421Z digest=sha256:662a9dbc709a823a7bf102d832987f815833feb460875adbb4a8ab0cc5a9607c

Observation 14ecd0dd-a6f4-472e-961b-460215eba19a · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.306590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.884976Z digest=sha256:2f64791fc684ab15c09299b1ed0708e69278cf7d4175c8633cb6d0908d887bc6

Observation 23b1f492-326d-4dd8-b1b0-c5a864c906f6 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.295544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.888324Z digest=sha256:aef67f25ef49b6cbbd4f1184d91751b2856c68f44d22df92a161d18a801ddb2e

Observation 598ba009-f5cd-4530-8272-9e2c58842eb2 · outbound

This paper cites H.; Sadler, B.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting H.; Sadler, B

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.283311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.891605Z digest=sha256:ec52144a720b00054bd241144b14e978904c2b76b64e096fbeaaf10432c3d6e2

Observation 5c6ef65e-f41b-4872-84ed-b9c6a83fd178 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.270981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.895322Z digest=sha256:dfd9b776c9a47336a6721a1d2c6ac73130d69c3cca8db6ca5c5cbad410f2e146

Observation 9124f636-ed3c-4f6e-ba19-9b26190776b6 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.258180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.899103Z digest=sha256:b17061eb96e65d7fbfe3fedf15fea5c78ebf7074184c7728c03ee74b1940b90c

Observation dd4fa7a6-4250-42e6-ac1d-289d73622c29 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.902451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.902451Z digest=sha256:f6f98d00feb01851311490d85f8a16b0c64c54631f5bd80890032cc5ccbbd682

Observation fc121910-51f0-486a-a7f3-67ca2a509bc0 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.244619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.905627Z digest=sha256:997397241871c543c0290e71323e990deb56bd9b097fa1390242b2b4a0f4d1a7

Observation 263817a5-8948-426c-bbda-94e44b92593e · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting ModelScope Text-to-Video Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.909066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.909066Z digest=sha256:1d6fd161e3e03bd68533bd3e83a8f072b5d772a3384428e65e4c68beb3356ad4

Observation 2deac228-7cb7-4aa9-9c5f-88a36854967d · outbound

This paper cites Z.; Ge, Y.; Wang, X.; Lei, S.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Z.; Ge, Y.; Wang, X.; Lei, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.233322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.912657Z digest=sha256:5844446825f84ff1f89e50903f17dece48e823b5ba09ff77798569b7613b37f0

Observation 3a5dc640-8b66-4a9b-a0ea-dc86b1d8a70f · outbound

This paper cites H.; Miech, A.; Pont - Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting H.; Miech, A.; Pont - Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.220251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.915869Z digest=sha256:3003e8e25b641084a8b30ff99f12483fa306f23d202ddd0085f5ce680cb2c48b

Observation 39972daa-76e7-4027-bbe8-784006803196 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.206886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.919591Z digest=sha256:e971d0082b1a3d6f33c4f6b46e3aa9f5f7a727a77ffc13f4e9d2481eb57ea8f8

Observation 784b6a82-36b1-4ffe-9c22-1619ecd18db8 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.923237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.923237Z digest=sha256:60524bbaa12192239cf46e25a0f253c3e2f92c096dccaff9eef90c0ce50a8ed5

Observation c17ebd86-dcb9-49a2-8308-8f94be3ee6bc · outbound

This paper cites G.; Wildes, R.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Wildes, R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.193535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.927426Z digest=sha256:b63f1ba417f14a648e1e09e4976b502496d0e4b6f65882e2c8e9b83e0990ae9f

Observation 9820123e-4482-43f6-bbaf-d247a33215a4 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.931033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.931033Z digest=sha256:0dcd92a1a960481cc9ef908b4b175d669228201e69526efd33cc255d50cc3f30

Observation 10658eb0-4a2c-472f-9925-e6ed54ea9942 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.180515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.935214Z digest=sha256:720eeffb001f41471ce95326b23f116939135b228ae6faac52c6e15f7fba1d7f

Observation 6dcd6dfa-435f-49f7-aabc-98b375f1bcb8 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.167551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.938944Z digest=sha256:116e964bab5290982066613d2fb12cfa56e4354e29bf3b8b21526de03d827fd4

Observation 1e4c5177-b20d-49d4-92e9-a7b8df3ffaf0 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.154446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.942721Z digest=sha256:e6d949490e6b71ad856280c20fb10f47c039b77a797071b23aa7300600dde12b

Observation ad2dbf87-0579-46ce-bd2d-1804976dbdf5 · outbound

This paper cites G.; Fouhey, D.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Fouhey, D

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.140752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.946423Z digest=sha256:34b61f2ad953ea4c5137715d8b6dbcab69a2a2eceb787f069fe071f1185673ed

Observation 63549908-fe39-45f7-ae68-14993bd8d06a · outbound

This paper cites , " * write output.state after.block = add.period write newline.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting , " * write output.state after.block = add.period write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.950286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.950286Z digest=sha256:b1899f5fa8c92f33d990a3bacec4573be38a707583ef3cf5a7c8bf0e78077fb1

Observation b0daf325-8131-4ca7-8b22-a331bb7f1e31 · outbound

This paper cites write newline.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.954450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.954450Z digest=sha256:67494c6cc87518a8d75e359c9c969d53260ae70b420c8c206e4d9d2c3191eb70

Pith citing papers

No inbound Pith citation observations are available.