Pith. sign in

Paper Citation Record · LEDGER

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2412.01987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01987 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:04:09.060651Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-23T00:30:55.729900Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:32:18.226311Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbfc89c9-4c74-41e5-acbe-0ae8d7cc9fc5 · outbound

This paper cites GPT-4 Technical Report.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.631013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.631013Z digest=sha256:eb6db6f568116f4f61f1a44975ba3cb74b2f857d3db00462806b5dbf2fa3990d

Observation 0474e1f2-d6b2-4e26-a975-3bf24cdf8c0b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ht-step: Aligning instructional articles with how-to videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.169513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.639434Z digest=sha256:576f5e8ddf2178e9396c64ff93ab0857308cb956bed535c18d698cf301429f87

Observation 8f61ac1d-c42a-4270-a247-83f6c736d298 · outbound

This paper cites Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.154558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.644101Z digest=sha256:96002d8e33cd46abc55273c81538384de036eff0e4b2919b28cbc4afa2d8fa6a

Observation 875724f0-374e-43f8-b6f3-ef6ab842f263 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.140176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.648560Z digest=sha256:cc94a1f019906b8f2f260f14e40495e3b8dc26bb4253186559791e1f646bfaaf

Observation 8ab9f193-3f26-4f9d-a311-62485c7168ad · outbound

This paper cites Whisperx: Time-accurate speech transcription of long- form audio.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Whisperx: Time-accurate speech transcription of long- form audio

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.126539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.652991Z digest=sha256:54abbc798067a07758f5943e8999e6685d6478492a85d843b04dfbeca5d1ec24

Observation cb136022-b1d8-437c-a9af-85786097d1b9 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.657233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.657233Z digest=sha256:15a4f1b0baeadaef52db5bb707762f87f2e15d11dd7f38bdf274898ebc519c71

Observation 2228cb1a-90c3-4c2f-af15-114b78151792 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.661872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.661872Z digest=sha256:e75ca55b427f058107026284955e369baade4ba24a493a2987302119ed75b3c3

Observation 628aaaac-5f4e-46f2-ac6e-73cfb7c3be57 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.670093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.670093Z digest=sha256:a91976a481d7b832fd8258a07073bd27c1606b10b5f3918920d55717aabab336

Observation 746ffd17-70f9-4af0-a183-6a1a1a6b9442 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.675581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.675581Z digest=sha256:c3f2f943ff03936a35b8bb7101276a7bca6e881e130a20ef17111166e0763b80

Observation f136592c-ea7d-406d-809b-5d5afe95275f · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.685044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.685044Z digest=sha256:3be11129deed24c638c978d267718bdbb3f38727dea10b0f1a1155bee3a72c01

Observation 06fae0db-32da-468c-a6bd-38a1976d1f43 · outbound

This paper cites Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.430891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.697307Z digest=sha256:62992bb77a66d172dcfcc9baade28022666f641c8752b67dbdd6d7b9b53a9053

Observation 72bfeceb-e09b-43f6-b98c-ac53ebc28a20 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions In- structpix2pix: Learning to follow image editing instructions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.101887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.704361Z digest=sha256:c46d93f862e3ce4896ef792670ff1b9a00ac71c26615bd4b79f7016df7afecb0

Observation 38ec8158-98fa-4360-8dc1-1a764f8a2779 · outbound

This paper cites Video generation models as world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video generation models as world simulators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.709987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.709987Z digest=sha256:c04cacd74a36b46f8219cdf0b51f14f94caee4ce49338ca47cbd1c15baac2fd6

Observation fb6aa962-a943-45e2-9180-a2aec0997299 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.714910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.714910Z digest=sha256:ae79f13f92e3534f4124ed357cf1312fb7f3398cd48f232d09c5bb2c1408ae09

Observation 9f601e5c-8a61-48e9-9bff-bd37a1e99ac1 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.724998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.724998Z digest=sha256:79f9d40bec673cfe72a564460f0dfb074170c96734123c6d751e8f5886827a33

Observation 3be16d03-b94e-4d79-aace-587c08d7e986 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.730901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.730901Z digest=sha256:074166991369532d826ea8693368d037b5e03497137d2802e394080bf5e8faf6

Observation de56ccef-0ea9-4e96-a12d-695c9042cc4d · outbound

This paper cites Palm: Scaling language modeling with pathways.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Palm: Scaling language modeling with pathways

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.058051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.735941Z digest=sha256:8432e4f07db3ee50769e6bf09228a385da5c54462a97e7e5018c0c6837907b4a

Observation d8a4ef4d-ce20-4221-a470-e5b41932c615 · outbound

This paper cites Learning universal policies via text-guided video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning universal policies via text-guided video generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.742913Z digest=sha256:04218b76cbdda4defe23951a86bcda6b7895192b466851cb599b042e0280cff2

Observation 63ad7054-90a2-4ab5-abfb-c809dea7d116 · outbound

This paper cites The Llama 3 Herd of Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.747773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.747773Z digest=sha256:047ce14f309e837a166539e9540ebad468eb61e3412db42b26bad1b0baac8e74

Observation efc81969-1507-4b68-bb28-67ee39e1c0c9 · outbound

This paper cites Step- former: Self-supervised step discovery and localization in instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Step- former: Self-supervised step discovery and localization in instructional videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.028020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.753963Z digest=sha256:ddf0a852ebd25c4a7c53509d7b87f8b4cabc4cf8b6f53ac085dc55ba14064a00

Observation f12d4f78-d0a0-4581-a0f7-9b5029f6f0f1 · outbound

This paper cites Data Filtering Networks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Data Filtering Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.758855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.758855Z digest=sha256:5a32f1418199e0388e8191f339045f3d5a2ad4ae09937ad785256875bccf4bdb

Observation 2df30a7a-d2c4-4f11-bc5e-911df71131bd · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Preserve your own correlation: A noise prior for video diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.763399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.763399Z digest=sha256:3722e8e1c4df4442f3571821854d38a69c094aef8d800c24744d4ff04aa48204

Observation 71c850b4-401d-4bf5-9ee1-348ffdf0f099 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.767955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.767955Z digest=sha256:b6ed3be74f191567e96f2dcd7737320f7a5b138a4747abcd41a56944f2e9f011

Observation 087efbe0-b9a4-4ab9-9432-f7c3f38a3c80 · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Photorealistic Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.772500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.772500Z digest=sha256:d8abd5ec3338dfcbaee7d67e45aef2706b1dcbfcb4d7b61d21944659e56f4ab7

Observation 01f158c6-5bd0-4186-8ecd-12540da6f884 · outbound

This paper cites Temporal alignment networks for long-term video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Temporal alignment networks for long-term video

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.001952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.778529Z digest=sha256:9ab0ff89e678e9734e5a11efb0c411c8b2899f3262fe166fa590758e0f056060

Observation d77c6347-6bce-4269-bd9d-e48011e07a49 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.783762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.783762Z digest=sha256:47fb75a0f57af54ed6ac9984d5888433e30fd881d98548757ac661415def0be6

Observation 2b027e12-2a04-4346-9ad4-79bb40dc6edc · outbound

This paper cites Denoising diffu- sion probabilistic models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Denoising diffu- sion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.789267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.789267Z digest=sha256:868c5cc423d04fa3574241e6abe6980ad6da4155f04afe7508976a102f9171ab

Observation 3abf3e7a-6051-4e66-9f39-dc79b95c8b6b · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Imagen Video: High Definition Video Generation with Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.793860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.793860Z digest=sha256:9e2b2db1486ac058a42b873641e93015858eaf0abb8f5d56d2fcef006b3aa4d7

Observation 26c9e021-b692-4deb-88b7-002fa62d32bd · outbound

This paper cites Video dif- fusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video dif- fusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.800354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.800354Z digest=sha256:87eb4025c1f35cd431bbd81f891b1151c814b58d49c617c84912a692c79a4c4d

Observation b396e86c-3622-448c-8055-ac46ab252cf8 · outbound

This paper cites Make it move: controllable image-to-video generation with text de- scriptions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Make it move: controllable image-to-video generation with text de- scriptions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.969266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.806210Z digest=sha256:b89c01a9396778a1795ca404fe0094d5578c389df7fb86b763efdc0ebfbbd0de

Observation 9b7a549b-ad21-4a77-8b66-2143d34c7663 · outbound

This paper cites An edit friendly ddpm noise space: Inversion and manipulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions An edit friendly ddpm noise space: Inversion and manipulations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.953814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.811584Z digest=sha256:0385df5f35b04877163b811c5ef0ab2d2409701a61db4af443fe7c03fa2911e0

Observation 7324c0f5-d6b0-43d8-ade9-fd63baecc2ec · outbound

This paper cites Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.817423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.817423Z digest=sha256:1c4d5893b4ccbee259049c9c5951bef57cdfce05393f06d42fea42f4f087a8fc

Observation 55037516-e7ce-47bd-bed1-a644e7532620 · outbound

This paper cites Large language models are zero-shot reasoners.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Large language models are zero-shot reasoners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.940558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.828636Z digest=sha256:bad01f0b8985dd92092e5b50c65db65fb3cc93033125b6bc078eb76a88251883

Observation 04de4a20-014e-45dd-9ebe-c5fb5bb9f704 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.833207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.833207Z digest=sha256:7e8350d7101654e3df35a45fa71fbd4add400b9f2b43959a3801ce56d8a22dc6

Observation 784a9e7e-f9ad-483e-b297-44665d9fe130 · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.838693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.838693Z digest=sha256:4ff2447cfecb2fa5f833776ab6490e2d2b1fed8a36fa1ca7860c44bfd0d742c3

Observation 97fb824f-cc38-46a4-880d-f76455855370 · outbound

This paper cites LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.843887Z digest=sha256:e54fef5bf3d7fe172a273d52aa659887d58a3f9fb1288abc4d5e09fb602900c8

Observation 6cf761fe-5126-447f-9cd1-21b14f35bc65 · outbound

This paper cites Multi-sentence grounding for long- term instructional video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-sentence grounding for long- term instructional video

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.926952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.848623Z digest=sha256:ea1e2d324bae25c1534ed2525ce21191e57b189d721e59d065c6018f49f56835

Observation e15e0d37-d20b-425f-9fbd-a0e04fa81b58 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.852974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.852974Z digest=sha256:d5160990a568126764ee31fef585753bf4cfb6655d3aeac3b8644a5adb741f5a

Observation 2e6c8c9a-86e6-4774-bc4c-faed00a04766 · outbound

This paper cites Learning to ground instructional articles in videos through narrations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning to ground instructional articles in videos through narrations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.913479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.857693Z digest=sha256:40a5d00274763226bb7e41bbec7a3c8c7580944a3ae2f98ba78d6eace9d6b6ae

Observation d71a7704-a7bb-4419-9c33-a5e13bd09666 · outbound

This paper cites Vidm: Video implicit diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Vidm: Video implicit diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.899081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.861703Z digest=sha256:cdcea0695d85d403bb690b4f10b1ad181bbbaa7cade7c30ce6b9534d595b5e83

Observation b36f24c2-8541-4af3-847b-8808db6542bb · outbound

This paper cites Generating illustrated instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating illustrated instructions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.886483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.865718Z digest=sha256:1895664582363227861a6c4df389f4defa33b597f76e7d2f97d28058ee3ea440

Observation 3997a520-f7c9-4b35-8c17-146d1d48f8a8 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.873064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.884360Z digest=sha256:2e65dfed5355937f605f7f28eac9f8b406cbf3863db0efd442bf74e61f7b9997

Observation 75ab12e5-d43e-4af6-bc3b-3f0068300426 · outbound

This paper cites Visual reinforcement learn- ing with imagined goals.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual reinforcement learn- ing with imagined goals

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.858339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.889614Z digest=sha256:601b8c41b89ded8f1cfb543e640b7002bd74318da9adefbe1e84133dcb5bd6cc

Observation 3159ebe7-26d7-4e29-a8c7-7056b9841d72 · outbound

This paper cites Dinov2: Learning robust visual features without su- pervision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dinov2: Learning robust visual features without su- pervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.844241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.894028Z digest=sha256:651ea57f6011bbceaa2db9152472f38a1a23d3cdca0fe42aa14c90bf724ca190

Observation 91487550-b785-4f9b-98c3-6241dcfd339d · outbound

This paper cites Coherent Zero-Shot Visual Instruction Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coherent Zero-Shot Visual Instruction Generation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.220057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.900527Z digest=sha256:637f7cf8518163d60c64b226da5ddb45d5e1fd080ea9e2e5b3617580ee9f254f

Observation 60140c13-e07b-4231-85d3-a63fc6abda67 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Movie Gen: A Cast of Media Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.907577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.907577Z digest=sha256:9c93812b63455bb4021f357bbf35e199d4cebb7a31741694d40c92913d784054

Observation 4619a875-76d8-4c97-b9cc-687d634fb7a6 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing transferable visual models from natural language super- vision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.829895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.913694Z digest=sha256:878bad38d893db2a1e504ddbd3afd06237cb07c49e4d00565c7befa7bc8003f9

Observation e2eb7c0b-4895-4727-8a31-4d58a0bef8ac · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions High-resolution image syn- thesis with latent diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.815433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.922999Z digest=sha256:703ed7276bf22cab8c4b4d88ac1f57604c90ac7b1e0669e303cbb3d100812aa4

Observation ff5beb7d-edc1-49b3-b812-878915810340 · outbound

This paper cites Gen-3 alpha.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen-3 alpha

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.800306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.929037Z digest=sha256:4545e39e48b1d4ea5be2b690492c5d25d2838d4a67df251b0a07f3bc06944224

Observation abd3eec9-2225-48e8-a13b-a20ccc259fa1 · outbound

This paper cites Howtocap- tion: Prompting llms to transform video annotations at scale.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howtocap- tion: Prompting llms to transform video annotations at scale

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.786673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.935087Z digest=sha256:0e99d6bc08872ea02fd2c8acae9fd6dd9f3b8358f56e8394387262911dafbec3

Observation d89dc2ff-12b3-47c0-9b6e-d4c0834ef508 · outbound

This paper cites Ego4d goal-step: Toward hierarchical understanding of procedural activities.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ego4d goal-step: Toward hierarchical understanding of procedural activities

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.772416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.941594Z digest=sha256:c8ba7b40f47d2edc62f9e3e242080a0424134d5f3f1d92f4a497cdc4e517b905

Observation f4c24efa-4fec-44fd-8d0f-eb90fd4941d9 · outbound

This paper cites Multi-task learning of object states and state-modifying actions from web videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-task learning of object states and state-modifying actions from web videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.755919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.947814Z digest=sha256:6f08d66f5c74e1f187b86e268139c75458ff5c52b717476499e0ed491862e8ec

Observation 91449512-da44-444c-9e3c-6da5db8539b9 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.954781Z digest=sha256:57b21b3b30967a32a77c6efa91324a76b2aa3fbcdfb05052d259a12e708c4a68

Observation a16c4a41-8d5a-4f5c-8f2b-4e07dd8542a1 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.707756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.967511Z digest=sha256:6cd8e2c37d7b33ea20547fe04118feb9a4bf7c62495cc2cb0900b234c099d240

Observation 23aa6fd0-2760-465d-912d-8f5aed6eb5e4 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocomposer: Compositional video synthesis with motion controllability

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.688504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.971997Z digest=sha256:cbbe651b080969f6737304de18bc68e7b34cdefe001e650f65be66ee73fa5425

Observation 3205a1c7-a4b1-4f96-8dd6-048908a636d5 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.978979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.978979Z digest=sha256:db964675ddca0517528d9a7bc24589f0dcdbb9264d98c1212ccd4acde167bdc1

Observation dd22c36a-6f25-4b61-93fb-ffc4df674270 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Flow as the Cross-Domain Manipulation Interface

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.983607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.983607Z digest=sha256:fdc1e93154b710fb4a482263336e7ea4e3e495d84d296e7bee1d41f02f1b0020

Observation ba9f1114-b91f-457a-b153-a3a3db55453b · outbound

This paper cites Learn- ing object state changes in videos: An open-world perspec- tive.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing object state changes in videos: An open-world perspec- tive

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.660696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.988887Z digest=sha256:4a4a051cf5debeda62a61c11ec96b1e3ea865912c4bbfdda2b237313b92695b2

Observation 1845b683-1324-44c6-bc0d-25ebc34d4a7d · outbound

This paper cites Unloc: A unified framework for video localization tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unloc: A unified framework for video localization tasks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.642664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.993253Z digest=sha256:48e16f8a59cd4a28d7ea54f8e751a8f6751fac81692cca44d0e037cbed5d7cc8

Observation a2b2db66-34d9-4f6f-b6e8-e32ddb6aa599 · outbound

This paper cites Dif- fusion probabilistic modeling for video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dif- fusion probabilistic modeling for video generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.624772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.997859Z digest=sha256:508ae159785055bdfaf4e2afafce36d4d3ed285a4b30df1de729191e82a618a2

Observation 3c955992-aea2-4564-a178-a26a2d839d06 · outbound

This paper cites Learning interactive real-world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning interactive real-world simulators

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.608529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:09.002378Z digest=sha256:668ed0bbeadc3327e2a1a19300c7a16830b446f389ded7548fa37c7ce8790c43

Observation bef5f001-9e54-445e-b330-9f7a4fb563b9 · outbound

This paper cites Visual goal-step inference using wikihow.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual goal-step inference using wikihow

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.592663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:09.006957Z digest=sha256:b1d0cda3c0878aa1854f459790369503f16593b9cd6e6140024210498f86b428

Observation c7c52e77-656e-4c1e-a25e-74deb2d6c8bf · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.011524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.011524Z digest=sha256:e099ab5c9312cf0a7ba5e4cec3ca6cab8e22a29dbd196c6e13823e82ac04a71b

Observation a711e4fc-19a1-400b-b73e-8c4cd7eef45c · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video probabilistic diffusion models in projected latent space

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.575661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:09.016495Z digest=sha256:9e7a870e36180fb1cd8925efd36609d98598e1d184fe1542488052f405d6695a

Observation c6ab4b64-56a5-4d2e-8b74-27373b76e153 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Scaling Robot Learning with Semantically Imagined Experience

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.020787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.020787Z digest=sha256:e6111c27afdd4599ea85e479554b81785bdabd8055dcfe226f04e6fcc9edc23a

Observation c37e19e9-9711-4bac-ab38-727aa59b8c3d · outbound

This paper cites Sigmoid loss for language image pre-training.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Sigmoid loss for language image pre-training

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.558855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:09.025635Z digest=sha256:c409df0f6edb334b53c79fdf76edeacacc87356610f6613419872c5a42aca5f4

Observation 80dc2485-ed8d-4389-a1e3-0753cf4468de · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.045663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.045663Z digest=sha256:11b86c95cffdb10f945d3f1f292bf33aa3a6a00d3c42d75256e6c7d29ffd101e

Observation 73f4869a-77b4-4c89-baac-b08d423997fd · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.050387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.050387Z digest=sha256:c42179577bfe46d9d8af0da575e7e85a791a2040fad5f6a34cd6e3bd52969e6b

Observation 85dda75d-30b6-47d7-ac36-27a93e83fce2 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.055356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.055356Z digest=sha256:29afb22c7e547bccf71b6babbd4aa1722f3c0d5c1a0c1c1627a677bb154bd052

Observation 10e75bc3-40ce-4afc-a621-b5d893492669 · outbound

This paper cites Put some aluminum foil in there.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Put some aluminum foil in there

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:04:09.531802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:09.060651Z digest=sha256:441bb810d9d727f726a6ebf4eecbf399184488e9cf97445277f5f746c2eab69b

Observation cef7d800-01de-4556-9f7b-c2b20a34323f · outbound

This paper cites an unresolved cited work.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:04:09.723153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:04:08.960237Z digest=sha256:444307fc99b04b56b1070d60b60b600e9b73eb85e9e2b5b3874227cee8d85318

Pith citing papers

Observation 273928a8-b33a-4626-bce9-409575dfca99 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.229135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:3f874fad40eeb436b3aac482a030703792b6fb955875fb9274c112d1e058dd12

Observation ab65c9fd-487b-4646-9eea-77765388e187 · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:25.991899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:27fbb592e0487c1eb207a0d569b1a8028f556a15f0881db9572c5b1831d42d5e