Pith. sign in

Paper Citation Record · LEDGER

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2412.01987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01987 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:04:09.060651Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-23T00:30:55.729900Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:32:18.226311Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbfc89c9-4c74-41e5-acbe-0ae8d7cc9fc5 · outbound

This paper cites GPT-4 Technical Report.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.631013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.631013Z digest=sha256:eb6db6f568116f4f61f1a44975ba3cb74b2f857d3db00462806b5dbf2fa3990d

Observation 0474e1f2-d6b2-4e26-a975-3bf24cdf8c0b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ht-step: Aligning instructional articles with how-to videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.169513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.639434Z digest=sha256:d1f30bb4fc3a469d3530b8ef7539d73288155ec1b335a546f79586e8c1e90db9

Observation 8f61ac1d-c42a-4270-a247-83f6c736d298 · outbound

This paper cites Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.154558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.644101Z digest=sha256:133484d8cf0b53423d1cb7ed6e7910fbfe241d5f517ac49fb8a0d4f63028cf9d

Observation 875724f0-374e-43f8-b6f3-ef6ab842f263 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.140176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.648560Z digest=sha256:9384bab92aa365ac82454bc7c1c65097e1f1d06e2f2965ddfb545aa518b7642a

Observation 8ab9f193-3f26-4f9d-a311-62485c7168ad · outbound

This paper cites Whisperx: Time-accurate speech transcription of long- form audio.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Whisperx: Time-accurate speech transcription of long- form audio

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.126539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.652991Z digest=sha256:7fbbccf15f3ffe12de9ba166c6d835b89e9c74f6951ae1a52953b100cb2a8964

Observation cb136022-b1d8-437c-a9af-85786097d1b9 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.657233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.657233Z digest=sha256:15a4f1b0baeadaef52db5bb707762f87f2e15d11dd7f38bdf274898ebc519c71

Observation 2228cb1a-90c3-4c2f-af15-114b78151792 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.661872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.661872Z digest=sha256:e75ca55b427f058107026284955e369baade4ba24a493a2987302119ed75b3c3

Observation 628aaaac-5f4e-46f2-ac6e-73cfb7c3be57 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.670093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.670093Z digest=sha256:a91976a481d7b832fd8258a07073bd27c1606b10b5f3918920d55717aabab336

Observation 746ffd17-70f9-4af0-a183-6a1a1a6b9442 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.675581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.675581Z digest=sha256:c3f2f943ff03936a35b8bb7101276a7bca6e881e130a20ef17111166e0763b80

Observation f136592c-ea7d-406d-809b-5d5afe95275f · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.685044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.685044Z digest=sha256:3be11129deed24c638c978d267718bdbb3f38727dea10b0f1a1155bee3a72c01

Observation 06fae0db-32da-468c-a6bd-38a1976d1f43 · outbound

This paper cites Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.430891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.697307Z digest=sha256:50292a4c7b4d93c45e67eea2f28152d926964e8ab146977144e3e8fdcf221e11

Observation 72bfeceb-e09b-43f6-b98c-ac53ebc28a20 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions In- structpix2pix: Learning to follow image editing instructions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.101887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.704361Z digest=sha256:35a6a56e5e56bcf8a18736ef8e89ce7d443d44ea201835cbdef3ceaab44ab9b7

Observation 38ec8158-98fa-4360-8dc1-1a764f8a2779 · outbound

This paper cites Video generation models as world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video generation models as world simulators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.709987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.709987Z digest=sha256:c04cacd74a36b46f8219cdf0b51f14f94caee4ce49338ca47cbd1c15baac2fd6

Observation fb6aa962-a943-45e2-9180-a2aec0997299 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.714910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.714910Z digest=sha256:ae79f13f92e3534f4124ed357cf1312fb7f3398cd48f232d09c5bb2c1408ae09

Observation 9f601e5c-8a61-48e9-9bff-bd37a1e99ac1 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.724998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.724998Z digest=sha256:79f9d40bec673cfe72a564460f0dfb074170c96734123c6d751e8f5886827a33

Observation 3be16d03-b94e-4d79-aace-587c08d7e986 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.730901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.730901Z digest=sha256:074166991369532d826ea8693368d037b5e03497137d2802e394080bf5e8faf6

Observation de56ccef-0ea9-4e96-a12d-695c9042cc4d · outbound

This paper cites Palm: Scaling language modeling with pathways.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Palm: Scaling language modeling with pathways

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.058051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.735941Z digest=sha256:fe6ae3a122c7d3900b49cb00f71bc847ccbcca03b4cb65b4ed2263b2ac9d95a0

Observation d8a4ef4d-ce20-4221-a470-e5b41932c615 · outbound

This paper cites Learning universal policies via text-guided video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning universal policies via text-guided video generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.742913Z digest=sha256:6164e5d78e058df648e20a08b615816efaf99304dc53bfc1531c3b844c099ad4

Observation 63ad7054-90a2-4ab5-abfb-c809dea7d116 · outbound

This paper cites The Llama 3 Herd of Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.747773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.747773Z digest=sha256:047ce14f309e837a166539e9540ebad468eb61e3412db42b26bad1b0baac8e74

Observation efc81969-1507-4b68-bb28-67ee39e1c0c9 · outbound

This paper cites Step- former: Self-supervised step discovery and localization in instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Step- former: Self-supervised step discovery and localization in instructional videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.028020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.753963Z digest=sha256:c08e524e3ba8e302c89ac13b291411c877105c10b1a00ed54335be153b73231f

Observation f12d4f78-d0a0-4581-a0f7-9b5029f6f0f1 · outbound

This paper cites Data Filtering Networks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Data Filtering Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.758855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.758855Z digest=sha256:5a32f1418199e0388e8191f339045f3d5a2ad4ae09937ad785256875bccf4bdb

Observation 2df30a7a-d2c4-4f11-bc5e-911df71131bd · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Preserve your own correlation: A noise prior for video diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.763399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.763399Z digest=sha256:3722e8e1c4df4442f3571821854d38a69c094aef8d800c24744d4ff04aa48204

Observation 71c850b4-401d-4bf5-9ee1-348ffdf0f099 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.767955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.767955Z digest=sha256:b6ed3be74f191567e96f2dcd7737320f7a5b138a4747abcd41a56944f2e9f011

Observation 087efbe0-b9a4-4ab9-9432-f7c3f38a3c80 · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Photorealistic Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.772500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.772500Z digest=sha256:d8abd5ec3338dfcbaee7d67e45aef2706b1dcbfcb4d7b61d21944659e56f4ab7

Observation 01f158c6-5bd0-4186-8ecd-12540da6f884 · outbound

This paper cites Temporal alignment networks for long-term video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Temporal alignment networks for long-term video

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.001952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.778529Z digest=sha256:1e79a8e155f81b724ab055ef38c0f88aa1970cd600a64eb01c7a4890b20ac8a7

Observation d77c6347-6bce-4269-bd9d-e48011e07a49 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.783762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.783762Z digest=sha256:47fb75a0f57af54ed6ac9984d5888433e30fd881d98548757ac661415def0be6

Observation 2b027e12-2a04-4346-9ad4-79bb40dc6edc · outbound

This paper cites Denoising diffu- sion probabilistic models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Denoising diffu- sion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.789267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.789267Z digest=sha256:868c5cc423d04fa3574241e6abe6980ad6da4155f04afe7508976a102f9171ab

Observation 3abf3e7a-6051-4e66-9f39-dc79b95c8b6b · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Imagen Video: High Definition Video Generation with Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.793860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.793860Z digest=sha256:9e2b2db1486ac058a42b873641e93015858eaf0abb8f5d56d2fcef006b3aa4d7

Observation 26c9e021-b692-4deb-88b7-002fa62d32bd · outbound

This paper cites Video dif- fusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video dif- fusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.800354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.800354Z digest=sha256:87eb4025c1f35cd431bbd81f891b1151c814b58d49c617c84912a692c79a4c4d

Observation b396e86c-3622-448c-8055-ac46ab252cf8 · outbound

This paper cites Make it move: controllable image-to-video generation with text de- scriptions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Make it move: controllable image-to-video generation with text de- scriptions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.969266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.806210Z digest=sha256:e2d88b1b23288fe204b37ee4660769b12b1d6825e9e5e23488077ad3dbd6ed7a

Observation 9b7a549b-ad21-4a77-8b66-2143d34c7663 · outbound

This paper cites An edit friendly ddpm noise space: Inversion and manipulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions An edit friendly ddpm noise space: Inversion and manipulations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.953814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.811584Z digest=sha256:aea4c144db4e7a00e7ad961ca92d9f45fe18ba7a24cfb0c90c681db9b0deb80e

Observation 7324c0f5-d6b0-43d8-ade9-fd63baecc2ec · outbound

This paper cites Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.817423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.817423Z digest=sha256:1c4d5893b4ccbee259049c9c5951bef57cdfce05393f06d42fea42f4f087a8fc

Observation 55037516-e7ce-47bd-bed1-a644e7532620 · outbound

This paper cites Large language models are zero-shot reasoners.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Large language models are zero-shot reasoners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.940558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.828636Z digest=sha256:7c35ec56d0a87c024a72118d50fb37da6039bea95097fc0e3ca3a6a27e6634ca

Observation 04de4a20-014e-45dd-9ebe-c5fb5bb9f704 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.833207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.833207Z digest=sha256:7e8350d7101654e3df35a45fa71fbd4add400b9f2b43959a3801ce56d8a22dc6

Observation 784a9e7e-f9ad-483e-b297-44665d9fe130 · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.838693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.838693Z digest=sha256:816dcf758f4688f985b3ab26bd81c67fb8dc7a3d1a7bf238abd900972bc9ab02

Observation 97fb824f-cc38-46a4-880d-f76455855370 · outbound

This paper cites LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.843887Z digest=sha256:e54fef5bf3d7fe172a273d52aa659887d58a3f9fb1288abc4d5e09fb602900c8

Observation 6cf761fe-5126-447f-9cd1-21b14f35bc65 · outbound

This paper cites Multi-sentence grounding for long- term instructional video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-sentence grounding for long- term instructional video

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.926952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.848623Z digest=sha256:565896fe3500d4df3bda105ec8a4d5f4ed73915354b6ceface3f1e71f024a982

Observation e15e0d37-d20b-425f-9fbd-a0e04fa81b58 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.852974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.852974Z digest=sha256:d5160990a568126764ee31fef585753bf4cfb6655d3aeac3b8644a5adb741f5a

Observation 2e6c8c9a-86e6-4774-bc4c-faed00a04766 · outbound

This paper cites Learning to ground instructional articles in videos through narrations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning to ground instructional articles in videos through narrations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.913479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.857693Z digest=sha256:f143ae1b695263878075aa9c098e39ffe4f14c8440e2b79ad47531613116eefa

Observation d71a7704-a7bb-4419-9c33-a5e13bd09666 · outbound

This paper cites Vidm: Video implicit diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Vidm: Video implicit diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.899081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.861703Z digest=sha256:19334bb8e10b7c8696664d050c64a789d4b23e270e5cc19e8b3cb152a9262828

Observation b36f24c2-8541-4af3-847b-8808db6542bb · outbound

This paper cites Generating illustrated instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating illustrated instructions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.886483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.865718Z digest=sha256:7ba9f73285f174a8fa1a6ec4f142427707211e1a94bc769b5a973092bd1876b2

Observation 3997a520-f7c9-4b35-8c17-146d1d48f8a8 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.873064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.884360Z digest=sha256:339c2e93ca7691fecdc26717dc503025cbd91d8ef2e82a2a71380f96e8204b5e

Observation 75ab12e5-d43e-4af6-bc3b-3f0068300426 · outbound

This paper cites Visual reinforcement learn- ing with imagined goals.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual reinforcement learn- ing with imagined goals

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.858339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.889614Z digest=sha256:05010fb5c39f0684535080618efd073c232a931bed3baee8e929014f73e3cbc3

Observation 3159ebe7-26d7-4e29-a8c7-7056b9841d72 · outbound

This paper cites Dinov2: Learning robust visual features without su- pervision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dinov2: Learning robust visual features without su- pervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.844241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.894028Z digest=sha256:b97cc347a5e3c2e08b3dca09e87b9175eca5cf5e628c9fd20de40cf1c8629a91

Observation 91487550-b785-4f9b-98c3-6241dcfd339d · outbound

This paper cites Coherent Zero-Shot Visual Instruction Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coherent Zero-Shot Visual Instruction Generation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.220057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.900527Z digest=sha256:9fe4adc23c37cdd1e908782394a930bd8586878ed774f76d0e27310efc324432

Observation 60140c13-e07b-4231-85d3-a63fc6abda67 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Movie Gen: A Cast of Media Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.907577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.907577Z digest=sha256:9c93812b63455bb4021f357bbf35e199d4cebb7a31741694d40c92913d784054

Observation 4619a875-76d8-4c97-b9cc-687d634fb7a6 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing transferable visual models from natural language super- vision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.829895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.913694Z digest=sha256:c40759dc659c7037bb6e618e2444fdf4c1dfde5a0b56186c763d30246ea86029

Observation e2eb7c0b-4895-4727-8a31-4d58a0bef8ac · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions High-resolution image syn- thesis with latent diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.815433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.922999Z digest=sha256:18fe172f17428e966bb5fc47f57982f1a8e497fd649188975fa96907a8f11645

Observation ff5beb7d-edc1-49b3-b812-878915810340 · outbound

This paper cites Gen-3 alpha.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen-3 alpha

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.800306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.929037Z digest=sha256:49b39beb58f9fdea0260865b67624b38c8fe794ce239e32dd4d236a9eebe09da

Observation abd3eec9-2225-48e8-a13b-a20ccc259fa1 · outbound

This paper cites Howtocap- tion: Prompting llms to transform video annotations at scale.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howtocap- tion: Prompting llms to transform video annotations at scale

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.786673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.935087Z digest=sha256:8f38b7313f4c7354d113127128b159ca266773a1cbbb3c364edbe16119f1af87

Observation d89dc2ff-12b3-47c0-9b6e-d4c0834ef508 · outbound

This paper cites Ego4d goal-step: Toward hierarchical understanding of procedural activities.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ego4d goal-step: Toward hierarchical understanding of procedural activities

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.772416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.941594Z digest=sha256:25786e7f6ea2ead296199af7328a918cfbaa28d7decd8aff5b540502498307d8

Observation f4c24efa-4fec-44fd-8d0f-eb90fd4941d9 · outbound

This paper cites Multi-task learning of object states and state-modifying actions from web videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-task learning of object states and state-modifying actions from web videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.755919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.947814Z digest=sha256:6c58cd05a63b4e8b1af8903707f0418be2a60c5f64b31a31d46163fb91f4782e

Observation 91449512-da44-444c-9e3c-6da5db8539b9 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.954781Z digest=sha256:6771e3cd460da910173021cb8710c12ebabcd7c43eb61df8be6e94f2063d5b98

Observation a16c4a41-8d5a-4f5c-8f2b-4e07dd8542a1 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.707756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.967511Z digest=sha256:66ce268dffd0c93660546d2c30e1a84707d64c7b06a7615b164332bd42980a88

Observation 23aa6fd0-2760-465d-912d-8f5aed6eb5e4 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocomposer: Compositional video synthesis with motion controllability

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.688504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.971997Z digest=sha256:0bee8ab9bb1804290d5d63d160650a62cfb99025e1b291663c10f283f19a41ff

Observation 3205a1c7-a4b1-4f96-8dd6-048908a636d5 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.978979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.978979Z digest=sha256:db964675ddca0517528d9a7bc24589f0dcdbb9264d98c1212ccd4acde167bdc1

Observation dd22c36a-6f25-4b61-93fb-ffc4df674270 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Flow as the Cross-Domain Manipulation Interface

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.983607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.983607Z digest=sha256:fdc1e93154b710fb4a482263336e7ea4e3e495d84d296e7bee1d41f02f1b0020

Observation ba9f1114-b91f-457a-b153-a3a3db55453b · outbound

This paper cites Learn- ing object state changes in videos: An open-world perspec- tive.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing object state changes in videos: An open-world perspec- tive

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.660696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.988887Z digest=sha256:e565793417c9be770c4b738fc95768e528f746ea854c4546b1f70021ef7e0b16

Observation 1845b683-1324-44c6-bc0d-25ebc34d4a7d · outbound

This paper cites Unloc: A unified framework for video localization tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unloc: A unified framework for video localization tasks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.642664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.993253Z digest=sha256:47858deb627ff728aa888a9a3c84f7464870b1edb485b9c52d062486ba803918

Observation a2b2db66-34d9-4f6f-b6e8-e32ddb6aa599 · outbound

This paper cites Dif- fusion probabilistic modeling for video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dif- fusion probabilistic modeling for video generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.624772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.997859Z digest=sha256:59084ad9436e0774552206f8aab8ce192f3c1f33618afc8ca04136b8aba07a36

Observation 3c955992-aea2-4564-a178-a26a2d839d06 · outbound

This paper cites Learning interactive real-world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning interactive real-world simulators

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.608529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:09.002378Z digest=sha256:fd465c6ad60aae7dc6d933e8e20833c70ac65899c6c9f9f544ba30cc1742b97e

Observation bef5f001-9e54-445e-b330-9f7a4fb563b9 · outbound

This paper cites Visual goal-step inference using wikihow.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual goal-step inference using wikihow

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.592663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:09.006957Z digest=sha256:3362524d69f20e22b316daf6c1abf12afba5994f4a8189a51e4dcee4737ce3e7

Observation c7c52e77-656e-4c1e-a25e-74deb2d6c8bf · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.011524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.011524Z digest=sha256:e099ab5c9312cf0a7ba5e4cec3ca6cab8e22a29dbd196c6e13823e82ac04a71b

Observation a711e4fc-19a1-400b-b73e-8c4cd7eef45c · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video probabilistic diffusion models in projected latent space

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.575661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:09.016495Z digest=sha256:ea110f3aad5fe9a7ba602efe27273abc56c8575ce438f90a561f3e8e424386a3

Observation c6ab4b64-56a5-4d2e-8b74-27373b76e153 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Scaling Robot Learning with Semantically Imagined Experience

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.020787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.020787Z digest=sha256:e6111c27afdd4599ea85e479554b81785bdabd8055dcfe226f04e6fcc9edc23a

Observation c37e19e9-9711-4bac-ab38-727aa59b8c3d · outbound

This paper cites Sigmoid loss for language image pre-training.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Sigmoid loss for language image pre-training

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.558855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:09.025635Z digest=sha256:4e09b234ec82d2688ca7646fda5bcacf93c25c7c89a8e7375c58fb31fcca0a94

Observation 80dc2485-ed8d-4389-a1e3-0753cf4468de · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.045663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.045663Z digest=sha256:11b86c95cffdb10f945d3f1f292bf33aa3a6a00d3c42d75256e6c7d29ffd101e

Observation 73f4869a-77b4-4c89-baac-b08d423997fd · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.050387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.050387Z digest=sha256:c42179577bfe46d9d8af0da575e7e85a791a2040fad5f6a34cd6e3bd52969e6b

Observation 85dda75d-30b6-47d7-ac36-27a93e83fce2 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.055356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.055356Z digest=sha256:29afb22c7e547bccf71b6babbd4aa1722f3c0d5c1a0c1c1627a677bb154bd052

Observation 10e75bc3-40ce-4afc-a621-b5d893492669 · outbound

This paper cites Put some aluminum foil in there.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Put some aluminum foil in there

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:04:09.531802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:09.060651Z digest=sha256:1be751c0a09f79fd321afee31066a862fed429a37cffdc9021f8452ffbd1df12

Observation cef7d800-01de-4556-9f7b-c2b20a34323f · outbound

This paper cites an unresolved cited work.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:04:09.723153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:04:08.960237Z digest=sha256:78df9c163c19cec29501e54281f2cb80b52bef8182e305328093ab19937453f2

Pith citing papers

Observation 273928a8-b33a-4626-bce9-409575dfca99 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.229135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:eb9aea981e7cf74c0e70ee2cf44baf740455c8c0148cad13083d13125b5473e4

Observation ab65c9fd-487b-4646-9eea-77765388e187 · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:25.991899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:a2d1d41125b129b3e3824c4ce89fa58bd7128bab58185a7b9b6949ccff14a70f