Pith. sign in

Paper Citation Record · LEDGER

Learning Human Skill Generators at Key-Step Levels

As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2502.08234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08234 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:57:41.808405Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b15a81f-0f91-46a7-98ce-ffb83ee5f24b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024.

Learning Human Skill Generators at Key-Step Levels Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.450088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.582945Z digest=sha256:afcf1e8a2ccfe31eebbce2d11a15e841144a81c70e5af40117a8574884bc464c

Observation b28e4f32-db7d-4f7a-b3a9-d2eb4aef30d7 · outbound

This paper cites Character region awareness for text detec- tion.

Learning Human Skill Generators at Key-Step Levels Character region awareness for text detec- tion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.440236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.587959Z digest=sha256:dbd096cddf63b9d12dd4c61e414ce4c554a78ccb0cd2c4b7bb2cd27d1b66231f

Observation aeac7dab-3087-4cf5-b4b9-10ad7510f05a · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

Learning Human Skill Generators at Key-Step Levels Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.592358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.592358Z digest=sha256:2045bdbab47ac6bf35cd22cf12ca8df7cf7068d0f1aa9377d1782ce58cc33ae5

Observation 4d18419e-e528-4449-a31e-e3c679f8c782 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Learning Human Skill Generators at Key-Step Levels Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.596848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.596848Z digest=sha256:3fabb0a6f17f1475a2ba2bef3185e20824af91ad4767d2c192fbeed79a541b0d

Observation 8142fb59-4c3a-426d-af2b-84b93943b4e2 · outbound

This paper cites Video generation models as world simulators.

Learning Human Skill Generators at Key-Step Levels Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.601151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.601151Z digest=sha256:d931b9beed5a1bf98b21aed7dcea0d5de5c2b327e02ef6f28d0d536959af38e5

Observation a031fa14-c750-4464-889c-4aa6b4090bfb · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Learning Human Skill Generators at Key-Step Levels Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.424727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.605099Z digest=sha256:93e49c0dede9c5e8144a79e2b417f4fe13b4ea6b8eca8bc187bcd7b28c101e51

Observation f9238a44-384b-42bb-bdda-b63a9b332c3a · outbound

This paper cites Procedure planning in in- 12 structional videos.

Learning Human Skill Generators at Key-Step Levels Procedure planning in in- 12 structional videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.414968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.609286Z digest=sha256:01565653d69b642b9668f8932df10e7bf874d186dff085bf0a574f1d6b9404e4

Observation 862b9e94-22ad-4014-ac79-6a043492211d · outbound

This paper cites V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset.

Learning Human Skill Generators at Key-Step Levels V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.404619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.613181Z digest=sha256:5e71ee4ef8a3031c12f73fce4d818949444018d49983a80a795f637c3c175ca4

Observation 6ccd12c6-14e7-48fb-b16c-615589d8bb65 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Learning Human Skill Generators at Key-Step Levels Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.394083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.616878Z digest=sha256:c9db43b99deb1ef782dacd7a26b460406e38e0575fd2a3aac6c93be3c0f176b3

Observation 0bcdc815-534b-429d-956a-eb1bff33fad6 · outbound

This paper cites Claude 3.5 sonnet.

Learning Human Skill Generators at Key-Step Levels Claude 3.5 sonnet

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.384134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.620810Z digest=sha256:4b1e3ddab8f7b00e7778c9a5dc393513d2e473eb056f4e1b02d0c404ee019542

Observation cf4c67af-e8a1-4625-84e8-6b747e2bef53 · outbound

This paper cites Animateanything: Fine- grained open domain image animation with motion guid- ance.

Learning Human Skill Generators at Key-Step Levels Animateanything: Fine- grained open domain image animation with motion guid- ance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.373948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.624605Z digest=sha256:50460ced9630a874fa972e2430d52b9540f177f85f21d1fda8f2b7fcc36f4d0b

Observation c3f30dec-35df-4485-90e2-84578e514b3b · outbound

This paper cites The EPIC-KITCHENS dataset: Collection, challenges and baselines.

Learning Human Skill Generators at Key-Step Levels The EPIC-KITCHENS dataset: Collection, challenges and baselines

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.362984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.628361Z digest=sha256:9e2875a79046f5933fba255db0b6dac7b5e7049708e55cf4823c5d00ba79388e

Observation b14b932e-06f4-42b7-8332-5745980dbeae · outbound

This paper cites Doell, and Jason J.

Learning Human Skill Generators at Key-Step Levels Doell, and Jason J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.352953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.632029Z digest=sha256:7c486abb958d83572f4aa48b66c4c708f2f3ea06f531c067632287c230eb7110

Observation 0b8563f8-1c56-43d5-bb27-2e7741c19b0a · outbound

This paper cites Video Language Planning.

Learning Human Skill Generators at Key-Step Levels Video Language Planning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.635876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.635876Z digest=sha256:b2af36167a0eab254c5128348a37bc741b68ac372ed7203a13383ca80e9daad9

Observation a6160406-5171-4117-ac40-9ff0e852b1f5 · outbound

This paper cites Learning universal policies via text-guided video generation.

Learning Human Skill Generators at Key-Step Levels Learning universal policies via text-guided video generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.343045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.639969Z digest=sha256:7402d98bf9d79d19aaa6a8ca2252b9fb74fdc6be4cce2d7e0a0ef5f6b94ea389

Observation 2d3d93ea-c6aa-46e8-9439-4ca85b970f95 · outbound

This paper cites The Llama 3 Herd of Models.

Learning Human Skill Generators at Key-Step Levels The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.643663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.643663Z digest=sha256:5b91be3e1b278e89fcf6f263afc96f11a556450ae9a2fee099acf0692b7d5d2d

Observation 4be22ee2-836a-4ea0-94fd-fcf691ac8854 · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense.

Learning Human Skill Generators at Key-Step Levels The ”something something” video database for learning and evaluating visual common sense

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.332513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.647707Z digest=sha256:5856e108de7df9b641165be01fd7d3a044e3339be26cd36534a34c429dd1b189

Observation fa9757cf-f239-4d0f-91a3-31f77e69b607 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Learning Human Skill Generators at Key-Step Levels AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.651277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.651277Z digest=sha256:ad9c34035f2a50d27f492433670455cf2a070aa9ad17745fda7ad69a9c3983d0

Observation 8249a066-287e-4a1d-b67c-0565564cdcb0 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Learning Human Skill Generators at Key-Step Levels Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.321383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.655476Z digest=sha256:1c16b478e383733248de31e178dab540f785f4aacfba1019cf33f364cf89d512

Observation 7072952b-5400-490e-8403-9efcd254d6be · outbound

This paper cites Denoising diffu- sion probabilistic models.

Learning Human Skill Generators at Key-Step Levels Denoising diffu- sion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.659027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.659027Z digest=sha256:7e255955dc39c60f6b54de1cf143ca5ff2ddab2f88dfe5e81f6f706521372a67

Observation 7cd2a483-4297-40d8-bdfd-453ee4a091bb · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

Learning Human Skill Generators at Key-Step Levels Vbench: Comprehensive bench- mark suite for video generative models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.303866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.662629Z digest=sha256:c4206eb27faf3f45ae317bce18989b8fcaa9be641ac449e2b0834a5da1767ec5

Observation 8c33e2b3-915b-4fde-bc74-803c9be71552 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Learning Human Skill Generators at Key-Step Levels The Kinetics Human Action Video Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.666094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.666094Z digest=sha256:0a81faf183330be405f2872a465347d527bd892fb510d702329218fb41309022

Observation cb93cf69-3c11-430b-a6e1-70b817733680 · outbound

This paper cites The language of actions: Recovering the syntax and semantics of goal-directed human activities.

Learning Human Skill Generators at Key-Step Levels The language of actions: Recovering the syntax and semantics of goal-directed human activities

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.292234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.670181Z digest=sha256:516a5c7584aa1c8fefcb94c5aa825832320fde0492aeebb3efe4adf6f5e4e7c3

Observation e376017a-7f18-470f-a3ca-878bde00487b · outbound

This paper cites an unresolved cited work.

Learning Human Skill Generators at Key-Step Levels Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:57:42.281239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.674076Z digest=sha256:aac535e1c815cee90d586aac04784c18ebcf6b04b871cd16a9a265177c64764f

Observation c3193487-7874-4e52-905b-2b352f1e574b · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Learning Human Skill Generators at Key-Step Levels Latte: Latent Diffusion Transformer for Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.677889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.677889Z digest=sha256:d23a305f7fd6407a97a7fa37cf72301a56a02b624d0cef5cb6ab9412f1464137

Observation b2521d54-76d4-40e6-8d87-4a66524998e9 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Learning Human Skill Generators at Key-Step Levels Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.271516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.681886Z digest=sha256:92b82dd6799fc5c80b07c8a1530696690faf4f1e72d985ae3e398e8e8c5d8e03

Observation fd6799bf-3171-4a07-9213-93720205aa0e · outbound

This paper cites Gpt-4o release.

Learning Human Skill Generators at Key-Step Levels Gpt-4o release

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.261132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.685785Z digest=sha256:c843cd15ecdf72aa96b57d5dc2a60c364d9745066906c1d35395e7aa5c7ab63e

Observation c7c78915-af01-431a-909e-74beb5cab700 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

Learning Human Skill Generators at Key-Step Levels Gpt-4o mini: advancing cost-efficient intelligence

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.250388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.689473Z digest=sha256:5fdfb9f3d66780768f5bfa88c2323e5ae5cd71539ad35a95dd4d01641099f75d

Observation 9cf035b1-497b-4340-8e04-a50ab56bad56 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Learning Human Skill Generators at Key-Step Levels DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.693527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.693527Z digest=sha256:3c2e45c95ae57391e9ac294ade1e334beb47a01157e3241a46a3ec2e106032c7

Observation fc634a86-c32d-4597-bf55-c2c41e6e6cbd · outbound

This paper cites Scalable diffusion models with transformers.

Learning Human Skill Generators at Key-Step Levels Scalable diffusion models with transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.697503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.697503Z digest=sha256:77f9ea2eb16ce4f083402a1300dd282c728360ec6d1dd34e8928c81eccde95c9

Observation 5a0ac98a-0486-42db-af27-bfafb4e6f147 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Learning Human Skill Generators at Key-Step Levels SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.701183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.701183Z digest=sha256:d1293fd27fe8c87b7c0b99c2d356e5eb20eb981665e098e057ac6a1d0a856bec

Observation 78bc2d0a-a0ab-4266-8b2b-0bd9b7cbf693 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Learning Human Skill Generators at Key-Step Levels Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.705234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.705234Z digest=sha256:dcf741fa22ebe7d801239dd2eaac20d0f499aeabca616f711a656e5c7d503538

Observation 93597df1-d0db-4a22-8bd1-e615fd138797 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Learning Human Skill Generators at Key-Step Levels Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.708915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.708915Z digest=sha256:033862e1963eca0eec59a765df33f57e212cfb8326c37a61c9ca6d7f083230ae

Observation 9893da96-cba5-4395-8184-a28e4f6384e3 · outbound

This paper cites Skills, rules, and knowledge; signals, signs, and symbols, and other distinctions in human performance models.

Learning Human Skill Generators at Key-Step Levels Skills, rules, and knowledge; signals, signs, and symbols, and other distinctions in human performance models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.226149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.712854Z digest=sha256:c3e16f4e06f6a750489e43845d2a958cf1fb5eb7a3b2df08d6d0cd1fa26a9c84

Observation a59097f5-d700-4d84-80af-f497a6c64f2f · outbound

This paper cites A database for fine grained activity detec- tion of cooking activities.

Learning Human Skill Generators at Key-Step Levels A database for fine grained activity detec- tion of cooking activities

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.214252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.716913Z digest=sha256:bf6d9ce29607c27b0202a80417f5ccf83bfc3c19a45f09473de36c48fb5ea3d8

Observation 608336e7-a93f-4c8b-bc62-692ec82d47c6 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning Human Skill Generators at Key-Step Levels High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.720611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.720611Z digest=sha256:419ae2cb2f3733759c4f713a37f6b51f0e97448aea82d0e6439ff473fb04d08a

Observation 5f035334-60cc-4b52-8741-0381fff9c689 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning Human Skill Generators at Key-Step Levels High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.724335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.724335Z digest=sha256:d7741ec679a4ef178a1dc7ab0546ff38f8ec10d0e3c679909d4303528c867216

Observation d675b854-7876-4acc-a12a-a65c2c856792 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

Learning Human Skill Generators at Key-Step Levels As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.186665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.728176Z digest=sha256:be7d1909686dc62e94e39a06223a05b961f5be934e034d05f3bda7822967cadb

Observation c66fd543-5930-475c-826d-a41ab152b3f8 · outbound

This paper cites Tulyakov, and Mohamed Elhoseiny.

Learning Human Skill Generators at Key-Step Levels Tulyakov, and Mohamed Elhoseiny

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.174198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.731981Z digest=sha256:2359a012e3d34deea01d49cf062e3ae31992992b4f0012f56ea885a0b3b40a04

Observation 54abcc50-6b32-4a62-a212-86e1de9dfeff · outbound

This paper cites Look for the change: Learning object states and state-modifying actions from untrimmed web videos.

Learning Human Skill Generators at Key-Step Levels Look for the change: Learning object states and state-modifying actions from untrimmed web videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.161314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.735568Z digest=sha256:f4315a2f717ed680383dbd76bfab9c8452e6fe6ca5abdba19f53f3c6e6065a9e

Observation c1cbc740-8d1f-4de4-ad37-4a08d806a359 · outbound

This paper cites GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos.

Learning Human Skill Generators at Key-Step Levels GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:57:41.894755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.739141Z digest=sha256:e1b0fcc060fb4844360f435e43789db6f483c594c78d1abf675eedbdee23691a

Observation 72bd222b-d278-46d2-b63a-19fc12bb0af9 · outbound

This paper cites an unresolved cited work.

Learning Human Skill Generators at Key-Step Levels Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:57:42.149776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.742915Z digest=sha256:b103800ede6c5893c2cb72458d6ff24cf872167888b54b5f70c594d0f6ba3923

Observation 25f80926-b6c1-4afc-94c0-77eadcb5d26f · outbound

This paper cites Plate: Visually-grounded plan- ning with transformers in procedural tasks.

Learning Human Skill Generators at Key-Step Levels Plate: Visually-grounded plan- ning with transformers in procedural tasks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.139262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.746536Z digest=sha256:40012acfb1f4c3a1df693efacf617adc17493fc93be39a4080ae6896c444584d

Observation 4360b4cb-7909-4f79-b17a-d29e15164814 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Learning Human Skill Generators at Key-Step Levels EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.749692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.749692Z digest=sha256:8ff1a622185879115da27e87a84728a0b15061ace33413baf29049cd3eade3ca

Observation ac8de97c-418b-430b-b092-0926d9099a4d · outbound

This paper cites A comprehensive survey of procedural video datasets.

Learning Human Skill Generators at Key-Step Levels A comprehensive survey of procedural video datasets

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.127835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.753580Z digest=sha256:3aecf91d2260b21175df675355e57b378898b35499fd2ec357b8a1414cdfd6f3

Observation ff684aa7-86e8-4264-893d-24f821bd067b · outbound

This paper cites COIN: A large-scale dataset for comprehensive instructional video analysis.

Learning Human Skill Generators at Key-Step Levels COIN: A large-scale dataset for comprehensive instructional video analysis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.117106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.756920Z digest=sha256:6da9565a637085874f33ee4488a0682b74b1e06a07dccd2686916e04ae4d2c88

Observation a756a502-1e39-4bb7-9d06-b7c3d29fb6fb · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

Learning Human Skill Generators at Key-Step Levels Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.760392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.760392Z digest=sha256:ad9591f1ff6cc678a805c1eaa6f36311767e6101a4fb234ee69a6efb5441b8be

Observation cc08896b-20f1-41f3-ab43-2aae1aba9e4b · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Learning Human Skill Generators at Key-Step Levels Raft: Recurrent all-pairs field transforms for optical flow

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.763710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.763710Z digest=sha256:7b3fa5ecfd4ba73d1caffc0b0b1e590be26d42e5a33672fab1a5201fff9b96f6

Observation 183fe31c-5cbb-41fe-9458-c97d75655936 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Learning Human Skill Generators at Key-Step Levels Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.089029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.767223Z digest=sha256:5b9063dc7c2b3d7ed411da2ffe82cf87120b82688da180a0d4ecfcf313e62ec7

Observation dccf7034-e184-4de5-9156-2f4c2c9647d5 · outbound

This paper cites FVD: A new metric for video generation.

Learning Human Skill Generators at Key-Step Levels FVD: A new metric for video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.076814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.770757Z digest=sha256:9dcd51cda1966d39f4724b0824764f0f714a75f6f070ed024afd959c519f4152

Observation 05673831-b1ff-4c37-bb59-0f5778a4b0b2 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

Learning Human Skill Generators at Key-Step Levels Event-guided procedure planning from in- structional videos with text supervision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.064809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.774451Z digest=sha256:ba8764804db88ba80095638b1305070428e92a240f6c75d097113864027d4f0d

Observation 38c10197-f20b-4356-933c-132a379a5dd5 · outbound

This paper cites PDPP: projected diffusion for procedure planning in instructional videos.

Learning Human Skill Generators at Key-Step Levels PDPP: projected diffusion for procedure planning in instructional videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.052088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.777811Z digest=sha256:583e1bdd0999eda4e9e823233fb7b4928d25f36fe7f2d240eed4fc1280f94aaa

Observation 88bb9576-8025-4cc8-8808-b03d213121ff · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

Learning Human Skill Generators at Key-Step Levels Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.781154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.781154Z digest=sha256:e99865cc5a0fe19e788bcf1f23ea163f5d90c14c6cbaac4cc70062b00af3df0f

Observation 5f3a08e6-a538-4f44-9983-8f1710ca49e9 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

Learning Human Skill Generators at Key-Step Levels Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.041272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.784886Z digest=sha256:962a848262665857a83ab325e43a6f18da9632f43ad444c414d4073654fb4637

Observation 2d8b7ea6-9b50-416e-a032-a55cceb2672f · outbound

This paper cites Learning Interactive Real-World Simulators.

Learning Human Skill Generators at Key-Step Levels Learning Interactive Real-World Simulators

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.788494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.788494Z digest=sha256:dbb3afba588bce56389fb896247b430a41e7961db53d0cf7bd651c6b2e7f2142

Observation 0c749702-373f-475b-aef7-0596903adbdc · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Learning Human Skill Generators at Key-Step Levels CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.792380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.792380Z digest=sha256:b335c544871bd1d4676ba537550621f6f3ccfc4819449f46410a0c959a9f0e1f

Observation c36ef58e-a876-4025-ad90-5cb9765f8d34 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Learning Human Skill Generators at Key-Step Levels IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.796675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.796675Z digest=sha256:da280cfe161a25113f347f790eff50bff6453abd61c9e539734aa4009da5efa4

Observation 2dd1794c-ee9e-4dad-943d-9887aa7b1de3 · outbound

This paper cites Der- panis, Richard P.

Learning Human Skill Generators at Key-Step Levels Der- panis, Richard P

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.030861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.800838Z digest=sha256:ca99dda662b2ccea8dc7aa899c76b10d4e750e00336ff88ae4b1686d5309c3f0

Observation d4a441bc-6f7f-4f98-9e02-644a6ec06356 · outbound

This paper cites an unresolved cited work.

Learning Human Skill Generators at Key-Step Levels Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:57:42.017712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.804814Z digest=sha256:72d07be126f01f17c8c0a347edcce6ad07582cb92736303caffbcb9e30019ede

Observation a15cf893-8100-4a22-913f-e85defcf0df0 · outbound

This paper cites Fouhey, Ivan Laptev, and Josef Sivic.

Learning Human Skill Generators at Key-Step Levels Fouhey, Ivan Laptev, and Josef Sivic

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.004808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:57:41.808405Z digest=sha256:a1cfc44d9b0b2f5b0d1ed99ca57bebc3e292e77d011151850a2fbc2df8afd14d

Pith citing papers

No inbound Pith citation observations are available.