Pith. sign in

Paper Citation Record · LEDGER

Learning Human Skill Generators at Key-Step Levels

As of 20 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2502.08234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08234 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:57:41.808405Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b15a81f-0f91-46a7-98ce-ffb83ee5f24b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024.

Learning Human Skill Generators at Key-Step Levels Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.450088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.582945Z digest=sha256:ea5fa335ddf7efff1801fffcf274e29806fc7014a36d3ab849421723ce423dc1

Observation b28e4f32-db7d-4f7a-b3a9-d2eb4aef30d7 · outbound

This paper cites Character region awareness for text detec- tion.

Learning Human Skill Generators at Key-Step Levels Character region awareness for text detec- tion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.440236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.587959Z digest=sha256:885fa183e7d1c47c636730bc636bd465affe637b62bb17fa21f8c09f84901b1b

Observation aeac7dab-3087-4cf5-b4b9-10ad7510f05a · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

Learning Human Skill Generators at Key-Step Levels Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.592358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.592358Z digest=sha256:96b2a118dcc508102898374af8bcdfa62b3e6b30ec743fb6096e83ae7c33d334

Observation 4d18419e-e528-4449-a31e-e3c679f8c782 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Learning Human Skill Generators at Key-Step Levels Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.596848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.596848Z digest=sha256:e052f9cfcc21d28433fdb3a40bd70ba86c51176773821d88b404affe5739c4af

Observation 8142fb59-4c3a-426d-af2b-84b93943b4e2 · outbound

This paper cites Video generation models as world simulators.

Learning Human Skill Generators at Key-Step Levels Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.601151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.601151Z digest=sha256:08a8923dfd2c9dfdc2dc8ab73363aeceee9e0f61d7b1c707f0e12e3ec7593c36

Observation a031fa14-c750-4464-889c-4aa6b4090bfb · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Learning Human Skill Generators at Key-Step Levels Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.424727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.605099Z digest=sha256:d6b057a0cff873b16408cdeabf336a534fba4bfd9134413f9a506075e7f4e6a1

Observation f9238a44-384b-42bb-bdda-b63a9b332c3a · outbound

This paper cites Procedure planning in in- 12 structional videos.

Learning Human Skill Generators at Key-Step Levels Procedure planning in in- 12 structional videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.414968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.609286Z digest=sha256:6d0810b1a7f3b92cd34ac07be4d41e88ffe53dee82c07c2bfa2e76189ac8328f

Observation 862b9e94-22ad-4014-ac79-6a043492211d · outbound

This paper cites V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset.

Learning Human Skill Generators at Key-Step Levels V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.404619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.613181Z digest=sha256:345a3125cc5c9e06e836c19a06a82f06cb64b7881b0feecc6ea5ca9b7a84dbb4

Observation 6ccd12c6-14e7-48fb-b16c-615589d8bb65 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Learning Human Skill Generators at Key-Step Levels Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.394083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.616878Z digest=sha256:db9221230463afbcb45ebbee0f06202fbd4c0a24500165b989e7d54ad1f3a229

Observation 0bcdc815-534b-429d-956a-eb1bff33fad6 · outbound

This paper cites Claude 3.5 sonnet.

Learning Human Skill Generators at Key-Step Levels Claude 3.5 sonnet

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.384134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.620810Z digest=sha256:6f577585526293f01deb9dffb3e3ad50438dd3b3b13cdc923e962a21688ba058

Observation cf4c67af-e8a1-4625-84e8-6b747e2bef53 · outbound

This paper cites Animateanything: Fine- grained open domain image animation with motion guid- ance.

Learning Human Skill Generators at Key-Step Levels Animateanything: Fine- grained open domain image animation with motion guid- ance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.373948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.624605Z digest=sha256:d0d8be242475c11e73d99a0f5c569801f253fe9f17910e62192a08eb8b12d4d2

Observation c3f30dec-35df-4485-90e2-84578e514b3b · outbound

This paper cites The EPIC-KITCHENS dataset: Collection, challenges and baselines.

Learning Human Skill Generators at Key-Step Levels The EPIC-KITCHENS dataset: Collection, challenges and baselines

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.362984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.628361Z digest=sha256:aff72cd9635c5a50ebfc8108125e3fbfc15ceba4486da76f0c4956de69553dc3

Observation b14b932e-06f4-42b7-8332-5745980dbeae · outbound

This paper cites Doell, and Jason J.

Learning Human Skill Generators at Key-Step Levels Doell, and Jason J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.352953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.632029Z digest=sha256:316171775f93540c2345b60b93355b47f536b4028263b3142aa1b0d7a83969f9

Observation 0b8563f8-1c56-43d5-bb27-2e7741c19b0a · outbound

This paper cites Video Language Planning.

Learning Human Skill Generators at Key-Step Levels Video Language Planning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.635876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.635876Z digest=sha256:a47233bb3bdc22cb12f45e530c674a2e61af93338c4a64e5ebc86259e596c702

Observation a6160406-5171-4117-ac40-9ff0e852b1f5 · outbound

This paper cites Learning universal policies via text-guided video generation.

Learning Human Skill Generators at Key-Step Levels Learning universal policies via text-guided video generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.343045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.639969Z digest=sha256:c5151ddbdd50653e3ee2e43fbe8a19ffe72bd1f17e0b2ec9a0aa014ee9ff3fa4

Observation 2d3d93ea-c6aa-46e8-9439-4ca85b970f95 · outbound

This paper cites The Llama 3 Herd of Models.

Learning Human Skill Generators at Key-Step Levels The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.643663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.643663Z digest=sha256:a51f0a769620541dc9e5cb7f096f31cf85cb5b8421758684863f24dbfcc632aa

Observation 4be22ee2-836a-4ea0-94fd-fcf691ac8854 · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense.

Learning Human Skill Generators at Key-Step Levels The ”something something” video database for learning and evaluating visual common sense

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.332513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.647707Z digest=sha256:8c33553fc979fbeb0ea8df3fe1fda7de58314d2c5d1501432d50417fab60eb4d

Observation fa9757cf-f239-4d0f-91a3-31f77e69b607 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Learning Human Skill Generators at Key-Step Levels AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.651277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.651277Z digest=sha256:8b5454478e91b7d65f7aa3aaa44b6b01026fc695f4e9c570637c03b777a874ba

Observation 8249a066-287e-4a1d-b67c-0565564cdcb0 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Learning Human Skill Generators at Key-Step Levels Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.321383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.655476Z digest=sha256:67974aa1b5d0c72ec50623fa4f8a41c81aba2cd484c4e64048790af3f3ba2fba

Observation 7072952b-5400-490e-8403-9efcd254d6be · outbound

This paper cites Denoising diffu- sion probabilistic models.

Learning Human Skill Generators at Key-Step Levels Denoising diffu- sion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.659027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.659027Z digest=sha256:ae54823623c07bb7a55d7d7075bb2f8f560d5aa85224bb51710b0124ff717434

Observation 7cd2a483-4297-40d8-bdfd-453ee4a091bb · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

Learning Human Skill Generators at Key-Step Levels Vbench: Comprehensive bench- mark suite for video generative models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.303866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.662629Z digest=sha256:e2178468c277baecdd3272fd91770b4a1062f80b9e7b30458f1d742f5dc8ade8

Observation 8c33e2b3-915b-4fde-bc74-803c9be71552 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Learning Human Skill Generators at Key-Step Levels The Kinetics Human Action Video Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.666094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.666094Z digest=sha256:4495b2e922e5cf536cf9132e0bbd0c17a59997904de7f142b56530d2e6feb5c9

Observation cb93cf69-3c11-430b-a6e1-70b817733680 · outbound

This paper cites The language of actions: Recovering the syntax and semantics of goal-directed human activities.

Learning Human Skill Generators at Key-Step Levels The language of actions: Recovering the syntax and semantics of goal-directed human activities

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.292234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.670181Z digest=sha256:abf079ed37086a48961c8db5b62b5bb18648181809e0c1f16760723177027ae9

Observation e376017a-7f18-470f-a3ca-878bde00487b · outbound

This paper cites an unresolved cited work.

Learning Human Skill Generators at Key-Step Levels Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:57:42.281239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.674076Z digest=sha256:e0b862ee5ebb878e7743e66fc83f1f7ed5b9d9610261c102ea841530e4e0519d

Observation c3193487-7874-4e52-905b-2b352f1e574b · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Learning Human Skill Generators at Key-Step Levels Latte: Latent Diffusion Transformer for Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.677889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.677889Z digest=sha256:c0ca9b58e4dd28f08679c0e2491d23cdb8e3618296ea711e2b070eb2445d46bb

Observation b2521d54-76d4-40e6-8d87-4a66524998e9 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Learning Human Skill Generators at Key-Step Levels Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.271516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.681886Z digest=sha256:02ee6178815f1a3966e5d6c9b4d640b95183eb3c86dfcdb328bfd997c91c37e5

Observation fd6799bf-3171-4a07-9213-93720205aa0e · outbound

This paper cites Gpt-4o release.

Learning Human Skill Generators at Key-Step Levels Gpt-4o release

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.261132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.685785Z digest=sha256:2780ff47e2a83ae68f93cd1bd60c7456ffb7667ddc42fb4ed41f3abe3a58637c

Observation c7c78915-af01-431a-909e-74beb5cab700 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

Learning Human Skill Generators at Key-Step Levels Gpt-4o mini: advancing cost-efficient intelligence

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.250388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.689473Z digest=sha256:f4ee15204380e1ae466c32407cf93c856e09d463cd82df58eeed97f4866aa729

Observation 9cf035b1-497b-4340-8e04-a50ab56bad56 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Learning Human Skill Generators at Key-Step Levels DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.693527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.693527Z digest=sha256:35532f7fff05884ab23df9bb327396cfe28a77f40355ce675e8be1788800963a

Observation fc634a86-c32d-4597-bf55-c2c41e6e6cbd · outbound

This paper cites Scalable diffusion models with transformers.

Learning Human Skill Generators at Key-Step Levels Scalable diffusion models with transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.697503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.697503Z digest=sha256:1b47372467c30aa61b529a562cc1490b2da63edafdfa5cf4115d3c2ed07f562e

Observation 5a0ac98a-0486-42db-af27-bfafb4e6f147 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Learning Human Skill Generators at Key-Step Levels SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.701183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.701183Z digest=sha256:7f8dac618a7d274c18e157e9884ab87de6b9b9e5aad9ddf774b18d91b6b0024a

Observation 78bc2d0a-a0ab-4266-8b2b-0bd9b7cbf693 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Learning Human Skill Generators at Key-Step Levels Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.705234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.705234Z digest=sha256:40da3bc9c00ea030560b369266fd53b47130b1f3bdc8aefb9b568dc359178459

Observation 93597df1-d0db-4a22-8bd1-e615fd138797 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Learning Human Skill Generators at Key-Step Levels Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.708915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.708915Z digest=sha256:dcdad365a0a443da2c8a071add6f5992cd265b0c53d9d66c7a8a933eab7ebbcf

Observation 9893da96-cba5-4395-8184-a28e4f6384e3 · outbound

This paper cites Skills, rules, and knowledge; signals, signs, and symbols, and other distinctions in human performance models.

Learning Human Skill Generators at Key-Step Levels Skills, rules, and knowledge; signals, signs, and symbols, and other distinctions in human performance models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.226149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.712854Z digest=sha256:3bbaa02f29ce87d5422d070c1577ecdeacab2fd9c6b1d004fead6c9fdc3bca21

Observation a59097f5-d700-4d84-80af-f497a6c64f2f · outbound

This paper cites A database for fine grained activity detec- tion of cooking activities.

Learning Human Skill Generators at Key-Step Levels A database for fine grained activity detec- tion of cooking activities

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.214252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.716913Z digest=sha256:3a46ab078d084701f1d50529b7a984f25d0c48ec24666c74b042cfc704744e0a

Observation 608336e7-a93f-4c8b-bc62-692ec82d47c6 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning Human Skill Generators at Key-Step Levels High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.720611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.720611Z digest=sha256:f1fd61524f6d777cce4c09a7726d072b34aba0b0ea5902f776ec4ab8e6b313b6

Observation 5f035334-60cc-4b52-8741-0381fff9c689 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning Human Skill Generators at Key-Step Levels High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.724335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.724335Z digest=sha256:abc7d6a911dc3796d3128068077da16a69a53f20ef6cd213e3105a02ff00a877

Observation d675b854-7876-4acc-a12a-a65c2c856792 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

Learning Human Skill Generators at Key-Step Levels As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.186665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.728176Z digest=sha256:20bfcf49c705bf1533acdc7e5142446b54d7ae25948abe98189f8ec6db121f29

Observation c66fd543-5930-475c-826d-a41ab152b3f8 · outbound

This paper cites Tulyakov, and Mohamed Elhoseiny.

Learning Human Skill Generators at Key-Step Levels Tulyakov, and Mohamed Elhoseiny

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.174198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.731981Z digest=sha256:18290b88d72792c3d40157eebbc3138d2ac185ab16e3e3187e41c0f389bdabcf

Observation 54abcc50-6b32-4a62-a212-86e1de9dfeff · outbound

This paper cites Look for the change: Learning object states and state-modifying actions from untrimmed web videos.

Learning Human Skill Generators at Key-Step Levels Look for the change: Learning object states and state-modifying actions from untrimmed web videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.161314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.735568Z digest=sha256:bc1109fd644a0a7e41f13102eaedf33666488bdc5595c3ff2a8bbdc3c6325b8e

Observation c1cbc740-8d1f-4de4-ad37-4a08d806a359 · outbound

This paper cites GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos.

Learning Human Skill Generators at Key-Step Levels GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:57:41.894755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.739141Z digest=sha256:d153481487561d82a0356c2aa52c4867c5c29145b460210b4a72251d294664f1

Observation 72bd222b-d278-46d2-b63a-19fc12bb0af9 · outbound

This paper cites an unresolved cited work.

Learning Human Skill Generators at Key-Step Levels Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:57:42.149776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.742915Z digest=sha256:73730f3bc7e43cedecb07df5e47359e7bee946a97fe731560019cc8dbc97f197

Observation 25f80926-b6c1-4afc-94c0-77eadcb5d26f · outbound

This paper cites Plate: Visually-grounded plan- ning with transformers in procedural tasks.

Learning Human Skill Generators at Key-Step Levels Plate: Visually-grounded plan- ning with transformers in procedural tasks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.139262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.746536Z digest=sha256:b1a909fe4ea81c37899dc31dcc983b1f8e711b7dce738112600b1437057576e2

Observation 4360b4cb-7909-4f79-b17a-d29e15164814 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Learning Human Skill Generators at Key-Step Levels EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.749692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.749692Z digest=sha256:bd5f22d7d6c77eb378eee402c218f4005097610dde41259f59b9c44c36ba3004

Observation ac8de97c-418b-430b-b092-0926d9099a4d · outbound

This paper cites A comprehensive survey of procedural video datasets.

Learning Human Skill Generators at Key-Step Levels A comprehensive survey of procedural video datasets

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.127835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.753580Z digest=sha256:8614857cc777b9d3cb9fdee8ddc392b033b2b970a703e867dc8d47d39c17d234

Observation ff684aa7-86e8-4264-893d-24f821bd067b · outbound

This paper cites COIN: A large-scale dataset for comprehensive instructional video analysis.

Learning Human Skill Generators at Key-Step Levels COIN: A large-scale dataset for comprehensive instructional video analysis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.117106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.756920Z digest=sha256:0e62ece3809da45213dec2d6d6cc6f00756b4ca85fffe96cc921efd824ccb6cd

Observation a756a502-1e39-4bb7-9d06-b7c3d29fb6fb · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

Learning Human Skill Generators at Key-Step Levels Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.760392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.760392Z digest=sha256:7dfd0f2a6a496e7ba854dea8366a5ae50c93b7523d500ceabb4c66ef54cf8957

Observation cc08896b-20f1-41f3-ab43-2aae1aba9e4b · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Learning Human Skill Generators at Key-Step Levels Raft: Recurrent all-pairs field transforms for optical flow

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.763710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.763710Z digest=sha256:1158108d61c78b8d9d23be46d8679e032ce3b8b3ed4938766f1bad1c53fb9360

Observation 183fe31c-5cbb-41fe-9458-c97d75655936 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Learning Human Skill Generators at Key-Step Levels Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.089029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.767223Z digest=sha256:845682206307b42ce62d6f0a20b3af6480ed464097faa0cbb030aaff9bf07a27

Observation dccf7034-e184-4de5-9156-2f4c2c9647d5 · outbound

This paper cites FVD: A new metric for video generation.

Learning Human Skill Generators at Key-Step Levels FVD: A new metric for video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.076814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.770757Z digest=sha256:f891966cd682f1770b78fb16398d911dcfc53325cb9aecd6a90aa7163ee78dc5

Observation 05673831-b1ff-4c37-bb59-0f5778a4b0b2 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

Learning Human Skill Generators at Key-Step Levels Event-guided procedure planning from in- structional videos with text supervision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.064809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.774451Z digest=sha256:49f03508039862feeed1bc5e58ca5b543b2c0330bc406f1347394f285056e288

Observation 38c10197-f20b-4356-933c-132a379a5dd5 · outbound

This paper cites PDPP: projected diffusion for procedure planning in instructional videos.

Learning Human Skill Generators at Key-Step Levels PDPP: projected diffusion for procedure planning in instructional videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.052088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.777811Z digest=sha256:b9fb8b007dc62d658bb3f6815b947f9de4ea4a3bee5c58f8edd4757c278b987b

Observation 88bb9576-8025-4cc8-8808-b03d213121ff · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

Learning Human Skill Generators at Key-Step Levels Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.781154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.781154Z digest=sha256:ec9e65cf5ad4ed355c95d9a57920e6fb88ae9fca11271b3ad88e8b1ffbf908d0

Observation 5f3a08e6-a538-4f44-9983-8f1710ca49e9 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

Learning Human Skill Generators at Key-Step Levels Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.041272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.784886Z digest=sha256:60f6b0eefe314f827269f7bbc64860b34b42a3be2c60c40fc3afa6ca4a2a9e3f

Observation 2d8b7ea6-9b50-416e-a032-a55cceb2672f · outbound

This paper cites Learning Interactive Real-World Simulators.

Learning Human Skill Generators at Key-Step Levels Learning Interactive Real-World Simulators

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.788494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.788494Z digest=sha256:54b4513af72cf93bbea7946d49592117a50c81754b3b7fd1b1d73ccb518c0cc0

Observation 0c749702-373f-475b-aef7-0596903adbdc · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Learning Human Skill Generators at Key-Step Levels CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.792380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.792380Z digest=sha256:0b353559ba03024960cfb6a6e74042412a7a78326f0997b3d2a5e53744edfbc4

Observation c36ef58e-a876-4025-ad90-5cb9765f8d34 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Learning Human Skill Generators at Key-Step Levels IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T05:57:41.796675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:57:41.796675Z digest=sha256:701785ac355bd2dd656fea7c80f616d1d2b5e5237a07cf8e2fa61c46c6650c40

Observation 2dd1794c-ee9e-4dad-943d-9887aa7b1de3 · outbound

This paper cites Der- panis, Richard P.

Learning Human Skill Generators at Key-Step Levels Der- panis, Richard P

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.030861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.800838Z digest=sha256:06e11ebca0641df5ef5580d28ec33785cd752427e32069a79c289b258baf3e93

Observation d4a441bc-6f7f-4f98-9e02-644a6ec06356 · outbound

This paper cites an unresolved cited work.

Learning Human Skill Generators at Key-Step Levels Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:57:42.017712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.804814Z digest=sha256:5d3bc829f36d112a097affa7727f507f0f6593d2def60e1b1ad88ba5f721ff86

Observation a15cf893-8100-4a22-913f-e85defcf0df0 · outbound

This paper cites Fouhey, Ivan Laptev, and Josef Sivic.

Learning Human Skill Generators at Key-Step Levels Fouhey, Ivan Laptev, and Josef Sivic

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:57:42.004808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T05:57:41.808405Z digest=sha256:b2f0a7fae2922f3dfd29163271b7f8265ee2fb2a92d9dd8199397bfb8c3a81cb

Pith citing papers

No inbound Pith citation observations are available.