Pith. sign in

Paper Citation Record · LEDGER

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

As of 3 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2605.03637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.03637 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T15:44:25.021995Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:28:10.101325Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact31
  • verified fuzzy22
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 827612a1-0079-43b8-8226-b35270fc0466 · outbound

This paper cites GPT-4 Technical Report.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing GPT-4 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:01:31.548286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:5f612ccdf1e742ace40e1c37e903f410f3a0ed97d5d7a18d65b6cd63e92aba51

Observation 7d317eb4-1997-4ba4-b875-06c7fa58895d · outbound

This paper cites ScrewMimic: Bimanual Imitation from Human Videos with Screw Space Projection.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing ScrewMimic: Bimanual Imitation from Human Videos with Screw Space Projection

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.539086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:dd53e79c23ad0d8edcb2b1410c07ad6f2c37a10b7e3a9cec1b2a44c73a44e0d3

Observation 096874db-3c42-467f-8d73-433da61f372e · outbound

This paper cites HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.542416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:b47f1d9ac1f5828a8f122ab4d867d8b71d77f189186b19df76eace31bd885ea4

Observation 07c6e875-c634-4ac8-ace1-057e06b9bade · outbound

This paper cites DeepSeek-V3 Technical Report.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing DeepSeek-V3 Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:01:31.545329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:6244c1ae6be04ce6547e2123c72082913e55069f1e0bd62ffa37b054722d3eae

Observation 5cc62dec-6b38-4355-8bd6-3a1d4c7c73d6 · outbound

This paper cites Improving image generation with better captions.Computer Sci- ence.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Improving image generation with better captions.Computer Sci- ence

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.959542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:bca39853cdbab32ac345310f2b61d91255bd26ea7a66edf88a4cad67ac6c7e48

Observation cffcf19f-b6ec-4a10-94d3-bfe2b32d95f9 · outbound

This paper cites GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.558785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:b7f18aed8f816862ae97e9b806fe27c81dab7630d1286fa5f1f8fd117213848e

Observation d29b6301-b25f-4edf-b952-99846598dd74 · outbound

This paper cites Video depth anything: Consistent depth estimation for super-long videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Video depth anything: Consistent depth estimation for super-long videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.962284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:0163d3ec988e2deb444c31517cabe8b7d589c93358fec783ceb20c0b5cbf6dd1

Observation 623ab139-3a95-4e11-9e51-e459127bd83e · outbound

This paper cites EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:40:29.471447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:ca0df7423591fbdae1abfaec4b51b742b74f82283e4df29a264c83f7bf52fee4

Observation 8aaa76c3-8fc9-45ad-a5d7-9b1988302356 · outbound

This paper cites Club: a contrastive log-ratio upper bound of mutual information.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Club: a contrastive log-ratio upper bound of mutual information

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.964871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:5e3a7791e970fddb338cece43a9ed3c3913677a94d390108d04b20ecaf3f08fc

Observation 411e13ff-b86e-4dee-97d6-c190cf691409 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:38:11.644840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:7d405acf2b7884eaebae7479f9ba5e44ee00ccfbb998cc0945d67e7f7edb104a

Observation 623f8c52-0de6-41f0-9a16-8f588d2a1b11 · outbound

This paper cites The trimmed iterative closest point algorithm.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing The trimmed iterative closest point algorithm

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.967763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:dc432ff808082278a009bc9605c6f49b93d45db2cab305e6917e528f9c9131df

Observation 100584e7-f53a-43c4-b823-e3208d564c15 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing VACE: All-in-One Video Creation and Editing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:53:54.408026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:a614854f71a36460a1fbfc6e877bad80d96550b3ca1afbfd1958723a717d9c55

Observation 4d214dac-365b-4072-8430-673be4fb1ef5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for lan- guage understanding.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Bert: Pre-training of deep bidirectional transformers for lan- guage understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.970453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:9c34437f8bb4875615988be0f840d86dc360a3ba78bbee6f8cbc5fc30126988d

Observation 7db13b20-5d7f-4f27-b68f-8cc2646410e3 · outbound

This paper cites EgoMimic: Scaling Imitation Learning via Egocentric Video.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing EgoMimic: Scaling Imitation Learning via Egocentric Video

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.535838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:45ca555a8d1e4b6935c6f5a396975d09b5ce4f944ff2247040b7c4bbd37bfe83

Observation 8349e0f5-826a-4ad8-b1f8-953621fa6c1c · outbound

This paper cites J., and Hilliges, O.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing J., and Hilliges, O

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.972827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:a9258a9d2035a457126670fb0f743aefb1f16e88509475f35dea2251729dfebc

Observation 9bf3b0d6-1ab6-4762-80b9-821ca37e7586 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Learning to Act from Actionless Videos through Dense Correspondences

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.562222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:3484257fe9c78890d3fb2bd673c7f42449b51a851dcd553b1e5361ce45c9d149

Observation 03333a73-8ddf-4dd5-8d8c-4b42124d0a8f · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:01:31.515721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:84937d77fcc6a9a146ac52a6591d876b6607458667353d57cb9065c2f54e7961

Observation 636ac2b2-f3c5-4423-b6b2-75871dd435ce · outbound

This paper cites The” something something” video database for learning and evaluating visual com- mon sense.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing The” something something” video database for learning and evaluating visual com- mon sense

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.975268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:e74d3cd6e3c7bc698a4cbb941ed0280081d68c74089194115ca3db20e8fc2709

Observation b8b5d8cb-2743-44c8-84ae-91ff6f5d2232 · outbound

This paper cites Decoupled Weight Decay Regularization.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Decoupled Weight Decay Regularization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:31.518757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:5061336f9fc90bef181d97047222e77512934237944fd41cf84d44ce0b523106

Observation 2fa89b3d-51f9-4cc8-a196-616e2effa82b · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Ego4d: Around the world in 3,000 hours of egocentric video

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.977664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:9ae3f77665cd00e88cea756878e92def62527f78e40be38f103229660b0ead6e

Observation c18e144a-27e8-429f-8c08-7b656dee87de · outbound

This paper cites RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.522365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:ceddcb314d80c1c9631fdb26aa0e84d32512ed02e153f64f6d0dd872991d1dc3

Observation 2f36abad-7ae0-4429-ac31-17aa112b0a56 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Representation Learning with Contrastive Predictive Coding

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:31.528392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:0fe2ac54bbf7685bdb594fdcfa47337fdc3b8d5ef74f546a89a76d6ea0f1ac1e

Observation 10cf7beb-7f86-4f23-af33-dd66878d335c · outbound

This paper cites R+X: Retrieval and Execution from Everyday Human Videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing R+X: Retrieval and Execution from Everyday Human Videos

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T11:01:31.504658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:972e2676f411534f0efd9cd6da65806592af8ac2236ebf87801d423fcce88a76

Observation 42ae427a-7feb-4e1e-bb3b-37633b241f87 · outbound

This paper cites Vbench: Compre- hensive benchmark suite for video generative models.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Vbench: Compre- hensive benchmark suite for video generative models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.987051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:81ce8009f828ded793176b8d86f3f60faafce7065553ff2363a04c6d614a27ea

Observation 736088d1-9945-47de-bda1-9ec40d26d688 · outbound

This paper cites Learning to Transfer Human Hand Skills for Robot Manipulations.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Learning to Transfer Human Hand Skills for Robot Manipulations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.492828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:56b55702b4779da471e984683bfe491453785cef196c8d0c975d85535558e1be

Observation c6a306ff-7d2c-4ca1-8b3d-88957381655b · outbound

This paper cites Dreamgen: Unlocking generalization in robot learning through neural trajectories.arXiv e-prints, pp.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Dreamgen: Unlocking generalization in robot learning through neural trajectories.arXiv e-prints, pp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.989579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:69d9cd254950acd347bcf3ffe5ffd287937525fdbf4ab0ed533c4648caabd3d8

Observation dbb3ee31-5762-4b22-bfd2-af6614e1c1e6 · outbound

This paper cites Yoon, Ryan Hoque, Lars Paulsen, Ge Yang, Jian Zhang, Sha Yi, Guanya Shi, and Xiaolong Wang.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Yoon, Ryan Hoque, Lars Paulsen, Ge Yang, Jian Zhang, Sha Yi, Guanya Shi, and Xiaolong Wang

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.453293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:c17f815ca1f8360640fe2cef443aa351a325f743e138444b697937bab20af07e

Observation 81bcc662-f9a8-4a3d-9aed-7edcbaf56393 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing SAM 2: Segment Anything in Images and Videos

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:31.489255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:7f3afebe07e246f65768c0b98ee9f08ec125c92555bfffdc3dd51dc5d7417263

Observation d545aff9-9179-4a9a-936d-54882d989ab8 · outbound

This paper cites VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.496416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:a3437e069a018df0a7c9e729e0eab99005731c4a97bfaac62ba5c300e2047554

Observation 036f130d-abd2-4855-a9b9-b8912290057f · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:31.508531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:6351284d5bce09deea8bdf751bf9a51c5cbd5058cf58d252b1e474b6d43ac6f8

Observation 69589ecd-2561-4e99-ac06-124e7fbb4dc2 · outbound

This paper cites J., and Lee, Y.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing J., and Lee, Y

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.441833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:e449212ae6f468b488ecbb9ef258c87bda546a407ff4c29eb264a3b778a882df

Observation 21202640-44db-4b99-aee8-10a6bf19853b · outbound

This paper cites MimicPlay: Long-Horizon Imitation Learning by Watching Human Play.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.577514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:790741cfafaa5b334f1297276cfcfc656758326ea711946522e92b8e001570f7

Observation 3fca940b-5fe7-474f-87b6-03e3dd9a9085 · outbound

This paper cites Unianimate-dit: Human image animation with large-scale video diffusion transformer.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Unianimate-dit: Human image animation with large-scale video diffusion transformer

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.511904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:10e48af44cff4343acd3cc89fd05ee8e88eade179e348080dbe3c3574482347a

Observation 8cad6fa4-5c1d-4154-807e-78a46ee80e90 · outbound

This paper cites Masquerade: Learning from In-the-wild Human Videos using Data-Editing.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Masquerade: Learning from In-the-wild Human Videos using Data-Editing

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:55.330788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:1dc07cbb557c0050f69840ff8039ba9a4f643a66e3ae8b94ff54110e55e39c1a

Observation d9e6393c-8aad-44a7-8673-c1e43c96417b · outbound

This paper cites $\mathcal{D(R,O)}$ Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing $\mathcal{D(R,O)}$ Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T11:01:31.479570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:441e878724156e01e9dd0517ee4ff7620eacbab328f7c1818d83600768e6ada3

Observation 3da35b47-e06f-4896-b2a5-015047bfdfda · outbound

This paper cites Qwen-Image Technical Report.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Qwen-Image Technical Report

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:01:31.482358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:5c5297d92ef6f4ec13915f023fcd55b0a3d2e9693f0778348ffcfcd448676b53

Observation 7850f2e8-e6fa-44b3-9e2d-9bda9a76888c · outbound

This paper cites Phantom: Training Robots Without Robots Using Only Human Videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Phantom: Training Robots Without Robots Using Only Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:52.652914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:6dd323c69cb63f5a0824ad0aec1affd9884a2e72d6b2f502a02f04dbca8e671e

Observation 1e55a747-8acf-4d87-8b01-d7bd2a6e5e57 · outbound

This paper cites H2r: A human-to-robot data augmentation for robot pre- training from videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing H2r: A human-to-robot data augmentation for robot pre- training from videos

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.457390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:a837851571d186c35c0d834491cfcf3f8c9e5601cf20880b9e564354db64411f

Observation 29029250-12f1-40d2-9148-8a76c2a07711 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:31.472008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:9d3e0a79b49ddfa4e8d0acc5c40326854fa09b28cf4c73e2cc2416c976f36ae8

Observation b40b883f-1a5c-4c35-b122-38c03108dc4c · outbound

This paper cites Unified Video Action Model.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Unified Video Action Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:50:29.857677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:516a14b846093a668bce4f1734c7feb93f79d0944d3942da79efdbefb45d5a27

Observation fc4ada06-80e7-4344-95ab-7ce18c2cf58b · outbound

This paper cites Latent Action Pretraining from Videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Latent Action Pretraining from Videos

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.461201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:6c4c0324548eaf3f18a0c364f20354a9f45e565bf8e4090b5d44d4005957bf3d

Observation 314c1b96-1330-46a8-9ef4-a9993e959e87 · outbound

This paper cites Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.468932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:4e8aa8518a2735caecc533c00a117b095fd4bb9ae36750ca7f8b1c0961d48dde

Observation 964f7a6c-8252-4f81-98cd-6ce121d33c6f · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:31.554848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:1012ec80bfe6501d977309000c4702978abc4a43c35e638e858cb9d2ecc45dcf

Observation 4d166979-b203-430d-acb4-3871f233bc3f · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:31.551664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:f91985e39936a1af785d48f4b8d12c63fbc64cc63d8c7d2fedf3344c7ef3d1ea

Observation cb32dd4d-7976-460d-813c-e831a0b4e9fe · outbound

This paper cites DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.521687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:aa14589a5458deb4ca384e04bee06bbb86c551426a373e60cfdf19686b4d4420

Observation e61044ea-c82c-4080-8904-4a688f594a79 · outbound

This paper cites One-Shot Imitation Learning with Invariance Matching for Robotic Manipulation.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing One-Shot Imitation Learning with Invariance Matching for Robotic Manipulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.485970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:e96771552e287779ed26412ef46aa408c63ddec35ea0cccba40123a9b931229e

Observation a68571d5-8c32-4e15-aa19-b2cc69440840 · outbound

This paper cites Follow your pose: Pose-guided text-to-video generation using pose-free videos.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Follow your pose: Pose-guided text-to-video generation using pose-free videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.935200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:773c4428c5207b41c44bd08ca2f769f5451193eff9d88b9ecc8a838831cbd972

Observation d0045015-5519-4f35-a577-c466e0bbcfce · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.500911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:e147cda6f93b769260072b9fc094072d133e05a0825c1672dcc14e91049b5da0

Observation 6b89b11e-1263-486e-bc8b-4969fac48dad · outbound

This paper cites G., and Li, Z.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing G., and Li, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.937784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:cf2d0c216a748a131d6457b816bd3faa70e52a49e7aed4c8cec5cd5170eea42b

Observation d4245804-3427-4f29-8c1c-0f0d35f82237 · outbound

This paper cites Z., Zhang, D.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Z., Zhang, D

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:58:44.678388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:30fc2c12285d27a0fe7db40a77acb7e7bc74a820b2683f428b31d1b1fd6541a8

Observation 9f33d108-dee8-4b43-b290-7d9ab262ed30 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:01:31.525402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:e3e5169c5fa7fe2365878e68b120f8eb4bb89d6f220be1755432d157798e44c6

Observation 5c904740-9e62-4a7e-a43b-15029303366b · outbound

This paper cites Vision-based Manipulation from Single Human Video with Open-World Object Graphs.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Vision-based Manipulation from Single Human Video with Open-World Object Graphs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.464839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:94d23650d7142fb46633f8282fd1e0b61e1211c23ea311dff69439b5111100a2

Observation 76706d66-0676-402e-919a-f7092e49246e · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collab- oration 0.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collab- oration 0

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.940302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:926f839250c90312eb5681ab5f9d14ddb1fe7910e5bb42a1e7efc66ca289374b

Observation 7a039b3e-5c35-49c5-b407-100df565ae67 · outbound

This paper cites an unresolved cited work.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Unresolved cited work

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-05-27T02:58:44.675151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:e1b7cf87d556d56d53b25740867fc3325ba88defd419bf42ca71be9d97d44021

Observation 8b645966-d423-4007-966d-86f995824606 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.956907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:d81375e1f5e48130de54924aa45d9e366a5a8087b6a325e5f12f642b042754ec

Observation 9654e2cd-4e7f-4b7c-b623-4ca09b21d72f · outbound

This paper cites an unresolved cited work.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-27T12:24:03.980101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:4bc2d43558d774e209489ef4d31d3336e61943dc4bf4bf0ff0f877dc506820ba

Observation 4d9d527e-8879-4105-924b-e74dd4919dda · outbound

This paper cites W., Myers, V ., Kim, M.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing W., Myers, V ., Kim, M

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.982573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:89a67d052dc111a24d47c16aa812086c34569a9c08320c6a0cf5a135060fcf35

Observation c65e1b88-d8f4-4c6f-b98a-7eea2f1ec8f4 · outbound

This paper cites C., Sheikh, H.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing C., Sheikh, H

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.984754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:7a932e681850d528d5ea446657b926d339f8edd5a7320820f4f158f17e808761

Observation 0dfd7665-5bbb-47df-afff-8ceeed8d4c18 · outbound

This paper cites Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.954062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:2b4bbc48e63b164661cf8aaab6117c50c028a11869bd292815339544d4347f06

Observation 9e74d60c-f91e-4f57-9884-1d921ace1e41 · outbound

This paper cites J., Shou, M.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing J., Shou, M

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.951345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:05aa9383eda3a59310428ab27f2a5968258982721bf70bd7934d3eafda8a7882

Observation ecf6162b-5a82-4843-9a4f-9b7d718cdcea · outbound

This paper cites A., Shechtman, E., and Wang, O.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing A., Shechtman, E., and Wang, O

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.948766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:7583ce84aef0f19be963c7bfafd727aebbc9071475ff7e60a1086f1967f44d1c

Observation 1d34e2b8-7fa6-459b-bd82-de384ba70d00 · outbound

This paper cites Motionpro: A precise motion controller for image-to-video generation.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Motionpro: A precise motion controller for image-to-video generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.945780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:16fc646001a3b14b9ec6f85063f6f8d153f3fd318803c23705e314602ffa26a3

Observation e4785666-54a4-48c6-ad65-576622b35d19 · outbound

This paper cites Z., Zhang, D.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Z., Zhang, D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T12:24:03.943044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T00:50:06.906904Z digest=sha256:bede46530918042e42f50663ce9d1b56b3445f667a931c418a9346e335e6fab4

Pith citing papers

Observation 75a50c3e-30a1-4e3c-b583-61baa9cbd6f3 · inbound

AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning cites this paper.

AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T06:16:27.546219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:16:27.546219Z digest=sha256:8c5b2f6970f240172b998601250c0afe81cba7325888bcb591f60357047fe8f4

Observation 7f584892-e9a6-4ada-ba96-5c55b2067267 · inbound

AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning cites this paper.

AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T15:28:10.101325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:28:10.101325Z digest=sha256:f9e697099f38b5862a8b19d5a6bb51c860ecfb5015991a8c388255be6bfa85d9