Pith. sign in

Paper Citation Record · LEDGER

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

As of 10 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2507.17462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17462 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:38.496401Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T00:17:39.409452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:25:09.923270Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ca5a907-6541-4edd-be98-86e26c94919c · outbound

This paper cites CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.294528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.294528Z digest=sha256:674ebe1b4c29106b8532dc48fff9164b4a28f0de36568ef3642843447be0b427

Observation f0a1a531-f8d7-473f-a58b-97a808b6eff5 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Scaling Robot Learning with Semantically Imagined Experience,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.735035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:31.361436Z digest=sha256:393893056bd6e5585e2c75b43767bc39d5e8fe6b5fb4135388569de2766a95cb

Observation 8baeeaea-cb0f-4687-bc98-7a2727894b97 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.519794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.519794Z digest=sha256:5e29435374612b8f171cb0be09237c7c79e7765efc1c6de196dd829d7e502106

Observation 4be03cb1-ac15-4dd4-bb8c-6a463f158267 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.611898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.611898Z digest=sha256:f7c3428a0f7476fe04e578c48c01ea773ff38c61c0353f59cc3fcab450f98242

Observation 04a71937-6bfb-4ee6-a0e0-869bd17dd796 · outbound

This paper cites MagicDrive: Street view generation with diverse 3d geometry control,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents MagicDrive: Street view generation with diverse 3d geometry control,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.492290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:31.715847Z digest=sha256:0e612c41095a64e882aab207f810e2a3230b1be558077e022f00bd48c0e2f388

Observation ad4d3104-934c-424c-a9f3-e97ebffce656 · outbound

This paper cites BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.818402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.818402Z digest=sha256:eed9657f813417b68f893bcc6c190ebb26abfaca54cf41068318ddbea4a7cde4

Observation 35592f22-efe7-47ff-92fd-c461a4bc71ed · outbound

This paper cites Street-view image generation from a bird’s-eye view layout,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Street-view image generation from a bird’s-eye view layout,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.224327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:31.939060Z digest=sha256:3dd27f744ef16af621d9cfadbd7635bd9b9e6c6cf6e17410a16724fd2fa102ba

Observation 5ab518d0-e587-4160-9616-96361742f5ca · outbound

This paper cites Dragvideo: Interactive drag-style video editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dragvideo: Interactive drag-style video editing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.962585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:32.047960Z digest=sha256:8f467f08989f9e01c70b06cd691891505b32298c1660b228042d4e95d1591239

Observation 26ed3f74-1906-4583-8acb-be8ce7bf7b73 · outbound

This paper cites Video-p2p: Video editing with cross-attention control,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Video-p2p: Video editing with cross-attention control,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.700618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:32.210605Z digest=sha256:36e0bb5c9c509340a64b910d8a82a3cf6c6e384780b6068bc7ec126d7f659d08

Observation 6c75b870-d27c-432c-bd5f-a7f2bfce9af3 · outbound

This paper cites Visual commonsense-aware representation network for video captioning,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Visual commonsense-aware representation network for video captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.383791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:32.355751Z digest=sha256:4db71303c2d643fa1ea787eb67e94f14fea1fb2659f6e02411f5e897fafecf08

Observation a17b54cc-3fe0-44e2-bea7-a313f0067b13 · outbound

This paper cites Diffusion model-based image editing: A survey,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Diffusion model-based image editing: A survey,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.135486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:32.484985Z digest=sha256:90774f26e8f8b9486c062be2a387d8680525af2714291531476e5a5dcb9485a8

Observation 858904f2-5e64-41b3-bb0c-63489df87bf6 · outbound

This paper cites Feditnet++: Few- shot editing of latent semantics in gan spaces with correlated attribute disentanglement,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Feditnet++: Few- shot editing of latent semantics in gan spaces with correlated attribute disentanglement,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.811434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:32.605859Z digest=sha256:80a305823b80fd0144ce23990a73f8cf80369ab757f63d5d8715625a98244af6

Observation de2cc8c5-818d-4812-82d7-c105f1d3360d · outbound

This paper cites Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting edit- ing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting edit- ing,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:32.743292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:32.743292Z digest=sha256:49b94043d380ff8411aaae3927db035fb022da3332fcbb2be96c417a0eeba714

Observation f33ea19d-8b14-464f-98c4-669954140c4f · outbound

This paper cites Efficient dynamic scene editing via 4d gaussian-based static-dynamic separation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Efficient dynamic scene editing via 4d gaussian-based static-dynamic separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.429108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:32.872349Z digest=sha256:110737f1b0435c3bf41fb78c9d732a3b10fcf6db50e809fbc0c2bff8f3790016

Observation 240dac32-42fa-4a83-95bf-d3b4d47eab96 · outbound

This paper cites Generating long videos of dynamic scenes,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Generating long videos of dynamic scenes,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.173267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:33.031861Z digest=sha256:64877a3ca2376828de82b733879e859223712ca92135540f1cd1c01fe466bf5a

Observation b2344980-f9e9-4149-aac0-0f5ced496595 · outbound

This paper cites Storydiffusion: Consistent self-attention for long-range image and video generation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Storydiffusion: Consistent self-attention for long-range image and video generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.868743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:33.186072Z digest=sha256:a09c424086de4fbf0ad0820ec0c6313237e2e151b349de2af8efb0dd0fc4e4f0

Observation fef81db5-5e53-49eb-85f1-025222e676f7 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Evalcrafter: Benchmarking and evaluating large video generation models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.588423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:33.288535Z digest=sha256:7439e06554a7087e7bd8cd884dfab2785b002124bd832a928933437411d87419

Observation d6105e95-80cd-4038-9ef7-ce23d0e3fe6c · outbound

This paper cites Towards long video understanding via fine-detailed video story generation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Towards long video understanding via fine-detailed video story generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.394123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:33.414870Z digest=sha256:9f14ec4140c1ff41d931fdf989e911886148074cf4228d2da3319c2d810821ac

Observation b30c98a2-6965-4090-9e4b-9012ea9e36a0 · outbound

This paper cites Maskgwm: A generalizable driving world model with video mask reconstruction,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Maskgwm: A generalizable driving world model with video mask reconstruction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.094305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:33.564446Z digest=sha256:deb39eb8de47a40bb37835528f00158d05819ea49bfd71918e3d9d2a33adeb58

Observation 537a1492-6b2a-408f-880c-16d2b2a1bf8c · outbound

This paper cites A Survey on Long Video Generation: Challenges, Methods, and Prospects.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents A Survey on Long Video Generation: Challenges, Methods, and Prospects

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:33.702516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:33.702516Z digest=sha256:5d38eb9d03c154feb2a53efb50e183bcfd20ceb0ff8e1fefe30464d1d24904ae

Observation 64896346-6477-4c74-a9cf-a30b6a39df53 · outbound

This paper cites Dall-e-bot: Introducing web- scale diffusion models to robotics,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dall-e-bot: Introducing web- scale diffusion models to robotics,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.734812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:33.900587Z digest=sha256:0cf0fe9dd762d4153641f20679ae2df091add77aaa463e23645d7ad6ae69bf13

Observation ec05318f-4a89-4cac-a4b0-b51e7bc9083a · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.065833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.065833Z digest=sha256:79e3be87567ac9e2e6e9f20fbdf8a325b7fcd4a7c25078c81cf4ccb78f67d88b

Observation 7fe0b49a-3265-49fe-a6aa-54abf355c307 · outbound

This paper cites GenAug: Retargeting behaviors to unseen situations via Generative Augmentation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents GenAug: Retargeting behaviors to unseen situations via Generative Augmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.194140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.194140Z digest=sha256:5a7ded996fd6c7c8359df289180cbf596c3134c364d50eee6468bf3e42a35dff

Observation e7569c55-ebc9-4a05-9667-47468ebe9be1 · outbound

This paper cites Semantically controllable augmentations for generalizable robot learning,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Semantically controllable augmentations for generalizable robot learning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.333520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.333520Z digest=sha256:aa81bf010e01c9bb4738698ed912b89b27b7f7dc0ee7a4ecf611241f751dd3fa

Observation 3f9cb659-3d2d-4a89-9357-d4a00e51e36a · outbound

This paper cites Roboagent: Generalization and efficiency in robot manipu- lation via semantic augmentations and action chunking,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Roboagent: Generalization and efficiency in robot manipu- lation via semantic augmentations and action chunking,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.481792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:34.459950Z digest=sha256:ebc80365a07087e33546e3466089619c21b97c06b046bead69275a7145b9a4d8

Observation f48df0bf-9d3b-4967-bb78-03bd870c2454 · outbound

This paper cites Segment anything,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Segment anything,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.586243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.586243Z digest=sha256:9a53116ffdf324d80a484f7d3cc7f25e2bda1eafae378b463972a5661077ada3

Observation e8b228d8-3a7f-428f-bc33-44f5c8b892d9 · outbound

This paper cites EnerVerse-AC: Envisioning Embodied Environments with Action Condition.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents EnerVerse-AC: Envisioning Embodied Environments with Action Condition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.708735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.708735Z digest=sha256:f0972df38803f8fc37c62dc4391cbbd1fd49c05612098843f6d2dad953bbd20f

Observation 27c73efe-2b8f-4912-8279-17ef1c9a34df · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Zero-1-to-3: Zero-shot one image to 3d object,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.247498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:34.849354Z digest=sha256:9d2d8fdd80eafb6a08426f3825da68731ac0bd49caabe12dd3414e72278e7bf1

Observation 0bce7733-0683-4bd4-b8a0-e45affa1d359 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.980027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.980027Z digest=sha256:8fe1b94fb9ae300b7b53b420ad93b333b8cf64a73cc0e684af4b5bdb43c6d3a9

Observation 1c2321a9-0043-42e0-8703-03a03eeb6977 · outbound

This paper cites 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:35.168365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:35.168365Z digest=sha256:615e708f2af866191f5c1720f252ff90c11162e4de45e84c7bf6876effa707cd

Observation 2d9265b2-33db-4031-a976-f83f19887450 · outbound

This paper cites Dge: Direct gaussian 3d editing by consistent multi-view editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dge: Direct gaussian 3d editing by consistent multi-view editing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.001632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:35.339800Z digest=sha256:3b0c2ec85e50c066bfa56faa709e11eb5b19a21c003e1a73f32372b22f7abf81

Observation e8c810d4-e30d-4347-ae5c-2f106ca3aa20 · outbound

This paper cites Imfine: 3d inpainting via geometry-guided multi-view refinement,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Imfine: 3d inpainting via geometry-guided multi-view refinement,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.681720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:35.479138Z digest=sha256:1179ec2926acc6324ed3a5a900f458359117bc0a4e596a684b7fb6712c2bb077

Observation 822d2eb3-7671-400a-b957-cd6d7ef32a5d · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents High- resolution image synthesis with latent diffusion models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:35.647926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:35.647926Z digest=sha256:a98438dd75a929f3b459e0a84e1847736bb445d3610de0b946759ad4fbbb0ecb

Observation efaa2062-6f9e-416b-8b88-5f1e572d0c7a · outbound

This paper cites Towards language-driven video inpainting via multimodal large language models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Towards language-driven video inpainting via multimodal large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.373678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:35.820351Z digest=sha256:c021264632b8754a428945ddf1c0c642431143a5c28bb905e44517049056aa8f

Observation 0d35a874-59eb-41ee-8366-83b33dab9d1c · outbound

This paper cites Brush2prompt: Contextual prompt generator for object inpainting,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Brush2prompt: Contextual prompt generator for object inpainting,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.065267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:35.926503Z digest=sha256:b4af475edf2ccb992dab8345caa6926486e63c4d85f174a9ec9b039595050ba3

Observation 9e760139-0a9d-4a6d-9ac5-19d643a221d6 · outbound

This paper cites Sketch- guided image inpainting with partial discrete diffusion process,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Sketch- guided image inpainting with partial discrete diffusion process,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.819891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:36.052743Z digest=sha256:80eab0cb857a5748ecfdbd2cafe844d387f176602a929d15dc1e0e9a0edcfa4e

Observation 337a5171-982a-4350-842d-46761fed50a2 · outbound

This paper cites Few-shot image generation via style adaptation and content preservation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Few-shot image generation via style adaptation and content preservation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.501753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:36.256605Z digest=sha256:918a0ddacef0609dd2c47cee74e37b963b0a78d48273fcf986dbde028f0a2cc8

Observation 2b5dfbe2-4686-4566-ac14-412938512fcb · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Learning transferable visual models from natural language supervision,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.417013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.417013Z digest=sha256:24ce9b237f7e7a57b5fed12595220a6e299a8843226dfd90f3e67e175348264f

Observation c1f65cef-63aa-4d4d-b954-2bf34b686539 · outbound

This paper cites Video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Video diffusion models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.559772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.559772Z digest=sha256:27d4767975af705035fc04977a7b29f2bb8ea83dd5ed21b343f2a64cf81d3c4b

Observation e59b2166-dec2-49ad-b81d-6e467a3a7ffb · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Enerverse: Envisioning embodied future space for robotics manipulation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.744744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.744744Z digest=sha256:a2391fbfc4356cca84e0f5ecf2237cfc1385e90139b093732bbdced3f56a009e

Observation 02991ebf-d39f-4ce1-b936-b379fb0b1826 · outbound

This paper cites Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.178023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:36.907756Z digest=sha256:65d484d9fa2d9989d53afacc0cca213e668559e50ec3ac6d53222b97b8ff64bf

Observation 4e7cb55a-f64f-4b26-b156-4c97ebcb9b65 · outbound

This paper cites Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.812003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:37.094974Z digest=sha256:c8a4bea2d73b5210463a8e167ee5a6531feff9e2f7dc21487614b615bba01315

Observation 475c8523-758a-435a-a4cc-b3631bffabf6 · outbound

This paper cites Progressive autoregressive video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Progressive autoregressive video diffusion models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.489380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:37.295899Z digest=sha256:1b788c22b25726e9f339194e2ba38cfed4d106315fe70560929a0055480feb56

Observation 8a3c1759-7c0a-4c46-8052-9405e6541aa0 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents From slow bidirectional to fast autoregressive video diffusion models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.457637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.457637Z digest=sha256:9babc67a9625020629d0520f4ad2c9521479a2d2dabc13d204373be1b341c3c7

Observation a207e586-b1bb-4667-bdb0-8ad1045082c9 · outbound

This paper cites Qwen2.5-VL Technical Report.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Qwen2.5-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.599945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.599945Z digest=sha256:b17ec4a043ddbd8e7f76700a59d9546a83c12414f65660dd5b75e23dd4ddcc07

Observation 72315267-80ee-42a1-b6e5-0598d6458ecc · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Step1X-Edit: A Practical Framework for General Image Editing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.812216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.812216Z digest=sha256:606d6e343692ed8efc71c34543880eb8aeef20ea4740beed9a0714d89339301b

Observation 80f9a6bd-903f-46d4-82d4-275ddf05a4f6 · outbound

This paper cites Robotwin: Dual-arm robot benchmark with generative digital twins (early version),.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Robotwin: Dual-arm robot benchmark with generative digital twins (early version),

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.242888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:37.961048Z digest=sha256:d3eaf52490ad13dafeaafc53b554c0157aa3b3b61444a788a16b79cc78e4c0f5

Observation 4b404dc2-4398-4132-8db1-459fd9fb345d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via ac- tion diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Diffusion policy: Visuomotor policy learning via ac- tion diffusion,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:38.099075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:38.099075Z digest=sha256:81014367030d5e55552f0d257b0ef7d10c7e8ae0b7c36f1a74e4b931fa9943de

Observation cd782825-b3e2-4a52-a651-73a0e9b715f1 · outbound

This paper cites Schedule your edit: A simple yet effective diffu- sion noise schedule for image editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Schedule your edit: A simple yet effective diffu- sion noise schedule for image editing,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:38.933412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:53:38.294649Z digest=sha256:583ccb9e74ccb210f7b338452979edc4adf7658d83c02e8dd7f83dd254e165c9

Observation b0ffbd48-09d1-4a61-98ee-25730bfa6101 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:38.496401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:38.496401Z digest=sha256:c905ae241f467eeef05f400b31fb7cfc61ec5ffbcf574f0e59f1ef387039695b

Pith citing papers

Observation 6fb304ea-b10f-4f1d-8fde-8c04a951d874 · inbound

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling cites this paper.

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:32.152811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T19:16:50.004359Z digest=sha256:0bb425e9bb7e517bf07fb44164e6bdc32407268152231179be77386d3ee3bcc9

Observation 23da8182-92d7-4d43-a694-e877fb20b4a1 · inbound

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling cites this paper.

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.924677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T00:17:39.409452Z digest=sha256:2046442cb4bcf466a682cc4126f5e1d8a8851546068434a0c23bebd4f483da78