Pith. sign in

Paper Citation Record · LEDGER

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning

As of 20 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2607.09825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09825 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:16:18.444295Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d21c3f39-ff4c-46c0-9377-1f9ed95c6293 · outbound

This paper cites an unresolved cited work.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:f98a0a15a524d80193071313378c40c84663a01d95528374610ab71937c53cdf

Observation 1a9e1690-711b-4d77-bd24-dc28788264d8 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Emerging properties in self-supervised vision transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:82a26cf03d0a645ddadffb184ac55cbe945fc5c958873cdca33a3ac59a1f4d69

Observation 7c6c0035-98c9-452b-8e11-ab9342987ca1 · outbound

This paper cites Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:0afded39a3a528b6da52a2b93e611c03dc1ccab851c08d97b405987dfa5e2875

Observation 5769fcdf-36d2-4dd9-b8f2-5a501c0880a8 · outbound

This paper cites Dif- fusion policy: Visuomotor policy learning via action diffusion.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Dif- fusion policy: Visuomotor policy learning via action diffusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:d9779b9bf59f1974eebd1b61ba5311049fe96f363ef392b0cd1c7b5993d6877e

Observation e6c20047-e6d0-4a8c-accd-46a412e1df54 · outbound

This paper cites Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Georg Heigold, and Thomas Kipf.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Georg Heigold, and Thomas Kipf

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:2180dd0898633df035e3018838c9264be3b90592e61ca6d3c2989d640556ae63

Observation cbe21540-d6f7-44d4-9164-b0532f9a1f18 · outbound

This paper cites SPOT: Self-training with patch-order permutation for object-centric learning with autoregressive transformers.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning SPOT: Self-training with patch-order permutation for object-centric learning with autoregressive transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:36669f3a4b1f17f619dd1fc53e2a8810d499ed26604fd614c66f6d895a322f85

Observation 6b0de0dc-0086-47d3-b7dc-de6be97c2b15 · outbound

This paper cites Elsayed, Aravindh Ma- hendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Gr- eff.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Elsayed, Aravindh Ma- hendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Gr- eff

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:af181afc77dbd55c691403943c4e46a72449267fac98f54547cc7d96f3bd89a5

Observation 4da1e232-2bc9-421a-bcae-8f28171c86d6 · outbound

This paper cites Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:07d9738cdcc411f285e244674d4300232eb65c591b3f4305fef099852bf03bd0

Observation df1fab8f-408d-4515-9c0f-28ea657caf90 · outbound

This paper cites Object-centric learning with slot attention.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Object-centric learning with slot attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:f5ba97a94cec7ad5ec1c24878ea6cf50d4c7b01f908ab5e5aede26e0b1d07e91

Observation 61571298-dec9-4341-b172-d642a6e0da31 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning R3M: A Universal Visual Representation for Robot Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:1133c3a4ffa575bd2c1d50159d08a8725eb096cb6464cabf06a0fc7d37109c5e

Observation 1fcd72bf-0c41-490c-a563-164046556eb9 · outbound

This paper cites UniGaze: Towards universal gaze estimation via large- scale pre-training.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning UniGaze: Towards universal gaze estimation via large- scale pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:00b4850e8e7a14bf677911e4c63c399a7ff50594095912c65912a8d82d1c8a35

Observation d605f260-4107-4637-bc09-db7c4425f4b9 · outbound

This paper cites Real-world robot learning with masked visual pre-training.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Real-world robot learning with masked visual pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:5bbf6898e12bef3b122d1fcbaa0dd4ac2b95187c4cbdc03dece5bf7394b93db7

Observation 5c483134-84b9-4cc7-b440-9996ad2feb10 · outbound

This paper cites Bridging the gap to real-world object-centric learning.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Bridging the gap to real-world object-centric learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:eae16509808d4c6b82b45edac344d71db3479205fcb84f0897347b5a00032f61

Observation e6c6b9df-c8a5-45c5-8dbe-e6cfd0759729 · outbound

This paper cites Vision trans- formers need more than registers.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Vision trans- formers need more than registers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:5dad73b2fb7727800e90ebd4eafa411bf69bd1c2cf30609ca5b6eb24b212dbfd

Observation 8cc58953-8bec-40d8-b3e8-9ad5779c8318 · outbound

This paper cites ManiSkill3: GPU parallelized robotics simulation and rendering for generalizable embodied AI.Robotics: Science and Systems, 2025.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning ManiSkill3: GPU parallelized robotics simulation and rendering for generalizable embodied AI.Robotics: Science and Systems, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:6e470b9b969464d5da6d58629ef83ab72a30074badcc18edcd1d0a040210813c

Observation 914fc1cb-d5ec-48ba-bc94-af4a3d423731 · outbound

This paper cites SlotDiffusion: Object-centric generative modeling with diffusion models.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning SlotDiffusion: Object-centric generative modeling with diffusion models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:2d3431491bde5155b47bcb74d69dbbb2e72f628bd8fc81e24b8177415e84d06d

Observation 4ec39184-f587-406e-ad99-e5921c56af02 · outbound

This paper cites Utonia: Toward one encoder for all point clouds.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Utonia: Toward one encoder for all point clouds

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:f08b7d436e77ac0ffb4f3506fab5d4aa7d152e91a5341e91fed2364ed80baead

Observation 2bc5bf99-b5d1-4e6f-8beb-6a4449b77384 · outbound

This paper cites Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:26ba5de37df8f407a5435c39af0b3f397d7de9e12c294a8cf85fcaedec28b2ed

Pith citing papers

No inbound Pith citation observations are available.