Pith. sign in

Paper Citation Record · LEDGER

Theia: Distilling Diverse Vision Foundation Models for Robot Learning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2407.20179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.20179 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:03:27.121387Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.580056Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e6bc9326-0a31-4aa2-baf8-79d69a5f684c · inbound

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation cites this paper.

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:13.141072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:13.141072Z digest=sha256:08bbcce4bb72a711d826fb035a78716036c3373f5c5fcdede04a077fc835c286

Observation 08f83ba3-34d6-4866-a1cf-ebdd5e21bba4 · inbound

SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation cites this paper.

SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T23:04:14.561926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:04:14.561926Z digest=sha256:c5748b5c5391ad474204d0fc69015645e9fe1388ad4769bdd6456632101c6718

Observation 54a0ef54-462a-492b-a145-9f1a8e825c62 · inbound

Object-Centric Representations Improve Policy Generalization in Robot Manipulation cites this paper.

Object-Centric Representations Improve Policy Generalization in Robot Manipulation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:27.121387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:27.121387Z digest=sha256:23cf0258e455047990183341e716ec5d7d8b872e8230ba10857bc1c5f37e75b1

Observation 5cec6f46-9634-477d-81a8-74278cc13a85 · inbound

Is an object-centric representation beneficial for robotic manipulation ? cites this paper.

Is an object-centric representation beneficial for robotic manipulation ? Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:50.078534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:50.078534Z digest=sha256:6afb5b4bd3ce780d898a5893228e61f0ff4865e25c1d2a883e26cf226825bc31

Observation a9cb98dc-1036-4994-8d70-8f919ae0bd3c · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.660719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:85189c3f113506f10565d8d6673d5a022a12c5e5ed984890aad85342a5cdc538

Observation 26caf29b-1180-4d8a-aa97-7c020eed7bff · inbound

Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models cites this paper.

Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:22:58.808101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:22:58.808101Z digest=sha256:c0667a1ae24d02194bc6910a465111e8d2a1e56419d8edc6ab074d781405b31f

Observation 3c575d66-de10-419b-9a9a-87d97cf82a59 · inbound

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation cites this paper.

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:19.552910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:26:19.552910Z digest=sha256:1496cea866b2a5f43e17da9de012fceb79948db6186301d23745b0ccab6fba9e

Observation 890b4ed2-661c-42f5-88be-4dcaf068aabd · inbound

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models cites this paper.

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:24:04.765662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T05:23:11.261860Z digest=sha256:f7b818cceb5b28550e26c3ec722eed25ed6b41b9b98186af57c9d46271bfdcc3

Observation 319d9c38-1d2e-40b0-b065-2741b47c2424 · inbound

SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models cites this paper.

SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:23:23.731777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T20:21:47.430865Z digest=sha256:adf60a5920ea911c13fe01741bf922d12ce214015923dae7d5f2f9e747e47acd

Observation 72cff02c-9387-4abf-9791-edfbdb9ef77d · inbound

Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation cites this paper.

Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T07:04:44.923245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:04:44.923245Z digest=sha256:9ddf98657b698e710ad5d114aab21c306f7609b5faaf32000f5ad70deebe57d2

Observation 4923df3b-62c9-4fc4-8bfb-0c93c0e8c268 · inbound

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors cites this paper.

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:55.177952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T09:37:02.168898Z digest=sha256:f562463a144b21904dc334029ea4c657a40cb9919e9b9b25b56aa796bad8eead

Observation 181c552a-0578-4513-8992-5c66a83521f3 · inbound

Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation cites this paper.

Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:07.173665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:30:21.475036Z digest=sha256:72724581a32eddbca5f39e929a9baebc4583fe9a388190126ed655cb15a821f1

Observation 5129e6c7-3db3-46b5-a89c-b4c1535e013f · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.613444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:c7c7992dfb608c01b7b1de6ff4e745dc0e2d3cee60cfc7518b2fb46637ee4d23

Observation e36865a4-f893-4107-be7b-acd44ab738e4 · inbound

LACE: Latent Visual Representation for Cross-Embodiment Learning cites this paper.

LACE: Latent Visual Representation for Cross-Embodiment Learning Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:42:48.316408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T21:38:33.937153Z digest=sha256:c53ce767f74d55384603f2f6ac3d39b3817aaf8c187aab09d1346b92a7cc3d1e

Observation c1c81271-8e74-4de3-a5e8-286af5926ef0 · inbound

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization cites this paper.

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.241421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T10:47:33.591670Z digest=sha256:e1929dc51035e0649a749bf9aa3d40aeddce3b75e046a6fff07e31ba8078bfac

Observation f9a1dc47-02e1-47ac-a0b8-0e32d2a9388e · inbound

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation cites this paper.

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:29.281950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:28:41.952330Z digest=sha256:df94f470480bc60433307596d46e49703b526b5de7a473718ccda1e9292f728f

Observation 0b1cabc2-4a61-4189-a742-e4820622a1cd · inbound

Action-Effect Memory Pretraining for Robot Manipulation cites this paper.

Action-Effect Memory Pretraining for Robot Manipulation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:48:04.840744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:22:17.435427Z digest=sha256:dda592476a5600cf63116806398964dac4b3687123066815e11ca5773eaed999

Observation ba4c843e-e302-4175-beb8-53ae8518e5cb · inbound

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning cites this paper.

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:29.582106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T18:25:13.494219Z digest=sha256:5c0e4066eeacf23fe75b534111ef52f482b5bd9ac4e46d52cd837e3ecdee1d48

Observation f41f797f-ca3e-441a-9b29-09198ee0bab1 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.210298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.210298Z digest=sha256:6a2b4278c99123136a665d63b93e3ab60bfb53c085f26b11d63ca49f82381db2