Pith. sign in

Paper Citation Record · LEDGER

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2605.31234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.31234 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T22:04:09.296855Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:31.334140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact25
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35ba7f4b-76ac-43b1-8817-eabe773a0b30 · outbound

This paper cites Latent Action Pretraining from Videos.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Latent Action Pretraining from Videos

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.862338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:8699ce907c127d30a2aebbf5a09c089850c8c3b460dcc69fb53a07e2ad0689f2

Observation 32aa5940-6bd1-4895-baeb-e3d3a16511a7 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.880387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:9169015f3bc24e3b74977a8d311e105db2e14b02abde40b5aa44ab7ca50d9549

Observation 7798bbd5-02ca-4d48-a471-b698704aeccd · outbound

This paper cites J., and Lee, Y.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model J., and Lee, Y

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.857254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:05aae2c5c84897366a69c189ced535e3472bf193d345af92a5014c4d78208809

Observation 43bdcb58-c2db-4c92-8681-e5411df3c993 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:a79999a02468ad3d18d1aa14c9f5efe6dc025f52a7f8c13e56bea6a0dfcee43b

Observation 064d0ce0-74bc-4ff7-8b8e-abfaca8ea110 · outbound

This paper cites Kareer, D.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Kareer, D

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:eb47bdf1c7123074ad2543860e191a10bf1d3087bd266d2bfd3b42e935255480

Observation fd2b0dbb-915d-462e-a9a2-8558caa50aa8 · outbound

This paper cites arXiv preprint arXiv:2509.22199 (2025).

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model arXiv preprint arXiv:2509.22199 (2025)

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.860090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:e8895c60abc8fbd3ae1b58d6cbe9e48c77111b4f44177fd8602ebb4630994e3c

Observation 6a093d09-7fb5-4210-a04e-d5fdfc8e3699 · outbound

This paper cites Dexumi: Using human hand as the universal manipulation in- terface for dexterous manipulation.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Dexumi: Using human hand as the universal manipulation in- terface for dexterous manipulation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.864905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:8024192c27a24e57f2947baa1274bf7c64b7207cce844c3a80c19c5807843c5f

Observation 42fc5a09-bc12-4092-b6b8-cc41e2272440 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:456a8f776f323cac1a9bfd7448afd51e246d6a5ddd7a85958ce0d1534c17b24d

Observation d194bc2e-f932-440a-a3fe-6bbec9073162 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model OpenVLA: An Open-Source Vision-Language-Action Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.864485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:657ef51e33d42be61f52264c73525c2ed060076c0f080f3f9494da4efc87d9ee

Observation a03984bb-8b93-466d-b361-620c6c29afe8 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.876832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:bc71a9d7e59ffb50933aaab892971259a102c81d7c3f879082bc38476ab9a4ea

Observation e01e983c-4254-4ae8-a537-c8858f3b617c · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.861914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:179b68a159d396ffe320045f909b38b83aec06dab57860b49edf1a96bd739b18

Observation 1c0df870-7193-4d33-bc0d-30c4d0db2665 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:5833b05facb22b6e50fa7e3453d5d5fabec533cfb66561eb97c77194f411383f

Observation 93e23b68-3d6b-4038-a72b-1225fb19e32c · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Octo: An Open-Source Generalist Robot Policy

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.859283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:a44761619a098d71189d75533e2202982bd854809a3653b894ba1c09dc88637d

Observation dc4e691f-b7f0-493f-a413-67f31826560a · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.878008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:1758a8f1ce6f9c16d77e1e316725ac21dc2eb6b00c73cead1e2c2544b2a94b67

Observation 0ea0dfb9-827f-4003-ba3f-9db05ebc6cad · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.869549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:b244178aa5de1694346bb7450e5e5d997ba501e4f9cd884fa38d2aabf048a59c

Observation 5879ffed-fee5-4786-8183-47c259e5ac3f · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.872735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:84f044e73382da37a33e7c9fab422d5033186aad803e0641d3ddd193d65e825b

Observation 70717396-c379-4ad4-be3c-847929f6a063 · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.828518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:362bb6f149513d68f6a563fc47cb87b304e53a4c465901ede96ee8427de569d3

Observation 36f83a56-0072-4c87-b74b-9317167cb482 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:b64450af280828e179db5feee57cfe62b29e1ecf0901408af5620815a8debf30

Observation 66491bf7-482b-402a-8085-bb8907fa9795 · outbound

This paper cites Xiong, Q.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Xiong, Q

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:670aa7c9abdac2bdc04e5d389c3f7adee37c0940ecad17adc3fae6b82331efbe

Observation ad161016-7bb7-4fe6-ad2b-bb91d39a2f9b · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Flow as the Cross-Domain Manipulation Interface

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.870093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:b945b133c4828dbeb32b90d0fdda907f2e4705c93641180fdc55e4ad7fc1d8eb

Observation 053c5aec-d184-4b12-89d2-c6a7fffd8c09 · outbound

This paper cites Dream2flow: Bridging video generation and open-world manipulation with 3d object flow.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Dream2flow: Bridging video generation and open-world manipulation with 3d object flow

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.879323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:c85f91dfc439d1e195e1059fac7f7a3f4bd56b6364e74d1a72fd641147184917

Observation a5149324-bddb-46b0-8e8f-b2ba96f0aca5 · outbound

This paper cites RoVi-Aug: Robot and Viewpoint Augmentation for Cross-Embodiment Robot Learning.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model RoVi-Aug: Robot and Viewpoint Augmentation for Cross-Embodiment Robot Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.875390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:6a91e6401ec249c14674d573a9222d928ee301a7e4b3838844a98346814a246e

Observation 0c48d54d-9fdb-4815-a340-8eb285135d4a · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:dbda4306a330e886a6b5583065f5d0c9df26d911a8ca18490a7e44cbde1b45ed

Observation 6f7be1f5-8fa4-4301-92af-9add02daba39 · outbound

This paper cites OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.853926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:03a6e35960512d7e25da0c8b89c0b053ba1b5c595bda9141a86e78529a097286

Observation d98d2698-de49-45ff-85cd-5f29bd1ba444 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:8984e8c235648bb667345610f22d672e46ee6d078948c5e551db191936e778bc

Observation dc173fc5-0de5-45e0-b9d2-7a24665d9ae0 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.866879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:f813e469f245d480697c9bd33ea22b55f1ac617da4a29d54d8a24b50ef3b9604

Observation 0528053c-90e5-4e9a-a3be-2f2609f06edc · outbound

This paper cites Human2robot: Learning robot actions from paired human-robot videos.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Human2robot: Learning robot actions from paired human-robot videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.840796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:4bdcdc791bd99a540f544213e534244a0d23a5c41805969673bbb39bed5ab68d

Observation 3d1d54bc-cab0-4e2a-8457-312c3ba68754 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:812901fc465f8361f19815091fa8cb59c92a5eed73dc1a333ff91c392a58ac95

Observation bcd5b816-09a3-4b20-9c41-c3958f8bc118 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:58a48615e5ad83512174a889a105c1dc924df4105e28c1c1d2fda53895596eb8

Observation f5e5bddf-bb26-453c-aaaf-b6070511d466 · outbound

This paper cites James, Z.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model James, Z

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:e4c636f132ae0b83635dd0729eb13a19a3828f7de5231d5de968dac720046002

Observation b9d80d99-ee8b-42f0-8c7a-ff17d73a253a · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.874408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:cd3d4b1e7fcd8afc0dc7b7346e57449ead83276d6589b563c2ff27d0643a8d7f

Observation 2fdecf46-b87f-4a73-8443-9f2a24b1aaad · outbound

This paper cites Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.871892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:986db8a78356e9f3aba77cb5aa316969bcbbe0d947732362d20fd752a23c3f5d

Observation cd3b6cb3-e617-48f6-9a23-d167f338a5e0 · outbound

This paper cites Qwen3 Technical Report.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Qwen3 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.828901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:7e3c91f2fe195e6b484153641e0ad8909a77b68d1294ae9524db560328fc368f

Observation 3d7ee676-4f8c-4a4e-9f65-577cb22c9551 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:075a5377a6320d0f89023a0f29d15e72128a45da2b3ae7147baa193e1127c07c

Observation 13b58a9c-ff64-4a76-9aa0-11893ee77d72 · outbound

This paper cites Doersch, Y.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Doersch, Y

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:02f30556dc4c013b0ea4496bec546e2f04f2f459657bc55d78bd12416b896b7d

Observation 9c1cc4fc-6f01-4c0e-9c69-abcbfed54641 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:fec95b0a0222029664e84397a22d6265b526a786fd1fba765b7075d57a516a7c

Observation e63e87b2-d476-41d7-898f-25fd56ec8bec · outbound

This paper cites Embodied Hands: Modeling and Capturing Hands and Bodies Together.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Embodied Hands: Modeling and Capturing Hands and Bodies Together

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.842693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:6ad913580beda564531385550f90ac99e479ebe835b8791d5c028a198e945588

Observation e9d5f882-8369-41c2-812e-47a56fd1ab28 · outbound

This paper cites Yadav and M.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Yadav and M

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:eedddde87a14b2fc9c50dda524d10beae57de9dfe53170e1758feb46046ab3fe

Observation aa901d56-61d9-4944-afe0-b4583f30dfaa · outbound

This paper cites Karamcheti, S.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Karamcheti, S

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:498e7b8e193c4ea261840342277e4f7e562bafd67763ff8bfdf4d867190d6560

Observation 47cfa813-6f33-4c28-a66f-e58eb02e9f15 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model DINOv2: Learning Robust Visual Features without Supervision

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.850895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:75d3d4e8edd38f6ef516a9a67117b33007b903e96b940c274265006b960c2f07

Observation 2fa1af2e-1e3d-43e8-83c9-0860a7a943a2 · outbound

This paper cites an unresolved cited work.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T22:04:09.296855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:386f8227f85ed3349f2bb6f2485411ed975b50d3d34ddc30e02866273b24e409

Observation 4a3a4bea-dcba-45ba-a093-82660c00d795 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.834458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:04:09.296855Z digest=sha256:12341e26916ca7cea35b8f389ca91052629da8ff86af9303946536209909e349

Pith citing papers

Observation 074e1e22-2d11-4cd4-b40b-49822b0b9199 · inbound

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models cites this paper.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.334140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.334140Z digest=sha256:b42e63546a9438458a64357497f77766c9bf95843d2118c996ac331115f2694d