Pith. sign in

Paper Citation Record · LEDGER

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2602.19710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.19710 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:37:16.250864Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:11:17.634768Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:39:17.424396Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7836cd2-72d2-4240-9beb-46ff0e3673c0 · outbound

This paper cites Objectron: A large scale dataset of object-centric videos in the wild with pose annotations.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Objectron: A large scale dataset of object-centric videos in the wild with pose annotations

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:09.276954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:09.276954Z digest=sha256:ec3385d7018ecc0ed11a5816c6d722c51488a909ab5a8fc6e9b9f3d2ae56fb54

Observation 13787cf2-8a07-4f77-812b-a6d664dcbc45 · outbound

This paper cites On the representation degradation in vision- language-action models.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies On the representation degradation in vision- language-action models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:09.389547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:09.389547Z digest=sha256:e18b33175f54ab1ec56ea1fa40ff09f3d7fabcdbcbb5b139b0383a14405f2c05

Observation dc1b9b76-3585-4d13-ad50-241b1b6d04a8 · outbound

This paper cites Qwen2.5-VL Technical Report.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:09.569609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:09.569609Z digest=sha256:4214a959b7535c0c01e0e3c7593b2ad10371335f88fb0c064e8840830174dba0

Observation 1f14761d-9d52-456b-9e60-da1e6a74cf26 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:09.716828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:09.716828Z digest=sha256:e907b92ea03a585fccfd008e49dc045376e5cb633657ce25c434159a43abd637

Observation e4356ceb-fefb-4b66-918b-6a8d9d2f8b4d · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:09.848953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:09.848953Z digest=sha256:7761345280cd98002a4d1309639a86ead2bedf2f11ee69fe25d714893bda7a19

Observation 50fa0fdd-f25a-4d27-9cdf-b650b250858a · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:09.982605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:09.982605Z digest=sha256:1af1d236a7f9ce56d54cf2007117c623ffb66c19d757a6d61b57d5726d84ca4a

Observation 32de1248-beff-414b-9a66-056fb6ba3064 · outbound

This paper cites In9th Annual Conference on Robot Learning, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies In9th Annual Conference on Robot Learning, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.171748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.171748Z digest=sha256:b909cee796b1c5e8abf0ecf093af6fb37fe87b1a0ad335501e225006c01a865b

Observation 4af5f446-d60d-4fdd-8490-c16efd2d18e7 · outbound

This paper cites Omni3d: A large benchmark and model for 3d object detection in the wild.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Omni3d: A large benchmark and model for 3d object detection in the wild

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.355369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.355369Z digest=sha256:bde3cf54cabfb816369cc3b31d996f38b7aa9d75cb879151c5ef12ce364562ca

Observation 3665d143-5566-4892-aacd-59f730cda58a · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies WorldVLA: Towards Autoregressive Action World Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.505000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.505000Z digest=sha256:b4963a65e334eb5537592383dda4048a9c1421d5035c196ce262cc141c8c5042

Observation 4e7fb85d-047b-4cca-a206-12cb859cff1b · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.660577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.660577Z digest=sha256:ac074d87e16ca6a0e311fa357e0b3b1ed23f3b480aa1e55c2dc35393a7de3722

Observation 695d346a-885e-4bbe-ab8f-f360d97c2d60 · outbound

This paper cites Training Strategies for Efficient Embodied Reasoning.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Training Strategies for Efficient Embodied Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.787322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.787322Z digest=sha256:69496c040366075266d87f8825ee84b087ba8f46efdeba085dc3d6a8d5dc55a2

Observation a825d5c6-361f-4032-84ec-2cb6bbbc3e42 · outbound

This paper cites Language-Image Models with 3D Understanding.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Language-Image Models with 3D Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.906144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.906144Z digest=sha256:97779430d9e9b5c41d26d01f0cca7f076aa3002c6213561e6a7b47522684d84f

Observation f395ef63-03ef-440d-867d-58d13819b45d · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.048220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.048220Z digest=sha256:0af450369fbe4059cfd97b47541a4566567455ae7cff7f9d0cbd25c8c8eb542b

Observation 332ec85a-a121-4926-8f00-0b3157c5210a · outbound

This paper cites Vla-0: Building state-of-the-art vlas with zero modification.arXiv preprint arXiv:2510.13054, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Vla-0: Building state-of-the-art vlas with zero modification.arXiv preprint arXiv:2510.13054, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.241225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.241225Z digest=sha256:020bc7fc0bc380deb097cbc9ce589fe018e56f7872b80ef0894788f867feaa02

Observation 9158f500-ae05-4265-8f33-ca70f58be589 · outbound

This paper cites Seed1.5-VL Technical Report.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.438697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.438697Z digest=sha256:f37280fd5624d9c2480c5c42ebe1ea6a86dd428080b7c99eeb770bfd746af98b

Observation d6bc258a-8622-4d21-b897-b3622abba245 · outbound

This paper cites Pow3r: Empow- ering unconstrained 3d reconstruction with camera and scene priors.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Pow3r: Empow- ering unconstrained 3d reconstruction with camera and scene priors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.571557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.571557Z digest=sha256:86c5299c179ba5758d506dc62fd90bde22b2ced69bb663b0ac2718f84a731dcf

Observation 36e77b2e-5337-40a6-a07c-f381598307c7 · outbound

This paper cites Don’t blind your vla: Aligning visual representations for ood generalization.arXiv preprint arXiv:2510.25616, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Don’t blind your vla: Aligning visual representations for ood generalization.arXiv preprint arXiv:2510.25616, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.691309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.691309Z digest=sha256:b682412063241e0b5fb02b4f7d6179df5fa6c644f98e0f85b7004a5da1416d07

Observation ab137981-81d6-46f0-8282-1c87ec346ec8 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies OpenVLA: An Open-Source Vision-Language-Action Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.867178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.867178Z digest=sha256:94d1656de1c975a44aab3709994ff134fb4ac76339fa06408b872ec2d8ecbc53

Observation 608225c9-c9d3-437c-9ee3-100846ec7415 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies MolmoAct: Action Reasoning Models that can Reason in Space

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.038367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.038367Z digest=sha256:a6bffc13a79d3fab19aed432360b595e1d88d43118267e60b78abfd50938a171

Observation 8e19ce90-9856-4c74-8aa1-fab6df6129f2 · outbound

This paper cites Spatial forcing: Implicit spatial representation align- ment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Spatial forcing: Implicit spatial representation align- ment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.262166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.262166Z digest=sha256:582fabd9ebdc17b0a9c441a5de1392711a8653e75cc638f5d13922c1810c6380

Observation ff1c7ab5-0453-4562-ab5d-224f51eef893 · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.445617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.445617Z digest=sha256:cf557faa8414c77596bfb65e1e5e45a140f00e6bfce9914aa32f2bd4ca435361

Observation f9490399-d3bd-4063-bc71-c22b33f310db · outbound

This paper cites Onetwovla: A unified vision-language-action model with adaptive reasoning.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Onetwovla: A unified vision-language-action model with adaptive reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.615594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.615594Z digest=sha256:1980b17e0f746a778e707e79ba9ccd7b795a5828bf859bcd6281c0c5cf2c707b

Observation 6855142e-b978-4985-9626-8d6b5cb7039e · outbound

This paper cites Flow Matching for Generative Modeling.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Flow Matching for Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.715620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.715620Z digest=sha256:6ba2f5a39f6a0740a2f403daf17227a558988dda0b76c85f77ade0464f6d90c8

Observation a6781868-2bad-440d-a316-42773c16ae41 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36, 2024.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.837474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.837474Z digest=sha256:480003c2d3f0b8c8fdad54aa8ba8db59ab40ae3f79de56e76c9c8fc95877ec9f

Observation 00ce6743-9a5f-41dc-8a83-b2bb6d4a9225 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.899050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.899050Z digest=sha256:55e2f7a8d2fb217e0690d88775558f85913640faff562e79bcefaa85c4e5b891

Observation 4041a377-8649-4145-baa2-65e6f5f33f11 · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:12.941124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:12.941124Z digest=sha256:af936d5b9a6e6680747d01b9bd1a37dbc7ab28fe119f8a5bfdb9adbe719c1148

Observation 89b1aa89-e22d-4ce3-a247-bbb8a6ffa1cd · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:13.125283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:13.125283Z digest=sha256:4dfc469943882119abd36fe3ab19c444ed19eca8671c409acabc02678ebc4407

Observation a0d59919-a6a7-49dc-8105-3f2c08e04de7 · outbound

This paper cites SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:13.241240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:13.241240Z digest=sha256:3bb62671b9b975fbd504ea7cbfcb8d2f38b26dc300757ab3ff4ff105c726b925

Observation 50fa8582-1f65-4fa2-8cc4-bab0e4de3690 · outbound

This paper cites Locateanything3d: Vision- language 3d detection with chain-of-sight.arXiv preprint arXiv:2511.20648, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Locateanything3d: Vision- language 3d detection with chain-of-sight.arXiv preprint arXiv:2511.20648, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:13.473981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:13.473981Z digest=sha256:7fa70dce8dcf76476225728e4e600d93d64f2bd0430d1ccf5721840f2ad25679

Observation b02e4c95-c70a-488d-b79e-fd43c711c706 · outbound

This paper cites Spa- tiallm: Training large language models for structured in- door modeling.arXiv preprint arXiv:2506.07491, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Spa- tiallm: Training large language models for structured in- door modeling.arXiv preprint arXiv:2506.07491, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:13.627124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:13.627124Z digest=sha256:7636fff2c9e6ca9c6da94c0e0fc56d53891bba9610b73dfca878c7088fbd47bb

Observation 58dc68ba-727b-4dc8-a9e4-3b0ed97a74f7 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:13.747288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:13.747288Z digest=sha256:f38170d2a8554bc0e80f1686dd318b5a025a1d39bcbd1bc2ab41667ef871c084

Observation e2e84639-41aa-45db-95b8-40ce6cf16bc0 · outbound

This paper cites Eo-1: Interleaved vision- text-action pretraining for general robot control.arXiv preprint arXiv:2508.21112, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Eo-1: Interleaved vision- text-action pretraining for general robot control.arXiv preprint arXiv:2508.21112, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:13.975339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:13.975339Z digest=sha256:95c425e2e7a271e191150e28682b5b454dcdd555331cda1ebee1db6034cadd0c

Observation 2a9f7163-7f55-49c6-b515-1e2d44d64994 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.096872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.096872Z digest=sha256:2589942d8e808b747f97814ddcd62ee481c2f52f3f89853eb43f5b673baed1db

Observation 0d7908e9-4133-462c-8987-902a9cf46eea · outbound

This paper cites Qwen3-vl: A frontier multimodal large lan- guage model.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Qwen3-vl: A frontier multimodal large lan- guage model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.213063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.213063Z digest=sha256:cb5ca3e6cfe025eb2ce7f5dfabd7c8058f3a5c66e2d61d8eba30f62f6abc26cc

Observation 70d727fc-7eb0-4d66-a788-1f2b3873e63e · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.453522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.453522Z digest=sha256:3ca379c9f63fcb08c062500679d2ae66c773aaa9f0ab9416eb069b36dc1f0d11

Observation cbb2ba4e-6064-464e-9287-655ed5bf425e · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.611612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.611612Z digest=sha256:9b2f5abd2c20eefb6e8212a0d2c2a45303fa6cba396c6873ff62b379b377575e

Observation 1ec095cf-0eec-4824-bdbb-0def141888c1 · outbound

This paper cites Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial rea- soning.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial rea- soning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.725557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.725557Z digest=sha256:385c4700910b5043d31665dd5d49c9150023f50616831886d20fc6d938910eec

Observation 1212fad9-76b7-4c03-8e71-0e9887b77dc2 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Gemini Robotics: Bringing AI into the Physical World

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.891831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.891831Z digest=sha256:7891893ffe48b80bea0936c3699df344fc7dc2451f03f80b0b0ae4779d886f09

Observation 99aa2b4a-aeeb-4118-bade-36b97fa5adf2 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Octo: An Open-Source Generalist Robot Policy

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.956373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.956373Z digest=sha256:443c614b494ecb176d39e625815e2dde15d8cf20b4f55f0921bb7008de6adbd4

Observation 408a266b-fe64-4d6d-a97c-0ced98d15a1f · outbound

This paper cites VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.978308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.978308Z digest=sha256:ae346bb56031595bf80e2b4bf7c6e676fa0af574c3002132a3b9fd3f560a8ac6

Observation 5a9ff308-f506-4843-ada3-067a6187d278 · outbound

This paper cites N3d-vlm: Native 3d grounding enables accu- rate spatial reasoning in vision-language models.arXiv preprint arXiv:2512.16561, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies N3d-vlm: Native 3d grounding enables accu- rate spatial reasoning in vision-language models.arXiv preprint arXiv:2512.16561, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.087357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.087357Z digest=sha256:4b882846c02fd0acc97a5df8787cd3b09ba7a9a11cd32284047ad04664daf849

Observation 3354e406-cad3-44e4-a168-aa6be5881022 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.172508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.172508Z digest=sha256:1e6631f829098e6147f8516c228460a1317519bcd5aa5b2cab5950aa1cb9df46

Observation 0b496158-cb66-4d09-bcd6-0b6d266b69c2 · outbound

This paper cites Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.251491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.251491Z digest=sha256:7e793181b12a8aa178a54dc3ccbe9509694a1d82257f81ff9bd935052a696c2e

Observation e8597b77-d3c7-4985-b394-26b413f23c41 · outbound

This paper cites Instructvla: Vision-language-action instruction tuning from understanding to manipulation.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Instructvla: Vision-language-action instruction tuning from understanding to manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.332694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.332694Z digest=sha256:d5feee0127371215c5c1324cc6b13379c2a904e321a8e5eca02579b6ba8497bb

Observation 1c0049f5-38d9-4087-ba75-7af585525d4e · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.434643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.434643Z digest=sha256:be5bde18f0a6956e8cc5d307be863f0996dfc182583418007558afce778b3dc4

Observation 98f4fbad-b8c3-4acc-9277-2b09dd81eb7c · outbound

This paper cites Sigmoid loss for language image pre- training.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Sigmoid loss for language image pre- training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.598305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.598305Z digest=sha256:556bddabf4a1c14f47f51d9a7aeaa9faf24aa07538a289858c66cc4a1b009fdb

Observation 18d1f9c5-29de-4ab9-b147-807980952787 · outbound

This paper cites VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.741346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.741346Z digest=sha256:90b611ef0ec6679268afd4dafc5e41c763bf6fa3671ef850c5b29ae833058e42

Observation 5045d1ff-c7a7-44e5-9a3c-0d2322629782 · outbound

This paper cites Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.903815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.903815Z digest=sha256:33cdd13fb7afab75d585915e1797bfafef589f1aad842bef2345733dec4bcdf3

Observation a7774f0c-7958-4a18-9b98-868e92ec8d31 · outbound

This paper cites Cot-vla: Visual chain- of-thought reasoning for vision-language-action models.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Cot-vla: Visual chain- of-thought reasoning for vision-language-action models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:16.012864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:16.012864Z digest=sha256:5b2499930ba14e22b7a05b882334ca0563a9d6908effcdfaab4233ce370257a5

Observation 48770eb8-b9d6-4007-856f-b98651a35fdc · outbound

This paper cites Chatvla: Unified multimodal understanding and robot control with vision- language-action model.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Chatvla: Unified multimodal understanding and robot control with vision- language-action model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:16.137591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:16.137591Z digest=sha256:3fc018fe4e1e64435280e4f55c0558d6d8eabbe03ee9b732a1847af88ab5bee1

Observation 47a1e215-21a2-49a5-982f-4a1c8d8ce4a7 · outbound

This paper cites Rt-2: Vision-language- action models transfer web knowledge to robotic control.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Rt-2: Vision-language- action models transfer web knowledge to robotic control

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:16.250864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:16.250864Z digest=sha256:1eb19f34224590d57e7eeecaa1a6851c740ddc9a232950d92b2e8f6301901fe5

Observation 6cf70152-9102-4e65-b5fd-273de537d6d3 · outbound

This paper cites an unresolved cited work.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Unresolved cited work

Reference 2025

Resolution
parse uncertain
no resolver link, observed 2026-08-02T21:37:14.293501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.293501Z digest=sha256:144d6a5de7cf3ea746e19cb672cc215f996175e46e0937ccc132e6b55c25eac9

Pith citing papers

Observation 14adc3dd-c331-4945-844f-7b8981d089e0 · inbound

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models cites this paper.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.870836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:439f67a9d06605b31395c833adca35c94a699dc51371da254984085a71dac0c8

Observation db88cc04-c30a-4c28-864c-8cfed2513b62 · inbound

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? cites this paper.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.425754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:e5fed64fc0756ed8668d45c7a0b853edba61b613bc2b5b7a84c8b527904ec884

Observation debb4867-1ec5-4d65-ac42-eeedc4ffc602 · inbound

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions cites this paper.

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:17.634768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:17.634768Z digest=sha256:a8731730d249dac22fb92992ba1e54deedac95c886e7fb0a9ac827eed82094cd