Pith. sign in

Paper Citation Record · LEDGER

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

As of 13 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07314 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:29:55.365544Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22fd29ac-b83c-409e-8d90-27297fe18a37 · outbound

This paper cites PaLM-E: An embodied multimodal language model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models PaLM-E: An embodied multimodal language model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.195731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.231185Z digest=sha256:d123f67995f39925305b9e3c10639f3e6875ea00775ff8d6ffd7e75dc1709b3a

Observation 3fa02969-224b-4744-a1ae-5da4257e0daa · outbound

This paper cites RT-1: Robotics transformer for real-world control at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.186044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.235069Z digest=sha256:d201c36922a958d9b8b95a8a724f83b45a501f99705d0a48a7eba8a91d580af8

Observation 29a27b32-044a-4c8a-903a-48e1d3ac7059 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.175689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.238673Z digest=sha256:8160795229cd74e0faaa120d7dc87fe47c60922f85cd83ec49cde46f65585b34

Observation ea4da3e9-37b1-4529-9654-dc7e9ac20716 · outbound

This paper cites Open X-Embodiment: Robotic learning datasets and RT-X models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Open X-Embodiment: Robotic learning datasets and RT-X models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.165594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.242091Z digest=sha256:29cadd8ffe29e8c2596f755221d7681a65800c2366deca5ae335fe211b3b7c50

Observation 56230c39-9f3d-4b74-b647-a0e692aa832e · outbound

This paper cites Octo: An open-source generalist robot policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Octo: An open-source generalist robot policy,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.155897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.245579Z digest=sha256:60698df60b9492938051be5995cd1e07de19b839406578786887a226d13c12a3

Observation 12848eb0-1192-4339-b31f-2d29fe98b861 · outbound

This paper cites OpenVLA: An open-source vision- language-action model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models OpenVLA: An open-source vision- language-action model,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.147186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.248925Z digest=sha256:fbfdfe8eb7fe1c4a223815a4a60ca96e064cd7d997d0e997905475ed178f9f52

Observation bf801935-81ea-416f-8f26-aec8e029f257 · outbound

This paper cites Fine-tuning vision-language-action models: Optimizing speed and success,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Fine-tuning vision-language-action models: Optimizing speed and success,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.138028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.252088Z digest=sha256:bdd499da1f3e4ac8e31387da882ade7e3b1f76822f3713828834f437b7fbe106

Observation 0f01f29c-4545-49d7-95f6-b4fc8b577c71 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.127108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.254959Z digest=sha256:6c9622a66e149cc2fca00930799627196ab8ba542a79ee9dd01735f3c6b14bca

Observation 5476859f-65e1-4d2b-adcf-a043a72076f6 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.257629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.257629Z digest=sha256:aecd3e789bca430d8a05329f31e7c8d2d04d61a4d74cc2569c10b5e2606918d7

Observation f73d1b42-aa20-4d12-8ef7-f7e48bd7899f · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models SimpleVLA-RL: Scaling VLA training via reinforcement learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.116999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.260771Z digest=sha256:f99dc952973b0dabf916e7ad90fced54ab29c608fbe19c54b78e6cbb77681446

Observation 34301185-637a-412e-ab6e-6c5c09477126 · outbound

This paper cites RL Token: Bootstrapping Online RL with Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.263923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.263923Z digest=sha256:2f6d95ab5ec7b423b958717ded0709ad00b6ce85f6fcfc101b00bee51099c6fd

Observation 9f211459-f700-4493-a8cb-9098f4574656 · outbound

This paper cites Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.106712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.266959Z digest=sha256:19c4a3e39deba07c0d3d048d9546102c8da39d8653ddec351b345c457ab8394a

Observation edfee9b4-2bab-447b-8495-e59005959870 · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.270093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.270093Z digest=sha256:48478f967820c6c9522dd999b4f821a370eb4ae3fd3cda878ffdd18c8c28f6cd

Observation a5f798a9-de3c-464c-b2f5-7123f70a7771 · outbound

This paper cites Addressing function approximation error in actor-critic methods,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Addressing function approximation error in actor-critic methods,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.096260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.273617Z digest=sha256:b0ccc9894ebe9c0134c224d60a1700a4158a11fede61e945569ae15404636a5a

Observation 7668e1bc-5b3f-4710-82f3-f085ce17c84f · outbound

This paper cites BridgeData V2: A dataset for robot learning at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models BridgeData V2: A dataset for robot learning at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.086289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.276679Z digest=sha256:85457a1b6b3fccb850b25d07a350b4c4de94354182e31ff421254b5aeed727c9

Observation 29ca5594-d6bd-4964-94df-c05bdf3801a0 · outbound

This paper cites Vision-language foundation models as effective robot imitators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Vision-language foundation models as effective robot imitators,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.279752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.279752Z digest=sha256:70f2b4e80c9a74ae1f75e47758a2fb9514067dc773dd115244112d5863afc9e6

Observation 4b7bdbb6-6c16-4379-8ffe-24ff1d0f6af1 · outbound

This paper cites Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.915698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.283014Z digest=sha256:c84045cce84cdcb796b157cccacbc0c32bc8696dec1f5b24aabd093a9ee22e50

Observation 6dadfcca-bde9-4ada-9e09-db4045287d65 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π 0: A vision-language-action flow model for general robot control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.070341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.286220Z digest=sha256:814cb28bfe92f9bd9da975a9deae5a7c28074d1728db7bd4e6fb3b5a23b8c0b0

Observation 26618d8a-a725-4ff9-845f-2dafac8509ee · outbound

This paper cites π0.5: A vision-language-action model with open-world generalization,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π0.5: A vision-language-action model with open-world generalization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.060558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.289517Z digest=sha256:18508b0e1288dc80c94cac60155658aa578627573234ea359976c964fb6b9b56

Observation 87607237-d798-43ab-8d8a-065063327d4b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.296092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.296092Z digest=sha256:b06e134cddd0638511dcd41b3896c2df895d17e3f775d54fcda09f96b46c52e8

Observation 3066997d-de2d-47fe-ade4-6bcea7246a0e · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Interactive Post-Training for Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.299516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.299516Z digest=sha256:5023d6221a6e66b364a19749e3c08bad6a322487e0eca087f55507cf303ce247

Observation eec1aa47-6d89-4fbe-bd06-e161450369d4 · outbound

This paper cites VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.040569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.302904Z digest=sha256:f0c1b7f9a81ad048784e91aab362f395d64212ee394fb2e102f7a25932b6a371

Observation 8599fa7e-bda9-4196-9164-a095a7ad8298 · outbound

This paper cites ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.030298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.309572Z digest=sha256:b105eaa466a095bec1eb913d7720191c8dcd82dacd45ba23518f76aa9dd97ab2

Observation ef126e3e-aeae-4b5e-8752-2cedf4b0b831 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.312519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.312519Z digest=sha256:2b0d23ee8a73a5c33f7920b441b2864a34cbaee545688650403859e329a3d50b

Observation 6eb70623-a93a-45c2-8578-e2e368f4b264 · outbound

This paper cites Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.660050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.316233Z digest=sha256:43799c4c4815536764290b2a389f1fa6b4b9c5bc2637ec6af6978173366eda19

Observation bff91253-b426-414e-92fa-ed83fc42e6f1 · outbound

This paper cites DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.319520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.319520Z digest=sha256:449e4e37f93d1b87033599a580de21a54893236109e488fccb1a33bf76489d16

Observation 1595bda8-f98a-4546-a9fd-5b91cd3c17be · outbound

This paper cites RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.322806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.322806Z digest=sha256:0351ae07e6d243ae7a68d337a2415a8a367c9f8cfc2c998917fcca33a095547f

Observation c9190653-2b3a-4c89-be44-49d47b8d0d05 · outbound

This paper cites FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.020826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.326214Z digest=sha256:1990e4389dc85f3910b02b443bfc8bfb269ca797b6b62ad8f922df4e0463fbe7

Observation df257f01-d68e-4020-86cb-20df0eadc475 · outbound

This paper cites CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.010337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.329247Z digest=sha256:947a1919146b004136915607ce19d8a4f713c31782d660c24df1c61e6058464e

Observation 070e97bc-b8ca-414e-b4a4-7b62c0b62d05 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.332304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.332304Z digest=sha256:9853ed36bd0af0ce3d10d281b5ca1351f7a8df27ae6e60d27516e0775ea524b6

Observation c0d5ca24-f249-4b16-aee9-1bbd471e85db · outbound

This paper cites Unified Vision-Language-Action Model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.335648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.335648Z digest=sha256:0f46b1a89a077d1f3769ff01bfc6a2f606b88f7db1a013f4c27e7191fc93ac56

Observation bcf3a1ed-5608-4cf1-85df-33d4ca959e10 · outbound

This paper cites Disentangled robot learning via separate forward and inverse dynamics pretraining,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Disentangled robot learning via separate forward and inverse dynamics pretraining,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.999192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.339006Z digest=sha256:c59e43ebe0fe8dd090fedf820eccc841f7edfb7985db412d4d54c6f009acff8b

Observation 494b723d-3fc4-4c09-be34-5b9b3e44242f · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.342197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.342197Z digest=sha256:95e1013220430c457c6476e07cc0c8cf62ba7d2bee31e2b5f73d35127b5fb96d

Observation d3f1a906-5701-4748-890a-26e4d2cd3052 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.345462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.345462Z digest=sha256:4c3907590a5e321d6ca3624208e93333c053d1cff2d577c52637990f59c00faa

Observation 151c7e52-33d6-4715-9660-055547a9ecc6 · outbound

This paper cites Closed-loop visuomotor control with generative expectation for robotic manipulation,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Closed-loop visuomotor control with generative expectation for robotic manipulation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.988461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.348290Z digest=sha256:865aa47b7523e59ebbea1221b41d5c77e548b504953c2d414baf17c7eaa20eaa

Observation 57f9af1d-d3b1-49dc-8040-fe607b8cfead · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.351046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.351046Z digest=sha256:0e5a820e6e65668d410ae35837f262c3bbfef9b4f0ac54085ed859f8a2f076b6

Observation 329ece0c-ce06-44bb-bb75-991b50ca45d5 · outbound

This paper cites UP-VLA: A unified understanding and prediction model for embodied agent,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A unified understanding and prediction model for embodied agent,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.977851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.353821Z digest=sha256:f9716cea3295e241835ab55a7b2243e303a320a78370c7454182a7adde61056e

Observation 995e8f22-1a8e-46e0-84bf-941c89e41238 · outbound

This paper cites What matters in building vision-language- action models for generalist robots,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models What matters in building vision-language- action models for generalist robots,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.966734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.359366Z digest=sha256:647a05237f892331fc5f71dae9cecbbc2e8d0ddf7a056c8cf8af9da709732605

Observation b1c35929-4741-4869-a527-3e1902b19ff7 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.362007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.362007Z digest=sha256:8bda3adcf48385c5558a273fffe0cc998029569574d3af124efd5b95c16ec566

Observation 76e909f4-f93f-4e90-991a-6a937ec19a6b · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.356467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.356467Z digest=sha256:43b6b9600619b0a6b2c9834048967d41905987f504135297f3140f8aabf58c36

Observation 07711bc4-1625-48fd-b1e5-c6201cd13b93 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Video prediction policy: A generalist robot policy with predictive visual representations,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.955746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.365544Z digest=sha256:fada7752225d3d9146b87d28a2068f7c700f77eff2b758edc28e984e449d9ff3

Observation 7294401d-902b-4fdd-bf7c-f2e75b9fb205 · outbound

This paper cites an unresolved cited work.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unresolved cited work

Reference 305

Resolution
unresolved
raw_fallback, observed 2026-08-10T10:29:56.049969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.292911Z digest=sha256:ea7e50852c925c6964d070300d7a71c00fc81f93ddae83c561ba47216e647b9d

Observation c57bf128-f41f-4721-82f8-47351af4de48 · outbound

This paper cites Available: https://arxiv.org/abs/2510.00406.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Available: https://arxiv.org/abs/2510.00406

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.306328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.306328Z digest=sha256:955cc2ef68371581362dff0c7fd9f12091298d68c9286efd84e62cb1f9376861

Pith citing papers

No inbound Pith citation observations are available.