Pith. sign in

Paper Citation Record · LEDGER

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

As of 14 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07314 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:29:55.365544Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22fd29ac-b83c-409e-8d90-27297fe18a37 · outbound

This paper cites PaLM-E: An embodied multimodal language model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models PaLM-E: An embodied multimodal language model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.195731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.231185Z digest=sha256:5099e61d066b0cd2b5d7d7e5ec6da1ae665e4d95ad3909a7214b45d4c8ca2b26

Observation 3fa02969-224b-4744-a1ae-5da4257e0daa · outbound

This paper cites RT-1: Robotics transformer for real-world control at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.186044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.235069Z digest=sha256:a12cc118fca9cc37e79d52af1eebd9ca75a8c52e1b2f9b1920d6496846b3e0ff

Observation 29a27b32-044a-4c8a-903a-48e1d3ac7059 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.175689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.238673Z digest=sha256:01c3b164de6da340fc1c7ae12831e6353b190bbeb32591f2ee06b52d4d04a69b

Observation ea4da3e9-37b1-4529-9654-dc7e9ac20716 · outbound

This paper cites Open X-Embodiment: Robotic learning datasets and RT-X models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Open X-Embodiment: Robotic learning datasets and RT-X models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.165594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.242091Z digest=sha256:cfb1c33f649fb38abe36a1977d53ef1cf4b761084346cfb858d166531a5687d4

Observation 56230c39-9f3d-4b74-b647-a0e692aa832e · outbound

This paper cites Octo: An open-source generalist robot policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Octo: An open-source generalist robot policy,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.155897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.245579Z digest=sha256:32a4c999025153129f3c763573ddf3ef61a80de95a36e5e8b52b13bef066b376

Observation 12848eb0-1192-4339-b31f-2d29fe98b861 · outbound

This paper cites OpenVLA: An open-source vision- language-action model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models OpenVLA: An open-source vision- language-action model,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.147186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.248925Z digest=sha256:b5f5e714f93a5a347c3e87fbca2ff37424b9b671a858c7047b392de41120ae9b

Observation bf801935-81ea-416f-8f26-aec8e029f257 · outbound

This paper cites Fine-tuning vision-language-action models: Optimizing speed and success,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Fine-tuning vision-language-action models: Optimizing speed and success,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.138028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.252088Z digest=sha256:691f9c9507f6e84277ee60fdb20fc1bfe5f45a3cce8891fa125937533bbda555

Observation 0f01f29c-4545-49d7-95f6-b4fc8b577c71 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.127108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.254959Z digest=sha256:81d2afe0efb2047988775a9619aa9f4d4544a0bc2e7d3eeaa977532bd0bbaa3e

Observation 5476859f-65e1-4d2b-adcf-a043a72076f6 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.257629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.257629Z digest=sha256:a053b1ae9c95f5a65ea884e995928726bc097cfe87f60eeea7d05604de479b14

Observation f73d1b42-aa20-4d12-8ef7-f7e48bd7899f · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models SimpleVLA-RL: Scaling VLA training via reinforcement learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.116999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.260771Z digest=sha256:4718ee959aeee6138124f49e24a4a07b6cbe94a2ac573a1fa3cd7c68968dc205

Observation 34301185-637a-412e-ab6e-6c5c09477126 · outbound

This paper cites RL Token: Bootstrapping Online RL with Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.263923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.263923Z digest=sha256:72ff5fb5c3779460fda2c8c5c20981f53d5bdfb084fd93df3f22910acc6a9348

Observation 9f211459-f700-4493-a8cb-9098f4574656 · outbound

This paper cites Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.106712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.266959Z digest=sha256:1e9a56182af23fdfe6cdd96769336d85957e68df79f0fe4fcac17a3a71061420

Observation edfee9b4-2bab-447b-8495-e59005959870 · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.270093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.270093Z digest=sha256:7f19ec01bc9e7b18168ba3ac1eda56618d205c98b250b16800459ed392c931b0

Observation a5f798a9-de3c-464c-b2f5-7123f70a7771 · outbound

This paper cites Addressing function approximation error in actor-critic methods,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Addressing function approximation error in actor-critic methods,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.096260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.273617Z digest=sha256:b09cd6f9a6080a521fc33e815a1954d99ac4127b0b1f62e5de8edacacff06866

Observation 7668e1bc-5b3f-4710-82f3-f085ce17c84f · outbound

This paper cites BridgeData V2: A dataset for robot learning at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models BridgeData V2: A dataset for robot learning at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.086289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.276679Z digest=sha256:79b3f1d9f595d76c16a25336d0cbbc534b189126dcc00d63886505ea3b2eb136

Observation 29ca5594-d6bd-4964-94df-c05bdf3801a0 · outbound

This paper cites Vision-language foundation models as effective robot imitators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Vision-language foundation models as effective robot imitators,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.279752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.279752Z digest=sha256:c0e0f4783bd05f0f0ee066cdc3e5ccaf6fac7fc7d7478c6570631538c4769867

Observation 4b7bdbb6-6c16-4379-8ffe-24ff1d0f6af1 · outbound

This paper cites Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.915698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.283014Z digest=sha256:b0a15d3f44a858cd2dd2133fe70e35e4182443f2de07eb2eb43c6043f533924e

Observation 6dadfcca-bde9-4ada-9e09-db4045287d65 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π 0: A vision-language-action flow model for general robot control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.070341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.286220Z digest=sha256:aa43e237ac996db8bff023b1d402f48a7effa4b0f568bf6755a8629626e107d5

Observation 26618d8a-a725-4ff9-845f-2dafac8509ee · outbound

This paper cites π0.5: A vision-language-action model with open-world generalization,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π0.5: A vision-language-action model with open-world generalization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.060558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.289517Z digest=sha256:734cf563643d049dd14b25bce9f5091f3f4b95284cf692c615046bfec9a2f431

Observation 87607237-d798-43ab-8d8a-065063327d4b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.296092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.296092Z digest=sha256:5b2ba7347d2aa358e0bc4a92ee6568dae591dad313252b75ffccc7195da8d08f

Observation 3066997d-de2d-47fe-ade4-6bcea7246a0e · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Interactive Post-Training for Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.299516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.299516Z digest=sha256:bc8c4d20b259e545b7b9f3be2c3ad6f6859b0c9271ec7815857ce67f607c8713

Observation eec1aa47-6d89-4fbe-bd06-e161450369d4 · outbound

This paper cites VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.040569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.302904Z digest=sha256:7c7bb4e52d90e27f36e58114d7963e236bf0349cd4e3cf25b300c7c89c6d1eff

Observation 8599fa7e-bda9-4196-9164-a095a7ad8298 · outbound

This paper cites ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.030298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.309572Z digest=sha256:6c239188d67bb717c13044d8700772a0a9e19b66d37948e50d485d55490eff3f

Observation ef126e3e-aeae-4b5e-8752-2cedf4b0b831 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.312519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.312519Z digest=sha256:e5f8ae9cede005f05abd747cb6a759ad68cf3b05f4b94a477831ac601ce9a88a

Observation 6eb70623-a93a-45c2-8578-e2e368f4b264 · outbound

This paper cites Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.660050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.316233Z digest=sha256:9cdc2ace56275cb154d5bff183fdd23e5255a2d07ffe7f9f2bfdc8d6f8d5db87

Observation bff91253-b426-414e-92fa-ed83fc42e6f1 · outbound

This paper cites DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.319520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.319520Z digest=sha256:638a0e077b1171d3f26c451c233c0380ea4f7f0af2d627a91b391b079850979b

Observation 1595bda8-f98a-4546-a9fd-5b91cd3c17be · outbound

This paper cites RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.322806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.322806Z digest=sha256:c2194292085b77a8d38f2552bc2f8654e1eab82f3bf04b88079e33f7d1c5ef5d

Observation c9190653-2b3a-4c89-be44-49d47b8d0d05 · outbound

This paper cites FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.020826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.326214Z digest=sha256:a094cf9c751de0f5da0f09cc65572c15eb12f3b3c690db71f1a14c84872b33f8

Observation df257f01-d68e-4020-86cb-20df0eadc475 · outbound

This paper cites CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.010337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.329247Z digest=sha256:b74adc67ba8e25654ac9eaf6fc687bd2f71059247d2e6a51c6f85a1d7e6f524a

Observation 070e97bc-b8ca-414e-b4a4-7b62c0b62d05 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.332304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.332304Z digest=sha256:6f288ceefec85066f11f5985d5b3292e3d26d007ac9562ac6c978be0b974d3ab

Observation c0d5ca24-f249-4b16-aee9-1bbd471e85db · outbound

This paper cites Unified Vision-Language-Action Model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.335648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.335648Z digest=sha256:b307c230f9f8f30ce236f404b60be729ebcf302bcef68c986d776eca1eb938d1

Observation bcf3a1ed-5608-4cf1-85df-33d4ca959e10 · outbound

This paper cites Disentangled robot learning via separate forward and inverse dynamics pretraining,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Disentangled robot learning via separate forward and inverse dynamics pretraining,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.999192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.339006Z digest=sha256:518468e8caeee88d333d2ec2c3ef077bc884d9645ebebc020beb37a8620eca64

Observation 494b723d-3fc4-4c09-be34-5b9b3e44242f · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.342197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.342197Z digest=sha256:9d62b1c36134a39199dac858afc585d396473c3a5e504cc89d29f8f30d46b400

Observation d3f1a906-5701-4748-890a-26e4d2cd3052 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.345462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.345462Z digest=sha256:6c63a873bedf54a5de70b95c144c08ada84fbbe71f645b1e8de1d538a99df672

Observation 151c7e52-33d6-4715-9660-055547a9ecc6 · outbound

This paper cites Closed-loop visuomotor control with generative expectation for robotic manipulation,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Closed-loop visuomotor control with generative expectation for robotic manipulation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.988461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.348290Z digest=sha256:a11b2daaa408f032c54740b0f26bcde9541f384f5814e56613ef9634f6a9dc2f

Observation 57f9af1d-d3b1-49dc-8040-fe607b8cfead · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.351046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.351046Z digest=sha256:95e2c2577017a18bdf50cc2f2cbf39d965e507a83f2adb34b5bf12f3e3cdd73d

Observation 329ece0c-ce06-44bb-bb75-991b50ca45d5 · outbound

This paper cites UP-VLA: A unified understanding and prediction model for embodied agent,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A unified understanding and prediction model for embodied agent,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.977851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.353821Z digest=sha256:029119fd2156c81d3cf09dbf38c9661d8314f400db04f3acf73ad0c1fa0e5515

Observation 995e8f22-1a8e-46e0-84bf-941c89e41238 · outbound

This paper cites What matters in building vision-language- action models for generalist robots,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models What matters in building vision-language- action models for generalist robots,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.966734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.359366Z digest=sha256:3bd2760023c86ae78eaf77d252ad03338bdf2b6c9592cd1918ab5bdefb51a2b6

Observation b1c35929-4741-4869-a527-3e1902b19ff7 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.362007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.362007Z digest=sha256:374b2427d1e9ebc814db6c276334c946882d3ef848413356851cc42b1a76bb51

Observation 76e909f4-f93f-4e90-991a-6a937ec19a6b · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.356467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.356467Z digest=sha256:b98b7121b93ebe891a8bf93e9f49cd907132cc4ff5123a75f219f5dff0663268

Observation 07711bc4-1625-48fd-b1e5-c6201cd13b93 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Video prediction policy: A generalist robot policy with predictive visual representations,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.955746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.365544Z digest=sha256:0c58acb0d6ecafd7248de5fb0c0673b137f6ac02dbbe3b34ffd2b1a16f964a68

Observation 7294401d-902b-4fdd-bf7c-f2e75b9fb205 · outbound

This paper cites an unresolved cited work.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unresolved cited work

Reference 305

Resolution
unresolved
raw_fallback, observed 2026-08-10T10:29:56.049969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T10:29:55.292911Z digest=sha256:68f32c9e615fd92c8dcc36031187b4b30e9c78e27c790c8fc6991a6d305de6ec

Observation c57bf128-f41f-4721-82f8-47351af4de48 · outbound

This paper cites Available: https://arxiv.org/abs/2510.00406.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Available: https://arxiv.org/abs/2510.00406

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.306328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.306328Z digest=sha256:eabd714c645034fcc6089b414616781bb41fe3fd8f23f8ad9898fff5e4201f2b

Pith citing papers

No inbound Pith citation observations are available.