Pith. sign in

Paper Citation Record · LEDGER

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

As of 12 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07314 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:29:55.365544Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22fd29ac-b83c-409e-8d90-27297fe18a37 · outbound

This paper cites PaLM-E: An embodied multimodal language model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models PaLM-E: An embodied multimodal language model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.195731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.231185Z digest=sha256:dfd52c05429f87a823420d2e41c51b1df3be752d0c08c8c2a63cf07ddeb45654

Observation 3fa02969-224b-4744-a1ae-5da4257e0daa · outbound

This paper cites RT-1: Robotics transformer for real-world control at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.186044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.235069Z digest=sha256:908660a765832a1f560531dfe8fd1674cce4e9fb6f16abc121671aeade49d2a9

Observation 29a27b32-044a-4c8a-903a-48e1d3ac7059 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.175689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.238673Z digest=sha256:dd0b66ddd95f568e941c1e31ce6821b8014bffc70ebaece90b105d8a13d2df54

Observation ea4da3e9-37b1-4529-9654-dc7e9ac20716 · outbound

This paper cites Open X-Embodiment: Robotic learning datasets and RT-X models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Open X-Embodiment: Robotic learning datasets and RT-X models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.165594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.242091Z digest=sha256:e914446f6de595fdc569c3f63bb6d5daddee599eaf3c5071df60a7d2e27099cb

Observation 56230c39-9f3d-4b74-b647-a0e692aa832e · outbound

This paper cites Octo: An open-source generalist robot policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Octo: An open-source generalist robot policy,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.155897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.245579Z digest=sha256:1f669fa43907a7b6e51502a42f2a70b34d75db13113fb0ec1682eac5c85679d9

Observation 12848eb0-1192-4339-b31f-2d29fe98b861 · outbound

This paper cites OpenVLA: An open-source vision- language-action model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models OpenVLA: An open-source vision- language-action model,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.147186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.248925Z digest=sha256:716eba8ae6153dc11d7359872cbb3fcf4fd8c15cf95698d1f2a5dff76586011b

Observation bf801935-81ea-416f-8f26-aec8e029f257 · outbound

This paper cites Fine-tuning vision-language-action models: Optimizing speed and success,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Fine-tuning vision-language-action models: Optimizing speed and success,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.138028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.252088Z digest=sha256:c0387991ef91fe7317d199e34ce6bfc229862eb639dc4c930f5a6f06140f393d

Observation 0f01f29c-4545-49d7-95f6-b4fc8b577c71 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.127108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.254959Z digest=sha256:84a5a88690cc60cd1eb14e3d361a1b2f09ed58739c4a922ec328bfc827878912

Observation 5476859f-65e1-4d2b-adcf-a043a72076f6 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.257629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.257629Z digest=sha256:aeb024e844e5e3f9ebee024107c67ff2bdbe2d4065739bb094c2b5d91dc4d507

Observation f73d1b42-aa20-4d12-8ef7-f7e48bd7899f · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models SimpleVLA-RL: Scaling VLA training via reinforcement learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.116999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.260771Z digest=sha256:987697048d7f76eb8c06522c4e1a95d5db9c047358a2cfb5e79e59af1967eba1

Observation 34301185-637a-412e-ab6e-6c5c09477126 · outbound

This paper cites RL Token: Bootstrapping Online RL with Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.263923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.263923Z digest=sha256:521bc0f7f055843fd5ae5e8c0cc62cb8466c353bbbc20e114b687fbd4e818834

Observation 9f211459-f700-4493-a8cb-9098f4574656 · outbound

This paper cites Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.106712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.266959Z digest=sha256:62e7d6a749a62a4c13e540421df66bd34a2f5d58f8f9e6958389f8f4c2f362e4

Observation edfee9b4-2bab-447b-8495-e59005959870 · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.270093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.270093Z digest=sha256:cd2f6a41df85df9cdb59d03af4502b60c8acf0f1ec7c8f475cbd79c54213c561

Observation a5f798a9-de3c-464c-b2f5-7123f70a7771 · outbound

This paper cites Addressing function approximation error in actor-critic methods,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Addressing function approximation error in actor-critic methods,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.096260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.273617Z digest=sha256:5605361d3a885be14957f95354cca445001cdbd20ae9b975c26051c9a3f7502c

Observation 7668e1bc-5b3f-4710-82f3-f085ce17c84f · outbound

This paper cites BridgeData V2: A dataset for robot learning at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models BridgeData V2: A dataset for robot learning at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.086289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.276679Z digest=sha256:43e47f8177f1a0069fab9f12d0062d48be3fc82398dc0995912701ed80ea0e49

Observation 29ca5594-d6bd-4964-94df-c05bdf3801a0 · outbound

This paper cites Vision-language foundation models as effective robot imitators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Vision-language foundation models as effective robot imitators,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.279752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.279752Z digest=sha256:b6c458981bd81e3d93979b691634f578b62b7907f87cd3e852cabeed5208f75c

Observation 4b7bdbb6-6c16-4379-8ffe-24ff1d0f6af1 · outbound

This paper cites Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.915698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.283014Z digest=sha256:530bc37e801b34b9647811d999f1fbe376df1380ff5c6a77dc6a84119f415059

Observation 6dadfcca-bde9-4ada-9e09-db4045287d65 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π 0: A vision-language-action flow model for general robot control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.070341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.286220Z digest=sha256:c6324291bc8a579f28947169f2fbe8b0d1cc83c6326eda6a4c74f6dd6b91dbd4

Observation 26618d8a-a725-4ff9-845f-2dafac8509ee · outbound

This paper cites π0.5: A vision-language-action model with open-world generalization,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π0.5: A vision-language-action model with open-world generalization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.060558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.289517Z digest=sha256:a6a556790831c74dade1f83b4a7782ee8994ba62e9cdcf083b879abdc92bf80c

Observation 87607237-d798-43ab-8d8a-065063327d4b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.296092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.296092Z digest=sha256:ecdd4900cd790dcd3fc43c688a51a5c9a4379725d3621e84ba7137b6e579a227

Observation 3066997d-de2d-47fe-ade4-6bcea7246a0e · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Interactive Post-Training for Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.299516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.299516Z digest=sha256:e3a47dac4e5fad147acb2644b006e659e5ef3cb086d7dc9130e84618350342dd

Observation eec1aa47-6d89-4fbe-bd06-e161450369d4 · outbound

This paper cites VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.040569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.302904Z digest=sha256:a5ce2cb323d691565be5e64baf6d149adadb77bff2d031f17f1e7eba60b373f5

Observation 8599fa7e-bda9-4196-9164-a095a7ad8298 · outbound

This paper cites ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.030298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.309572Z digest=sha256:dceef96cecc2729d7801d515b622d82593a108ab8ad1f35c82c7f9cc453612ff

Observation ef126e3e-aeae-4b5e-8752-2cedf4b0b831 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.312519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.312519Z digest=sha256:2083b5620e01d0c415d389d8a994dfc8f8200cfc8ec4c678cb76fea43508c9f3

Observation 6eb70623-a93a-45c2-8578-e2e368f4b264 · outbound

This paper cites Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.660050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.316233Z digest=sha256:1f33a701d5bc6ed22b9f66ef53928c785fabd9941c3eb6f32b19ac05bea1bf40

Observation bff91253-b426-414e-92fa-ed83fc42e6f1 · outbound

This paper cites DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.319520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.319520Z digest=sha256:c03d4a4b51d402c64514ad4612c9c0933f26d4c27be42c63365295d6944c06d7

Observation 1595bda8-f98a-4546-a9fd-5b91cd3c17be · outbound

This paper cites RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.322806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.322806Z digest=sha256:0959ac236cc6ded26270f16eb6ed3c190c67427a67714def3cc60c0b03d1bfae

Observation c9190653-2b3a-4c89-be44-49d47b8d0d05 · outbound

This paper cites FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.020826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.326214Z digest=sha256:e675c5da0290ee7553a7ebc6eb2ecd70d851bbebc8ace78ec35075a3d5d8f8fb

Observation df257f01-d68e-4020-86cb-20df0eadc475 · outbound

This paper cites CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.010337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.329247Z digest=sha256:b86ddeca4058013bc18af43502c7b99938fc75bc2264b82cdcc84660f9cc6de3

Observation 070e97bc-b8ca-414e-b4a4-7b62c0b62d05 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.332304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.332304Z digest=sha256:aba7f6388b6a56cce67a190f348edb2112d6726e3e76596f28f5be40f39716de

Observation c0d5ca24-f249-4b16-aee9-1bbd471e85db · outbound

This paper cites Unified Vision-Language-Action Model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.335648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.335648Z digest=sha256:f70df64f8282e773013b436642857646d1a8834160f66b42fbf263eccab83ed5

Observation bcf3a1ed-5608-4cf1-85df-33d4ca959e10 · outbound

This paper cites Disentangled robot learning via separate forward and inverse dynamics pretraining,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Disentangled robot learning via separate forward and inverse dynamics pretraining,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.999192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.339006Z digest=sha256:c7560a60ffa6ec622dc53e4e4fe839404537ec4affffe4a4a776c85afa37903b

Observation 494b723d-3fc4-4c09-be34-5b9b3e44242f · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.342197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.342197Z digest=sha256:ac5c9a721353b60c774fd24980162e4292f077b483dda76e914550039fb89c0a

Observation d3f1a906-5701-4748-890a-26e4d2cd3052 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.345462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.345462Z digest=sha256:76b9e1dee0611526e2ce0ae847633b18d7d1e3a82436d4c3e3b465aa8bb28b0c

Observation 151c7e52-33d6-4715-9660-055547a9ecc6 · outbound

This paper cites Closed-loop visuomotor control with generative expectation for robotic manipulation,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Closed-loop visuomotor control with generative expectation for robotic manipulation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.988461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.348290Z digest=sha256:c214df7d978b0b5bb145ad97982fe9714d187d7d3d161f81788e7b4439c40671

Observation 57f9af1d-d3b1-49dc-8040-fe607b8cfead · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.351046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.351046Z digest=sha256:2e03e95bef9f62c9520cfefef04331a6f1cc198ed661c868c7930f362c4cef35

Observation 329ece0c-ce06-44bb-bb75-991b50ca45d5 · outbound

This paper cites UP-VLA: A unified understanding and prediction model for embodied agent,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A unified understanding and prediction model for embodied agent,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.977851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.353821Z digest=sha256:475aa8c48ba0929b3783e366690de1517162456106352c625ddea2888095328c

Observation 995e8f22-1a8e-46e0-84bf-941c89e41238 · outbound

This paper cites What matters in building vision-language- action models for generalist robots,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models What matters in building vision-language- action models for generalist robots,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.966734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.359366Z digest=sha256:e791739de828302272620572a4ee820d22b641d69d5ab51c2c48252628fd976a

Observation b1c35929-4741-4869-a527-3e1902b19ff7 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.362007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.362007Z digest=sha256:ea97466cb29e396690a07a82f77ea8f28207d5d1a12416fe270f7446f5762ce9

Observation 76e909f4-f93f-4e90-991a-6a937ec19a6b · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.356467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.356467Z digest=sha256:c52a014a4c4b45d83922d57d70e5233ea956fea3b4c026d0f137e0ba4eae113d

Observation 07711bc4-1625-48fd-b1e5-c6201cd13b93 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Video prediction policy: A generalist robot policy with predictive visual representations,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.955746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.365544Z digest=sha256:63ad3a5096475efc29ee9342ee31157e637b4a162f06f6b65160f7c4151c344d

Observation 7294401d-902b-4fdd-bf7c-f2e75b9fb205 · outbound

This paper cites an unresolved cited work.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unresolved cited work

Reference 305

Resolution
unresolved
raw_fallback, observed 2026-08-10T10:29:56.049969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:29:55.292911Z digest=sha256:56dbc1c5bfe952f52f62d075da3347d0aa407325881493b73d973cc1022be26a

Observation c57bf128-f41f-4721-82f8-47351af4de48 · outbound

This paper cites Available: https://arxiv.org/abs/2510.00406.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Available: https://arxiv.org/abs/2510.00406

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.306328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.306328Z digest=sha256:42cc1fa937edaf6573f54b1ffc1152c9d2d6d8de58f3d29508efe22326f846e9

Pith citing papers

No inbound Pith citation observations are available.