Pith. sign in

Paper Citation Record · LEDGER

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2505.19767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19767 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:10:14.710517Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84d7d996-2113-49df-b533-066e031b2686 · outbound

This paper cites GPT-4 Technical Report.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:10.441248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:10.441248Z digest=sha256:4e3fa93ff03b66a0bc91eebd91536192e8a93a461d4b66cdeb47828d1609645e

Observation 0cefab2a-604e-4025-b595-1d9023d64e15 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback The claude 3 model family: Opus, sonnet, haiku

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:10.514798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:10.514798Z digest=sha256:c46a5fb7b0d4027329ab8d5d3b7b2e5427a98e0183517e7fccc7c98f73c1371c

Observation 25d58bb7-f8a8-4b61-b200-6d99f679d5d5 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback PaliGemma: A versatile 3B VLM for transfer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:10.646523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:10.646523Z digest=sha256:461bae5d649d02861a943b278b691ba58b6222d146b1da4911794b144d6fc1d7

Observation 26cd55c1-25b9-4c35-89ce-385b705f08d3 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:10.773196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:10.773196Z digest=sha256:b4db08ec41b0d41bbd5317413314cc0d38afe8811255ee1992b2d409c6777836

Observation 10297ea3-b1a0-45fb-8768-1687b41cc52f · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Do as i can, not as i say: Grounding language in robotic affordances

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:17.989060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:10.884953Z digest=sha256:2624ab3bbf2c08e506674109427f6838c01d3c7405976ae15190676800baa6ce

Observation ca446ca0-4380-40a2-b952-0367c84dc081 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.005810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.005810Z digest=sha256:dbad77fcfb50d3d996d81ed31be01d13fac665acef18ac1ed3e662436d0d439e

Observation 2f5a51c2-fe4b-4098-a172-7fc68a980344 · outbound

This paper cites Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.153669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.153669Z digest=sha256:428da846654c30e45334212a5a4a273d2cc153eb807842700c6c445029c9726f

Observation 32791648-9a9f-474d-a85c-0e52b1a44714 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.309613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.309613Z digest=sha256:7dc0669919edc183db075a01bf735afd766b2a3e0698ccfe943797446674de87

Observation d2a96252-7877-47d8-aade-4f0386510bf9 · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.420265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.420265Z digest=sha256:33cc7d6b37f2691bf1776a17b590ac05d3a46dc33e037a98f310bbea189cad64

Observation 132f271d-e618-43d2-8cfd-7213a697e4ad · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:17.849191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:11.532371Z digest=sha256:4f63d256f82c2681e7c54315b2abd573ed23eeb0b66724c0d495e4f4b287a692

Observation deb26e8c-db32-4229-b4a8-e404067fa895 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.616045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.616045Z digest=sha256:c0f86bf68ad5468497a7038b8234b85bac3e5727c62b8505c2988b769069adcb

Observation a27cf8f7-cea0-43b4-b7c8-1ec0799ea99a · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.693083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.693083Z digest=sha256:3b7ecfc80a7cdc468390d9800554e77e49c1c022a249563e1ef716347415648d

Observation 9601e9e1-6c59-4e3b-982d-2d31622ec4e8 · outbound

This paper cites Diffusion Transformer Policy.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Diffusion Transformer Policy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.796953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.796953Z digest=sha256:e850814f06bcf84854f6913d49235c9497063095b9b7a557a7d7cf8ceeccc2d7

Observation 595a48b4-583f-4508-961f-4d32ac9e6a73 · outbound

This paper cites FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.936652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.936652Z digest=sha256:3d4b2316500057b7089d9bb49563a717d6fb7ad5a0b86d45267157579a57bad9

Observation f5ed7ac3-7079-4303-a82c-676d6c11c2f8 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.043431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.043431Z digest=sha256:b5f7b9c1349bbea3be6995c440e17b19d36d46bfc573fe859756e31d92179f8a

Observation 48acf3a8-a8c5-495d-bb47-c09580192076 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.160489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.160489Z digest=sha256:05c1c68b35395b412e4ab6d48e4a35aa0a1ccdc06d16a8850c80d484d5f631c1

Observation f5da6f77-aa0a-4f7a-aff8-73cba822e6c9 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.271482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.271482Z digest=sha256:578d221103d4decc295909989ea0c92435f1b4347cfe01f132a5517b96be6144

Observation f2593529-d885-4a0c-88e6-041bbb5838f6 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:17.705508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:12.340763Z digest=sha256:68c08f67cfa69af1c467864b84c9521ceb7014aa708d4e5a4967a92373fb5eee

Observation fba6cf7e-a5ab-4ff0-bad7-344f4cb4c0e4 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.407220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.407220Z digest=sha256:a3c830686392b5b5084374593d0f7a4b17b1a7fe627b7ddedcf4c94c07b3cf7c

Observation 170d07f2-10c9-4a20-b293-700f65812c2f · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.477348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.477348Z digest=sha256:834beccd84ac902ef41f1dfa39f77b2f61cb167de91a5721ad7aba68272181ca

Observation 662bd5c9-a28e-4bb5-9d50-0f00bdfe42f0 · outbound

This paper cites Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:17.499237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:12.561753Z digest=sha256:9353e251f6f4bafa85f9e2605b04c7104426498aaa0674767dc9732499f06d46

Observation 829083fe-0ec1-4bfc-8589-afd904549977 · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Manipllm: Embodied multimodal large language model for object-centric robotic manipulation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:17.366697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:12.646019Z digest=sha256:1b57518c5c3a943f903878d1437d191dd36a84ee27721a58f9504ffef046a10e

Observation 4cb70d16-4877-422c-80bd-9af4a09c4f5d · outbound

This paper cites Visual instruction tuning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:17.205490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:12.702322Z digest=sha256:7216533bf06029743cabe1e95b8ac530960498e1cb2fa4863513e03fd8aa8783

Observation a091da66-7b2d-410c-bd73-9b4b9b49537e · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.795591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.795591Z digest=sha256:8b332d088c8c4cb331a8078d492cb2ff63336519b10965538a30603ba4b3e831

Observation 8b0aca1d-5856-4d18-8f72-b1d6733a5134 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.871855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.871855Z digest=sha256:b172333f2c8e6f54ea015f4fcf2741a348ff7710c24f2f54568f4d5c9609fa6e

Observation 04b7eaec-cacb-4cbb-85b3-b00ad11efe56 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.939491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.939491Z digest=sha256:acacbd09f3912bfd2a8283a6fe6819e767463c6955c1bafca5932ca87a39784b

Observation 718535f4-56c7-42e8-9b0a-1fe475137173 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback A Survey on Vision-Language-Action Models for Embodied AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:13.032121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:13.032121Z digest=sha256:4c753f4148f37e873949c67262292e0d1c1444bff73cf566815e264c0e6fd046

Observation dbe37bec-c507-46d9-954d-854dce1def88 · outbound

This paper cites RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:13.120606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:13.120606Z digest=sha256:4ebfc5f3222ee9891168a258ec1a1e756c7fc4fccb95306d4a119d288a2134d8

Observation fd506478-a35b-4b75-8882-5beed27e6043 · outbound

This paper cites Genrl: Multimodal-foundation world models for generalization in embodied agents.Neural Information Processing Systems (NeurIPS), 2024.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Genrl: Multimodal-foundation world models for generalization in embodied agents.Neural Information Processing Systems (NeurIPS), 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:17.076094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:13.226413Z digest=sha256:9f734fb77e070154ccf23465b857452662106858e97c96ff7f7b8613c791de7a

Observation a926ba9d-a3ac-48e9-a991-a4c0054f2623 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:16.922107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:13.313734Z digest=sha256:50ac1f5147b246d38cb9d4b4a95f37d9244ccb6d9cb88fc8334de24f2504bc31

Observation c1af770b-cf65-43ff-bb77-4e464b3dd087 · outbound

This paper cites Policy invariance under reward transfor- mations: Theory and application to reward shaping.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Policy invariance under reward transfor- mations: Theory and application to reward shaping

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:16.789038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:13.386140Z digest=sha256:389d36a3cda9dc331a4fd552f8dcd83979e8b023d7361c19f9fb0e5af41082ab

Observation 789a1e12-6441-4c4f-8265-60bea975c1c6 · outbound

This paper cites Training language models to follow instructions with human feedback.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Training language models to follow instructions with human feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:16.634960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:13.451288Z digest=sha256:e3a98a3b8511aa6226525e777d8e2d8343f5a3eba1d554aec94aadbb8391da11

Observation 44169c1d-9b73-4100-9a24-cad41620917f · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:16.458385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:13.533614Z digest=sha256:4a25cf8af70c757f543800ed6140126266b4fb6d9f9a5ff6c6d8747104b47bb2

Observation cfce2e45-5d38-48af-a08b-d652d6b928e6 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:16.298773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:13.603390Z digest=sha256:d8b2b8154abe34876538c55e3e4f6b8e8b6c956fe74bf23ed30c8fcd0173c88b

Observation a51e8918-7449-4f2a-9cc9-b6d474515a9d · outbound

This paper cites Diffusion Policy Policy Optimization.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Diffusion Policy Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:13.679384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:13.679384Z digest=sha256:8c6997959a4d44d55652ef146eec75d4628fcb280d2ff6d35660eb8cf100e8c7

Observation 82e608ab-dec4-43f3-b186-d49f41d56a23 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:13.757826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:13.757826Z digest=sha256:38ad925d5e40ff8da01ae32f29f4d5599093d8bfd0c2246d0127fd112195ab05

Observation 25990dc4-a7e6-43a9-a577-72cf3780da04 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:13.833341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:13.833341Z digest=sha256:7ee745cab4a9bf605c8f48f0f00a0af624323d87bb83346b2347c10291aeb913

Observation 278423af-23a4-4733-9b88-b9827ebb80bc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:13.912863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:13.912863Z digest=sha256:6785641785363716eb97c9ddc53d0413275662d9361d7d795f3ac0527b496e79

Observation 9fc9c671-1819-4976-bc23-e2e8caef6cd0 · outbound

This paper cites Large language models as general- izable policies for embodied tasks.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Large language models as general- izable policies for embodied tasks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:16.163700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:14.007026Z digest=sha256:70bea2a3ff88ef8e6a33b42e0cbaaa9076f0cf9dc626c4e5840da2c7410c437f

Observation 26c447e8-a739-4411-a434-4c72a1249eed · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:14.083349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:14.083349Z digest=sha256:9aa04a3649f6f0b544a2aa187c2cac50c15be04b03088ddd119cbd18308db4c7

Observation f2b7222c-2b51-4f65-ae09-0e8342420b22 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:14.171540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:14.171540Z digest=sha256:b95a8b72fb709efa041845751b7013f665f1e437bcf73a2887c5b3ab6b43182e

Observation 398732cb-7f57-4aaa-af4b-a73d90fe4d7a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:14.238253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:14.238253Z digest=sha256:c9ed0bfdf9732f5cc6ee3eff2c967c193940852d1995ff5fc6ec6f1f3c9a765c

Observation f1bad90d-60d1-4953-ba4d-71d0bc6ef7e4 · outbound

This paper cites Reft: Rea- soning with reinforced fine-tuning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Reft: Rea- soning with reinforced fine-tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:16.006898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:14.284220Z digest=sha256:e702da2cac955e73288ab9bb576a84663357da43f5f11fa02cd6156bf2e30b8a

Observation 1dc21c5f-2644-4847-a832-8a11beb6bd49 · outbound

This paper cites Embodied Task Planning with Large Language Models.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Embodied Task Planning with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:14.451539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:14.451539Z digest=sha256:7c798b3ba7bc499abbf42070f86df952d0ad17dba03fda346476af5a9c189bae

Observation 0f1717da-f9b7-49bc-a6e7-40611aea8a24 · outbound

This paper cites Qwen2.5 Technical Report.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Qwen2.5 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:14.543919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:14.543919Z digest=sha256:7e6152bd26e49bd8eb6e7a11d0648c7e3679f61c6dfd757475c7f26fc98f62cb

Observation 19ca6b3e-7f11-4af9-84f2-e563aad0994b · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:15.877510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:14.623426Z digest=sha256:70c43299a4afb8cb74a0c1df3e2baf0a7f3608805221ba411a4839d175af3e63

Observation 0f7a4f8f-8980-45f4-91f8-fbca89ebc6f3 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:10:15.785766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:10:14.710517Z digest=sha256:a5145b52b54b6d0b901c0d0c4d6f710d2cd482a5d162e47066583039e94dd400

Pith citing papers

No inbound Pith citation observations are available.