Pith. sign in

Paper Citation Record · LEDGER

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.14739.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14739 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:15:51.066275Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 132e003a-25e7-4069-b575-e94390c2bef3 · outbound

This paper cites an unresolved cited work.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.391508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.391508Z digest=sha256:af3f08a45952483c94d6b5d13d3f7fe8b8b447c94e0b2bcb47ec75b0cbf77e3d

Observation 5a498e8a-2f24-4fe6-bea1-ef450fa1ff8d · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.666023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.666023Z digest=sha256:5582fd2d3ee26679efdb56dd7bee7ab9d45d5cdac1ab107e71a17e5b816e22ff

Observation 616e8db5-5661-41ba-a695-54049e221387 · outbound

This paper cites Point Tracking Improves World Action Models.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models Point Tracking Improves World Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.737286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.737286Z digest=sha256:ae216901965d728fb74af76c8224b807af286a76883117df67d4369ac94e1f13

Observation 5cdf3b85-eb4a-44af-b3c6-91a7ee53a099 · outbound

This paper cites arXiv:2602.09849.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models arXiv:2602.09849

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.806929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.806929Z digest=sha256:52a836c6335b38453108e79174ef78a2934a0b3997cb6909ed65929f71bd7f4d

Observation 76d049eb-b639-4378-89e3-6acf16dbbe5e · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.909928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.909928Z digest=sha256:20fbb3563d3c067f3e16800c0f29a1bb144f571ed6b8888f39d1ae6e197d75e0

Observation 537e55ff-7ed2-4d86-93a6-e21c567a26e2 · outbound

This paper cites Causal World Modeling for Robot Control.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models Causal World Modeling for Robot Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.992652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.992652Z digest=sha256:079aeeecb9a853df5efc503d1df9e2e2f6993e0cd3d28809519594ccd95f07f4

Observation 0d6220d7-d705-40b7-bcd1-c9ee34888c7f · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.034343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.034343Z digest=sha256:fb2cf08a694b79d5b02b4a2ad89fec319c5cbeb213e9885fe47a3b63cba06b66

Observation 50d238a3-05e0-4284-b603-5ca86d95bed1 · outbound

This paper cites HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.037508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.037508Z digest=sha256:f29c57e204515df06c674d510608424a4102a4face634468206902ac73e00c5c

Observation abc843e6-083d-4b10-b2d2-5bbecf27ef98 · outbound

This paper cites arXiv:2511.19859.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models arXiv:2511.19859

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.040696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.040696Z digest=sha256:2cc6c11f90784d7a47d6ab65e1bfa34d5e55be3bb8b6c70987240239188718f0

Observation 9e49428c-fea9-4777-b44d-2bf49e7839a9 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.043766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.043766Z digest=sha256:c3e9c4026621162fbae452830f614fe304ee1545db0efc85e812ea008731af88

Observation b22597d3-32b6-4406-a776-296dcc84a852 · outbound

This paper cites arXiv:2602.22010.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models arXiv:2602.22010

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.047215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.047215Z digest=sha256:5c3eff16140cf2e46b31869e2af3db8d629bd99dafd2f16202a456da47d9cb23

Observation d3e75dcf-7d87-4411-b6ff-09419e0934f9 · outbound

This paper cites arXiv:2602.10098.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models arXiv:2602.10098

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.051011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.051011Z digest=sha256:c82c81769b0ac3604b394fbc01a6455db05b73746a9252ec1be81a2e1f754050

Observation bbd8565c-1912-4596-b50c-34c12e46b91a · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.055527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.055527Z digest=sha256:1d19d9cfa541740058b47e8444889b436f1f5b78ef3cd6f81d76667a5f5ee23a

Observation 8f7472ba-a6b0-4525-81a8-148cd714ab7d · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.059693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.059693Z digest=sha256:c44ef25004a90e2d68cfe313b4a79f517757b50cfc2a49df3460433a2e6e45fe

Observation 99a72d23-3379-48fb-bdb2-4f58a3fe219c · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.062743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.062743Z digest=sha256:12abb9be1c5ccede10d08ad919926690ccce941a64136b7639e3b7b5502cc6ae

Observation e153b636-b722-4c4a-82b6-3398db214c54 · outbound

This paper cites arXiv:2508.18269.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models arXiv:2508.18269

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:51.066275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:51.066275Z digest=sha256:d900cea424c8cca113f6dea9296bf1f27c8077305a8f5c2f31c0d86f2a717933

Observation 1384c583-f5c3-4d73-9227-21fc6fb3e954 · outbound

This paper cites an unresolved cited work.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.595450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.595450Z digest=sha256:949987a1dc4590098ab8e4e2babf29da0ee5d9f277e4ce0979e6026d9471f2ce

Observation 0d5a7a9e-bb7e-4f93-93be-d4fe7e35c839 · outbound

This paper cites CoTracker: It is Better to Track Together.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models CoTracker: It is Better to Track Together

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.868951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.868951Z digest=sha256:05f094a927bfb4537bbf894891adabcc301bb63d9f80733a51ef1d5915bebd4a

Observation 08e1ce28-db67-4252-8660-5d0b55c5f279 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.469381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.469381Z digest=sha256:fd5262da6302ee71c7844644dc73368f5951e47bea044fa453a23275ab261f75

Observation da1d927b-27dd-4fa1-b366-d0c334512116 · outbound

This paper cites arXiv:2601.02456.

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models arXiv:2601.02456

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:50.533339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:50.533339Z digest=sha256:9bfbe2863179b163e8eac1f4e70d70b156b121b5888e42fcac80755578075273

Pith citing papers

No inbound Pith citation observations are available.