Pith. sign in

Paper Citation Record · LEDGER

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

As of 17 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2608.08839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08839 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:27:01.485780Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c4cdf7b-882d-406c-b4ef-208ea7513bda · outbound

This paper cites Motus: A Unified Latent Action World Model.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Motus: A Unified Latent Action World Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.187503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.187503Z digest=sha256:beffde940df2f66617e02f5d76413bf722cfaf873da6a0bb983808f9ac985478

Observation 45bff85b-e0d3-4f0d-a057-d9580aacb76a · outbound

This paper cites π0.5: a vision-language-action model with open-world generalization.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models π0.5: a vision-language-action model with open-world generalization

Reference 3

Resolution
verified exact
doi, observed 2026-08-14T04:27:01.535013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:27:01.213845Z digest=sha256:b2d3aec88a430fe196c9c4a4e93ef0dc30db5520e3d57d2bd4aea34963da4671

Observation 94a0bf8e-e965-4e35-bea9-f4996e576e80 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models WorldVLA: Towards Autoregressive Action World Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.246456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.246456Z digest=sha256:9844d40c8685c6f4837da26a81c056ee86dc62bca1fe7a851bbdeb5fa95b32e8

Observation 45301073-61e7-4430-a514-1d0e169dd403 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.271186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.271186Z digest=sha256:5de9344cd85dad6818d2d2c1458678458cb671c608f14f7f284ee7a7b7ff587b

Observation 41ef3d8a-b30a-4df4-a51e-6c1d3f278632 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models LTX-Video: Realtime Video Latent Diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.307886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.307886Z digest=sha256:43c2d9568743b83a56f95210a8642f00ab945747b5d507f52fe09a4c0c803d36

Observation 65d62c5e-809e-4f91-b6a6-4f2e9e8bfd07 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.313720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.313720Z digest=sha256:037ae7568396a66b363d0e9ceb019ab638c1aabdf09d880317248e17bb040dd3

Observation 6c097634-f4f9-47eb-a59c-167333b337e8 · outbound

This paper cites Rynnvla-001: Using human demonstrations to improve robot manipulation.arXiv preprint arXiv:2509.15212,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Rynnvla-001: Using human demonstrations to improve robot manipulation.arXiv preprint arXiv:2509.15212,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.318910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.318910Z digest=sha256:9dad86d416163e0ebb6d375fc7f73a37c53963286863d57856be8d3a2a0a0163

Observation 5647a053-0c1a-4c39-bb99-f593781aed6e · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-14T04:27:01.322000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.322000Z digest=sha256:5278cccd0c4e7dd055eac0ed55a856a24ccc4caded64517337c8caf32a1ab7c2

Observation 6f83a792-6cb8-4ded-8659-693c7593dfbd · outbound

This paper cites Unified Video Action Model.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Unified Video Action Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.332559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.332559Z digest=sha256:2e95765b2f4a5d05d1312bccef6f12cb4fde29e3893524465ed36fe6a6f76321

Observation f5fa2be7-2580-4e88-9baa-cdbd2005f8e1 · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.336990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.336990Z digest=sha256:08d80a82d6d4e0359d43f6adf35c5a028bf89ec340bee98fcfcdf50c3658e950

Observation d0d199e1-abd0-4e4c-80c8-412f57e1ed56 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Depth Anything 3: Recovering the Visual Space from Any Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.346136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.346136Z digest=sha256:e6b4df99d2fb6dd06e18bf2c23d9391f2f204b251ba6d68e01f5a469bcafa0cc

Observation 73a90e82-f850-41f5-a373-264846f5716c · outbound

This paper cites OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.356892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.356892Z digest=sha256:e0c9493da88371984307456fa457bae785343353bfb202d8af56c34e0c2f8abb

Observation 416ca311-f861-4115-88b2-4b45c2ff0c0c · outbound

This paper cites Mask World Model: Predicting What Matters for Robust Robot Policy Learning.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Mask World Model: Predicting What Matters for Robust Robot Policy Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.365387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.365387Z digest=sha256:4363a1d9c3c8dae86dfa46e1fa37bb51e37b8925bebc1c196802a961aca6b779

Observation 4f851e50-8e62-45e1-a400-2bce99d00a32 · outbound

This paper cites Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.378434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.378434Z digest=sha256:b1f4f2b29b77c8c445b85f17b455499ac90bf5b472ad8dad199b1d7fd497f138

Observation df5a5daf-c291-4eb0-8c32-af1a286e8981 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.382735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.382735Z digest=sha256:0869c28678ddc448fda4297a17091780f3705e6c6f0ca5371bd583286e985724

Observation 98d7f6f8-d680-4a1a-a310-45b741db355f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.415201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.415201Z digest=sha256:80c996cc4cfdca2a9af1051a64edf7d13888f6edb3090430dcee7a5d5b5e0532

Observation 27bcc08b-6877-4a20-ad7d-935a2ad0db66 · outbound

This paper cites S-VAM: Shortcut video-action model by self-distilling geometric and semantic foresight.arXiv preprint arXiv:2603.16195,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models S-VAM: Shortcut video-action model by self-distilling geometric and semantic foresight.arXiv preprint arXiv:2603.16195,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.436760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.436760Z digest=sha256:11db7e6987be5aa3acddeee2c5f44cbe776f875a77048a80b286bb969c4c0455

Observation 0c20bc1a-e422-464d-878c-7855c0698104 · outbound

This paper cites MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.450632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.450632Z digest=sha256:8881d98d3f9e7724ff13f980e04c9e8e96664a8f5cd61623e88fb59d82343d4e

Observation f19e0cf9-3b5d-42c8-8b9a-e98d39b885b3 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.455192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.455192Z digest=sha256:3d625fc41b326217d6af0f9213cdbac97389ba42cf86ef356e6334c6a9f97f31

Observation ea6ecb71-d1cd-4df9-b9cb-08a5ea330434 · outbound

This paper cites FlowVLA: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models FlowVLA: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.462371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.462371Z digest=sha256:f9d7a62b1c613b8d52ae8e123b39ed6f40e5b3da15b6644138b702e60daf6509

Observation 01a340b4-28fb-446a-a0ba-9d9f01afd41b · outbound

This paper cites DualCoT-VLA: Visual-linguistic chain of thought via parallel reasoning for vision-language-action models.arXiv preprint arXiv:2603.22280,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models DualCoT-VLA: Visual-linguistic chain of thought via parallel reasoning for vision-language-action models.arXiv preprint arXiv:2603.22280,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.467787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.467787Z digest=sha256:a06d6c30607f58c43c121b4eee28a481673ec5c953bb850d798b3c13e6abf4e7

Observation c3d68a9d-5b10-45b2-81f9-802e6f4f655e · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.479963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.479963Z digest=sha256:d847e4b5d4e5024b761d388db1ee3ae0c4e19d7b2342685b975159dc28151205

Observation e357fa73-6940-4db3-bb28-3a4ac4759dd9 · outbound

This paper cites DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:27:01.558768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:27:01.485780Z digest=sha256:fc38d47f0e98fd999190e3c8371380359c27c41aeadc884a98b500d07c324267

Observation 6a50ebe8-ed3b-405d-baff-8aefda09b651 · outbound

This paper cites ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.402929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.402929Z digest=sha256:e7a16999ed6d4b1ff08d7da9c9412d5a0b9d0117489ac9968959055a55bcf094

Observation 7cca0c34-2788-4da8-bdce-8b05e6dbe381 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.238908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.238908Z digest=sha256:fe08f5663d158d4a07d0788882202a3c7ccb665a02093770abe72d6c774b7a04

Observation b1b0234c-48e2-4d0c-8b5f-f201e595275f · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.290804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.290804Z digest=sha256:418c5b82908d5575bd10c560523f2980198faf29a6f5f5082e80a173dfb1ff4c

Observation 5e761f15-8856-4ccf-af6c-764da0415a20 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.206906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.206906Z digest=sha256:37210d0ce842ec0ada46ba87ae29bdc6e405217294ba4897d9f630a27ce5fcca

Observation 3c97199d-61b9-4aae-a5e6-e6f88464c71c · outbound

This paper cites Causal World Modeling for Robot Control.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Causal World Modeling for Robot Control

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.327148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.327148Z digest=sha256:77a85143155c7846dd326ccc0e263c53da4ee7768036e3b479b7b92d62bdb88a

Pith citing papers

No inbound Pith citation observations are available.