Pith. sign in

Paper Citation Record · LEDGER

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

As of 9 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2608.00502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00502 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:54:48.841652Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c5d8d37-3172-49eb-a030-76bbe9e0695d · outbound

This paper cites Qwen3-VL Technical Report.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.786424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.786424Z digest=sha256:bb7fb5fbcca513c0ded1c8831738f57a4ef0462dae4530a9edf1e2eaa2df32c8

Observation fe9811cd-c017-4fd8-bc46-6ae4eec4e4b8 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.790316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.790316Z digest=sha256:04058e3c4b07c061d21ed3f185490344a01da981514645f7b2763fa132e96932

Observation 1d74063a-18b6-467c-9346-ec035d883fc5 · outbound

This paper cites One-Shot Affordance Detection.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance One-Shot Affordance Detection

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:54:49.129302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T00:54:48.811578Z digest=sha256:645a5b172a85272f364ee1a527840245dd0088e1da74802f69c013391b3febe0

Observation aa4483fb-ef56-478a-89eb-4ce8a5f5083c · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.817378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.817378Z digest=sha256:748f6179d88b244be760abba4a41b0f9fd5d464ade41d1381734d89966e858d0

Observation 8ceeb62a-a73c-4ed5-9df3-906602311fc8 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.823565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.823565Z digest=sha256:54685fd1c5efff5fda3dc3a7e262dca46bfdf41a57382a63dfc47552fbe33db8

Observation b4127c37-5561-4372-bae1-2f28034ef8cf · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Aligning large multimodal models with factually augmented rlhf

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:54:49.198559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T00:54:48.826852Z digest=sha256:d618678d368512183f0ddf1d5912a007a0be658a68dfb7c173c0db3f286eba22

Observation a08943d3-b8cf-46ad-87ba-8b807a0a6bdb · outbound

This paper cites RoboBrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352,.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RoboBrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.829656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.829656Z digest=sha256:bdb310462b389caf798502c0e7c8cc9eef0ef4063919e7ffe7925172ec623fb4

Observation ce239547-b0fe-49b5-8120-2637e6fedde0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.835533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.835533Z digest=sha256:ec5e5dc20297ba17729fbefb5c971258bb78d0d394439b4e14fd7d9d63c7e183

Observation 47da41f5-3b1f-4c2d-bce4-0c8231d80e48 · outbound

This paper cites PartAfford: Part-level Affordance Discovery from 3D Objects.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance PartAfford: Part-level Affordance Discovery from 3D Objects

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.838732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.838732Z digest=sha256:1cde27257840305bd6c39cbdf1b9f1f4e449f8d2c95ac8d3a1a28c31b31bec5f

Observation 603f4ed2-4932-46af-8a9f-68303b365f65 · outbound

This paper cites Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.841652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.841652Z digest=sha256:70241244efd8d32300a86d51bf72a3cc0747a9d41ce1a8f717278a0191249508

Observation c4d9ff90-75a3-45c6-b2d8-b90ff19a7653 · outbound

This paper cites Mgpo: Thinking with images via multi- turn grounding-based reinforcement learning.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Mgpo: Thinking with images via multi- turn grounding-based reinforcement learning

Reference 1979

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:54:49.218206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T00:54:48.797363Z digest=sha256:9714fc4b45e1ac3cf739eef2beaeea70c972012af7c6703e73e45610ef9c953d

Observation 34304f19-c3d3-4887-9ef7-3aadb0dfa4fd · outbound

This paper cites Do MLLMs really see it: Reinforcing visual attention in multimodal LLMs.arXiv preprint arXiv:2602.08241,.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Do MLLMs really see it: Reinforcing visual attention in multimodal LLMs.arXiv preprint arXiv:2602.08241,

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-05T00:54:49.115128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T00:54:48.814582Z digest=sha256:b52872d32be72c4efb0b38e6e0a9d397d9a2a20ab9b2ebeb1e476e1f1571b3ac

Observation 9d870beb-a6c7-4cc2-947e-3c64e26ccd69 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.820591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.820591Z digest=sha256:884e388df1f3e0b7597376e778640e942b61e875624eb94b166e144c26cdb888

Observation bfc3d2c6-72be-4554-a5b9-3b0d19f72fe9 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance PaLM-E: An Embodied Multimodal Language Model

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.793928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.793928Z digest=sha256:ec9f1b57a3facc6c48a061e667ff549b13ac80838086d4d29592d08a54636681

Observation 97d9db3f-cfeb-4faf-b572-984c4dfa35f4 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.800997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.800997Z digest=sha256:a21c58c36fb29050db34db238baa3cc2e857ca7913c9fe42b8665d5b9f8a9ce4

Observation 08652df0-1ecb-4470-b233-cf5c238f4dd9 · outbound

This paper cites Improved baselines with visual instruction tuning.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Improved baselines with visual instruction tuning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:54:49.208446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T00:54:48.808322Z digest=sha256:b292aa1262f722686a5642c799cbbfe5f974ada65977e44bedad9085e0ef1015

Observation 7f743f60-859f-4db7-a84e-11729fbe09f1 · outbound

This paper cites Token-Based Affordance Grounding with Large Vision-Language Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Token-Based Affordance Grounding with Large Vision-Language Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:54:49.142536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T00:54:48.804739Z digest=sha256:b5b38723228c2f9c367ab907521d667f1a52a65ff3df8582f33420b8a69a3e0d

Observation 29606749-d338-4124-8d6e-841411590730 · outbound

This paper cites RoboBrain 2.0 Technical Report.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RoboBrain 2.0 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.781637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.781637Z digest=sha256:956468f377fa8019e5a2110c16358d40b548729799fb44014ab04fdd821284e5

Observation 9323f021-c714-4397-a258-cbc5f52e805d · outbound

This paper cites Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.832527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.832527Z digest=sha256:0a3b568063e8c4f5526822f4598e79fb4a5817850f7c4c3e6c15f4b35e93910d

Pith citing papers

No inbound Pith citation observations are available.