Pith. sign in

Paper Citation Record · LEDGER

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

As of 13 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.14635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14635 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:33.493545Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 168c138b-5b28-48a2-9e18-51b61348cba2 · outbound

This paper cites RT-1: Robotics transformer for real-world control at scale,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.433041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.433041Z digest=sha256:05b256770a15485e83e1de5e887c9ea0535fdf06ce20da3bcad56ef8bf0fab44

Observation 9d41ccb1-28fa-4344-8721-eb0f3c1270f8 · outbound

This paper cites OpenVLA: An open-source vision-language-action model,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models OpenVLA: An open-source vision-language-action model,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.498298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.498298Z digest=sha256:0b5ff4f46e4642091a7e9024994a608f6d834015717ffb5fe71b80cf276b4915

Observation f4340d76-8cb1-40bd-9926-c82e604cc9ac · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.563364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.563364Z digest=sha256:c673999b2c4958c567d61ccda60579c43008409d77253ecedba1a7545692f99f

Observation 15d08ad3-b1d1-4906-bbc5-4ace3c238c75 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.631505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.631505Z digest=sha256:9513fcc537860eb6fe8f854f6e302dee78c5b5f3e5185358b9c1651d3db08faf

Observation 7b41b2c2-a51e-4c51-959b-28a852d27df9 · outbound

This paper cites π 0.5: A vision-language-action model with open-world generalization,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models π 0.5: A vision-language-action model with open-world generalization,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.689972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.689972Z digest=sha256:e6588e5fad9db6d7eb085dc092aae34102bae971b2c0147625966472f7a6b8b5

Observation 05b9df07-540b-435f-9d44-3ad02a2d2a90 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.776432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.776432Z digest=sha256:1e3a20e5ff233044e51359899911226c8c34c95e74e5c1025464747dfbe1e5f5

Observation 77c6d70d-6cb4-45f1-ba0e-04c1d9600072 · outbound

This paper cites VQ-VLA: Improving vision-language-action models via scaling vector-quantized action tokenizers,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VQ-VLA: Improving vision-language-action models via scaling vector-quantized action tokenizers,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.861393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.861393Z digest=sha256:8848f907e20f6c494f9f015c1744e662e2154ffeb95439b436278eb5184b1755

Observation b1533588-0a9b-4320-a324-d7079b8c57bb · outbound

This paper cites Latent action pretraining from videos,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Latent action pretraining from videos,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.942388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.942388Z digest=sha256:74b85216a0f563b8486f2bc86b319a4d63deb68466deb4575096f169f51b8ae7

Observation 476587b7-83dd-4d1d-9db3-723908f040ae · outbound

This paper cites Learning to act anywhere with task-centric latent actions,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Learning to act anywhere with task-centric latent actions,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.060679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.060679Z digest=sha256:5a540243b303bd4dbea46863282e991f3656cf7000452af7b50da50b491ddfdc

Observation b3014be4-9765-487d-8d81-2ee225112ea3 · outbound

This paper cites RotVLA: Rotational Latent Action for Vision-Language-Action Model.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RotVLA: Rotational Latent Action for Vision-Language-Action Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.181563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.181563Z digest=sha256:4683747d5f1f86ff7efdff0fd4e6df9390bf6238c28d20752022e31f753f8cf7

Observation ef5d9e8a-f341-4bfa-91ca-2e9dbd3a5b62 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Learning transferable visual models from natural language supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.319349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.319349Z digest=sha256:62b679a496add12d30db50f5e0b43cf0426f5e4e67e51c562bd41664824269e6

Observation 2a7b6a1c-7f0c-4f82-805e-08098bf12975 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Align before fuse: Vision and language representation learning with momentum distillation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.461696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.461696Z digest=sha256:04055c69f8c20a9270b89df4dffd05e200febb3b2d0b2632bb79649c9e383b3e

Observation 0e83d065-144b-446d-aecb-f7e79013b687 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.599740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.599740Z digest=sha256:731cf5980a7d3f61103e3cf4d7fdd34b3dbcf9380d9909c5064da178a573d2ab

Observation e9de61cc-a106-44b7-9630-e6e93b987c5e · outbound

This paper cites Visual instruction tuning,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Visual instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.685381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.685381Z digest=sha256:536db428f4489cbb59719f0a66d2630951343045e8deaa05c3328a7fb76b7660

Observation ed05aa5d-8654-43c9-8fe2-d5c4f95ef435 · outbound

This paper cites Perceiver: General perception with iterative attention,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Perceiver: General perception with iterative attention,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.777303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.777303Z digest=sha256:30ac85c22c2e08a93b11e72762eee7351b011631eeda1e388613924fa50c31e9

Observation d7716786-d9b1-4637-94f3-fc2a37d4d9fe · outbound

This paper cites Flamingo: A visual language model for few-shot learning,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Flamingo: A visual language model for few-shot learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.870154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.870154Z digest=sha256:2b65b0d72dd7237672f2ed9a1c8948aa83a2b88d4acca443a1ad529942b7ad65

Observation 923415a7-9bd2-49a4-88fe-a6d3bd47c635 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.960688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.960688Z digest=sha256:a5f2e85ede6ffda2dc5e21cdfb7d0adeef2f085b36a2a683f9666524b44964ce

Observation 1c49c74d-9986-4f96-b0a7-db0f3047c11a · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.053667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.053667Z digest=sha256:4e174fc8a0bd3719ab18ac6ca28f29308f970c5101494ebb6d56a62025af3cf3

Observation 9637b991-567d-4564-8175-23d732540a7d · outbound

This paper cites VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.166730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.166730Z digest=sha256:8a48bb438b2c0e19dc3a8c254680e3830663e424ee4e4e3a921d9d17818a6be1

Observation ead814d6-4973-4ae1-9688-f1eda41eb078 · outbound

This paper cites DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.269443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.269443Z digest=sha256:6dc4862622d124e0b1cdebd0a49c6275ffcdfa50fc18d6a2f33e5a9558147fde

Observation 074e1e22-2d11-4cd4-b40b-49822b0b9199 · outbound

This paper cites HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.334140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.334140Z digest=sha256:54914f781c4156cfa6925595f4ae214a391a262cf1e7205f26107cba77275ff8

Observation 0bcd7638-d956-4d7a-9b0c-18e40d3d4ef2 · outbound

This paper cites LARA: Latent Action Representation Alignment for Vision-Language-Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.411913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.411913Z digest=sha256:492b8d0577a634aad212f303af2fa68dda4993956a95d67400ff0eea45c7b67a

Observation 56aa65d7-e7da-4bf5-a083-24c74d54eef3 · outbound

This paper cites Making Foresight Actionable: Repurposing Representation Alignment in World Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.476784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.476784Z digest=sha256:e7a209d41fb3bb9e76131225f529d005e9c48e251080c11d62818e02997d4113

Observation 8a87a132-008f-48dd-9d21-a3f5767226a7 · outbound

This paper cites A formal basis for the heuristic determination of minimum cost paths,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models A formal basis for the heuristic determination of minimum cost paths,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.541114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.541114Z digest=sha256:222f43038c28c2b2aa7283c78e1ca2e4d3f8e423023b85f48fe49def86b5286f

Observation 9cec96d1-84a1-4269-8b49-73017f76d0ec · outbound

This paper cites The dynamic window approach to collision avoidance,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models The dynamic window approach to collision avoidance,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.602905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.602905Z digest=sha256:059592c33670f814fa2a9d5cc46e55999d2bb92ea73c9762fb0bafd96131b3e0

Observation 7acdcb17-7df6-4517-bd60-2c2202c9e3e3 · outbound

This paper cites ViKiNG: Vision-based kilometer-scale navigation with geographic hints,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ViKiNG: Vision-based kilometer-scale navigation with geographic hints,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.688754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.688754Z digest=sha256:cdccaff3ad75f3bb0e5dd7d3584c6457752f0f5b508d3b483a111a98b7f91343

Observation 45b6e0e3-aa74-46fc-a565-6864dce056a6 · outbound

This paper cites GNM: A general navigation model to drive any robot,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models GNM: A general navigation model to drive any robot,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.745987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.745987Z digest=sha256:3c96445455d55502f94d0bf2c71c6ab10f74fac64f2290d8f2375c074dc4b9db

Observation 60b2c79c-6854-4342-9f7e-59017e491a0f · outbound

This paper cites ViNT: A foundation model for visual navigation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ViNT: A foundation model for visual navigation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.799338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.799338Z digest=sha256:8bf270125312e2033f5088cba5e8020e4c91710f23d4c3b38fccd30af4dd370c

Observation 3a20c86d-0c4d-4318-afe6-1498b8d62cc5 · outbound

This paper cites NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.849984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.849984Z digest=sha256:9f7307974b20c861f47441be97fff60073dd372592ea28f1253b88e99a11f411

Observation c0e19cde-c6eb-4aca-8ce8-022e72b98cc1 · outbound

This paper cites LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.928705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.928705Z digest=sha256:53762dcb07bcacd5ad0afa6f320fd27e428646663189c8c8b2215321bf355e2e

Observation c3517109-29f0-42f2-8a4e-e7f4e012a460 · outbound

This paper cites NaVid: Video-based vlm plans the next step for vision- and-language navigation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models NaVid: Video-based vlm plans the next step for vision- and-language navigation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.989633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.989633Z digest=sha256:9c8dfb140ed2f78fee3e5b32d3767bab25839665adaf55cb3377fa108033d114

Observation 641a3b0e-511b-44fb-a13e-ee9f7b980fa6 · outbound

This paper cites Uni-NaVid: A video-based vision-language-action model for unifying embodied navigation tasks,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Uni-NaVid: A video-based vision-language-action model for unifying embodied navigation tasks,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.050463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.050463Z digest=sha256:4daf019e3c23d173f2731f104734b491af0273da8ad02db6f6aea6965c2112bc

Observation a8c0ee6e-9933-4b35-9cf4-7dee1f335325 · outbound

This paper cites Embodied navigation foundation model,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Embodied navigation foundation model,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.102790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.102790Z digest=sha256:fe93ad2ec055b6d32932a40b1fa798821ef039aa8cbeaa4caa79bcce82aaa94c

Observation d593800d-480f-4f4d-b9fd-8912113eed2b · outbound

This paper cites Navigation world models,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Navigation world models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.173003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.173003Z digest=sha256:ae327c9e23a8923dfd6d6529ace05655d008b88be5e2499a02707e63fdf71086

Observation 42860d4a-8a13-45c9-adb6-3eccb20c7d28 · outbound

This paper cites Habitat: A platform for embodied ai research,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Habitat: A platform for embodied ai research,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.225527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.225527Z digest=sha256:c34edcbad6626c68ba37c8124d50d60a55c328c3512f04b4769b3d963169fabe

Observation 5ed974af-8670-469b-bac0-d40b8af23c2b · outbound

This paper cites ObjectNav revisited: On evaluation of embodied agents navigating to objects,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ObjectNav revisited: On evaluation of embodied agents navigating to objects,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.327882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.327882Z digest=sha256:3ed9ca200e7304ee16c0109e781a2095c958134ddf4314b48a71aa2844b62a16

Observation 4d62678d-e558-41a8-b12a-c0ee58e35ac4 · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.403060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.403060Z digest=sha256:6bc12bb7811ea0b2c6da00ba068a8e3b8ff9772fe4ae5cac20f547c56ac92729

Observation 10b45c3a-54d0-4c01-9a52-5529cd98d10e · outbound

This paper cites In the main training setup, this parsing uses thespecial-token branchexclusively.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models In the main training setup, this parsing uses thespecial-token branchexclusively

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.482082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.482082Z digest=sha256:a8d017963548cd8d21d8f9dec98ac40ce99d8f968b43489017b6d204a6b3ae02

Observation 308a49be-7698-4501-bf4d-2d46507a8aa4 · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.565185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.565185Z digest=sha256:aef16ad3e742a87ce86bcf7342d12021bb7f1e29711a2dfcf5ff8daa34dccb76

Observation cd1d99b8-d009-4994-9894-fb5cacfddc32 · outbound

This paper cites make a left turn.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models make a left turn

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.642066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.642066Z digest=sha256:d541109b5d17a7c7c5ec836f59c77ef90e311fa3dbf6a33cf116da18ec1ec096

Observation 49c84480-0031-416f-b8f9-295d3eb852dc · outbound

This paper cites Thedirect-fusion baselineuses the minimal shared-context interface described in Sec.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Thedirect-fusion baselineuses the minimal shared-context interface described in Sec

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.746572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.746572Z digest=sha256:d116e526055eed4480ed07e6effa91f850f760955ef5efa8c52e2b2c81099d6a

Observation c118e976-ecdd-4dfc-83a7-1c27e862da10 · outbound

This paper cites These set- tings differ in which parts of the inherited multimodal pathway are exposed to action-loss gradients.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These set- tings differ in which parts of the inherited multimodal pathway are exposed to action-loss gradients

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.820960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.820960Z digest=sha256:cae13eb0a43922e5f2a156fc8fbc2d6915389deb68b8d46a06dbca3450320822

Observation 7554865d-e876-417b-9474-1fda0ca23fbb · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.876278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.876278Z digest=sha256:a9db772fac8e85858c2a29fcbda41e2beaee57b936558739ed0c83ad6ae904c9

Observation 520758be-a9a1-48c8-bd17-3eb81162391c · outbound

This paper cites Token-wise rewriting and rewriting-subspace analyses examine how strongly the inherited pathway is rewritten and whether that rewriting is broad or targeted.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Token-wise rewriting and rewriting-subspace analyses examine how strongly the inherited pathway is rewritten and whether that rewriting is broad or targeted

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.984898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.984898Z digest=sha256:27dbb3b40b06c3a54384ef7fb95167e99c958691481776c75433f1c616888e6d

Observation 1fc04309-d7d6-4a5f-ab9e-e3044dfd14de · outbound

This paper cites These statistics are supporting measurements: they are not intended to replace the token-level visualizations in Sec.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These statistics are supporting measurements: they are not intended to replace the token-level visualizations in Sec

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.064517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.064517Z digest=sha256:9fdd34d929f0d795619c8bdd576b43e2587a7de7292a9d96ebda762711d30872

Observation 1f877ce6-2ddb-4ece-971e-ff9a85ac6098 · outbound

This paper cites We partition tokens into boundary, control, spatial, and other groups, and report each group’s share of the total rewriting budget in Table VIII.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models We partition tokens into boundary, control, spatial, and other groups, and report each group’s share of the total rewriting budget in Table VIII

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.108432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.108432Z digest=sha256:a8b5fdc8707e8ec7e0c861edccf1ed2b97a0a649ba5e79b5d43893e86101b5ea

Observation 5a284d3e-0b81-4b44-bc9d-dee06831325f · outbound

This paper cites We next ask whether this selectivity is also reflected in how rewritten dimensions are organized.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models We next ask whether this selectivity is also reflected in how rewritten dimensions are organized

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.159920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.159920Z digest=sha256:0b47cc1e5d5e5a1223679261b95db943b1990aa8f2e5d13c6f838a3ae2556fd8

Observation e3a8b086-09c5-4be3-ad73-9e8c228cfc88 · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.236218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.236218Z digest=sha256:4797eedba2157d5ac7a3b13a450570ea97d0fbe13b1121daa9d4e0be7c2b350b

Observation 71385eb9-5416-4727-a888-a2b4a0614e02 · outbound

This paper cites These statistics provide additional views of how closely an action-loss-exposed atten- tion map remains aligned with its action-update-blocked refer- ence.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These statistics provide additional views of how closely an action-loss-exposed atten- tion map remains aligned with its action-update-blocked refer- ence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.281574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.281574Z digest=sha256:2848f8019237677782f430b3fedd15b3aec20926f69714c11c15bbb84755988e

Observation 6701a273-65d1-432f-8b5c-3b6f98e49d3a · outbound

This paper cites 18 provides complementary distributional views of attention stability across the full set of phrase-level comparisons.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models 18 provides complementary distributional views of attention stability across the full set of phrase-level comparisons

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.348531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.348531Z digest=sha256:2e0bbb3f2d62ac560ecc4a7cbaa2ef4043a3d4556a5ce15e07347f015c646f2d

Observation c6967e63-079c-47b5-b8e4-eb3ea5aecc2d · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.434174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.434174Z digest=sha256:5a2cda625ae37b5c50a426d0d60d0d3fddf5341ba12a04ec0ec5a01298de4c6b

Observation d4f04fbf-6567-4331-8616-3e12c1f68dd2 · outbound

This paper cites 20 shows that the all-head average can hide substantial head-level heterogeneity.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models 20 shows that the all-head average can hide substantial head-level heterogeneity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.493545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.493545Z digest=sha256:e133bc9368d6f19f8c5699a6c398d0a87fb7d5c87c3af51b6fb968e71b760686

Pith citing papers

No inbound Pith citation observations are available.