Pith. sign in

Paper Citation Record · LEDGER

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

As of 6 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2607.14187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14187 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:16:53.178921Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1483f19f-b5a1-418b-8d45-1f748282ee72 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.065229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.065229Z digest=sha256:3f1376fdce7f57a23d9b01336d13670b1030df82fb3f6ba1ef077ed1c9d14de5

Observation 3fe4b1d2-fd68-4961-9478-07f5cdd54d21 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.411477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.411477Z digest=sha256:dc148aebfa05882ec37c765f25abe9c74bf32e1dfca56c33c6e7a274ef8fbbc4

Observation 70ee18af-e177-4d1f-9c81-642953eb68a5 · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.624046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.624046Z digest=sha256:3ebb83606eaba107946ea72dc7f103f5dca63d9b57e1f04f2b17ad71b3ce48d7

Observation 86b44128-0ad2-4f01-bd20-3da95425aaf3 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.829087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.829087Z digest=sha256:55e224f10f8780a54940b12651d8d9484779db79dd3d5f320e4ff87db5a1d03b

Observation 441af304-fadf-4ee3-b89e-8502d1c10f96 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.902835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.902835Z digest=sha256:442af69d41e476d6504796821c4940d215351764318ae072b86bee77379ce011

Observation 8bb7f226-a638-4647-b3c9-cfb9cf2cc548 · outbound

This paper cites Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.966378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.966378Z digest=sha256:2b643b8b068570ee94217f11b456c700114c418145fc5e55e0dec5f12a20d2f1

Observation a61824be-2977-4df6-9356-7133a8dc1da0 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.984970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.984970Z digest=sha256:20d911d3a1133c356f33e5e1a5dde4bd7f30375d3334e8ac5a010ac2dbb9a748

Observation 68faf5a4-ee68-4e77-89b1-77a87726cd0d · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.989658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.989658Z digest=sha256:18f150b1fa35a60a82dfd270e8d99751ede454d91c41e24aa25e9ff935354300

Observation 6b8cd26f-09cc-4d5a-9b3c-577f5a6a43eb · outbound

This paper cites Rynnec: Bringing mllms into embodied world.arXiv preprint arXiv:2508.14160,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Rynnec: Bringing mllms into embodied world.arXiv preprint arXiv:2508.14160,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.993819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.993819Z digest=sha256:d79c5f1010078750097390ebcd7c8d1da36be5da90f0e802d07a85110b5a9beb

Observation f293a8c0-a783-42f7-aead-801df76ed549 · outbound

This paper cites Rynnbrain: Open embodied foundation models.arXiv preprint arXiv:2602.14979,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Rynnbrain: Open embodied foundation models.arXiv preprint arXiv:2602.14979,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.998413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.998413Z digest=sha256:441d8226516d2987ef546d34107cd981c1f65a29fb7d3903aadcb1ec75e4171e

Observation 1a2bd0dc-7117-4fdc-95d2-d04bc7d30eed · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Emerging Properties in Unified Multimodal Pretraining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.002710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.002710Z digest=sha256:c01082f32dd9ca8db0f7a9514dd64f51642b1bee8cdf563994c95f6dc4f77631

Observation d2b96bf6-d4d6-4a6f-94f3-38a77bbca105 · outbound

This paper cites Rethinking video generation model for the embodied world.arXiv preprint arXiv:2601.15282,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Rethinking video generation model for the embodied world.arXiv preprint arXiv:2601.15282,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.007351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.007351Z digest=sha256:e24f7415c7b803986db36544b423e4a4f878140bc4442511699fef6e6706132d

Observation 80c327f4-c8ec-48bb-b4fc-cc9e65c1ca7a · outbound

This paper cites EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.011642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.011642Z digest=sha256:563e4f790bbccde89d78785e64b17b545a8979737a0f3d5b0795b95e62d00002

Observation ae0e4585-db95-4289-9304-68fa09bea061 · outbound

This paper cites MolmoAct2: Action Reasoning Models for Real-world Deployment.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.016209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.016209Z digest=sha256:c2727d5d54566cb1d0829053ac8575ac23d3ef7b74393c3e100d093f0edcd1e4

Observation bdbeaf69-9a4f-4259-9131-c8080434fb10 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.020900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.020900Z digest=sha256:f9bafcde48ec5fb73e5b8f0b2f32ddba2214a57175d510dca0a1f295c3c3f295

Observation 0a3fb7e9-2d9b-463f-8bcc-297c6f6a4f2a · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.025061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.025061Z digest=sha256:5eb5b835daa7928d6d84cfac86bd9636ce9b21008d07e8313f18c1f7c45fa20e

Observation 62f09888-3c15-4368-ad0b-20dc82fe5cfb · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.029372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.029372Z digest=sha256:ba37de001644e01c3b003ca53f9f8e1d306cefa5ecbf4f08b9f7e4f8a2cf7429

Observation edd55c7f-8b3e-4ec8-8f42-5aa9b0665447 · outbound

This paper cites Ctrl-World: A Controllable Generative World Model for Robot Manipulation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Ctrl-World: A Controllable Generative World Model for Robot Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.033359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.033359Z digest=sha256:5c5a558441a0be5f89f7107cfa8ec36ec316be98b4cb22b1f21ce446842a860c

Observation f0a9d072-836c-442c-93ce-1a0ff3a5bab1 · outbound

This paper cites EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.041803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.041803Z digest=sha256:b832a314ed13cd4b27b40f6dff97fe6edeff73b81d212cfb34ccef6e4bc103b0

Observation b0aac4a1-4c52-4a07-8567-fe51c20ab1d0 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.046467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.046467Z digest=sha256:ef01f3e1039e86c5150b991f45e04040b1e6e8f91492fec0ed887d76ce2ef65e

Observation 99ef8783-cc27-4d87-884f-0ce71a4da012 · outbound

This paper cites an unresolved cited work.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.050538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.050538Z digest=sha256:51e462506ac6557b0bcf37279fd7679586be71f0d0562ff180d5b4cadd970af9

Observation 3a66bdc8-57aa-487b-9a80-58a44eb08f65 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.054573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.054573Z digest=sha256:66e4d8b6e5e0433268e07839d5f92067293e2fbc0712aa75df63f83193a2e721

Observation cbfd2e79-e0a3-4186-baae-1dd0db7044a4 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.058412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.058412Z digest=sha256:2d7a4360432b0a644fb04f0c2223bba96b201b491d38b12d88c244ffdd20ef8c

Observation 46844f05-3ffb-441e-b5fb-d811455e5487 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.706197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.706197Z digest=sha256:19e88cf087c6dd1e91e908b28d3be48570d6ea991dba0e7a6414cbaf9b607e93

Observation 532ddc54-a7fd-4dee-b38a-6a75134d972c · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.063299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.063299Z digest=sha256:df380bd6ce4fd3e7fc28682ea1c3b5515871f08cc272b75d552e5abbdc98f872

Observation 827494e7-2d97-42c6-a2e7-265c38e680a0 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.067819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.067819Z digest=sha256:e55895d51a1ffb2da2eb7fe36d25274d134cbc5c5347f0addb063be85676bc07

Observation e67efccf-9c77-47e4-b0c7-f19c09994fc3 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination MolmoAct: Action Reasoning Models that can Reason in Space

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.072266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.072266Z digest=sha256:8cd92ed5fc35b38a9c82ecb049d77700561d5bd0612ba871fedb6cbf5529c9d6

Observation bff5aa85-5343-4384-84ec-f98fa6fa5973 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.arXiv preprint arXiv:2505.21500,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.arXiv preprint arXiv:2505.21500,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.077240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.077240Z digest=sha256:b0254ce8f944e113f1f2e0bd66ff92cb608d2c4e1b0ba03601bf37de12b2163e

Observation 38d59eae-cda7-4910-88f3-e3c8c5314bd5 · outbound

This paper cites Causal World Modeling for Robot Control.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Causal World Modeling for Robot Control

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.081610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.081610Z digest=sha256:55f54141da97133ab4008dbdf23e4412625db11c6791bf8c17833ec73c5977c2

Observation b9ba8fe4-bfa8-403c-af0e-cca74d46f49d · outbound

This paper cites Mm-act: Learn from multimodal parallel generation to act.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Mm-act: Learn from multimodal parallel generation to act

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.086347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.086347Z digest=sha256:1c66c6eef2b8ddccd51d801a28d033d7c0f7ed9029040563e7da560ef59065d8

Observation e4dd2bde-4167-4e29-b1dc-c66c53065d69 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.090389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.090389Z digest=sha256:e37270825285986bb3606c59e56de4be35318d09d1f9db1f1e5701cbd367a308

Observation 6466870e-29bd-446b-90c4-0c726a26244c · outbound

This paper cites URL https://www.biorxiv.org/content/10.64898/2026.05.01.722168v1.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination URL https://www.biorxiv.org/content/10.64898/2026.05.01.722168v1

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.094833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.094833Z digest=sha256:2b3a4f98c6b7ea0c6a24f486f885eb5a43a4c1ad878764695c5101b136f2fdd4

Observation b23715cf-5277-4c16-9e62-edd407dc59f3 · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination 3dsrbench: A comprehensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.099478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.099478Z digest=sha256:ac32593d9b55d3c5e074fc77063328044f9b7d72dab6d7c01635317c88c9aea0

Observation 8f63dfd1-1926-48cd-a770-f80e3f759670 · outbound

This paper cites GPT-4o System Card.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination GPT-4o System Card

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.103400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.103400Z digest=sha256:5e0a65cfeeaa9ce22c07be9de4d81aaeaa162a5db43f07e2fcd5da5b0d6263ca

Observation f3ea2a64-7297-40c0-9390-389db261c6e5 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.107420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.107420Z digest=sha256:f30542e8358b2b7ca0cce533fd5ee50e69c37d1f9f98e61d7caf8ab8d35566db

Observation 37db1346-387e-4609-9120-c79a43eb0db7 · outbound

This paper cites EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.111912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.111912Z digest=sha256:a3566f454533a68715aa4f4b7ab200fc99aaa19252fb4cef51ded601f1deab73

Observation d4dd7ce6-2cbd-4d6f-a17e-5507dff19f90 · outbound

This paper cites RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.120452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.120452Z digest=sha256:ff032a38f62be47b0206a464a161e9142515dd3c967694ef20b10a8ef406a952

Observation c86811f8-6d5f-427d-a644-7354c8e4db39 · outbound

This paper cites OpenAI GPT-5 System Card.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination OpenAI GPT-5 System Card

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.124550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.124550Z digest=sha256:02b1493e0d56dc903b706b6ea45a3b2999b45aaa4087011f9e47f9d2c16a378a

Observation 42c0d821-160d-4137-aaf4-e782ffa90d8d · outbound

This paper cites BAAI RoboBrain Team.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination BAAI RoboBrain Team

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.128754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.128754Z digest=sha256:90203a27e04c49b7925b10039fad2947dafde666c1653cfb1529a3359b944bbe

Observation 64d1afd2-ab83-4882-81a2-d358a7b7d2ad · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Gemini Robotics: Bringing AI into the Physical World

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.133015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.133015Z digest=sha256:d8801038f319d3f2c3f72f6cd078a1f75101f05be11b735d18cb1bbeacccd45a

Observation 8b676f83-b117-40c3-ba69-e0484e85f341 · outbound

This paper cites HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.137725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.137725Z digest=sha256:e434be20afe4b8e7ffd9a901aac29e651a5b3c3d95e95860ad4875baf1f0c684

Observation 2512f4c5-febf-4f33-917f-de470cb3ce5c · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Wan: Open and Advanced Large-Scale Video Generative Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.142236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.142236Z digest=sha256:9789377e906944b56986ee9262e7921992b5ebea1b65334b299d35b1633059bd

Observation 9378d785-8a77-4277-9a84-06c3b3a4c87f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.146991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.146991Z digest=sha256:559ad982e8fd78208e02fda52aaba4d0f3c577921ee5e89aac196ab78fed4051

Observation d9cb534f-f35d-4b70-b06c-c97799f76cc4 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.151662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.151662Z digest=sha256:d4b8c097e5e1523b7e3f01b501d9e73fe8b978e47fe2c53fe8a659edebe09e74

Observation 3d8f654d-eeb6-4763-bd63-eea21781b577 · outbound

This paper cites A Pragmatic VLA Foundation Model.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination A Pragmatic VLA Foundation Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.156408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.156408Z digest=sha256:71a629da540eaf94bfce2ca3992478e2e713bbb27c43a04aba894e4bda3b38b7

Observation 76dec642-57e2-4091-aa9d-cf732843a2e0 · outbound

This paper cites World Action Models are Zero-shot Policies.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination World Action Models are Zero-shot Policies

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.160827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.160827Z digest=sha256:8107509a0d3677384d61e9b94893f127de8dfd5f43748e035a405d6e6d7a62f0

Observation 1eafb56b-242d-48e1-a279-b236400e7160 · outbound

This paper cites Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.165132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.165132Z digest=sha256:a47aff9695bf5f7ee8186cacf52c91395e6c3d63959622ba35a1d6180d30e16d

Observation 58d9205b-0127-4f60-bcdb-574a6b05700a · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.169515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.169515Z digest=sha256:856487e036c3a806ffb7f7f689d0d15e6053bdad897b46b7cc6b6c74456bf3a3

Observation e6a6cb12-87f5-4a87-933f-14bf1666022c · outbound

This paper cites Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.174546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.174546Z digest=sha256:f2bb64e5ce223315b9849691d4f4abe20c3e80b4b9f498639f5dc6299c6d6236

Observation db948e56-df8f-40b5-b381-53089309c917 · outbound

This paper cites Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.178921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.178921Z digest=sha256:8cdbc80de230c4245ac46907e3c9a76228ef2cc23cc44b63c7278b7a5ea5c5d8

Observation 6eeba8aa-c59b-42dd-b91c-2fd9e9c3d1bb · outbound

This paper cites Sat: Spatial aptitude training for multimodal language models.arXiv preprint arXiv:2412.07755, 3,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Sat: Spatial aptitude training for multimodal language models.arXiv preprint arXiv:2412.07755, 3,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.115877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.115877Z digest=sha256:90bdb7d5879305325838a14824657b9787c1076e310b0920f79a9ec24a19fbcb

Observation 037afa4d-f0de-4652-86ed-d45f604371b3 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.168989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.168989Z digest=sha256:4da22bc388ffe0b4489b9a7b5cc822a3e5138f53351e8abed23c3658f6394b8a

Observation 0af90b43-9b58-4c5a-bcc3-e10ae3acd14c · outbound

This paper cites Qwen3-VL Technical Report.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen3-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.544902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.544902Z digest=sha256:25d386c3344a372ba549fc1bc1052cc63f2ef8bfb8874d481c1ed8b0d3ac511e

Observation 73623032-e266-4c79-a3e4-206c17193e77 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.770586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.770586Z digest=sha256:fdb85b433df6fb0b8432ca6b083f3b43d19533cf0cd2905ddbf991c09d52a6f6

Observation d2d0bc79-c207-4f12-ab3d-3e6fbcf929af · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.277478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.277478Z digest=sha256:4abd4353ad12478d2cd86ba1121cb7c8e04d0c755fd4d793423e52f7df27bea5

Observation 44690fa4-1f37-466d-95ad-61a70d875c4f · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.980185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.980185Z digest=sha256:6221c45afa7e69845080f38b799c793163e8f1602fc1be2c99e0adb84559e11f

Pith citing papers

No inbound Pith citation observations are available.