Pith. sign in

Paper Citation Record · LEDGER

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

As of 20 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2607.14187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14187 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:16:53.178921Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:47.525182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T04:54:24.211818Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1483f19f-b5a1-418b-8d45-1f748282ee72 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.065229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.065229Z digest=sha256:161944cfca18d52469617d7d62b84dd42adfbb381815e67ceb8b463afc6820d4

Observation 3fe4b1d2-fd68-4961-9478-07f5cdd54d21 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.411477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.411477Z digest=sha256:3a9e2f76728c66c2bf69034f33745659435a91d11304f03072e5289b7e4c42c3

Observation 70ee18af-e177-4d1f-9c81-642953eb68a5 · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.624046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.624046Z digest=sha256:071cb2474aaa65b6c2e2981e652f407688e9efa8693245fc40c6d0402d6fb8d9

Observation 86b44128-0ad2-4f01-bd20-3da95425aaf3 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.829087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.829087Z digest=sha256:7bc31515c3409f8cb17b74ee841c78a980ff107bec8b21848f2269227d707b65

Observation 441af304-fadf-4ee3-b89e-8502d1c10f96 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.902835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.902835Z digest=sha256:d60049b0bc904cf487ed1815db8a0c5db128e5d07482a425d3c9b1c4285828fc

Observation 8bb7f226-a638-4647-b3c9-cfb9cf2cc548 · outbound

This paper cites Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.966378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.966378Z digest=sha256:dbbf2bb86c51a4608428c8e8d07aae7534280a003b355a3cc6a323304c9edf2b

Observation a61824be-2977-4df6-9356-7133a8dc1da0 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.984970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.984970Z digest=sha256:df16df9608f41ce3833604a5d39fba4918c6a2720473eb9b04128ab129588c11

Observation 68faf5a4-ee68-4e77-89b1-77a87726cd0d · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.989658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.989658Z digest=sha256:f44b25c8de74096448ae5386e9e75b29e7f390384785261c0792ccd4641bfa1c

Observation 6b8cd26f-09cc-4d5a-9b3c-577f5a6a43eb · outbound

This paper cites Rynnec: Bringing mllms into embodied world.arXiv preprint arXiv:2508.14160,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Rynnec: Bringing mllms into embodied world.arXiv preprint arXiv:2508.14160,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.993819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.993819Z digest=sha256:d551ba985b4422dbe87d397eb2d20dbe811ac62f849a3754005c677d528ccb94

Observation f293a8c0-a783-42f7-aead-801df76ed549 · outbound

This paper cites Rynnbrain: Open embodied foundation models.arXiv preprint arXiv:2602.14979,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Rynnbrain: Open embodied foundation models.arXiv preprint arXiv:2602.14979,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.998413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.998413Z digest=sha256:11176c6d3b25be3b17370270c09b18d0b31402256aa9652788965260f0a373aa

Observation 1a2bd0dc-7117-4fdc-95d2-d04bc7d30eed · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Emerging Properties in Unified Multimodal Pretraining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.002710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.002710Z digest=sha256:96206a0024110ebe841b3cfa44b2a03cc33132978b2f2a678b8853e3e4f58402

Observation d2b96bf6-d4d6-4a6f-94f3-38a77bbca105 · outbound

This paper cites Rethinking video generation model for the embodied world.arXiv preprint arXiv:2601.15282,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Rethinking video generation model for the embodied world.arXiv preprint arXiv:2601.15282,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.007351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.007351Z digest=sha256:8f79e5f295544252cd2309b2ef6b00df6799ab0c4e9a9c6c50fb3c70e5962810

Observation 80c327f4-c8ec-48bb-b4fc-cc9e65c1ca7a · outbound

This paper cites EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.011642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.011642Z digest=sha256:1491aacad3c1ed4bc4824a7b046427f24b19a90b83716bf19a15d1c0f5dd0a83

Observation ae0e4585-db95-4289-9304-68fa09bea061 · outbound

This paper cites MolmoAct2: Action Reasoning Models for Real-world Deployment.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.016209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.016209Z digest=sha256:0b4149bdd270346d7f6478506596238f7c81f7594621d10c5574d803427ce37d

Observation bdbeaf69-9a4f-4259-9131-c8080434fb10 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.020900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.020900Z digest=sha256:fc84cec071109937ad6235a1a0b998c6f990b754225f544182d9c9756c08bb74

Observation 0a3fb7e9-2d9b-463f-8bcc-297c6f6a4f2a · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.025061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.025061Z digest=sha256:36857fe69722b2096387a25c9b2457800eb0ec261ed7d74fb441bc88e9debb6b

Observation 62f09888-3c15-4368-ad0b-20dc82fe5cfb · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.029372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.029372Z digest=sha256:a7fec9bc0285f67f074a86ef0e58e4aaab0463f9616c22bee5d2d549574c37a6

Observation edd55c7f-8b3e-4ec8-8f42-5aa9b0665447 · outbound

This paper cites Ctrl-World: A Controllable Generative World Model for Robot Manipulation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Ctrl-World: A Controllable Generative World Model for Robot Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.033359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.033359Z digest=sha256:cd018b8856430f868710a10fc512cf341634f04afcc9b1e23e568d98217d5c49

Observation f0a9d072-836c-442c-93ce-1a0ff3a5bab1 · outbound

This paper cites EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.041803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.041803Z digest=sha256:afdf184d6a9367b4a75c3b05973c4f2b931e303aadc8f10bfa2e7018aa3caff4

Observation b0aac4a1-4c52-4a07-8567-fe51c20ab1d0 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.046467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.046467Z digest=sha256:1e1ef281d50e20d47f867b498bf41a592307b72a34fb98c5c55d7e94fac4aa8a

Observation 99ef8783-cc27-4d87-884f-0ce71a4da012 · outbound

This paper cites an unresolved cited work.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.050538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.050538Z digest=sha256:94355776443c8f491c85033a5effc552dae39b6b469862e46cc1349b9ac320ee

Observation 3a66bdc8-57aa-487b-9a80-58a44eb08f65 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.054573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.054573Z digest=sha256:32cb7da6dd64dbb5a0941ba730789a9efc66ded39ffc6d1aa81cb99be39a4c6f

Observation cbfd2e79-e0a3-4186-baae-1dd0db7044a4 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.058412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.058412Z digest=sha256:9d281d5a0c7298dbc3b9eb4b85724d8c0fba0102a10b23cacc1b97b1fc6aadd4

Observation 46844f05-3ffb-441e-b5fb-d811455e5487 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.706197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.706197Z digest=sha256:0f4643265f7b333e26b545eaf871c2be961c4709c46dfa357a7afd843106c660

Observation 532ddc54-a7fd-4dee-b38a-6a75134d972c · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.063299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.063299Z digest=sha256:b96ddb3991e462b2f3883a06ef5f53fc95365b1a9b3b14588432ce61a27f79b9

Observation 827494e7-2d97-42c6-a2e7-265c38e680a0 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.067819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.067819Z digest=sha256:ff3fca4f86420af0ed27bd5a1a505c47faf2a5cbcf66e4f3ef1eb8d185b85997

Observation e67efccf-9c77-47e4-b0c7-f19c09994fc3 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination MolmoAct: Action Reasoning Models that can Reason in Space

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.072266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.072266Z digest=sha256:1c05bb3c0f8212c3fd07efd21411acc589419eb6ec5a69458a8c9c1d1749fd89

Observation bff5aa85-5343-4384-84ec-f98fa6fa5973 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.arXiv preprint arXiv:2505.21500,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.arXiv preprint arXiv:2505.21500,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.077240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.077240Z digest=sha256:d8c2bc6763707c40b7a77c2c6ecdd6059055659a7725ea6f4972b1fbc04b190d

Observation 38d59eae-cda7-4910-88f3-e3c8c5314bd5 · outbound

This paper cites Causal World Modeling for Robot Control.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Causal World Modeling for Robot Control

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.081610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.081610Z digest=sha256:dd6bd346b9be1fc34eb5cfba2489af610f190e3b0080707d3c07fe9566c01c46

Observation b9ba8fe4-bfa8-403c-af0e-cca74d46f49d · outbound

This paper cites Mm-act: Learn from multimodal parallel generation to act.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Mm-act: Learn from multimodal parallel generation to act

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.086347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.086347Z digest=sha256:5a41df03669a15022a6941f8d1ebb43bcca284a2c4c61235fb923f84a0cd8a84

Observation e4dd2bde-4167-4e29-b1dc-c66c53065d69 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.090389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.090389Z digest=sha256:b0fac4896efaaba0a61db6a6ada3c51962d0908ac5728b6146685fe5df31fb00

Observation 6466870e-29bd-446b-90c4-0c726a26244c · outbound

This paper cites URL https://www.biorxiv.org/content/10.64898/2026.05.01.722168v1.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination URL https://www.biorxiv.org/content/10.64898/2026.05.01.722168v1

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.094833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.094833Z digest=sha256:8450f41a0c0d3c92243b613437311ad63f497a1d31ec93cfcd77a8d6cf8f013d

Observation b23715cf-5277-4c16-9e62-edd407dc59f3 · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination 3dsrbench: A comprehensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.099478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.099478Z digest=sha256:8c84f22daadb16f530f0dd31670bf23ba2af05a6c96906f266e565f23f3645ac

Observation 8f63dfd1-1926-48cd-a770-f80e3f759670 · outbound

This paper cites GPT-4o System Card.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination GPT-4o System Card

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.103400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.103400Z digest=sha256:4ec07f31a85455f131810ee639076b724a83ce384d4c7c1a33e8ff58407090bc

Observation f3ea2a64-7297-40c0-9390-389db261c6e5 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.107420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.107420Z digest=sha256:7f003440355c301f428708bec07e1d1435f6447c203aa2d1e5f85d021d83ffbd

Observation 37db1346-387e-4609-9120-c79a43eb0db7 · outbound

This paper cites EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.111912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.111912Z digest=sha256:3525ecc79487c322060b116af4661aa6ad3a105046bfc1fc8a5e2f2bdfdacddf

Observation d4dd7ce6-2cbd-4d6f-a17e-5507dff19f90 · outbound

This paper cites RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.120452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.120452Z digest=sha256:56f34a24f9f288516bcb91c3b48cadf435b2b4b52e41ea9a80caa0175c21ccaa

Observation c86811f8-6d5f-427d-a644-7354c8e4db39 · outbound

This paper cites OpenAI GPT-5 System Card.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination OpenAI GPT-5 System Card

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.124550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.124550Z digest=sha256:41c01a9110622d3f08d42ad19e7b3aff25f55160a4f74a8d75fd25526d371be1

Observation 42c0d821-160d-4137-aaf4-e782ffa90d8d · outbound

This paper cites BAAI RoboBrain Team.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination BAAI RoboBrain Team

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.128754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.128754Z digest=sha256:cefe0538ca3678c7189839f05863dc8d851ea978e63fcb50471602ddf721fa59

Observation 64d1afd2-ab83-4882-81a2-d358a7b7d2ad · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Gemini Robotics: Bringing AI into the Physical World

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.133015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.133015Z digest=sha256:a8e9a0512a1779cf690a0008ed11f21d5dbaadc40f6e29bca4a64cdd57f54d0d

Observation 8b676f83-b117-40c3-ba69-e0484e85f341 · outbound

This paper cites HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.137725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.137725Z digest=sha256:dc02ed2a371a31841a553b5a5eace4dfb838034753c4f5a69f92b0e94505a2ce

Observation 2512f4c5-febf-4f33-917f-de470cb3ce5c · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Wan: Open and Advanced Large-Scale Video Generative Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.142236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.142236Z digest=sha256:f7cafe0395a568e20fc19d353381b4d7661d3fcce59918a8d53b4331938e1ef0

Observation 9378d785-8a77-4277-9a84-06c3b3a4c87f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.146991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.146991Z digest=sha256:c476f7b4e9cbfc6fc24e6cd2aa73541bbd2f231800da378772992ea073dca07e

Observation d9cb534f-f35d-4b70-b06c-c97799f76cc4 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.151662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.151662Z digest=sha256:ea725a02697d146f14143a31b8008458b2cb7d8d152250efd06c9205f9cdbdf0

Observation 3d8f654d-eeb6-4763-bd63-eea21781b577 · outbound

This paper cites A Pragmatic VLA Foundation Model.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination A Pragmatic VLA Foundation Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.156408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.156408Z digest=sha256:e66e30c9d56d9cc6919c65db93c1a1af5497bc8b572889ce84a5ff4c28cdf4b0

Observation 76dec642-57e2-4091-aa9d-cf732843a2e0 · outbound

This paper cites World Action Models are Zero-shot Policies.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination World Action Models are Zero-shot Policies

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.160827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.160827Z digest=sha256:da5b744aa06d3f7896cf86ca0c0b2995bcc38a4d53ebb8164320c9be4307564f

Observation 1eafb56b-242d-48e1-a279-b236400e7160 · outbound

This paper cites Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.165132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.165132Z digest=sha256:9f900726a05a935545842aadb33895b4a8867d95be8d33123e2db0a6d43fc52b

Observation 58d9205b-0127-4f60-bcdb-574a6b05700a · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.169515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.169515Z digest=sha256:9704765ac53787be61cc7fe660117366e0e6ce776a9ed7258b4722e9ce409b58

Observation e6a6cb12-87f5-4a87-933f-14bf1666022c · outbound

This paper cites Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.174546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.174546Z digest=sha256:fe6dc6eaeb2f8e19df562b043f54b8214a8d190fe37452a4c7ec57c66e5c0a1e

Observation db948e56-df8f-40b5-b381-53089309c917 · outbound

This paper cites Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.178921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.178921Z digest=sha256:ce95157cf80844741c8788e07d4e1d88c82bbc3a50165f77596e9857def2edf9

Observation 6eeba8aa-c59b-42dd-b91c-2fd9e9c3d1bb · outbound

This paper cites Sat: Spatial aptitude training for multimodal language models.arXiv preprint arXiv:2412.07755, 3,.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Sat: Spatial aptitude training for multimodal language models.arXiv preprint arXiv:2412.07755, 3,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.115877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.115877Z digest=sha256:7b003051c160052df4c9ad1d7f4e4840696f97bf3be00dfa383d25b6a06b6cc8

Observation 037afa4d-f0de-4652-86ed-d45f604371b3 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.168989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.168989Z digest=sha256:6d603ec74f5a2c42501d631e42d9359913524f37968a8609e680f7e68ba019d8

Observation 0af90b43-9b58-4c5a-bcc3-e10ae3acd14c · outbound

This paper cites Qwen3-VL Technical Report.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen3-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.544902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.544902Z digest=sha256:0798a1b69bd39f2327e528c45a366fbb41a54c5d2e8edf163a22391e8abf8bbd

Observation 73623032-e266-4c79-a3e4-206c17193e77 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.770586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.770586Z digest=sha256:82b2a33896be452d892ec19f80494478c048a8432c30ecbd05fc69e972df9a3e

Observation d2d0bc79-c207-4f12-ab3d-3e6fbcf929af · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.277478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.277478Z digest=sha256:aab42f785cb62e81dbf61fd40dfa104678bc4849a92effe7737e13b85198e6fb

Observation 44690fa4-1f37-466d-95ad-61a70d875c4f · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.980185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.980185Z digest=sha256:130a6f24ac4a81b05c615404316770f16921b9eb2eda68ea10da2ef5328ac0dd

Pith citing papers

Observation 9521893e-d7e4-4956-8b02-fd3118d003c6 · inbound

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment cites this paper.

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:54:24.221679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T04:54:24.149323Z digest=sha256:2fe198b0699062ea7bddbbf297bd8714b0446558f96574d9f187411d79e21f9d

Observation b40867ae-b406-4ec6-903b-aa2c6ede4220 · inbound

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment cites this paper.

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:47.525182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:17:47.525182Z digest=sha256:6720611476920fe9c10c24e2dd8bdea485c1cdb917a38b4ff327350aacc1118c