Pith. sign in

Paper Citation Record · LEDGER

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

As of 12 August 2026, this Paper Citation Record lists 100 of 152 outbound references and 0 inbound Pith citation observations for arXiv:2608.06756.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06756 v1

Coverage vector

measured 100 of 152 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:34.990563Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 152 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86fa4e9e-bea2-4d0e-bafb-9151c55c7ede · outbound

This paper cites Learning transferable visual models from natural language supervision.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.427180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.427180Z digest=sha256:45651dbe975e6c14e7428eda5f917d748f00557a87330663581cfe5478a4c517

Observation 1b686d8d-8975-4482-80e0-a87d317f2e8d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems (NeurIPS), 2022.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems (NeurIPS), 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.508527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.508527Z digest=sha256:eff6f00637cbd48ece8cbe479407aebaee4800f14221414031aa98864cfefb90

Observation 9558b71e-665b-4a39-81e4-c7be360f62ec · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.550014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.550014Z digest=sha256:56112b3b61569153b7a4d597b47c622587606086fb9289e6441b02c99a80579b

Observation dd610574-b780-475c-8b93-737ddcbab4ae · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.555035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.555035Z digest=sha256:98b9b5d4e69caab38cbc2f4228bd7a82accf0d52b60de6edd1dbd561c29a56fb

Observation a1ce5d4f-fff2-4c7f-9659-dbf2835af667 · outbound

This paper cites Visual instruction tuning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.560003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.560003Z digest=sha256:bdd0bed3805b3ae911cacf7dd1c2010b35b45216d3ae2abc1fe8df09a3cbf354

Observation 8c9e1efe-c2f4-413e-8834-fd4ec95c0b55 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.564957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.564957Z digest=sha256:504655977d7e309dff313a8b2277a6bc412e7b6235aa54060c5be55f145aaa98

Observation 03b4107f-aab0-430a-a07c-e806c9b72ab6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.571389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.571389Z digest=sha256:766d8a37a68054abbb8aca96e40ef0a03a3b08845800c720e6fede83cca8f802

Observation ef7ecffc-478c-4a73-9003-862bf182fa39 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.575729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.575729Z digest=sha256:6f3338d6c7c045237c6a3ccc63f56213d50a75b2e76acb5e94415a0b03d11097

Observation 6128c3a3-393b-43dd-afa6-abb521a80b06 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.628607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.628607Z digest=sha256:045e3438145c8d727b0c5a5ce2c66efa502cd2328d27435cc7d10321967b1c9c

Observation 0aefddd6-7be9-4181-b512-9e4dea5b8264 · outbound

This paper cites Code as Policies: Language Model Programs for Embodied Control.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Code as Policies: Language Model Programs for Embodied Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.705999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.705999Z digest=sha256:6d0b7286cc83d670d45d928f2db336272233d182c9497526badfa8b23869b9e2

Observation cc13936a-121e-42ec-afe3-48b8ee51cafa · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.752781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.752781Z digest=sha256:5b8c07a302c0a979628c803224d26c6ad53cb613ac15144b24b3b9d74e048634

Observation 13aa88af-1dcc-432a-9a86-1f8f00e72e8d · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Gemini Robotics: Bringing AI into the Physical World

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.841005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.841005Z digest=sha256:adb2c04d7b759fee46eb2c815ebae41f3dda733815581d8901f37e57488cddba

Observation 6bb8a24e-72ac-4260-b6c5-6e26a666a9a2 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.953272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.953272Z digest=sha256:bb5c827d62f073eeb21309218e7aaf4d22eb8f39024578025a0e5a10159acdcc

Observation 5666ebd1-4296-40d0-b628-97027be56e0f · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.957799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.957799Z digest=sha256:86b984be28ecc67799e8354cc792c3f00ed6073b76fe662b4661489a378465ba

Observation d8e06185-f056-48cb-acf6-b044c6e7f7cd · outbound

This paper cites Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.962774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.962774Z digest=sha256:782390248c9b7d112c30387ee05ca15927b883fe428beb28d5598008b49ce653

Observation 6dbc5041-6eaf-442d-9b12-5593da0039ea · outbound

This paper cites Rynnbrain: Open embodied foundation models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Rynnbrain: Open embodied foundation models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.966568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.966568Z digest=sha256:ccea24fcc0cccf6cf42f62345f141e2a7e26fbd4ab3d14d0c6d63fe75afe790a

Observation 162a4f34-dbd6-4735-a7cb-c240bb768c4b · outbound

This paper cites Hy-Embodied-VLM-1.0: Efficient Physical-World Agents.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.637598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:32.978898Z digest=sha256:8e345f59ead8fbc654b9cb2c0ee0444e220ce7da2e09d1574c5baa3794a683c2

Observation f8b0690c-20b7-4e47-9f58-65448b0824d2 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Cosmos 3: Omnimodal World Models for Physical AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.991114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.991114Z digest=sha256:6bd3fb40f3f3157cc11905ce9e1f5812681a29c21180b03a75ff691038917ed9

Observation 144621a8-978e-4ce9-85a6-2f52b3fead21 · outbound

This paper cites MiMo-Embodied: X-Embodied Foundation Model Technical Report.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MiMo-Embodied: X-Embodied Foundation Model Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.096528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.096528Z digest=sha256:20fe85eababd4547700b88fd0232edb05ad292881a5ed548349831093fa19c41

Observation e1a0f103-d737-4b61-a4bb-243c9b4052bf · outbound

This paper cites Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.155569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.155569Z digest=sha256:93065305eaa20fd27b3ea81b3247b6cac6585d875d4620be357217397d40685a

Observation f05fbe3c-af8b-4870-84d2-0542e86cb1fb · outbound

This paper cites Vesta: A Generalist Embodied Reasoning Model.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Vesta: A Generalist Embodied Reasoning Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.541187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:33.160307Z digest=sha256:27358bf820636e5c9eae84baedf5a062a1e64b29da7dd9f9674dad58f97c831e

Observation 7f6474f4-f9c4-455f-9091-189374bf72e2 · outbound

This paper cites ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.517943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:33.164594Z digest=sha256:00f6da59ae78e389ab66ce3db7464285807246711303a7fc9f28c24ab0f93725

Observation 92858ee1-ad6a-492e-9752-bfa56571f301 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.169445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.169445Z digest=sha256:6ef769adff07fafea001c732bd5f4ccc98ee53bf052645de3f098a77311364dc

Observation c68e0c8e-7b94-4b0c-97be-348dbf7ff35e · outbound

This paper cites Towards embodied agentic ai: Review and classification of llm-and vlm-driven robot autonomy and interaction.arXiv preprint arXiv:2508.05294, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Towards embodied agentic ai: Review and classification of llm-and vlm-driven robot autonomy and interaction.arXiv preprint arXiv:2508.05294, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.249748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.249748Z digest=sha256:0c0314d917680567459e52d61cee22d64fcfbcd7b81f1508502c6e4a1e37a631

Observation 33640b2c-5173-4717-a65f-a80b96a4e93f · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.Conference on Robot Learning (CoRL), 2023.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.Conference on Robot Learning (CoRL), 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.279788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.279788Z digest=sha256:8c27ada41c527d5bbf04e74f43e9a176eef725e2bb663510fd9441f17ca12842

Observation d089a960-b2e9-4632-bf07-7fa6250fedd9 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.285071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.285071Z digest=sha256:e2786343ee2b8b7d57054819404b7a4e2dffbd41b8bc89fee6bbad9d058ce5f1

Observation fdf9a712-d270-4e65-9028-fbeea8192e53 · outbound

This paper cites To mix or to merge: Toward multi-domain reinforcement learning for large language models.arXiv preprint arXiv:2602.12566, 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence To mix or to merge: Toward multi-domain reinforcement learning for large language models.arXiv preprint arXiv:2602.12566, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.332402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.332402Z digest=sha256:03557c37402aee8205dd830c7a6f234d54e336c79bc1f731c8c403889204eebc

Observation d008e851-84b8-4f50-9a8f-ce16372ed5d0 · outbound

This paper cites TIES-Merging: Resolving Interference When Merging Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence TIES-Merging: Resolving Interference When Merging Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.474080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.474080Z digest=sha256:f93e61ef7dffacbf868b0f7bbea393470b36e1787d01e9b0ff2bb445fb28d029

Observation 91755aff-dbcc-4616-982a-abbdcaf6c7fe · outbound

This paper cites Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.478764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.478764Z digest=sha256:c481994afb1d073a816abc1ba3ade8ad82e0d89b35c4e30ac8218f01ac794c80

Observation c92e48b1-89bd-47a5-9d7d-306aaa57feb6 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen3.5: Towards native multimodal agents, February 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.484070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.484070Z digest=sha256:40625282c29b087a0c5145fd941d2ccc894e1379e6088e25e42a7b8122dc309e

Observation 4b70de19-d3b2-4d1b-a254-93af939ef572 · outbound

This paper cites Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:37.153623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:33.489323Z digest=sha256:017951d6e58a2d2f59c7ed7c7856c83ed18a60eed0228b9998082c87adeadd15

Observation 49450709-1715-47a5-92fb-5aa070fb0f64 · outbound

This paper cites Spatial intelligence in vision-language models: A comprehensive survey.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Spatial intelligence in vision-language models: A comprehensive survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.493556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.493556Z digest=sha256:09b85c5257d0a4890985c4b828e4d3c89d22662e513fb5d1c4e5316f857a71b4

Observation ac357ebb-62a5-4f01-a8bc-6f133836dfbb · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Scaling spatial intelligence with multimodal foundation models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.595378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.595378Z digest=sha256:ff3530e7ed0e387e4da6c564f70662bd017c11576ff38e2fb6c50c945ce6ed6f

Observation 0e8ac789-c608-44cd-9679-15df91d82e69 · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.681742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.681742Z digest=sha256:77ba9fb195374a5693a5a1aeaa96c6301dde40dd3edbd139785fdb147baa121f

Observation 48f5a9d4-5ba1-46ac-9242-143276a5bb77 · outbound

This paper cites Mindcube: Spatial mental modeling from limited views.arXiv preprint arXiv:2506.21458, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Mindcube: Spatial mental modeling from limited views.arXiv preprint arXiv:2506.21458, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.686415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.686415Z digest=sha256:df3fe314739ad49c5b6cf41ea3cce50d761a4303857bd7f8062d0e5457a520d9

Observation ed101fe9-12c4-4e51-9dc3-2a8c2c769b54 · outbound

This paper cites EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.691153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.691153Z digest=sha256:680a1a25a02e1f11ec2d89f3e4a32697313684e19e3d691c84b44733acdd9080

Observation 15588c70-3e8d-42c5-b43b-e8105dab5da8 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.697230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.697230Z digest=sha256:24067cfc07983c322279ae2e4ca9b798a9942b360d9e41808f3a8cdbee57b509

Observation 52943f18-0a9c-492e-a782-7b40631aba5b · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Next-qa: Next phase of question- answering to explaining temporal actions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.742687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.742687Z digest=sha256:d5a199f5ae61bb6e1aa1e0e4e67a42cb6f0b2e95baa1ddaacf5cd03558232bfb

Observation 7e34d14b-8972-44f2-8160-102057ebd74e · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36: 42748–42761, 2023.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36: 42748–42761, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.806130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.806130Z digest=sha256:91f64e9a0575a80d83660d3aeabc2110cf3e956f7bba138c8aa67564626406a3

Observation 21664cc2-0118-4b40-a07f-dec0158dd4de · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.811024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.811024Z digest=sha256:b2e8748fe7ec352b07a6a9909e9b7415a8b5b71755baa72f9b8163828fb11a2e

Observation 50b4cbe8-816b-4771-bd03-26e64f07a985 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.815978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.815978Z digest=sha256:5520939dd7cc1a7dce96494f71606a53af069779d6b8eac2d0761a0782028b19

Observation 9a915b03-7cc8-4b4c-9167-2d8e39b2fad8 · outbound

This paper cites Scaling rl to long videos.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Scaling rl to long videos

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.820656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.820656Z digest=sha256:f689a994df3f0426793ecad8ff3414a6b968948d096d9d48bc6a7888f98b6097

Observation 83b37b98-f0f9-4a27-b8fc-3630fc025293 · outbound

This paper cites Sigurdsson, Gul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Sigurdsson, Gul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.963434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.963434Z digest=sha256:d9ea66dbff811a6a6281e5380053563bcf4ce9c889bd172f767d272cc31847bc

Observation 2092ae7e-27ce-46de-8f73-cfb520dd1326 · outbound

This paper cites Localizing moments in video with natural language.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Localizing moments in video with natural language

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:33.967117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:33.967117Z digest=sha256:1f9c056d1c202f7f0c06405ad39644c32936466dea64f08bd6550ca52f2aa687

Observation e239d02d-f01f-4552-8eea-1acabe332c49 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Hierarchical video-moment retrieval and step-captioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.035589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.035589Z digest=sha256:7301e6f032cf1fa32255a59ece9c8866e4eb0c9c332f79a2d329e39a397b0895

Observation eee2537f-48ac-4f0c-a942-aee7a7f2a935 · outbound

This paper cites Queryd: A video dataset with high-quality text and audio narrations.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Queryd: A video dataset with high-quality text and audio narrations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.097731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.097731Z digest=sha256:fee14ada8a9c218d97b20ee8695fcd2473b7666aed34ed4412e2b5b57d1ad93a

Observation 4a3cba58-04d7-4264-a127-c786019291ed · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.103487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.103487Z digest=sha256:2936503461e1e6fcd9e41e959fee13ed843981ff8a84263760c76a3d62e05215

Observation 581d0e6d-922b-493d-957d-1ec240bfe891 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.108407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.108407Z digest=sha256:07fd14c21cd38ed50aa018fe6c3ce2070f36613c8e8e33945a1b73c5cb64da71

Observation a9271ed1-8430-4575-b124-4d6abe3a7614 · outbound

This paper cites Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.141816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.141816Z digest=sha256:594e6f99e6760a34eb8969fb5e831ba5eb3dd032bad0a5750f0933813e8721b5

Observation 5134f1df-a2ad-4d25-a522-232b9379943d · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.146624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.146624Z digest=sha256:3e1d11e7ad57136e977cc5cc144edb25faa289cd966dbea3d6f3e77d03b1c256

Observation c77499d7-845c-43eb-8b61-5b6eb9cc16a7 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.152165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.152165Z digest=sha256:b4c7c00729f074d631e3ec9f5caba29045ce29f780919b551d8414ea5f49adce

Observation a475d797-c9c9-496f-be8d-66ef66f0f527 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.157336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.157336Z digest=sha256:91bab01faa12b8787009e1c76ba1313a054aa179feb54882cfc90e9d04a03c36

Observation d7cec257-041a-42a8-b470-45aa21c02dc3 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.161372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.161372Z digest=sha256:736e1875a07a214a6097001593ca348c7be79286f6ceed42df540a4945e041e9

Observation 7685dd33-d366-4b69-aa92-7fa57addc942 · outbound

This paper cites Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv preprint arXiv:2512.24653, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv preprint arXiv:2512.24653, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.165396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.165396Z digest=sha256:77968e6f6ca35c9f78ee8963cf43a9d682725940daa659159d3ce477e85c4ef2

Observation cfeac7ad-05ee-473a-ba1f-c75afb3ad17e · outbound

This paper cites From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.169970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.169970Z digest=sha256:398e43a7f8dc93d0af8b2de7808f9bc8ea0fcc1cec892ba64ca5b23ebb5d45f8

Observation ea4dbfdc-93a9-4a2e-a591-a17284b2caad · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.174958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.174958Z digest=sha256:71bd6c20fdc1d78e91cc8b5743d21162a943ffb7cdb4c232f4d25e81bed64e17

Observation 2b0413ca-47bc-4422-a25a-fed4ca82253d · outbound

This paper cites Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.179159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.179159Z digest=sha256:96e3653ff5a0bb283f90cb0f105e80540da2ab90ef99fded4aba3a5ee750365f

Observation 28ab7762-b38e-4f95-b268-f111e381a36c · outbound

This paper cites Roboafford++: A generative ai-enhanced dataset for multimodal affordance learning in robotic manipulation and navigation.arXiv preprint arXiv:2511.12436, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Roboafford++: A generative ai-enhanced dataset for multimodal affordance learning in robotic manipulation and navigation.arXiv preprint arXiv:2511.12436, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.269327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.269327Z digest=sha256:b480b0dae56f79ad5d0fbac6a2392daedcbc9685389c9e1a850151741f07f981

Observation ae3bc33e-e223-41ac-bcf5-b6e32a498d83 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.348952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.348952Z digest=sha256:c564e350f07985c979f658f19aaae07d873ec029e1bd22b26924c8ed45600977

Observation 2ac471f3-3385-4990-b3f4-0f97783f9428 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence OneThinker: All-in-one Reasoning Model for Image and Video

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.360274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.360274Z digest=sha256:b1167fb0107c73d1d7f14417ec69eae34f707f341d67204f73133ee3100a54cc

Observation 7c7f7a20-bed4-4527-abc6-eecd1df2401c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.364746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.364746Z digest=sha256:6a0242527570cb0bf3f9d33fc30e2884f72a3048a10e79a69877235a5c4cc69d

Observation c7902afc-23c1-4703-971f-b8c6c196d616 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.371034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.371034Z digest=sha256:4df94f544c86f7de7844ed849753b76c76ab97551c4efbf059ddb460be1db488

Observation da04dddd-e6f9-4246-83c3-73792ffe8c6d · outbound

This paper cites Thinking with visual primitives.Technical report, 2026.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Thinking with visual primitives.Technical report, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.375958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.375958Z digest=sha256:b986b991c877aec5c032856cff7454e02c34082e8d08dfe5e7a629f21f709f58

Observation 16b43f36-a6d6-4399-b2f8-b37f383835c5 · outbound

This paper cites Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.422610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.422610Z digest=sha256:1dcc406ebe16a6e2b6815f68a53c78c3abd2eaccf862102c984a28018ee3fcc8

Observation c0505013-0f33-46c6-9863-ed86be8aca6b · outbound

This paper cites Computing discrete fréchet distance.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Computing discrete fréchet distance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.467999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.467999Z digest=sha256:f1e769322214034d90a7a3c3edc338919db6e2535aa54a967f6fef000e46506d

Observation 66a70d7f-1737-4421-a885-58eda448e264 · outbound

This paper cites Editing models with task arithmetic.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Editing models with task arithmetic

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.519320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.519320Z digest=sha256:58d88b4d448eb53e149cf5d3207afb8350c69c169c8f2788a4e89b82f6b67209

Observation b3e6ddeb-108c-43ba-8aaa-8f31591712ce · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.522970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.522970Z digest=sha256:55101357ad9e0d75ad511fc4ca8ee62a8ca1c1b593cf433513fc7309300b6e5c

Observation 9240bcc3-d422-4ac5-b089-cfa9a1000620 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.527373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.527373Z digest=sha256:e7094d04f93a1146bbcd107a6a84637ebeef9186d398a4487371938aeec2999e

Observation c2b34af9-5341-4fc5-b015-cbd3944e1baa · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135, 2025

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.531761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.531761Z digest=sha256:c9b289508486ffdb5a4e94918061fe91e8e3c5300fa54f7b1aadae6c2eaf3cbb

Observation 95287c8b-3854-4d03-bffc-13c3dcd6c0b9 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.arXiv preprint arXiv:2411.16537, 2024.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.arXiv preprint arXiv:2411.16537, 2024

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.535818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.535818Z digest=sha256:1f8b44015819546247e1527d56a34563f5f7be625c74c48357820c5fe6a7717e

Observation 919792d0-af6c-420c-908a-63c095ce76b0 · outbound

This paper cites Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.554776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.554776Z digest=sha256:424e42e5291c9578464a44b439b54063602f8e66ce3ed9c5dc7059e5cb377d21

Observation 398ecf40-eb65-45f4-905d-5072113e88ba · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Openeqa: Embodied question answering in the era of foundation models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.579247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.579247Z digest=sha256:aaccd50bec35e24b35538cc07449cbc63c3640f1b2fc452e7cd67ba1facefa89

Observation 4fd7c8d5-a0d7-4fa5-90c2-3f5c95f42a2b · outbound

This paper cites Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.626969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.626969Z digest=sha256:8033beeb5487b5c8c165c3fbbe6c3f6cb5e8936233aa1d1b384fd1d4a5b88201

Observation 8322d1ec-0c1b-41f0-9e61-bb31e74c71f9 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.631324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.631324Z digest=sha256:dfdaf84b3a46fe72dacc35c374595a81ce9094efb9a7a12e62153634aa948532

Observation bd18a110-790c-4eee-bddb-ff6638e71253 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.635521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.635521Z digest=sha256:ad7a7cea985c0c6b01362d5beb0e5109c87a1a746689fe0ed94f43c941e11832

Observation d88f57ac-3e11-4f7a-b937-26d411effd17 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.640153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.640153Z digest=sha256:8e2d0c90ea5802304e2ab892da5473543b52943444f29ee6eec0cfdebd9fcb84

Observation 9b8217a6-8b99-4e16-b871-02f201571e3e · outbound

This paper cites Timelens: Rethinking video temporal grounding with multimodal llms.arXiv preprint arXiv:2512.14698, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Timelens: Rethinking video temporal grounding with multimodal llms.arXiv preprint arXiv:2512.14698, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.644370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.644370Z digest=sha256:b4329a09f88ac37248434a2d9e154539f2552bb293b3ead8c04e9d8583dc50bc

Observation 40616467-1d5b-42e1-ad5a-19fd41f99973 · outbound

This paper cites VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:36.252207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:34.648497Z digest=sha256:9ece4f1fefc1051da60b16e8377443e862189732861b85c00c8f3e56d41b8669

Observation 4f568cba-cfae-46ec-905e-d54e13e39fab · outbound

This paper cites Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.676111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.676111Z digest=sha256:afb21cedd31c3e87fd53ed0c31f6e2895ea4d2a3ee2eb7e3b6629bf3a8011300

Observation 27cbe886-8340-4e8e-af68-7b37eaf1ebf1 · outbound

This paper cites Point-it-out: Benchmarking embodied reasoning for vision language models in multi-stage visual grounding.arXiv preprint arXiv:2509.25794, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Point-it-out: Benchmarking embodied reasoning for vision language models in multi-stage visual grounding.arXiv preprint arXiv:2509.25794, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.719268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.719268Z digest=sha256:b0738449525b3ead184069b30d4ae40a2a18041037b794434cabb45b85e99e12

Observation 71d4f895-b6d9-4a10-bc89-7c27ae3ce6bb · outbound

This paper cites Navitrace: Evaluating embodied navigation of vision- language models.arXiv preprint arXiv:2510.26909, 2025.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Navitrace: Evaluating embodied navigation of vision- language models.arXiv preprint arXiv:2510.26909, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.744188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.744188Z digest=sha256:6c0bf58801b69eaa11b350b1aed3525b01325b9b074280bb5d76d7aaad6e85cc

Observation d7558b24-56da-4a40-869e-fc62c065f7c0 · outbound

This paper cites an unresolved cited work.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.749228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.749228Z digest=sha256:b2e52c1ba17f3a97f36bce338314ae14f76299d22999289d082eecec5d792eac

Observation c5273e05-751c-4c42-bcde-3cca43a00d34 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.753811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.753811Z digest=sha256:395f40a61df85801b849ddf5de8f4cdb7c53b121278787c23789db0173c7af21

Observation ca55d428-8b33-4f4a-b722-7ec2ef0e057e · outbound

This paper cites Grok-1.5 vision preview and realworldqa.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Grok-1.5 vision preview and realworldqa

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.758326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.758326Z digest=sha256:515a8e32a3e07368948a5805166b762128131246de85a72f84d5811e5ec9c570

Observation 4e99aed2-a855-4603-85e5-e2604a61e274 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MMBench: Is Your Multi-modal Model an All-around Player?

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.762841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.762841Z digest=sha256:d26212c824b323c6fb0dab29861f81aeed8d1dc66ea44901d98031e38ca10b62

Observation 7f21dbfd-f9ab-4a8e-b851-747d9f020d6c · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Instruction-Following Evaluation for Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.767697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.767697Z digest=sha256:d261215d4069cd7419d7e439bfc77481186e8be7a2e2fae13a06b7ee86e41f2f

Observation 0d32e7c7-a8aa-43a4-967e-aa660d459f69 · outbound

This paper cites MMLU-Pro: A more robust and challenging multi-task language understanding benchmark.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MMLU-Pro: A more robust and challenging multi-task language understanding benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.773055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.773055Z digest=sha256:4c8a78f3def80e704a1cfe136bba394a383e91c98f32dd742559325c086dc761

Observation 085afeb7-cc38-4aa6-89b7-a271d3b69cc2 · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.777235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.777235Z digest=sha256:9d8145c1967b07d67ad33c3b91b97bcdeba5a60d62751b625593eb806e94c3ff

Observation 4fc0a1cc-53b9-41f7-be0a-3015c08d04e7 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.781389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.781389Z digest=sha256:421c5d68f25b26cc59a7e9486dfe583fed6a1da032c064b683675a7f9cd3df58

Observation 8f25946d-13b1-45f8-bec8-487fea92ac0d · outbound

This paper cites DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:18:35.961548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:34.803552Z digest=sha256:59777ed0d22ea5f7448cbf912836dc7dbcdbb504b2bf5122d9bf075a8b425f57

Observation 75ba6374-2fd7-4256-900b-3abf96f374ab · outbound

This paper cites Qwen3-VL Technical Report.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Qwen3-VL Technical Report

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.858442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.858442Z digest=sha256:0e5c747d1688a498aebf971ca770ea825b989966fc8b247f93a3b4d830322fdd

Observation 40d43281-44b5-4d7c-9c6e-d0f7b39db8b8 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.918689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.918689Z digest=sha256:bc80f62c8b8f3abc6195d0e579bbbf02ad834c9e0cfca26bd901f7b61ea43581

Observation 3a4aafdf-5e88-4e14-be9a-8312c51e3e0f · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.Advances in Neural Information Processing Systems (NeurIPS), 2022.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Procthor: Large-scale embodied ai using procedural generation.Advances in Neural Information Processing Systems (NeurIPS), 2022

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.954091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.954091Z digest=sha256:7ff0148151947073c392631953260fd7a48a77f56760f1cfc689d2f53bd2f6e7

Observation 245f033b-dc99-4955-b21a-bbbd6393cd3a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence HybridFlow: A Flexible and Efficient RLHF Framework

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.958615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.958615Z digest=sha256:362f231a68c6a89b26de9a168211de39831cd6860d4523313f5d1c3c8cd61478

Observation 143c72bf-dc07-436f-8035-c98e13d32153 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.963819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.963819Z digest=sha256:a6d41eff5c7eea8ca95c26d676eb71db48b3cadc72ecc311d73753a71187fcd8

Observation 20da74c3-02f5-482f-a03e-a66a0670e512 · outbound

This paper cites bbox_2d": [x1, y1, x2, y2],.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence bbox_2d": [x1, y1, x2, y2],

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.969726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.969726Z digest=sha256:d449b0ec7edc7bccbde039d87f19a365d41bfc70315142cbcd525b6498b83c47

Observation 6d543b52-9bfc-4277-9c69-ab14ab36e936 · outbound

This paper cites The left hand is holding a green lighter.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence The left hand is holding a green lighter

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.976635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.976635Z digest=sha256:a5e02deb7b9668e363766985fad424ba9b13acffa2a530250aacce738a39f9d9

Observation 8531c9ee-2c0b-4f8a-97d7-3d689f2ad37c · outbound

This paper cites The top of the lighter is now glowing red/orange, indicating it’s hot or just used, but crucially, there is no flame coming out of it.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence The top of the lighter is now glowing red/orange, indicating it’s hot or just used, but crucially, there is no flame coming out of it

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.981728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.981728Z digest=sha256:74567355a054620877e222910b9a1116c2005c22ad4b50af47d07ef3f273bfff

Observation 39cffba3-a8fb-4858-ae55-44c03298c32e · outbound

This paper cites Based on observations, is the lighter on or off?.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Based on observations, is the lighter on or off?

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.985997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.985997Z digest=sha256:aa9a37cd891196a7e29e8008a84149cfa08c1edf3d582854c690bc98b6a74ded

Observation e5b9aa84-f903-4067-a1e5-8c134fc50d35 · outbound

This paper cites * The club sandwich must be inside the packing box.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence * The club sandwich must be inside the packing box

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.990563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.990563Z digest=sha256:efcd9b69dd5a418a35588e68f3e14981b02060792ef21e3df9cbbf7abde3d8df

Pith citing papers

No inbound Pith citation observations are available.