Pith. sign in

Paper Citation Record · LEDGER

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2602.02533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.02533 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:43:04.627490Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:43:04.520409Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T15:43:04.746675Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b4e81357-9ff2-4195-97c7-84368156384e · outbound

This paper cites HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:43:04.752075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.520409Z digest=sha256:1df32e0590cc3cc4265ff15a797cd2994731106f4474761b9794b91fd4da1537

Observation 374d0762-650f-4ea5-904f-4e4928c967a9 · outbound

This paper cites Hyperbolic Semantic Alignment Our HMVLA framework is based on the Lorentz model in hyperbolic geometry, as shown in Figure 2.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Hyperbolic Semantic Alignment Our HMVLA framework is based on the Lorentz model in hyperbolic geometry, as shown in Figure 2

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.976355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.525749Z digest=sha256:f6e051a7360da1d3bf4d4e4b4eea3c19814d3492ed35a96815f125a4f3bbfde9

Observation 3b3459f3-0e28-4d79-8c9d-f5fd089ef3dc · outbound

This paper cites Specifically, we utilized four datasets: Spatial, Object, Goal, and LONG.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Specifically, we utilized four datasets: Spatial, Object, Goal, and LONG

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.966116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.530082Z digest=sha256:2375def28318f378ea77568f6c62e4a8ccf9362682313a2cbb06ff9996af9617

Observation 7e2b28c3-31f1-4c66-a8e1-e64a33f8b7b4 · outbound

This paper cites By embedding multimodal features into a hyperbolic space, our model effectively cap- tures the inherent hierarchical relationships within image-text data.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models By embedding multimodal features into a hyperbolic space, our model effectively cap- tures the inherent hierarchical relationships within image-text data

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.956450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.534320Z digest=sha256:1177cf73a48bcb83734874e9471d511d72b032bb290e138e964562f8c4d39ac5

Observation d351ea94-7afd-49cf-9086-d23d26538873 · outbound

This paper cites 62277011), National Key Research and Development Program of China (Grant No.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models 62277011), National Key Research and Development Program of China (Grant No

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.946258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.539446Z digest=sha256:55410e5a587d5669f1b8439190cd4afa3f9c4745b094bf6f2b928cfa07ca2634

Observation 0adcf0ee-9299-4181-839e-83cb85499538 · outbound

This paper cites GPT-4 Technical Report.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.543112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.543112Z digest=sha256:0f82c2edcbc7d2eeeb6b4c0c63f385e9ae91ab7011af91f49bf5ee78c3e26869

Observation ec1c130b-751d-47a7-8162-50061ddafd40 · outbound

This paper cites The llama 3 herd of models,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models The llama 3 herd of models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.936471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.547408Z digest=sha256:a1ed5b96abac96fc4c28adf036f40e18f07b848d99bfaa0f26c2cdb6cdaa52a7

Observation cecb79b9-8a35-4979-b63f-de2fe9d14aef · outbound

This paper cites Llapa: A vision-language model framework for counterfactual-aware procedural planning,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Llapa: A vision-language model framework for counterfactual-aware procedural planning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.926842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.550777Z digest=sha256:0f097c271a1003c9089b55ede35673ddd84944bb091c5371bc96143e2e0b742a

Observation 7d577ea6-d172-4e20-b863-7435f629a6bc · outbound

This paper cites Sage: A visual language model for anomaly detection via fact enhancement and entropy-aware alignment,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Sage: A visual language model for anomaly detection via fact enhancement and entropy-aware alignment,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.917189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.554343Z digest=sha256:5051c258ef2c61465ffc13a172ec5e1cfb9d4e13ef95c4afa0256ce427debc8c

Observation 2aa2faff-d727-46f1-92bf-0a2b1a174921 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.558110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.558110Z digest=sha256:757cd58c1af3daf41244ee4815b862251e3d285fe960489a6b921e6ee26f3b28

Observation 4232ef47-45e8-483f-b5ba-713a3ae03824 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models RT-H: Action Hierarchies Using Language

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.561594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.561594Z digest=sha256:14a7963e1e13493acb5b07e483a6fdeae7d195fe9dc50fce88811b1a8f9b9a8a

Observation c9324eb2-2f2d-4205-b7c6-30da002c49fb · outbound

This paper cites ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.565021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.565021Z digest=sha256:5094201464a20ea49b7ec053db549d58438ba65020438831c2fd3bc9b34f8dfa

Observation f3a935ef-f69e-49ff-9339-8997b4400539 · outbound

This paper cites 3d-vla: a 3d vision-language-action generative world model,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models 3d-vla: a 3d vision-language-action generative world model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.907320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.568497Z digest=sha256:35d0531664708fb9faeec734479db9354ef06b6df441128a1f34eaa5d3c2e664

Observation 55c8244d-1f2e-41fc-88ac-3851e635161c · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.572491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.572491Z digest=sha256:d4c3bcd82ebcf07380c1b91ba36f3604bfd63bbca0eee2607ec7b8f22632efd6

Observation 0d2e091f-1111-4eab-ab2d-adb679afec2d · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.575927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.575927Z digest=sha256:9f6858d6e4682608bc32b09edb2a60a82c6ffc92db563c958adb65ae6b8061f0

Observation c019518b-7f15-43c3-b7ec-ffe281ea1d04 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models A Survey on Vision-Language-Action Models for Embodied AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.579259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.579259Z digest=sha256:cc493df6b4c58b1a09386a62bd7d80a618a5c1635826dd95880db1a8f8d2b0e1

Observation ed3b8eff-bb83-406f-a334-799a3b209fc8 · outbound

This paper cites Vision- language models for vision tasks: A survey,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Vision- language models for vision tasks: A survey,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.897516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.582438Z digest=sha256:bf8349c147a195bc4465a4d3a0ebe7c9e6cedcbd2be1547f9a2720697465c195

Observation a349a12c-ef24-414c-8aea-1a20004be9cb · outbound

This paper cites Rt-2: Vision- language-action models transfer web knowledge to robotic control,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Rt-2: Vision- language-action models transfer web knowledge to robotic control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.887837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.585455Z digest=sha256:83670e5563360a56693cbe41b00c87f0f8f4ccbd36b716091ac326d78d4005fa

Observation 25bab865-bac0-4005-9ab4-1f1a06fc8f48 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.588623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.588623Z digest=sha256:9993d21fae07b2198c8533f0523de558730bf271c73ecee44e7004c6fcd5294f

Observation b47d593f-eb4c-453f-9fc5-d3cfe3cccebe · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Learn- ing transferable visual models from natural language super- vision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.877209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.592045Z digest=sha256:6287ff9484e9d2d7c1137389058ced40cebc8da561828960c3ac052b8f2e4480

Observation 0e8122c4-7dcd-42c5-b4ee-676f65a607ee · outbound

This paper cites An ex- tensive study on pre-trained models for program understanding and generation,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models An ex- tensive study on pre-trained models for program understanding and generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.866215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.595321Z digest=sha256:09c735dd4d53bc5f5ce277911edd50957458dd3be1de2ff9b152a92eedd5d669

Observation d262a5c5-f323-48e6-8588-3f9b9e7b8006 · outbound

This paper cites Hyperbolic spaces,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Hyperbolic spaces,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.854757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.598774Z digest=sha256:46be4431cfdf44aa1632551229ca393b77191d9bd5225a4c9810c0122ab5e01d

Observation f16ec933-ddd6-428d-8459-7595aac87099 · outbound

This paper cites Hyperbolic image-text representations,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Hyperbolic image-text representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.844027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.601744Z digest=sha256:10cbc4e110a07fa373af233db334c6bb15cde93d2e2a3cde0c698e726db701c5

Observation 266bde50-db5b-4ec9-bfb9-1f57b3fb541a · outbound

This paper cites Libero: Bench- marking knowledge transfer for lifelong robot learning,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Libero: Bench- marking knowledge transfer for lifelong robot learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.832486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.604742Z digest=sha256:965fef057ac22339366ef3490fa8eef2ff315ff1d1de6f1ae3d0ca167ae71484

Observation 14f30393-63f6-4484-be68-53fb55ef1699 · outbound

This paper cites Zur elektrodynamik bewegter k ¨orper,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Zur elektrodynamik bewegter k ¨orper,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.821653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.607850Z digest=sha256:80c9530c9abcecaeb33b28c726f6906eafa179ba3144e95965c2094a308ce00a

Observation 85e28ff1-6b6d-4fdf-9da2-a7318ab2f29a · outbound

This paper cites Dita: Scal- ing diffusion transformer for generalist vision-language-action policy,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Dita: Scal- ing diffusion transformer for generalist vision-language-action policy,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.810647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.610848Z digest=sha256:b83b59a2d07bf44002f64ac5f946c6803c9435e68d9b2962b218eeb7a422cedd

Observation fd6e3668-c86d-4c90-adab-f88e938af3f1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.798625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.613767Z digest=sha256:6814ac886b6c1c7afb11454b30ec2f0edf2628d1ed4008e019b0f016f9b70081

Observation b7d81815-f5f3-4ad1-a968-98e8a21057cc · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.616894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.616894Z digest=sha256:ae97a47034b0eb4738a123c5ee4f3b02be84fd8e85d872e89066bd6b323c7491

Observation 895c2bbf-5bee-4c93-9995-09c5864946e0 · outbound

This paper cites Tra-moe: Learning trajectory prediction model from multiple domains for adaptive policy conditioning,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Tra-moe: Learning trajectory prediction model from multiple domains for adaptive policy conditioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.787292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.620779Z digest=sha256:396a3accf1daf589ad775e969f8910c5373226018ad781a52568787a7e2e979a

Observation 0baec4dd-c9d3-4142-830b-0b86a2662837 · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.775644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.624055Z digest=sha256:7a340f5931fd3cbe757693d9eaae13cca29bc46b930965dcf41f50e047a280d5

Observation e2e7f64e-bc1c-4d6c-8fc8-398d7eb6489d · outbound

This paper cites OTTER: A vision-language-action model with text-aware visual feature extraction,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models OTTER: A vision-language-action model with text-aware visual feature extraction,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.763546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.627490Z digest=sha256:1cf7988db26c25e0c8f249f33f260768959a771dd392012fcadad7e1739e1955

Pith citing papers

Observation b4e81357-9ff2-4195-97c7-84368156384e · inbound

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models cites this paper.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:43:04.752075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:43:04.520409Z digest=sha256:1df32e0590cc3cc4265ff15a797cd2994731106f4474761b9794b91fd4da1537