Pith. sign in

Paper Citation Record · LEDGER

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

As of 23 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2602.02533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.02533 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:43:04.627490Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:43:04.520409Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T15:43:04.746675Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b4e81357-9ff2-4195-97c7-84368156384e · outbound

This paper cites HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:43:04.752075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.520409Z digest=sha256:be5d19c12082114fe12e66b4a348e3db181a7ed3f3c060e11699041ebc1fc5a4

Observation 374d0762-650f-4ea5-904f-4e4928c967a9 · outbound

This paper cites Hyperbolic Semantic Alignment Our HMVLA framework is based on the Lorentz model in hyperbolic geometry, as shown in Figure 2.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Hyperbolic Semantic Alignment Our HMVLA framework is based on the Lorentz model in hyperbolic geometry, as shown in Figure 2

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.976355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.525749Z digest=sha256:a475f0774b4b9327c6b11b02523f0ce28bf74bdee1ff0b5fd98c8d954b119ab3

Observation 3b3459f3-0e28-4d79-8c9d-f5fd089ef3dc · outbound

This paper cites Specifically, we utilized four datasets: Spatial, Object, Goal, and LONG.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Specifically, we utilized four datasets: Spatial, Object, Goal, and LONG

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.966116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.530082Z digest=sha256:e7babd9df0756e68ad81bec32023f5a6343ccf6a528a5d79089f4b46661fe8a2

Observation 7e2b28c3-31f1-4c66-a8e1-e64a33f8b7b4 · outbound

This paper cites By embedding multimodal features into a hyperbolic space, our model effectively cap- tures the inherent hierarchical relationships within image-text data.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models By embedding multimodal features into a hyperbolic space, our model effectively cap- tures the inherent hierarchical relationships within image-text data

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.956450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.534320Z digest=sha256:9d77b042a3020fc5354582aab1b153b664ded80a88c3d860f1f4874c1a7ea370

Observation d351ea94-7afd-49cf-9086-d23d26538873 · outbound

This paper cites 62277011), National Key Research and Development Program of China (Grant No.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models 62277011), National Key Research and Development Program of China (Grant No

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.946258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.539446Z digest=sha256:688db4dd85fa08219fc11ce64804a6d90f850957e10862a73ca52afb7af77b84

Observation 0adcf0ee-9299-4181-839e-83cb85499538 · outbound

This paper cites GPT-4 Technical Report.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.543112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.543112Z digest=sha256:fc3ec81942bbf1034e83c902958d2e5e00323de879d828841ebfb55e50601e6e

Observation ec1c130b-751d-47a7-8162-50061ddafd40 · outbound

This paper cites The llama 3 herd of models,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models The llama 3 herd of models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.936471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.547408Z digest=sha256:433c683dfdbdc3bf1fb2057db9ccbace264d81a69a10bd5349ac670aae82cbb0

Observation cecb79b9-8a35-4979-b63f-de2fe9d14aef · outbound

This paper cites Llapa: A vision-language model framework for counterfactual-aware procedural planning,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Llapa: A vision-language model framework for counterfactual-aware procedural planning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.926842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.550777Z digest=sha256:1efa41d9bfb9b0a01a5d0f39237ceff687cac19307d3a942733ef84933744825

Observation 7d577ea6-d172-4e20-b863-7435f629a6bc · outbound

This paper cites Sage: A visual language model for anomaly detection via fact enhancement and entropy-aware alignment,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Sage: A visual language model for anomaly detection via fact enhancement and entropy-aware alignment,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.917189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.554343Z digest=sha256:e2ab1034ebc37840a0a6e870bfaf48b6ad30b77e53561bb43ce683b811f8fce1

Observation 2aa2faff-d727-46f1-92bf-0a2b1a174921 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.558110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.558110Z digest=sha256:50edeae0aa314a9e86be19e50ebe31820dcf406e66543065461601c155bb249b

Observation 4232ef47-45e8-483f-b5ba-713a3ae03824 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models RT-H: Action Hierarchies Using Language

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.561594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.561594Z digest=sha256:2aa3aeb586fc4924b0026ed390e6a36bdfa24926add0d69bd9096cbf06a378d8

Observation c9324eb2-2f2d-4205-b7c6-30da002c49fb · outbound

This paper cites ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.565021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.565021Z digest=sha256:17957c22b5fc57475a784efb81515790dc8f7b54e73ddfc45ec62e8e40cacabb

Observation f3a935ef-f69e-49ff-9339-8997b4400539 · outbound

This paper cites 3d-vla: a 3d vision-language-action generative world model,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models 3d-vla: a 3d vision-language-action generative world model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.907320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.568497Z digest=sha256:64e6fea2a696a11729c9d35add2f3080f297187f91d7b09c0a45d3e1c657e76d

Observation 55c8244d-1f2e-41fc-88ac-3851e635161c · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.572491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.572491Z digest=sha256:1d8a32fb593ca0a5002e4082b0141e65353e0aaec96b67c10c7b3c2e1944ed9a

Observation 0d2e091f-1111-4eab-ab2d-adb679afec2d · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.575927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.575927Z digest=sha256:49c2c4e2eb9bf56d3cdeec7dea01d20d5f9c455d8875ea835931877af1610f3c

Observation c019518b-7f15-43c3-b7ec-ffe281ea1d04 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models A Survey on Vision-Language-Action Models for Embodied AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.579259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.579259Z digest=sha256:78df9f84dc37034f8c21a71ccfaf5e5bcc4d1775fbc0489b0c399e268208d02d

Observation ed3b8eff-bb83-406f-a334-799a3b209fc8 · outbound

This paper cites Vision- language models for vision tasks: A survey,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Vision- language models for vision tasks: A survey,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.897516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.582438Z digest=sha256:7f6a58187e0fa5ec0afc953a42ae05d4c0ebb89c935760c683f414ed993df459

Observation a349a12c-ef24-414c-8aea-1a20004be9cb · outbound

This paper cites Rt-2: Vision- language-action models transfer web knowledge to robotic control,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Rt-2: Vision- language-action models transfer web knowledge to robotic control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.887837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.585455Z digest=sha256:cf26201f9e3c3d853f17454e1c6715d6bac678d1c0d3f516b30a796a118f7b0d

Observation 25bab865-bac0-4005-9ab4-1f1a06fc8f48 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.588623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.588623Z digest=sha256:c73e90ef0128a8f7cf2f88441c5bfb270b7d02cf997de3de1354d81b8322e9d9

Observation b47d593f-eb4c-453f-9fc5-d3cfe3cccebe · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Learn- ing transferable visual models from natural language super- vision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.877209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.592045Z digest=sha256:e6a671097c87f4e55912303d5dc50dce4c7abcc15ef52ee29092ca8b4a60133a

Observation 0e8122c4-7dcd-42c5-b4ee-676f65a607ee · outbound

This paper cites An ex- tensive study on pre-trained models for program understanding and generation,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models An ex- tensive study on pre-trained models for program understanding and generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.866215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.595321Z digest=sha256:1f2cd564b02cb2d1499ee278f760c2617f080bd316d60497a888770eb2d5a348

Observation d262a5c5-f323-48e6-8588-3f9b9e7b8006 · outbound

This paper cites Hyperbolic spaces,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Hyperbolic spaces,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.854757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.598774Z digest=sha256:cca089e1d096a1b77ab59704e36062f72ef03206f9ccaf0299324f6822eeada4

Observation f16ec933-ddd6-428d-8459-7595aac87099 · outbound

This paper cites Hyperbolic image-text representations,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Hyperbolic image-text representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.844027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.601744Z digest=sha256:a931c826e27872f750ffed261015813b7b13fd041920281ac5d6772016165a7b

Observation 266bde50-db5b-4ec9-bfb9-1f57b3fb541a · outbound

This paper cites Libero: Bench- marking knowledge transfer for lifelong robot learning,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Libero: Bench- marking knowledge transfer for lifelong robot learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.832486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.604742Z digest=sha256:23f1147d385c3418170622d9fc8233bb092acf9a8d41242175c458ca7bf3deb5

Observation 14f30393-63f6-4484-be68-53fb55ef1699 · outbound

This paper cites Zur elektrodynamik bewegter k ¨orper,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Zur elektrodynamik bewegter k ¨orper,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.821653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.607850Z digest=sha256:8520deb1582afddd09eb2b507f6f95ee89ad1897f442ccce118f8dd59d10af95

Observation 85e28ff1-6b6d-4fdf-9da2-a7318ab2f29a · outbound

This paper cites Dita: Scal- ing diffusion transformer for generalist vision-language-action policy,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Dita: Scal- ing diffusion transformer for generalist vision-language-action policy,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.810647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.610848Z digest=sha256:00d5b1606666df03e70e438b300f18061566b6309d9562d6de9bb3591ee10e00

Observation fd6e3668-c86d-4c90-adab-f88e938af3f1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.798625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.613767Z digest=sha256:25666d742195044bc008cca6c1bff00c773b448a29b39f7e62ab5a88c18e7818

Observation b7d81815-f5f3-4ad1-a968-98e8a21057cc · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:04.616894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:43:04.616894Z digest=sha256:6190e87ada6b07986a882f15c2b1b519bc869be80831c94b09d88ef969bfa322

Observation 895c2bbf-5bee-4c93-9995-09c5864946e0 · outbound

This paper cites Tra-moe: Learning trajectory prediction model from multiple domains for adaptive policy conditioning,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Tra-moe: Learning trajectory prediction model from multiple domains for adaptive policy conditioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.787292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.620779Z digest=sha256:cb398087ef058f1610e98a98f43196840d91c197f23761c2d6a1dc7e0de2b475

Observation 0baec4dd-c9d3-4142-830b-0b86a2662837 · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.775644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.624055Z digest=sha256:0a74067e0177180a9b18e4519b5eba26840bbcec1f08a754a98f60c962f08183

Observation e2e7f64e-bc1c-4d6c-8fc8-398d7eb6489d · outbound

This paper cites OTTER: A vision-language-action model with text-aware visual feature extraction,.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models OTTER: A vision-language-action model with text-aware visual feature extraction,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:43:04.763546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.627490Z digest=sha256:a16beec64c868a0d96dacd8390f3330cb80b000215fe5ae02e272ba4508f0b1d

Pith citing papers

Observation b4e81357-9ff2-4195-97c7-84368156384e · inbound

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models cites this paper.

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:43:04.752075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:43:04.520409Z digest=sha256:be5d19c12082114fe12e66b4a348e3db181a7ed3f3c060e11699041ebc1fc5a4