Pith. sign in

Paper Citation Record · LEDGER

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.13807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.13807 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:38:43.950532Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:35:04.661111Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da1fcaa3-c76b-4cb6-a7b7-28e65f9900a1 · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:43.950532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:43.950532Z digest=sha256:a6e69ba1b7737d437a5e4c6581f5fff7e2dcb3aae0bc4d2e71713a739dbc25bc

Observation 295e2f7d-6221-4401-874a-24cfa57b4430 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.060746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.060746Z digest=sha256:e348e429b600deb1ac837beb2d88f3878e128c2623fcf310b0611f707710283b

Observation ea651a9d-1b27-4bff-8db3-46c22c15aa32 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.341363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.341363Z digest=sha256:5ba7fd44767093cebcbd69de1270f4ec54c2de2987e17c64a00d2825fe451c45

Observation e4e6362a-cf70-4623-b7e9-d766c9cd45bd · inbound

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection cites this paper.

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:13.433486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:31:28.091617Z digest=sha256:8c0964939c38e360ef04f2e19550c5cc9916f006bcc24f8b3f56b313d5d3fcd0

Observation 8a1deaec-c68e-47f1-987f-0db9b217f58b · inbound

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding cites this paper.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.662610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:29:27.063028Z digest=sha256:08b85f8bd32d43c70690aca2b8149eb0684a9ecf0884f250d71ad1e6f7e807fc

Observation 8f5c2a54-2c1c-46d1-a50e-8ff3b43a3dcf · inbound

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation cites this paper.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.186399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.186399Z digest=sha256:8dcad06d2404dfc18f999d781181f34b6e53e74fee1fca59f6383f78926f96d6