Pith. sign in

Paper Citation Record · LEDGER

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

As of 5 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.14497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14497 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:59:18.438342Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87102886-f169-4b84-a307-920d8b8cf547 · outbound

This paper cites Vqa: Visual question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Vqa: Visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.673764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.673764Z digest=sha256:407e2b43b7c91ba27b33292a2ecc29c906b14d08ed74eac06f3d9079b733fb49

Observation d1baaa2c-29cd-459b-8cd8-0bd917856433 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.709112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.709112Z digest=sha256:e3c38af65dcfa2f527607a40b23d9c796878c91bcd2325686d9a4811bdd0ac78

Observation c6563afa-6cfa-45d2-b7a5-fa893b1a44db · outbound

This paper cites Where did i leave my keys?- episodic-memory-based question answering on egocentric videos.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Where did i leave my keys?- episodic-memory-based question answering on egocentric videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.760465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.760465Z digest=sha256:4f3e0fe96df211c9f09f1ed3adf09f37c4e44ecfd15b53fba93b31087c3a034c

Observation ba5afb2f-81d0-4d25-bcfd-99e37fb528cc · outbound

This paper cites Ad- abins: Depth estimation using adaptive bins.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Ad- abins: Depth estimation using adaptive bins

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.813459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.813459Z digest=sha256:b8d812c438e019ee36431bf21f9837584dc0ea587c503c597386f21b6dbea5a5

Observation 53c3138e-a0d3-48de-b8b1-ae1269a4eec2 · outbound

This paper cites Local- bins: Improving depth estimation by learning local distributions.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Local- bins: Improving depth estimation by learning local distributions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.889620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.889620Z digest=sha256:294d11ea1925be5b55712bf2960b9638233ca32524e82f8733d2be7b14b4dab0

Observation c913e4c3-f841-42a2-a44e-5bb888d908cd · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.973132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.973132Z digest=sha256:165890ec3b030388684760e033f1ee133aa58e4a24f0633c63c65874dbf5ae5f

Observation b00dd320-b97d-4d9c-b48c-de6115cbe39c · outbound

This paper cites Scene text visual question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Scene text visual question answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.051371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.051371Z digest=sha256:ebb2f51013a2666105fc7602b52038f41fd6d54591668c2033c2835e5dff43de

Observation 0904920c-fad8-4d7d-861f-109b43a2e7b4 · outbound

This paper cites Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.110105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.110105Z digest=sha256:fed0940b2338eeb87c478fc05cfcb52c64d128bed1e4d3ff2b81f2c567cf69e0

Observation eb49d6b0-ea75-4ba5-ad42-13c9129251b5 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.160109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.160109Z digest=sha256:e36b9e62085a9d31c9a616c4ce051c687bce12be2c021f7337f588153b6ab724

Observation edc05bd7-0390-45f0-82c1-7042fad95dfe · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.211777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.211777Z digest=sha256:be40fcc7cc05cbbf00afe97a43a418a08fad07199952e1f742e8213ae1eeeb1b

Observation 4e551c8d-b69d-4ceb-a583-a5fa3b571b01 · outbound

This paper cites Egothink: Evaluating first-person per- spective thinking capability of vision-language models.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egothink: Evaluating first-person per- spective thinking capability of vision-language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.261373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.261373Z digest=sha256:3f8351d3418f5a667f44d8a10796ac45dffb9337dbdafa43820fcdb1eabbb725

Observation 7200e93a-5def-46fe-bcb1-3cace70790b4 · outbound

This paper cites In- structblip: Towards general-purpose vision-language models with in- struction tuning.Advances in neural information processing systems, 36:49250–49267, 2023.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation In- structblip: Towards general-purpose vision-language models with in- struction tuning.Advances in neural information processing systems, 36:49250–49267, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.307319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.307319Z digest=sha256:cd2122c58f7cf6104d372c1abfb3887d8d4c86900eb1021e02f04d588e2e1b6f

Observation 7c6fbf91-1078-4122-86ee-f7e220cad36c · outbound

This paper cites Egovqa-an egocentric video question answering benchmark dataset.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egovqa-an egocentric video question answering benchmark dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.366901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.366901Z digest=sha256:f8f908fd88ada665e642d58be0ddc1368e6fffcfa0800cf097ebb424f30407f9

Observation 64d09369-925b-4cb3-955f-c0ca4032c1ef · outbound

This paper cites Deep ordinal regression network for monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Deep ordinal regression network for monocular depth estimation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.370801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.370801Z digest=sha256:d2840312fb3a43b923ba5f0fb349a4b5d20004d3a4451621c901a0dc0c3b1b60

Observation a68823fd-3ddf-4389-a26b-3b300c8eedd2 · outbound

This paper cites Geowizard: Unleash- ing the diffusion priors for 3d geometry estimation from a single im- age.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Geowizard: Unleash- ing the diffusion priors for 3d geometry estimation from a single im- age

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.386825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.386825Z digest=sha256:0614728d042350e2f08c1d8ffca12462bf17ca8859f4e8b173daa25de4f6b0a9

Observation a0f02a9a-0189-407b-8ee2-f0ac16bb1249 · outbound

This paper cites Unsu- pervised monocular depth estimation with left-right consistency.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unsu- pervised monocular depth estimation with left-right consistency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.496125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.496125Z digest=sha256:87ad6de10e41b36e568ea9f5c43b5aaabb5e4177161bd73b00183e31fa5f6423

Observation 8dfe3c77-c5b6-4d0e-9487-613d8a9f0a94 · outbound

This paper cites Digging into self-supervised monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Digging into self-supervised monocular depth estimation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.635625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.635625Z digest=sha256:7bc1d19e657befa9a138a43cf9e459e4fc7648527cebbf8cef7a5f9f37e9073e

Observation bfd9642e-3138-4ad4-90fd-e6b1130a8b85 · outbound

This paper cites Depthfm: Fast generative monocular depth estimation with flow matching.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depthfm: Fast generative monocular depth estimation with flow matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.715378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.715378Z digest=sha256:7b1dee4d95d8d2b3c51a3951f97f19035e145be408da092d7853488de5061bfb

Observation b6a72b5d-c784-499e-b672-ae41c181adca · outbound

This paper cites Towards zero-shot scale-aware monocular depth es- timation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards zero-shot scale-aware monocular depth es- timation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.800762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.800762Z digest=sha256:5c3e7c1cb2035b4a4ed9738127f0c6c86a1d7c5ca360baad781fa2a3711fca50

Observation 1043def6-f760-4ee6-a3ee-efca78ed658f · outbound

This paper cites an unresolved cited work.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.915084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.915084Z digest=sha256:5af0d11913d719bd2f1fb3d11f72ad681f17f49787763deba5c20077ebf3d3e3

Observation c3d7ac6d-e563-47cd-95ef-b71a75f18254 · outbound

This paper cites Ego- taskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Ego- taskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.989903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.989903Z digest=sha256:c2b82bd95807fddb8525da41e625a21e34712dc36c50062d104faf8a49084810

Observation bedfa933-1617-41d8-be95-6468ee25ca17 · outbound

This paper cites Repurposing diffusion- based image generators for monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Repurposing diffusion- based image generators for monocular depth estimation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.054865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.054865Z digest=sha256:9c9c1753cbfbe248cfd89f44ab24f48ab07952b19857cd4e8b7e6fbc39a3f3d7

Observation aac8e7bc-8820-4dba-a97b-c697b76ec942 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.130687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.130687Z digest=sha256:e29c2919e7047a62efbd12e94b15a8ff8e0809c6778ccbb0ab4e960f1c3ef364

Observation f2837398-d58e-4ca5-9547-b1366468f428 · outbound

This paper cites Egocross: Benchmarking multimodal large language models for cross- domain egocentric video question answering.arXiv preprint arXiv:2508.10729, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egocross: Benchmarking multimodal large language models for cross- domain egocentric video question answering.arXiv preprint arXiv:2508.10729, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.237961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.237961Z digest=sha256:cc90b7fb762bc961c1a52fb13e399fce3359c1201349ed7e88c4c4fe9dcd208d

Observation 3ee1175e-1636-415f-a0b1-67519baae3d8 · outbound

This paper cites BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.314451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.314451Z digest=sha256:c08250525b7dd44a1f23354c5b621fd0e4cd2469d28299487d7ffd12d166bd17

Observation 8b4ab995-c0aa-4481-b14c-79666e8cd61d · outbound

This paper cites Patchfusion: An end-to-end tile-based framework for high-resolution monocular met- ric depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Patchfusion: An end-to-end tile-based framework for high-resolution monocular met- ric depth estimation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.396078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.396078Z digest=sha256:967c6ed42766ecee045b3b31ccecb60d80891574861c793868112bb6bf8d638c

Observation 4e457fda-1d07-4c99-af11-3dc7964ebf5b · outbound

This paper cites Unibind: Llm-augmented unified and balanced representation space to bind them all.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unibind: Llm-augmented unified and balanced representation space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.465821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.465821Z digest=sha256:cbd4f724d624a10519779429b48e8886d8bc66359b3d2a0853c83497eb681f94

Observation 63523dd7-515a-4f88-a91d-3a270c51eab6 · outbound

This paper cites Realrag: Retrieval- augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Realrag: Retrieval- augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.610317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.610317Z digest=sha256:e45310e0f9638f9d68d97632cf4b6afa337a492c959d0dbb045bdbe71d12491a

Observation f9a25bdb-33de-4010-8f4d-d7ec315278b4 · outbound

This paper cites Single image depth estimation: An overview.Digital Signal Processing, 123: 103441, 2022.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Single image depth estimation: An overview.Digital Signal Processing, 123: 103441, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.686243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.686243Z digest=sha256:c18d9fe24c517cbe583c99dcab842eab902413a9e3bab75172f207733c4a7b01

Observation 75867d58-7767-4496-b5a4-60dd58376bac · outbound

This paper cites UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.809240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.809240Z digest=sha256:9e026334bd560f9cafb59886b076f4b996a1d6a2c0d6ab94d1f56d8637e4628c

Observation 19aa8d47-db44-4ab3-8735-4d50f542d48e · outbound

This paper cites Towards robust monocular depth estimation: Mix- ing datasets for zero-shot cross-dataset transfer.IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637,.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards robust monocular depth estimation: Mix- ing datasets for zero-shot cross-dataset transfer.IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.895759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.895759Z digest=sha256:c569652e946b52ec0c532634fc923f22c5e86ad620e1816d69070e6f335effcf

Observation aa72a3f3-6776-49ce-a8db-4fe5854726ee · outbound

This paper cites Monocular depth estimation us- ing neural regression forest.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Monocular depth estimation us- ing neural regression forest

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.960864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.960864Z digest=sha256:6f083598f64380b4e96cf442c6859c67a04b14b14fe168a5b061e384d48d660d

Observation e09c48aa-acf9-4216-90f2-c1c3b39690db · outbound

This paper cites Towards vqa models that can read.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards vqa models that can read

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.077636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.077636Z digest=sha256:71b93cd2cc7e903858b5f5d4145bd3b556762c80f2a0d767f521314e50297688

Observation 8f5c2a54-2c1c-46d1-a50e-8ff3b43a3dcf · outbound

This paper cites AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.186399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.186399Z digest=sha256:8dcad06d2404dfc18f999d781181f34b6e53e74fee1fca59f6383f78926f96d6

Observation 58287337-48c1-4af8-bd02-c2ce6333dd5b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.304733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.304733Z digest=sha256:b60f31764b1a8f23ca676cf83625400691da2d8540297ed13f70a1a50b7a6623

Observation be175290-bc13-4f19-a30c-748ab95edd3e · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Emu3: Next-Token Prediction is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.406676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.406676Z digest=sha256:a4e68cf3acc81a85a9ee4aa527d030cfa5e49b1237d50ebc50a09297ddfdb017

Observation c0f015b3-0808-4241-8a97-b5a6fd790104 · outbound

This paper cites Fastdepth: Fast monocular depth estimation on em- bedded systems.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Fastdepth: Fast monocular depth estimation on em- bedded systems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.515648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.515648Z digest=sha256:541b0c73540501f8cd9f51fc9da16a79dbf06c3c6e3a1c0884aa095f1855160d

Observation 029520bc-4b9d-41fb-a7a1-c1ca069ee1b2 · outbound

This paper cites Visual question answering: A survey of methods and datasets.Computer Vision and Image Understanding, 163:21–40, 2017.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Visual question answering: A survey of methods and datasets.Computer Vision and Image Understanding, 163:21–40, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.647380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.647380Z digest=sha256:2a53564251646ecf304d03f9ca5a903966c987eef72beb57a612dc2862b5377e

Observation b3001eee-3042-4fda-be0e-89e8fe23942b · outbound

This paper cites Egolife: Towards egocentric life assistant.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egolife: Towards egocentric life assistant

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.731562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.731562Z digest=sha256:972267131ba14eb6ef35755e49fb0daf02d51a6d22d1a2cca6c955134943b939

Observation a0246b27-dfdb-4ffc-9442-702193a6e75b · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37:21875–21911, 2024.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depth anything v2.Advances in Neural Information Processing Systems, 37:21875–21911, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.802293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.802293Z digest=sha256:e0d59ad413aa3bd50e2240284f8a692f21a231260ea22fa0af29d48d936fb376

Observation 9a8cce54-98b9-488c-ad2a-3b209436c2fe · outbound

This paper cites MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.934979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.934979Z digest=sha256:fb028867a1353ec81a3d555368065f095a3535271838406e812b5e047c5f1ad0

Observation 967597eb-dbee-467e-b868-02cb603a993a · outbound

This paper cites Metric3d: Towards zero- shot metric 3d prediction from a single image.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Metric3d: Towards zero- shot metric 3d prediction from a single image

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.058550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.058550Z digest=sha256:e4fba9d326949556db9c018c116e6fbcbe5d55fc948db9c8233677873134306e

Observation 857d1fd3-24c7-4d7b-ad72-d2e06c774bbf · outbound

This paper cites NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.177460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.177460Z digest=sha256:73c76f3809269a593200b52550874ab79cadee5672db85ccde732e807d5efae8

Observation ace98f0a-0237-417a-a0ad-8f059242d5c8 · outbound

This paper cites Egonight: Towards egocentric vision under- standing at night with a challenging benchmark.arXiv preprint arXiv:2510.06218, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egonight: Towards egocentric vision under- standing at night with a challenging benchmark.arXiv preprint arXiv:2510.06218, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.266673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.266673Z digest=sha256:5d356a7423715e31a0bb7deecd0e1b3a4070f8d00a3b89bc1cc7f0641a1c74e6

Observation 9400028d-e145-47f0-830b-fc6eae73cfb5 · outbound

This paper cites Egotextvqa: Towards egocentric scene-text aware video question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egotextvqa: Towards egocentric scene-text aware video question answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.334406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.334406Z digest=sha256:d553e1f56d007a1224f5574bcf853d13a5ad8db24438abcd5139728a83bc18a3

Observation d5169957-d4f4-42c6-8280-353bf109e023 · outbound

This paper cites ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.438342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.438342Z digest=sha256:3a2e7c17744907baa8f8e71144258e3f26361d9b265743a60e87e7d4dc89f655

Pith citing papers

No inbound Pith citation observations are available.