OVO-S-Bench provides 1680 human-annotated questions on 348 videos to measure streaming spatial intelligence in MLLMs across instantaneous perception, spatiotemporal tracking, spatial simulation, and allocentric mapping.
hub
Aria Everyday Activities Dataset
15 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
abstract
We present Aria Everyday Activities (AEA) Dataset, an egocentric multimodal open dataset recorded using Project Aria glasses. AEA contains 143 daily activity sequences recorded by multiple wearers in five geographically diverse indoor locations. Each of the recording contains multimodal sensor data recorded through the Project Aria glasses. In addition, AEA provides machine perception data including high frequency globally aligned 3D trajectories, scene point cloud, per-frame 3D eye gaze vector and time aligned speech transcription. In this paper, we demonstrate a few exemplar research applications enabled by this dataset, including neural scene reconstruction and prompted segmentation. AEA is an open source dataset that can be downloaded from https://www.projectaria.com/datasets/aea/. We are also providing open-source implementations and examples of how to use the dataset in Project Aria Tools https://github.com/facebookresearch/projectaria_tools.
hub tools
citation-role summary
citation-polarity summary
polarities
background 3representative citing papers
SuperMemory-VQA provides 4,853 human-verified QA pairs from 52.9 hours of egocentric AI glasses recordings to benchmark AI systems on realistic long-horizon memory tasks including an unanswerable option.
EgoTraj is a new open multimodal dataset of 75 long-horizon egocentric human navigation sequences in urban environments with head pose, gaze, and scene data, plus benchmarks of trajectory prediction methods.
EgoTL provides a new egocentric dataset with think-aloud chains and metric labels that benchmarks VLMs on long-horizon tasks and improves their planning, reasoning, and spatial grounding after finetuning.
Introduces Personal VCL formalization and benchmark revealing LMM context gaps, plus an Agentic Context Bank baseline that boosts personalized visual reasoning.
Kaon-driven freeze-in of light dark matter at low reheating temperatures requires couplings large enough that NA62, KOTO, and KOTO II can probe the same operator.
EgoEverything is a 5,000+ question benchmark that uses gaze as a proxy for human attention to generate long-context egocentric video QA pairs, where the best VLM scores 63.1% vs humans at 83.5%.
SpeechLess enables micro-utterance AR interactions by binding prior interactions to personal spatial context for intent extrapolation.
Memento captures verbal queries with spatiotemporal contexts, discovers recurring interest patterns, and proactively delivers updated AR responses when matching situations are detected.
SAM 3D reconstructs 3D objects from single images with geometry, texture, and pose using human-model annotated data at scale and synthetic-to-real training, achieving 5:1 human preference wins.
Fusing stereo vision features with text prompts that include object class and approximate volume via a projection layer improves volume regression over vision-only baselines on public datasets.
The paper introduces World Action Models as a new paradigm unifying predictive world modeling with action generation in embodied foundation models and provides a taxonomy of existing approaches.
The Monado SLAM dataset supplies real egocentric visual-inertial sequences from VR headsets to fill gaps in existing VIO/SLAM benchmarks for difficult real-world scenarios.
citing papers explorer
-
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
OVO-S-Bench provides 1680 human-annotated questions on 348 videos to measure streaming spatial intelligence in MLLMs across instantaneous perception, spatiotemporal tracking, spatial simulation, and allocentric mapping.
-
SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
SuperMemory-VQA provides 4,853 human-verified QA pairs from 52.9 hours of egocentric AI glasses recordings to benchmark AI systems on realistic long-horizon memory tasks including an unanswerable option.
-
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
EgoTraj is a new open multimodal dataset of 75 long-horizon egocentric human navigation sequences in urban environments with head pose, gaze, and scene data, plus benchmarks of trajectory prediction methods.
-
EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks
EgoTL provides a new egocentric dataset with think-aloud chains and metric labels that benchmarks VLMs on long-horizon tasks and improves their planning, reasoning, and spatial grounding after finetuning.
-
Personal Visual Context Learning in Large Multimodal Models
Introduces Personal VCL formalization and benchmark revealing LMM context gaps, plus an Agentic Context Bank baseline that boosts personalized visual reasoning.
-
MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware
Kaon-driven freeze-in of light dark matter at low reheating temperatures requires couplings large enough that NA62, KOTO, and KOTO II can probe the same operator.
-
EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment
EgoEverything is a 5,000+ question benchmark that uses gaze as a proxy for human attention to generate long-context egocentric video QA pairs, where the best VLM scores 63.1% vs humans at 83.5%.
-
SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented Reality
SpeechLess enables micro-utterance AR interactions by binding prior interactions to personal spatial context for intent extrapolation.
-
Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant
Memento captures verbal queries with spatiotemporal contexts, discovers recurring interest patterns, and proactively delivers updated AR responses when matching situations are detected.
-
SAM 3D: 3Dfy Anything in Images
SAM 3D reconstructs 3D objects from single images with geometry, texture, and pose using human-model annotated data at scale and synthetic-to-real training, achieving 5:1 human preference wins.
-
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
Fusing stereo vision features with text prompts that include object class and approximate volume via a projection layer improves volume regression over vision-only baselines on public datasets.
-
World Action Models: The Next Frontier in Embodied AI
The paper introduces World Action Models as a new paradigm unifying predictive world modeling with action generation in embodied foundation models and provides a taxonomy of existing approaches.
-
The Monado SLAM Dataset for Egocentric Visual-Inertial Tracking
The Monado SLAM dataset supplies real egocentric visual-inertial sequences from VR headsets to fill gaps in existing VIO/SLAM benchmarks for difficult real-world scenarios.
- HeadRoom: Lightweight, Edge-deployable Pipeline for Adaptive Notification Routing
- HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control