Pith. sign in

REVIEW 4 cited by

EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.04174 v2 pith:ITZU7HXD submitted 2021-07-09 cs.SD cs.CVcs.LGeess.ASeess.SP

classification cs.SDcs.CVcs.LGeess.ASeess.SP
keywords speechdatasetalgorithmsaudioarrayaugmentedcocktailcontains
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Augmented Reality (AR) as a platform has the potential to facilitate the reduction of the cocktail party effect. Future AR headsets could potentially leverage information from an array of sensors spanning many different modalities. Training and testing signal processing and machine learning algorithms on tasks such as beam-forming and speech enhancement require high quality representative data. To the best of the author's knowledge, as of publication there are no available datasets that contain synchronized egocentric multi-channel audio and video with dynamic movement and conversations in a noisy environment. In this work, we describe, evaluate and release a dataset that contains over 5 hours of multi-modal data useful for training and testing algorithms for the application of improving conversations for an AR glasses wearer. We provide speech intelligibility, quality and signal-to-noise ratio improvement results for a baseline method and show improvements across all tested metrics. The dataset we are releasing contains AR glasses egocentric multi-channel microphone array audio, wide field-of-view RGB video, speech source pose, headset microphone audio, annotated voice activity, speech transcriptions, head bounding boxes, target of speech and source identification labels. We have created and are releasing this dataset to facilitate research in multi-modal AR solutions to the cocktail party problem.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AmbiDrop: Array-Agnostic Speech Enhancement Using Ambisonics Encoding and Dropout-Based Learning

    eess.AS 2025-09 conditional novelty 6.0 of 10

    A speech enhancement network trained only on simulated Ambisonics signals, with dropout, generalizes to unseen microphone arrays when inputs are encoded via Ambisonics signal matching.

  2. Ambisonics Encoder for Wearable Array with Improved Binaural Reproduction

    eess.AS 2025-07 reject novelty 5.0 of 10

    A joint Ambisonics-binaural encoder loss is proposed, but the claimed closed-form solution is not the actual minimizer, so the method reduces to a trivial filter interpolation.

  3. EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A joint distillation and policy-learning framework claims near-teacher accuracy on egocentric action recognition, active speaker localization, and behavior anticipation at a fraction of the compute.

  4. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

Pith tools