Pith. sign in

Paper Citation Record · LEDGER

EgoM2P: Egocentric Multimodal Multitask Pretraining

As of 7 August 2026, this Paper Citation Record lists 100 of 138 outbound references and 0 inbound Pith citation observations for arXiv:2506.07886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07886 v3

Coverage vector

measured 100 of 138 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:31:42.599141Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 138 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15894f6c-633e-4832-806d-3e43a6a26cda · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.346223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.346223Z digest=sha256:b6e203932f9e5dd79e84e48329bb5517537ad85dae6de89965f2e3514a94694c

Observation 0a8f9399-7488-4ef2-a179-2678ae232f49 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone.

EgoM2P: Egocentric Multimodal Multitask Pretraining Phi-3 technical report: A highly capable language model locally on your phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.349741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.349741Z digest=sha256:0ad7c0441999940b0d6401bf829cf571eab5f5660a353322f0335bbbb7b1745c

Observation cf5137c5-86e3-4333-9e18-23dcab4e0c56 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cosmos World Foundation Model Platform for Physical AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.352925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.352925Z digest=sha256:fca87add970997bcfa427fb246cbf1937e021ad0b6cc02a02bbfb094f38ce8b9

Observation f63069d3-06bb-4d13-8e40-4a4c0cd3b781 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

EgoM2P: Egocentric Multimodal Multitask Pretraining Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.356074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.356074Z digest=sha256:a19e7a20bcdabf24a72499571da3e280465372e326d637eaeff0aeb8c38dabb9

Observation 7fad7658-0c0b-409e-886d-56e0c4d08177 · outbound

This paper cites Scenescript: Reconstructing scenes with an autoregressive structured language model.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scenescript: Reconstructing scenes with an autoregressive structured language model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.358916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.358916Z digest=sha256:df0692b02f29e22c1b17090081949c1364d5eebfaab570f3d09f41af852a18c6

Observation 0617dc0d-0e34-4703-96c7-c1db19e69ccb · outbound

This paper cites Newcombe, and Vasileios Balntas.

EgoM2P: Egocentric Multimodal Multitask Pretraining Newcombe, and Vasileios Balntas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.361669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.361669Z digest=sha256:041b8c2c25c0a886ac51e0cfd5878fe1a91a4d12399846228ec0a27375c5b8f1

Observation bc8ce495-3b8b-4b19-bc39-a6519066f386 · outbound

This paper cites MultiMAE: Multi-modal multi-task masked autoencoders.

EgoM2P: Egocentric Multimodal Multitask Pretraining MultiMAE: Multi-modal multi-task masked autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.364505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.364505Z digest=sha256:93a2085a8daffa9c693d3cc23a5dd9f47a1b0b6fccc822398dfa0005555bae61

Observation 449cfd9e-580e-4616-901f-ab5e8a91cdd1 · outbound

This paper cites 4M-21: An any-to-any vision model for tens of tasks and modalities.

EgoM2P: Egocentric Multimodal Multitask Pretraining 4M-21: An any-to-any vision model for tens of tasks and modalities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.366966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.366966Z digest=sha256:74b00a958d82d7ab0441068c19450106e12c13a16485b68fb419059b9d14ba58

Observation 13c506a7-8635-4d99-9918-5431ab0404ce · outbound

This paper cites Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.369371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.369371Z digest=sha256:63e7476249c96a077c127430a01155a5c92336cc5e9dcb8f0ece0bc72d68a4cf

Observation f68b6e4d-f5d9-4ffe-b548-1bef9b959f7b · outbound

This paper cites Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking.

EgoM2P: Egocentric Multimodal Multitask Pretraining Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.372012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.372012Z digest=sha256:98858b0c6218c57b4659fe48bc8652b931a967843559f6e36a29acf8cbd53210

Observation b8f5cdf0-83c4-44f1-8530-eb19bc33e420 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

EgoM2P: Egocentric Multimodal Multitask Pretraining Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.375059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.375059Z digest=sha256:87354dd3c7234f66faceb4af4de43a22d737d2d4ceefe4eba296cdaabb17cf94

Observation 49ff6499-ee5d-430f-8f77-31886228d558 · outbound

This paper cites Align your latents: High-resolution video synthe- sis with latent diffusion models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Align your latents: High-resolution video synthe- sis with latent diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.377428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.377428Z digest=sha256:001fff54ae77dc221a892e6df4029ee6db42283053e3bf5589226a7acf6916ce

Observation 541dd731-2077-4ef5-b5a6-00c8efb30f53 · outbound

This paper cites Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.379761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.379761Z digest=sha256:d7f7d62b644a2bf4dddd81e5e2df8912543d07da56c8de39474b433cbc9b5fa3

Observation a852ea0f-f485-4732-8c3c-2aa14db15dee · outbound

This paper cites Genie: Generative interactive environments.

EgoM2P: Egocentric Multimodal Multitask Pretraining Genie: Generative interactive environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.382148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.382148Z digest=sha256:6887e77a82c1c234e200b56d32124200e59aa876ac00b42fa8431529853c122c

Observation 5c1d5162-9dbf-4d9a-b756-670d26569b51 · outbound

This paper cites End-to-end object detection with transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining End-to-end object detection with transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.384474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.384474Z digest=sha256:5dac29581b5d620c1d01d0f5a0701bc1f8ef4f228422431d794b04262b50e2df

Observation 3f9630f1-da1a-4b4d-9abd-2f82a4159002 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining Emerging properties in self-supervised vision transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.386991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.386991Z digest=sha256:8f6f655f209ed0cbec18932f7478dbb9e34fe0adcc7db6b75a7eabdf0fc22d9a

Observation 15fb6131-0fbc-44c4-a196-1853f71245d8 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.389243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.389243Z digest=sha256:7854dd5f4d72db3abb54a2f7556e34ac7bda45876b3c30cdbe561fda27807376

Observation f6e90f27-0f2f-4cdd-98fd-c557c5db6b1b · outbound

This paper cites Fleet, and Geoffrey Hinton.

EgoM2P: Egocentric Multimodal Multitask Pretraining Fleet, and Geoffrey Hinton

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.391946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.391946Z digest=sha256:5984bb7bac77cca6f01367dc89571e54ad755b2b028c452311e949824f49d5fa

Observation 8afff46f-d8e4-473d-8a61-84d53bd87593 · outbound

This paper cites Control- a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Control- a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.394269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.394269Z digest=sha256:e585430e811c5155dc776c14647c0cf1587609aaaf9aa044d6f05ee35c5665c2

Observation 2b910072-2163-4309-89dc-77bd1f85f945 · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scaling egocentric vision: The epic- kitchens dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.396601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.396601Z digest=sha256:983791dc685cb6f4cf006d7745b23112a2cf873525f50c0ad807f6e35edb819a

Observation 00a2618b-7fce-4b42-9bec-9b401e93b5e4 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EgoM2P: Egocentric Multimodal Multitask Pretraining An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.399010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.399010Z digest=sha256:62c3a7d1f4dd16ba57e6d097a1baa72c298883d6441f4e46ef9dfe753d3c0794

Observation 6d95ac89-690b-435f-af8c-9526af2a8c0b · outbound

This paper cites Structure and Content-Guided Video Synthesis with Diffusion Mod- els.

EgoM2P: Egocentric Multimodal Multitask Pretraining Structure and Content-Guided Video Synthesis with Diffusion Mod- els

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.401827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.401827Z digest=sha256:3520db64ec391ecc1a2c8980e655f96238e69d0833ca27be3e9ad8df428ecb76

Observation 34ffb50e-9f66-460b-8ce2-7ad3e4a53694 · outbound

This paper cites Black, and Otmar Hilliges.

EgoM2P: Egocentric Multimodal Multitask Pretraining Black, and Otmar Hilliges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.404240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.404240Z digest=sha256:96fd2a05b7c5f4168fe24f0271fce34341b34f8f235c7701b1c290e1f9796557

Observation 3dfc6897-3adf-4384-893a-18b9fd95211d · outbound

This paper cites HOLD: Category-agnostic 3d reconstruction of interacting hands and objects from video.

EgoM2P: Egocentric Multimodal Multitask Pretraining HOLD: Category-agnostic 3d reconstruction of interacting hands and objects from video

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.406558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.406558Z digest=sha256:b6f86adaf7fc7cb35fd2f591c24f80a0ccdf73872f4f3c11521a72f52a4d143a

Observation 505417fb-eae6-4fa0-b2de-e1c0b6e9f9d2 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

EgoM2P: Egocentric Multimodal Multitask Pretraining VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.408928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.408928Z digest=sha256:e338b3c77d7aa28faf813d2dd2b4ecccd72ca081b7f25327a0916b04f65c017b

Observation 55a2aeec-0259-4acf-8ae5-04c17ce47536 · outbound

This paper cites First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations.

EgoM2P: Egocentric Multimodal Multitask Pretraining First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.411701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.411701Z digest=sha256:1ba7e8a1bb818050dfb2e48762f17d3a8f31d018136fd04ea4d80b1532cc3863

Observation 5fd4f5f4-7520-4d04-a576-e2df90ba89ff · outbound

This paper cites Imagebind: One embedding space to bind them all.

EgoM2P: Egocentric Multimodal Multitask Pretraining Imagebind: One embedding space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.414302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.414302Z digest=sha256:18956658d20e8c2eea7cf554369bf668f0b8d2051c1213b81f3fb9f8d1894288

Observation 9e5a3b5b-3cb4-4816-a31b-6ad45b96c684 · outbound

This paper cites Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018.

EgoM2P: Egocentric Multimodal Multitask Pretraining Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.416754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.416754Z digest=sha256:ac0c14503eccfb6f5291d249a2bbeafa0879c7db65728439bf4441e5ddd97b00

Observation 0801895d-a503-4cb6-a12f-0c36348f3d69 · outbound

This paper cites Ego4d: Around the World in 3,000 Hours of Egocentric Video.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ego4d: Around the World in 3,000 Hours of Egocentric Video

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.419200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.419200Z digest=sha256:a4b2fe03bdb9587f80c0146063663ffec8fcc45b594e6d146ada40eb5c8541d0

Observation bcf61676-1751-4004-80ab-d9232c3a7592 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first- and third-person perspec- tives.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ego-exo4d: Understanding skilled human activity from first- and third-person perspec- tives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.421469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.421469Z digest=sha256:6397ffebe0700c5de93f1b4f284b945c3f9905936ff5ff322e247a621fcd6fc2

Observation 59fcde70-c33c-4575-9deb-69eab36e6483 · outbound

This paper cites Animatediff: Animate your personalized text- to-image diffusion models without specific tuning.

EgoM2P: Egocentric Multimodal Multitask Pretraining Animatediff: Animate your personalized text- to-image diffusion models without specific tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.423750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.423750Z digest=sha256:2924b72eb83229e900b5a2a1fc2459d5dc2bd305715d9a0d6d3e6018b173c642

Observation 9af1b819-83cf-4904-80fc-840d4704c4aa · outbound

This paper cites World Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining World Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.426124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.426124Z digest=sha256:72154458607efb121bb441bc99f6b80c71f6c8780265240bf7e642eb770ecef4

Observation 7c1575ea-1ee2-4ed0-8020-9e970c0fd46a · outbound

This paper cites Girshick.

EgoM2P: Egocentric Multimodal Multitask Pretraining Girshick

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.428661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.428661Z digest=sha256:889b3fa9930539901c1f9a86a8594f40c09650ff5ada6ed4002eefb15db2fcbe

Observation cfecc510-d9c4-4fa7-8696-949fad01d9f0 · outbound

This paper cites Classifier-free diffusion guidance.

EgoM2P: Egocentric Multimodal Multitask Pretraining Classifier-free diffusion guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.430965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.430965Z digest=sha256:766a4055491a21e2c06331b244b64bbf2f9b50d943b8c911b738a4205623c449

Observation c087ee39-de85-40d5-ab00-c371152241bb · outbound

This paper cites Video diffu- sion models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video diffu- sion models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.433709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.433709Z digest=sha256:85bbddb72cd23ee057416b143183a1a204d8b6d561998f294ee27c88a8b74098

Observation 4a4af980-0b14-40ca-abf4-c16d7b4263c4 · outbound

This paper cites The curious case of neural text degeneration.

EgoM2P: Egocentric Multimodal Multitask Pretraining The curious case of neural text degeneration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.436271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.436271Z digest=sha256:9a6374c4b34ae175ff5e79123bce5062b02c5af31f08bc29266417ca01c832db

Observation 7c459429-91fc-4ea3-a3cc-f48d486d3cf5 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to- video generation via transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cogvideo: Large-scale pretraining for text-to- video generation via transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.438690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.438690Z digest=sha256:d279bb4ad9a09078677d9f84bd8618693a1e91b116b10ecf7702d15a1d47ac8a

Observation 1e6a5851-42d7-4baf-be4c-4db19fd1e816 · outbound

This paper cites Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.441013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.441013Z digest=sha256:563967747b37e96a4a2e63be61203d43c76fb045d63dfbe1f45ea8685192e57f

Observation 0ff59b2a-d994-4462-9651-ffcf2ec61f47 · outbound

This paper cites Ross, and Alireza Fathi.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ross, and Alireza Fathi

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.443348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.443348Z digest=sha256:8ab7e8f264c8c006f4b552bdd0af2600f43502a0be0943a48ef7489a905b36b1

Observation d63a9227-59ab-40b8-b469-aae67ab99370 · outbound

This paper cites Predicting gaze in egocentric video by learning task- dependent attention transition.

EgoM2P: Egocentric Multimodal Multitask Pretraining Predicting gaze in egocentric video by learning task- dependent attention transition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.445693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.445693Z digest=sha256:0a19fa90c5effa2d0ed12d42d6ef2ee382c88b22203116a500ad594e57300db8

Observation 2a3c13d0-fb71-46ce-b967-3ad57c9ea5f2 · outbound

This paper cites GPT-4o System Card.

EgoM2P: Egocentric Multimodal Multitask Pretraining GPT-4o System Card

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.447923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.447923Z digest=sha256:da796dc879b4bfa02127a3d79b2bffd1c3229095d723e91c16d42538c9c3983b

Observation 400f5c80-0591-4a7f-a8ab-ffd96cbb14b8 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition.

EgoM2P: Egocentric Multimodal Multitask Pretraining Epic-fusion: Audio-visual temporal binding for egocentric action recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.450515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.450515Z digest=sha256:377f928c3cf6701d89dffeb98e1e93e08b90426cbca684541bfb349bd51442f8

Observation 04454c8a-dc38-4169-a86b-a8039f528166 · outbound

This paper cites Re- purposing diffusion-based image generators for monocular depth estimation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Re- purposing diffusion-based image generators for monocular depth estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.453671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.453671Z digest=sha256:fe45cf5df33e0fe3f0bed11430382c131f68cfcb9725158677e96db5a4635931

Observation c185bded-22cc-4cc7-bf12-3d5c9ceec635 · outbound

This paper cites Video depth without video models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video depth without video models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.456344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.456344Z digest=sha256:205b9700c581bac2fa181dca2ee9369a0053fdb7a134e04076f0d4c64d9824f2

Observation 36c03888-f351-40d8-834e-23e0d4fd5506 · outbound

This paper cites Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators.

EgoM2P: Egocentric Multimodal Multitask Pretraining Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.458614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.458614Z digest=sha256:26e8cf1fea867f295552d3b62c7a9230dd8ded4bb27e627d91de74e4d75d595a

Observation 306b8dd8-70c4-4016-8e9a-b90231c1cee4 · outbound

This paper cites Segment anything.

EgoM2P: Egocentric Multimodal Multitask Pretraining Segment anything

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.460971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.460971Z digest=sha256:de13031c42ebacdcc9a6ea1829acc012310ac6a7f7c52ae834476437360b68a8

Observation 85cfafe2-386a-4e0b-a8b3-b6856a206452 · outbound

This paper cites Harmsen, and Neil Houlsby.

EgoM2P: Egocentric Multimodal Multitask Pretraining Harmsen, and Neil Houlsby

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.463435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.463435Z digest=sha256:a7a8d4991f29d6694cbdba045d6db072a5f21ecfb0bd4bfb0d9a3fd32de4a3cb

Observation 56202ac3-0576-4504-b221-2e677d77f744 · outbound

This paper cites VideoPoet: A large language model for zero-shot video generation.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoPoet: A large language model for zero-shot video generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.465709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.465709Z digest=sha256:4e0ff33d9b486dd2145133d89291e0eb813bd998188a05a5f903fc318e4bef3b

Observation 8aaa8fa9-d631-4558-9fcf-3e60bf0ae668 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

EgoM2P: Egocentric Multimodal Multitask Pretraining H2o: Two hands manipulating objects for first person interaction recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.468298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.468298Z digest=sha256:e5e2c63860e545f51a88989d9cfc8cfb0b381cdd4497f87f8f71e714e7a196ae

Observation 4caa288a-432f-4f65-affa-3db175476b05 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.470764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.470764Z digest=sha256:ee7f67b96315ad18fa1d7b3044fbed4447c08c98d5d4e26955bdc3e21677b2a0

Observation a82b8a1d-6305-4cfa-93db-c0ffa7b3405a · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lisa: Reasoning segmenta- tion via large language model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.473203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.473203Z digest=sha256:237b4ff890f651f88846c50c960ab14bc56313f222c115daf753b901067b4838

Observation 31d3e8ba-650b-4fb4-992f-ca9afb0f6ad1 · outbound

This paper cites Egogen: An egocentric synthetic data generator.

EgoM2P: Egocentric Multimodal Multitask Pretraining Egogen: An egocentric synthetic data generator

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.475810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.475810Z digest=sha256:25cd3ad401427316cd49302c442f0c6bfa706c526a1924b769fcfcc6678cb986

Observation f4bd1b8c-c516-46d8-a784-fb266f7e1a4f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoChat: Chat-Centric Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.478451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.478451Z digest=sha256:e1498f42e93b1204eeb55267e68cf6b49fb6767c03570b7f384a778aff6a16d5

Observation aec46962-59cc-45c6-acbc-8457da029965 · outbound

This paper cites Megasam: Accurate, fast and robust structure and motion from casual dynamic videos.

EgoM2P: Egocentric Multimodal Multitask Pretraining Megasam: Accurate, fast and robust structure and motion from casual dynamic videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.481256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.481256Z digest=sha256:d9eed68d2953e8d1b620241d5f2e67544f4fc7ade0ec7648aa726bc5a2361f78

Observation 45d097af-9a90-49d3-91e0-60c25371e42e · outbound

This paper cites Video-LLaV A: Learning united visual representation by alignment before projection.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video-LLaV A: Learning united visual representation by alignment before projection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.483820Z digest=sha256:5df3b5f6985ba736726fcf6a351ffca11cf8f6d5ba868dc15fa5f21a5e92f8cb

Observation 3e263881-3e9a-4794-86af-d218551e21ae · outbound

This paper cites Cross-view exocentric to egocentric video synthesis.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cross-view exocentric to egocentric video synthesis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.486760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.486760Z digest=sha256:2e192b07c4aad5c5d212fe259fb11726f88e0c5c45411e46de4b085553c07568

Observation 01a27f75-5b00-4edf-8daa-383a5ad7a7c2 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.489037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.489037Z digest=sha256:6a5999f7459f8710465a5737330970a4b0e18635a31fb278e0e1d9524fb16ea4

Observation 6b3c69ab-e05a-4fbb-b93b-29013773e9f4 · outbound

This paper cites Exocentric-to-egocentric video gener- ation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Exocentric-to-egocentric video gener- ation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.491572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.491572Z digest=sha256:653bc407624403652df0d2661ecede2cbc0b760281273ce502e1a06b6899acd7

Observation 6f7ff5ef-cb26-4e70-a7fc-dc3429d9c4d3 · outbound

This paper cites Li, Ying Shan, and Ge Li.

EgoM2P: Egocentric Multimodal Multitask Pretraining Li, Ying Shan, and Ge Li

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.494202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.494202Z digest=sha256:b78cf38159782eb7842be68cbe79ae1f715c01e59530c76b48b0d35610521c90

Observation fee22a39-4c72-4ac1-8728-4142f468e909 · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

EgoM2P: Egocentric Multimodal Multitask Pretraining Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.499922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.499922Z digest=sha256:3a2b2ad24c3532d807e2de877ad7bc7ed7d8d9185e133a462ac08eed9431e1a3

Observation baa0c2cf-52c5-41ae-99fe-26b4bc392390 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.502461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.502461Z digest=sha256:09f551e2354730a44e70b00ef5810694b8db0ef78c11e3b5807b30ab799807a8

Observation 9064648b-65c1-4b38-a48f-fdd5d5acbc60 · outbound

This paper cites A convnet for the 2020s.

EgoM2P: Egocentric Multimodal Multitask Pretraining A convnet for the 2020s

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.505098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.505098Z digest=sha256:18429129901577a8dd80f6333d3d284a0c071dae252768190e831618c9eb4d85

Observation b95d2052-f35f-4041-8469-e7085267b8e4 · outbound

This paper cites Decoupled weight de- cay regularization.

EgoM2P: Egocentric Multimodal Multitask Pretraining Decoupled weight de- cay regularization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.507569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.507569Z digest=sha256:b83c88a3a56985c8ea7f21350f46152206d7d27be231fdaf85f47f9ec3814967

Observation d35445cd-14d2-407e-a5e3-3c65dc60ab61 · outbound

This paper cites Unified-io 2: Scaling autoregressive mul- timodal models with vision, language, audio, and action.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unified-io 2: Scaling autoregressive mul- timodal models with vision, language, audio, and action

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.509938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.509938Z digest=sha256:f7d8f83a870f7b103b64eff44a6bf88b994780247468cf8d28cf9cf3aec4926e

Observation 4ddd5a67-06cf-42cf-968e-3ddfee63a802 · outbound

This paper cites UNIFIED-IO: A uni- fied model for vision, language, and multi-modal tasks.

EgoM2P: Egocentric Multimodal Multitask Pretraining UNIFIED-IO: A uni- fied model for vision, language, and multi-modal tasks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.345152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.512343Z digest=sha256:1d71a282f991b43f6b4857ff93c74a677c8ab50b91eb61bb3812f69c411dbaa5

Observation a87fcc94-61c6-4fb1-a7f3-9790ca147d68 · outbound

This paper cites Align3r: Aligned monocular depth estimation for dynamic videos.

EgoM2P: Egocentric Multimodal Multitask Pretraining Align3r: Aligned monocular depth estimation for dynamic videos

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.337663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.515421Z digest=sha256:14b5192eaebb21521e2ef73c9a318befcc57c4b92efc009644f3a7617e7a7e54

Observation e196fdd4-630d-480a-8b97-1abbf6d199b6 · outbound

This paper cites Dream Machine.

EgoM2P: Egocentric Multimodal Multitask Pretraining Dream Machine

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.329825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.518003Z digest=sha256:c1aaa0b026338487d99e66e4215d66889287d899fbab0dab4605eaae3fd9f14e

Observation 85040bb0-72f9-4587-923d-e6ba6321b337 · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Videofusion: Decomposed diffusion models for high-quality video generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.322374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.520485Z digest=sha256:cdd4d84df7070017a6684ca4071be669e2502d673d07da2a2755691ef781dcc4

Observation b49d3759-826d-4058-912b-f96bc2c8be84 · outbound

This paper cites Aria Everyday Activities Dataset.

EgoM2P: Egocentric Multimodal Multitask Pretraining Aria Everyday Activities Dataset

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.522928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.522928Z digest=sha256:5f2f1a270d12e4a197ce23c3a3be2cded99fada75ea070d59f0af740507a5089

Observation b6e9ab21-5033-431c-80fe-6db7083e494f · outbound

This paper cites Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild.

EgoM2P: Egocentric Multimodal Multitask Pretraining Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.525404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.525404Z digest=sha256:11591632521947b214c9bf75fe146fb2a427b4e3430b157ff26231173944d846

Observation 1fd07221-c02f-4ec3-a438-49767fe5dfa2 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.315038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.527939Z digest=sha256:daa5cf96b0cc2ab7a719c6969d126883ff93b3e96e3b0de18aaea5eb539a2c2a

Observation 9977c314-da17-4ddf-b118-483416f1e3f9 · outbound

This paper cites Mm1: Methods, analysis & insights from multimodal llm pre-training, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Mm1: Methods, analysis & insights from multimodal llm pre-training, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.306917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.530377Z digest=sha256:20aee4965a89188f441c2ed2b3a067ad32175a56df9c5ce1756cbd6188466848

Observation d9deb7a6-65ed-4c3a-b7db-84c72c36390f · outbound

This paper cites Project Aria Glasses.

EgoM2P: Egocentric Multimodal Multitask Pretraining Project Aria Glasses

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.299588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.532866Z digest=sha256:3b7a182a6a1bf358e781093bf2339f32aa70e73eede9f8ac418be872298945b8

Observation 8ba322c7-646c-418d-9622-261ee4a5c97b · outbound

This paper cites Transformers are Sample-Efficient World Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Transformers are Sample-Efficient World Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.535714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.535714Z digest=sha256:e3f44825ff2ec89f1f3f05a80412bcb21e1fc1904d8bb3717e042dd083679c3b

Observation 5dc57102-cd1d-463d-b21a-8dcf0ada7d48 · outbound

This paper cites HoloLens 2.

EgoM2P: Egocentric Multimodal Multitask Pretraining HoloLens 2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.292402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.538301Z digest=sha256:dab9cea7adecf67d759588477ed5151d5f6693541c0b63d7673fbdb1f0b9cd61

Observation c137608e-1b3b-4edf-bd0f-d879c2a3dbf5 · outbound

This paper cites 4M: Massively multimodal masked modeling.

EgoM2P: Egocentric Multimodal Multitask Pretraining 4M: Massively multimodal masked modeling

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.285124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.540832Z digest=sha256:fe4abc8d56c00baa6d5c01c2cacd35ba35734a25e478fd4a4ed716faec5c621d

Observation b4420ba5-ac62-4b89-b5cb-6046902b1220 · outbound

This paper cites AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation.

EgoM2P: Egocentric Multimodal Multitask Pretraining AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.277532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.543229Z digest=sha256:a508b17b3362bf18f7182139f13113f3f2c7f5baf07a7ab679cdf1e707bcefe6

Observation c79c1e4d-4e9a-4db0-aae0-0eff5aeb866a · outbound

This paper cites Video generation models as world simula- tors.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video generation models as world simula- tors

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.269716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.545537Z digest=sha256:d07112070ba35335d08aad8ecc9645cf6f8176271575cf426d84054ca1cd01bc

Observation 374a7076-4d4a-4983-9a51-fd3c31220660 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:31:43.262258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.548214Z digest=sha256:f76c0e39a3a3df8f98809aa03a085d2d6688bcc4230e7bb608c8318d2d581232

Observation 3fb761d3-f991-4650-a924-fac389397dcf · outbound

This paper cites Aria digital twin: A new benchmark dataset for egocentric 3d machine perception.

EgoM2P: Egocentric Multimodal Multitask Pretraining Aria digital twin: A new benchmark dataset for egocentric 3d machine perception

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.254661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.550637Z digest=sha256:b3cc71158f78fbca9c49e5cbe3d794fdab267b3129f47fe6f73b2c77058498ac

Observation b0563564-1aa4-491d-a10b-6305dffedad5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Movie Gen: A Cast of Media Foundation Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.552979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.552979Z digest=sha256:e145a5bb5e38f147cd08507d378d38983e8917934a9d8d1a899c3112792d8de0

Observation 2a104cc6-49f3-486a-a341-1b1e0f6f3a57 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EgoM2P: Egocentric Multimodal Multitask Pretraining Learn- ing transferable visual models from natural language super- vision

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.245929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.555752Z digest=sha256:9cbbf72b2995ef29c22786850787baf316db755d1581377fdeb8acc15e7e3f32

Observation 43e157ec-730b-452b-8b52-c5821f2d7f52 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:31:43.237643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.558132Z digest=sha256:8c7a98d638f10b890df850a7f55c8b2abffa8a6a240c0b21089012e4b8d895e9

Observation 64ae6b91-6d2e-4268-bbab-dff2f7364b2d · outbound

This paper cites High-Resolution Im- age Synthesis with Latent Diffusion Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining High-Resolution Im- age Synthesis with Latent Diffusion Models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.229469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.560589Z digest=sha256:ca0ccad550730cfbdc9056057c2bcf804d8bc383d3c68db72ab3adb20187b791

Observation b3f70a23-1293-4c39-9ff0-f360abc10974 · outbound

This paper cites Gen-3 Alpha.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gen-3 Alpha

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.221152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.562991Z digest=sha256:9d27651c2f945961a0144d6be5cbdfeb9a82e2375a264f31b13757c5e127cb3b

Observation 66bddc3a-7e9c-4f03-bd41-b88c485106d3 · outbound

This paper cites Lamar: Bench- marking localization and mapping for augmented reality.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lamar: Bench- marking localization and mapping for augmented reality

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.212813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.565470Z digest=sha256:c14a78b518543d91b97b3860dcdc745af26798261c8af9a52b41352174fead75

Observation 07d1e775-b7d6-4c92-94dc-a1b1a51ed7b1 · outbound

This paper cites Sener, D.

EgoM2P: Egocentric Multimodal Multitask Pretraining Sener, D

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.203556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.567826Z digest=sha256:7d6152fd9c4d3aceb9a43976f858e9a7daf07b4a78cd908a00485699805f53e2

Observation ee2d64bc-3bc5-478e-af14-5ff57d27f4ce · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

EgoM2P: Egocentric Multimodal Multitask Pretraining Make-a-video: Text-to-video generation without text-video data

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.570245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.570245Z digest=sha256:ef72f8d9af87f31a5aae657677079e6177d11c6203bb2a8ccc3adb74e90df8bf

Observation 2706b4d7-8665-4f5e-9dc2-06dd67743634 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

EgoM2P: Egocentric Multimodal Multitask Pretraining The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.572653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.572653Z digest=sha256:2581cb1fb16d44bf1b422605a312526e14de7190eb93ac9f6b8f224dc133d285

Observation c84926b5-1768-451c-b056-ae70ede2ec88 · outbound

This paper cites Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon.

EgoM2P: Egocentric Multimodal Multitask Pretraining Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.191261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.575522Z digest=sha256:e92aa92254b4d54dc283357eca3a101b9391753197e46289a134426442f81db8

Observation c5e558a0-b2ed-4d3d-bf55-4fb269ab28b3 · outbound

This paper cites Emu: Generative pretraining in multimodality.

EgoM2P: Egocentric Multimodal Multitask Pretraining Emu: Generative pretraining in multimodality

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.183503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.577766Z digest=sha256:85d068aa88f8a776fc18be1dda05ff912cf0378fdc64c4845e25f5ca955a9f15

Observation 9ba77c0d-e496-42e5-a6c2-5147d5e6eb40 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gemini: A Family of Highly Capable Multimodal Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.580076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.580076Z digest=sha256:e7fdf89e90d5e6f2f88fa3b3f8c3d1abe1dde28aabd6f6126949d64653af9b98

Observation 522fd6c4-b46a-4fc0-b37e-a072ab425b44 · outbound

This paper cites Kling ai video generator.

EgoM2P: Egocentric Multimodal Multitask Pretraining Kling ai video generator

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.176365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.582501Z digest=sha256:ff40fec7b046c0a2a8c9780c4eb165b13dbf807fdf2a2ea56e511e71dfd83e1c

Observation e8559f55-245d-4006-a0d8-af30007aedce · outbound

This paper cites DROID-SLAM: Deep vi- sual SLAM for monocular, stereo, and RGB-d cameras.

EgoM2P: Egocentric Multimodal Multitask Pretraining DROID-SLAM: Deep vi- sual SLAM for monocular, stereo, and RGB-d cameras

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.168267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.584808Z digest=sha256:a7a09de376cebc7aae63ba5bd1a89234c73d46c61460bf3a40554a0d9d1039b7

Observation 2fa6a850-a295-4b97-a6bc-d0ddc6ddd58f · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoMAE: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.160061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.587075Z digest=sha256:23adfdf6ec0a7347b3e436557f70be4837d41dd2224658ac93e8f84c1108c072

Observation 2a8dbe85-a2d5-4cb7-a498-9d7c352ce84d · outbound

This paper cites Towards accurate generative models of video: A new metric & challenges, 2019.

EgoM2P: Egocentric Multimodal Multitask Pretraining Towards accurate generative models of video: A new metric & challenges, 2019

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.151538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.589235Z digest=sha256:9ef4aa17f2ced33f01ed3d6c058866e6f8d4988e9c4be33d7a6dfc405f3b4e97

Observation 1e2f2197-f1ed-4ec3-9ad4-e1979f876097 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

EgoM2P: Egocentric Multimodal Multitask Pretraining Diffusion Models Are Real-Time Game Engines

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.591660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.591660Z digest=sha256:85771588010161abb82ccbcd9a5ca22f16dd6ab0fb03cb55d4c68b06fe59ca97

Observation 7e673542-0d8e-4206-bd37-c9d6f93d6215 · outbound

This paper cites Neural discrete representation learn- ing.

EgoM2P: Egocentric Multimodal Multitask Pretraining Neural discrete representation learn- ing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.142981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.594555Z digest=sha256:5879c92bc7f95d1d7a79e158b58be7aea60682f212063f79692e51efee607b5e

Observation 3fc44bf5-10be-4363-8f68-d9a1423b43b9 · outbound

This paper cites Attention is all you need.

EgoM2P: Egocentric Multimodal Multitask Pretraining Attention is all you need

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.134665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.596902Z digest=sha256:b75498ff953beaf8110530c2a06efec8673d3d2f6d53fab7bd4d6d6bae40c09c

Observation f383cfda-5eff-479b-b21c-46daadaed112 · outbound

This paper cites Phenaki: Variable length video generation from open do- main textual descriptions.

EgoM2P: Egocentric Multimodal Multitask Pretraining Phenaki: Variable length video generation from open do- main textual descriptions

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.126418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.599141Z digest=sha256:98fbc600b6c9c654c61ee09997f6cdecb3a0e5cc8ef6470f8f61bb67c47a5b87

Pith citing papers

No inbound Pith citation observations are available.