Pith. sign in

Paper Citation Record · LEDGER

FIction: 4D Future Interaction Prediction from Video

As of 15 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 0 inbound Pith citation observations for arXiv:2412.00932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00932 v2

Coverage vector

measured 100 of 119 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:55:59.914437Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 119 outbound references displayed

  • verified exact4
  • verified fuzzy38
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c95a774e-e9a0-4766-a269-52b754da3f11 · outbound

This paper cites When will you do what?-anticipating temporal occurrences of activities.

FIction: 4D Future Interaction Prediction from Video When will you do what?-anticipating temporal occurrences of activities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.397059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.397059Z digest=sha256:26af4fda3f17a3ebb8db67c2c0c68e16e6f355ce5075b6dbaccf9d4365f54991

Observation 022228d3-b30f-4870-826d-c08cd4edc939 · outbound

This paper cites A spatio-temporal transformer for 3d human mo- tion prediction.

FIction: 4D Future Interaction Prediction from Video A spatio-temporal transformer for 3d human mo- tion prediction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.401821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.401821Z digest=sha256:17258bf9dd64121900be89f4fe008b975a26a18ef163d81ef49982564655c8fa

Observation b58bab80-8db7-48bc-9ab1-8c9ea6846ff9 · outbound

This paper cites Zero experience required: Plug & play modu- lar transfer learning for semantic visual navigation.

FIction: 4D Future Interaction Prediction from Video Zero experience required: Plug & play modu- lar transfer learning for semantic visual navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.406523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.406523Z digest=sha256:7398a271edfe1c2ccf9d98373f33f0a3017decb0d2f640d2950aeaf3c1caa891

Observation 7e2b7f73-dc70-429b-83f1-2e990df1e54e · outbound

This paper cites On Evaluation of Embodied Navigation Agents.

FIction: 4D Future Interaction Prediction from Video On Evaluation of Embodied Navigation Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.412778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.412778Z digest=sha256:265960f72da8acd20ae3a1256adfefaf9a440f417ce9f8d310b173c8cf4aeae4

Observation dc20e4b7-d67d-4887-9c40-3dd2704841e0 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

FIction: 4D Future Interaction Prediction from Video Hiervl: Learning hierarchical video- language embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.418626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.418626Z digest=sha256:29feba18bc0f03cac9b8c9a9717caffd46bc17d03b5ccea3091a0de35ef213c9

Observation 2542b0b9-ef66-40d1-ae14-91d0690dcb4c · outbound

This paper cites ExpertAF: Expert Actionable Feedback from Video.

FIction: 4D Future Interaction Prediction from Video ExpertAF: Expert Actionable Feedback from Video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.425152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.425152Z digest=sha256:742d37b29ef557f88c9523759c75ab7739413fd4f0cf45b30bb7ad6625b7c8e7

Observation 1a821230-2642-46a5-84fa-8a3633942ca3 · outbound

This paper cites Video-mined task graphs for keystep recognition in instructional videos.

FIction: 4D Future Interaction Prediction from Video Video-mined task graphs for keystep recognition in instructional videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.429721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.429721Z digest=sha256:4441ed79c502a27ab1aa2875617c9daffc70a441f956cd7072610e029d60d331

Observation f90a7c01-627c-42ef-a83f-4f677f48f99b · outbound

This paper cites Affordances from human videos as a versatile representation for robotics.

FIction: 4D Future Interaction Prediction from Video Affordances from human videos as a versatile representation for robotics

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.434263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.434263Z digest=sha256:ae1bc5ec4052db0b2a6b9cf0bfb6b678f5b19cbd1a86adc4b984f3dc7fd57241

Observation 95c4d2c7-7c85-46bb-8864-5b7294c1375c · outbound

This paper cites ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects.

FIction: 4D Future Interaction Prediction from Video ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.438939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.438939Z digest=sha256:d9ccd3cde14183981d1662aecd3bd738840552f35efb3b589f8c1dc97aa1f647

Observation 2297cef5-8674-44dc-b5a2-c97368b3ec24 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

FIction: 4D Future Interaction Prediction from Video Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.443996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.443996Z digest=sha256:6a37ea62e8d5a5725f60eb5b39f4f3eabd98b86c8f80375701aed255cbebfa0c

Observation f8935a8d-f960-4b47-913e-5e629a3e7a4c · outbound

This paper cites Procedure planning in instructional videos via contextual modeling and model- based policy learning.

FIction: 4D Future Interaction Prediction from Video Procedure planning in instructional videos via contextual modeling and model- based policy learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.450889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.450889Z digest=sha256:da6db82c6bcf95f0495b71a75706c509e3e60fdac5cfb1a8ca779dfe32066ac5

Observation f740a5e7-ad07-4b37-9228-1038b8bc3012 · outbound

This paper cites Long-term human mo- tion prediction with scene context.

FIction: 4D Future Interaction Prediction from Video Long-term human mo- tion prediction with scene context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.454637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.454637Z digest=sha256:c669df3e2b68b8b95bdbfc501047183e08a0e31983012a1617ff0445aef81913

Observation b9a306b7-83f1-4529-a98d-6b87fb4af34f · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

FIction: 4D Future Interaction Prediction from Video Quo vadis, action recognition? a new model and the kinetics dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.457842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.457842Z digest=sha256:6ef5de35f03cf422a72ea44cfd07905de656909db7e8854f7cd8ae4474ea82f3

Observation 4c1da243-0ef4-44e7-8edd-fb9aab1adbc5 · outbound

This paper cites Procedure planning in instructional videos.

FIction: 4D Future Interaction Prediction from Video Procedure planning in instructional videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.461437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.461437Z digest=sha256:4e602259df941d026c460d789093965d31fa5794183cf5fa9349b50a69e2a53b

Observation 5eb360d5-b888-47e2-9637-920da724e3ac · outbound

This paper cites Expressive Whole-Body Control for Humanoid Robots.

FIction: 4D Future Interaction Prediction from Video Expressive Whole-Body Control for Humanoid Robots

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.464865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.464865Z digest=sha256:cfb3be5395d03d40e5d54d92cc08f3d55be9f6a639ce6117fe7a3eb33c394abb

Observation 58b08415-f29b-4bf7-9c4e-ef45c24e6c0f · outbound

This paper cites Context-aware human motion prediction.

FIction: 4D Future Interaction Prediction from Video Context-aware human motion prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.468904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.468904Z digest=sha256:d76068881cbd429a58c6ff731ccc127efd8f141484e29316c3c252b089a09837

Observation 065ad2b4-b807-45b1-8c15-3fdb8eddb8c1 · outbound

This paper cites Enrichme: Per- ception and interaction of an assistive robot for the elderly at home.

FIction: 4D Future Interaction Prediction from Video Enrichme: Per- ception and interaction of an assistive robot for the elderly at home

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.473018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.473018Z digest=sha256:7feda8e692071433d43b26e95ea1e5063a780b2822e5ae0afe8123062d0d1a3c

Observation 4ac30229-5b29-4828-b673-33e479576cbd · outbound

This paper cites Rescaling egocentric vision: collection, pipeline and chal- lenges for epic-kitchens-100.

FIction: 4D Future Interaction Prediction from Video Rescaling egocentric vision: collection, pipeline and chal- lenges for epic-kitchens-100

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.477471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.477471Z digest=sha256:3ecb74f7b2133dc6857e7dddae8ab03f54106d3811ccc04e7cc3bb87f0d70970

Observation ceac0663-40ab-4778-b4be-0bc8c98e4ea9 · outbound

This paper cites 3d affordancenet: A benchmark for visual object affordance understanding.

FIction: 4D Future Interaction Prediction from Video 3d affordancenet: A benchmark for visual object affordance understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.481221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.481221Z digest=sha256:c1f07209c5c718bf46c5725b3089d0fa44bee2bd494ff907387265d48d810eb9

Observation 6dcfee24-9873-4266-ae4a-81bc264f9f9d · outbound

This paper cites Cg-hoi: Contact-guided 3d human-object interaction generation.

FIction: 4D Future Interaction Prediction from Video Cg-hoi: Contact-guided 3d human-object interaction generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.485891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.485891Z digest=sha256:6ce982b28546e64ab9af7ebdb8dfb86fabda51289e0d1ebc85a0b53356a24d6d

Observation 265f3001-992f-4a17-ba35-d57b79014bc1 · outbound

This paper cites The Llama 3 Herd of Models.

FIction: 4D Future Interaction Prediction from Video The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.489732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.489732Z digest=sha256:025a41dc399088d07e77fd05242691dd92ad452c92bf37b2c6e441a5f034da93

Observation 96efb0e0-9d9a-46de-89a9-b2aff306908b · outbound

This paper cites Simultaneous local- ization and mapping: part i.

FIction: 4D Future Interaction Prediction from Video Simultaneous local- ization and mapping: part i

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.494237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.494237Z digest=sha256:2321975800ef5fcdd1324cdc9226571182234301c3bb65096a84c7916cc0c959

Observation 32520ba7-dcf5-46d1-b2d5-95079b5c9a9c · outbound

This paper cites Flow graph to video grounding for weakly-supervised multi-step localization.

FIction: 4D Future Interaction Prediction from Video Flow graph to video grounding for weakly-supervised multi-step localization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.497741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.497741Z digest=sha256:523dbe83faa73d8b1c577d5aaf33160617bd5159114a2cccbdf9397c3e4519e1

Observation 2dc1a19b-e9a0-4a68-8f4c-96490736776f · outbound

This paper cites Tokenhmr: Advancing human mesh re- 9 covery with a tokenized pose representation.

FIction: 4D Future Interaction Prediction from Video Tokenhmr: Advancing human mesh re- 9 covery with a tokenized pose representation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.501481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.501481Z digest=sha256:569ac9f41acc91efffb25e88e6d1b6bd4568f6682026f151c5d4f75aee30ad2a

Observation 85643271-4cea-4c54-9799-21e1a05540db · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

FIction: 4D Future Interaction Prediction from Video Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.505259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.505259Z digest=sha256:e276336658281bf67bd68df6de42c1abd95ac5206db96166fb1a67c46b22cd25

Observation 2299be9d-93e7-4fdd-90b7-400eaaf71e2c · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

FIction: 4D Future Interaction Prediction from Video A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.509543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.509543Z digest=sha256:cd4b4a37261df4e4c71a08ee498e0e3b999284c686de518fd1e487bbddc05950

Observation 9650bfbe-83f2-4201-9ec7-c30b3fd4ba0c · outbound

This paper cites Slowfast networks for video recognition.

FIction: 4D Future Interaction Prediction from Video Slowfast networks for video recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.513886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.513886Z digest=sha256:5a4508c1d50315b1cf0cde027208e6974c93f7291b62746cae293098eec3479f

Observation 2ca6f636-f66d-4bf6-ae28-7b0be065eb78 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.517545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.517545Z digest=sha256:cf51d81435d082c120fd67e3a097dcf514cb12c43c83f388db9935a1c46fbbdb

Observation 3f76046c-4380-4c1c-9f2c-70af36c75545 · outbound

This paper cites Rolling- unrolling lstms for action anticipation from first-person video.

FIction: 4D Future Interaction Prediction from Video Rolling- unrolling lstms for action anticipation from first-person video

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.521528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.521528Z digest=sha256:5ed894551782aba6638ba98ffcdb712ec2274832b0fa6719ee0ac12cca92ab9c

Observation a70b580e-32ca-4e3e-beba-618376031e57 · outbound

This paper cites Next-active-object predic- tion from egocentric videos.

FIction: 4D Future Interaction Prediction from Video Next-active-object predic- tion from egocentric videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.526581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.526581Z digest=sha256:7d0acb138b7da7b3c10611d134c3243c71d19722700d4d1cc17e87efe56a127f

Observation 01abb968-9e33-48bd-b23d-55d9f11abb22 · outbound

This paper cites RED: Reinforced Encoder-Decoder Networks for Action Anticipation.

FIction: 4D Future Interaction Prediction from Video RED: Reinforced Encoder-Decoder Networks for Action Anticipation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.535082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.535082Z digest=sha256:3c5b7cbf13cec4f0540c289ae702a0e9312a41e38754b7b7b1f9cdf59abba1aa

Observation a6063b88-8054-4b1a-8174-13b6c88ae93a · outbound

This paper cites Anticipative video transformer.

FIction: 4D Future Interaction Prediction from Video Anticipative video transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.538490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.538490Z digest=sha256:d364836ac275796f290cc8fa751da61026d0414c3fa8806b021cc714980b41c0

Observation 004f3e4f-8e7e-4127-88d3-bf299f685a08 · outbound

This paper cites Omni- vore: A single model for many visual modalities.

FIction: 4D Future Interaction Prediction from Video Omni- vore: A single model for many visual modalities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.542273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.542273Z digest=sha256:1ee42a334162049ba8885cce60b0d1d437fffa462caace94ebb9384512cff261

Observation bf196040-dcb1-4cfb-af48-5c104814bb4f · outbound

This paper cites Humans in 4d: Re- constructing and tracking humans with transformers.

FIction: 4D Future Interaction Prediction from Video Humans in 4d: Re- constructing and tracking humans with transformers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.545672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.545672Z digest=sha256:dffb6e181831e247b7bcfb4f54deab46f18dcfe8267d540aee20446aa0c999f9

Observation 47955e9a-7cf5-4cc3-8590-d870b02ffe80 · outbound

This paper cites Con- tactopt: Optimizing contact to improve grasps.

FIction: 4D Future Interaction Prediction from Video Con- tactopt: Optimizing contact to improve grasps

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.549484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.549484Z digest=sha256:afae7c751268e511204c612a5208a3acf85cbc3c4310d7e2ad3d1450b404bb79

Observation cbe0b724-33af-4318-b876-7ed2bbe59832 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

FIction: 4D Future Interaction Prediction from Video Ego4d: Around the world in 3,000 hours of egocentric video

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.553880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.553880Z digest=sha256:3c74f4602c1e83fd067bef10fdbdfb8d945ef81339efa4aad82f12e5d6607cfb

Observation 818be1cb-7ff7-494a-9c1b-967c34453325 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

FIction: 4D Future Interaction Prediction from Video Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.558067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.558067Z digest=sha256:65651eab703002a982446924641ac14edcfc08fc1878e131c6b02f3a91284b11

Observation 649f042e-568f-453c-9609-b41fc3830ebc · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

FIction: 4D Future Interaction Prediction from Video Lvis: A dataset for large vocabulary instance segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.562199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.562199Z digest=sha256:b80eae8915d7fe7ee8eaa4dbe205fbfed56f2d5efbb3b62a03f3c820dacb6662

Observation 896641e9-732a-4cae-a3d5-7b1c5182b8a1 · outbound

This paper cites Resolving 3d human pose ambiguities with 3d scene constraints.

FIction: 4D Future Interaction Prediction from Video Resolving 3d human pose ambiguities with 3d scene constraints

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.566700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.566700Z digest=sha256:d892b71044dc73b9caccdb2ed3cce4b08ab35b154d4d4b8b1e6656053c32f9c9

Observation ed044fe1-a82f-448f-b79b-0247bdcd712b · outbound

This paper cites Stochas- tic scene-aware motion prediction.

FIction: 4D Future Interaction Prediction from Video Stochas- tic scene-aware motion prediction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.571042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.571042Z digest=sha256:05d32e9f01955472ab2cc601630ed5d4917e7308352101abb86f22a377fdc2a5

Observation 14dd2e39-6a5c-4c28-a458-dc4e16fee3a4 · outbound

This paper cites Synthesizing phys- ical character-scene interactions.

FIction: 4D Future Interaction Prediction from Video Synthesizing phys- ical character-scene interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.575363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.575363Z digest=sha256:8641cadfa5ce9e52896a7786a00bdbc3094bfb080ea50c4f7bc9eb64864de32d

Observation ccdee9e4-0959-4787-b098-6152f839eea3 · outbound

This paper cites Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation.

FIction: 4D Future Interaction Prediction from Video Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.579824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.579824Z digest=sha256:cd5c73d0aee3bf1536740e4f7324f134558f51fd1a95d408f9efa1dfa6dc4e3d

Observation 90595afb-d08a-4055-9a06-4a6c2fe5c670 · outbound

This paper cites Diffusion- based generation, optimization, and planning in 3d scenes.

FIction: 4D Future Interaction Prediction from Video Diffusion- based generation, optimization, and planning in 3d scenes

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.584042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.584042Z digest=sha256:e6486e47f9afcfab3c465ddd223b916f23f2e536a76dd08804d30b1a61b43b2b

Observation 7ce5930b-6c19-4f41-ad7e-6a8ee7bda2da · outbound

This paper cites Human3.6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments.

FIction: 4D Future Interaction Prediction from Video Human3.6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.587680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.587680Z digest=sha256:4252d94dae93d55e2ba009cbc680edab642731f2188db0b007745253069879ad

Observation 9f203ce7-0efa-4553-b54a-9303e78e8174 · outbound

This paper cites Hand-object contact consistency reasoning for hu- man grasps generation.

FIction: 4D Future Interaction Prediction from Video Hand-object contact consistency reasoning for hu- man grasps generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.591402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.591402Z digest=sha256:93309405ae8bd12ebc8e2c9f8f042fc47e37b4f98b139c8c2242bef46c37abc5

Observation f0a1d69a-e07a-47c5-8728-ec5dedaf5cb7 · outbound

This paper cites Sym- phonize 3d semantic scene completion with contextual in- stance queries, 2023.

FIction: 4D Future Interaction Prediction from Video Sym- phonize 3d semantic scene completion with contextual in- stance queries, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.595135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.595135Z digest=sha256:b5abf0e0ce732dd3807a5c8e2f073907d03dc415771b46e07d5b4e49b87005a5

Observation 07767f29-b574-4145-aaba-ea6861a5e29d · outbound

This paper cites R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection.

FIction: 4D Future Interaction Prediction from Video R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.181188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.598689Z digest=sha256:869e8c4760f2e1cfc79d3b06872230c08e3c9adb428685e067f2fec4b2c68c73

Observation e35a7de4-ea32-4cb1-9a81-21fade0c5aa1 · outbound

This paper cites Segment anything.

FIction: 4D Future Interaction Prediction from Video Segment anything

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.602804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.602804Z digest=sha256:161a095f8912afc083ffbf6036e082394315bb9bb70793a66e1b0535ca90c870

Observation b05ffc6d-7d7f-4961-9cd1-1a35dc569ff4 · outbound

This paper cites Interactive object segmentation in 3d point clouds.

FIction: 4D Future Interaction Prediction from Video Interactive object segmentation in 3d point clouds

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.969264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.606482Z digest=sha256:bdea961568c5f6ee577e7eea12458cfe59722e65c17c71fda917567ed6c1988f

Observation 63eef815-7046-4c14-94d7-2ce13dd3dd57 · outbound

This paper cites A hi- erarchical representation for future action prediction.

FIction: 4D Future Interaction Prediction from Video A hi- erarchical representation for future action prediction

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.956226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.609535Z digest=sha256:8895359b796e0077f00ef785670f877b908459ac2280b8d6a180bb1c0dfbed36

Observation 2d32b3ba-df97-467e-927c-13a462eb3b29 · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

FIction: 4D Future Interaction Prediction from Video LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.612463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.612463Z digest=sha256:b467dcecb25784402344d68538be27748c099ac5a51c99ffc42b92c20007d405

Observation 1e9820cb-45c1-4420-a34d-f00c8bad3a4b · outbound

This paper cites UniFormer: Unifying Convolution and Self-attention for Visual Recognition.

FIction: 4D Future Interaction Prediction from Video UniFormer: Unifying Convolution and Self-attention for Visual Recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.615541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.615541Z digest=sha256:0b6cd5223d27a31c599f09a87a0f10f09e52e5eee151b51f3cb1ecd17f120a86

Observation c0cb4845-502b-4372-a53e-19f406dbed80 · outbound

This paper cites Mvitv2: Improved multiscale vision transform- ers for classification and detection.

FIction: 4D Future Interaction Prediction from Video Mvitv2: Improved multiscale vision transform- ers for classification and detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.943815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.618770Z digest=sha256:854f59e5a537843a73b5771d6bcef42c7ce284dca5ee58e37985bbc2ed0bf03c

Observation a0c3ed5b-aad4-414a-99b3-572528567b07 · outbound

This paper cites Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar.

FIction: 4D Future Interaction Prediction from Video Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.930511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.622214Z digest=sha256:7a6cf718840e804097c152f7eab801b57ec994ce620118ed80deed83019381b5

Observation 0fb9c6e7-6b18-4534-be8f-4499916cc2e4 · outbound

This paper cites Egocen- tric video-language pretraining.

FIction: 4D Future Interaction Prediction from Video Egocen- tric video-language pretraining

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.915968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.625968Z digest=sha256:18859119555ce60afb5761f8242b1f205e41c903a302ab40699496e01d68dd3e

Observation 3c376914-a25f-4b3a-835a-282c30660cd0 · outbound

This paper cites Microsoft coco: Common objects in context.

FIction: 4D Future Interaction Prediction from Video Microsoft coco: Common objects in context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.629458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.629458Z digest=sha256:ee48039719ce752f243265bea90103a92d3c97903738c5c37468246f65eae30d

Observation d1798fd1-2fc4-40ee-bb42-e11d60d02a8e · outbound

This paper cites Learn- ing to recognize procedural activities with distant supervi- sion.

FIction: 4D Future Interaction Prediction from Video Learn- ing to recognize procedural activities with distant supervi- sion

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.893705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.633145Z digest=sha256:0de524031e5c533afb0eb469ff540c57ea199b59a5ce59baf7fb0e59f20f55ad

Observation 04161796-2915-456a-aabf-ab8359d75e6d · outbound

This paper cites Deep patch visual slam.

FIction: 4D Future Interaction Prediction from Video Deep patch visual slam

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.881706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.637095Z digest=sha256:45915b9315b303155050664821bb8f4ebec421e7d55f50d6eaea80aa4da375c6

Observation 87fed159-0559-4ad8-8d4b-9908c2664220 · outbound

This paper cites Visual Instruction Tuning.

FIction: 4D Future Interaction Prediction from Video Visual Instruction Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.641299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.641299Z digest=sha256:6d7fb2eb6709d4f87eaad58bfe23a576d95b05a7c11443364304c915291484f0

Observation cf3b1fc8-af49-4d67-9408-ca464e6ada64 · outbound

This paper cites Fore- casting human-object interaction: joint prediction of motor attention and actions in first person video.

FIction: 4D Future Interaction Prediction from Video Fore- casting human-object interaction: joint prediction of motor attention and actions in first person video

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.870748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.645564Z digest=sha256:96e9157ca017e4a39819cc2470def1f70c58d98d9def22fa4b0f4dd8c2ccff32

Observation 8c1cdd44-0956-481e-82dd-e18914533a92 · outbound

This paper cites Joint hand motion and interaction hotspots prediction from egocentric videos.

FIction: 4D Future Interaction Prediction from Video Joint hand motion and interaction hotspots prediction from egocentric videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.859489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.649079Z digest=sha256:4a3ce3f55cd6a842070f1268a36dedcc747c4436c6126919ea67bb977d8ddaea

Observation bf640659-1a84-4232-b088-3e464f2819e8 · outbound

This paper cites Joint hand motion and interaction hotspots prediction from egocentric videos.

FIction: 4D Future Interaction Prediction from Video Joint hand motion and interaction hotspots prediction from egocentric videos

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.847867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.652471Z digest=sha256:40091b2d39a8772e90ff2a80a0a41877182701218488934237b6763723721c09

Observation 2860e921-7892-4b06-9a23-8b62fe354f34 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:56:00.837304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.655916Z digest=sha256:29a63f7b2f5ac826f8dd32ac48fa4c6b85e99f09cd74d14016b798b9c625d36d

Observation 9cbe04d5-4a68-4396-b4fd-0ad035ca3fb5 · outbound

This paper cites Multimodal sense-informed forecasting of 3d human motions.

FIction: 4D Future Interaction Prediction from Video Multimodal sense-informed forecasting of 3d human motions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.826379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.659781Z digest=sha256:15caaf31a3f7f79316f87bbee391dcbb5d4b6677b36324531cf55481b444625e

Observation 5b87f850-56d8-4358-958d-e7f36ea74cf0 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

FIction: 4D Future Interaction Prediction from Video Amass: Archive of motion capture as surface shapes

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.814816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.663603Z digest=sha256:f76fcfe1732a273de56c1660c71f0e00202763fb49ff942ef41cf9cf70a3edee

Observation fee3d281-a9ba-4e46-96fd-245e975f7d4c · outbound

This paper cites Contact-aware human motion forecasting.

FIction: 4D Future Interaction Prediction from Video Contact-aware human motion forecasting

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.803216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.667609Z digest=sha256:c9ec1dad1be7b518275d8445d14225d5cbe0795bbe64e97d08ad5ab842515221

Observation f8dfe2d4-90ed-4baf-b6a4-7b73c28a4120 · outbound

This paper cites On human motion prediction using recurrent neural networks.

FIction: 4D Future Interaction Prediction from Video On human motion prediction using recurrent neural networks

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.791592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.671250Z digest=sha256:b8955c3e37e489057fd85e2cc8e6e09b9b0ba094b96e61e4db222e7b01c5b9c5

Observation 00f2291d-45aa-49fe-9888-a8acafb95bfb · outbound

This paper cites Intention-Conditioned Long-Term Human Egocentric Action Forecasting.

FIction: 4D Future Interaction Prediction from Video Intention-Conditioned Long-Term Human Egocentric Action Forecasting

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.133452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.675387Z digest=sha256:2ca463f63d68aff39a52fce586edaa09de764f74fe316e8150b62086c86eb486

Observation 4986b394-7980-4259-8baa-bb8b72142364 · outbound

This paper cites Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips.

FIction: 4D Future Interaction Prediction from Video Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.778937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.679252Z digest=sha256:8f6e819659894a91aa8198aa82aa498e8fb7f4ad2c161ecb6ecbd5c43e46e577

Observation 75393f18-c82b-49eb-aa56-cd909c9a897f · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

FIction: 4D Future Interaction Prediction from Video End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.767520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.682922Z digest=sha256:0f700e868be39d41d82c9183f665276303b2da6886c9f2ed2be9791a7a521115

Observation 877213ac-c196-423a-8cb6-5e895cf645af · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- 11 thesis.

FIction: 4D Future Interaction Prediction from Video Nerf: Representing scenes as neural radiance fields for view syn- 11 thesis

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.753669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.686550Z digest=sha256:f0da658539a05fb5d0031ba7b7f35bb7549353266289b4e7c32dcbc9ec750614

Observation 9ec69155-3b27-4297-94da-1192e797f937 · outbound

This paper cites Graspit! a versatile simulator for robotic grasping.

FIction: 4D Future Interaction Prediction from Video Graspit! a versatile simulator for robotic grasping

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.741183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.690285Z digest=sha256:7d5ba4baf33f9c4d221266fa4d023f7e8287c56a1eb834d5a26a3d8a33576e75

Observation 44846b5c-4c04-4323-b8fe-dcdc281ff0bb · outbound

This paper cites Where2act: From pixels to actions for articulated 3d objects.

FIction: 4D Future Interaction Prediction from Video Where2act: From pixels to actions for articulated 3d objects

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.729177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.694406Z digest=sha256:e701c19b22c01b4e8fe14a21cd1f34c2dd0ffe4c4d374357f534301a68b20d91

Observation 0b75ac41-22cb-45bc-9a2f-9a0068bcdd1a · outbound

This paper cites Orb-slam: a versatile and accurate monocular slam system.

FIction: 4D Future Interaction Prediction from Video Orb-slam: a versatile and accurate monocular slam system

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.697773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.697773Z digest=sha256:4c274df13636849b45b3d5a78430a1321439e910d614bc611d53e96b68e97708

Observation c7b528b3-e084-4f00-a9ea-934d83e910ef · outbound

This paper cites Multi-label affordance mapping from egocentric vision.

FIction: 4D Future Interaction Prediction from Video Multi-label affordance mapping from egocentric vision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.710600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.704470Z digest=sha256:f58e2778e4b2be01b6dd401e29cb26a579097fd21635812b29974030e37469d9

Observation e41b3955-af65-441b-aa9e-703b140f0ec7 · outbound

This paper cites AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation.

FIction: 4D Future Interaction Prediction from Video AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.707512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.707512Z digest=sha256:de6f0529d43a0de0beb7018e1d107021a9cee851fe0b7d2a4ef03fb178ee6f5d

Observation 0d4db6a3-d95d-45a6-bebc-b9feb5891460 · outbound

This paper cites Nagarajan and K.

FIction: 4D Future Interaction Prediction from Video Nagarajan and K

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.697433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.710676Z digest=sha256:9f0e37d0ec5659dbfd46772c8d29200ec4df5c55cc9926571f3b4eb88c399f53

Observation 7be45799-0727-413b-8734-fb27ef1dad26 · outbound

This paper cites Grounded human-object interaction hotspots from video.

FIction: 4D Future Interaction Prediction from Video Grounded human-object interaction hotspots from video

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.685731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.713776Z digest=sha256:15837fdd78c1e589f6a275d9e313260c92895b2430744b56976239d56f75897b

Observation 26792741-cff7-4591-b15b-46d282ad7b76 · outbound

This paper cites Future event prediction: If and when.

FIction: 4D Future Interaction Prediction from Video Future event prediction: If and when

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.674007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.717287Z digest=sha256:be8e7f55b684ba99947dda088338041ef90c0e1bccd6273dd889defc2c1e6b3f

Observation 2472ca63-6afd-48f3-bcc4-77ef094f42c3 · outbound

This paper cites Using Geometry to Detect Grasps in 3D Point Clouds.

FIction: 4D Future Interaction Prediction from Video Using Geometry to Detect Grasps in 3D Point Clouds

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.107862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.837998Z digest=sha256:3a62f9cc9658c9048e02921ab133c47ce0dfb1fbcf968470c18dc2928048120b

Observation 1306d568-892c-4691-b533-161300a63894 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.842388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.842388Z digest=sha256:3e8a3206dcbe36a4271b19e7937e0f0885d9083743143592dd24287e1a5a3bd1

Observation 7414532d-4fee-4cef-96e3-b58cb3017e6d · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

FIction: 4D Future Interaction Prediction from Video Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.655407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.846244Z digest=sha256:9f354ba62aac8dbca2fea533f3ced3897b6242ec7e84587b60f976b41a07273e

Observation 4e6dd8f6-41fa-4a1d-add9-743bed487856 · outbound

This paper cites Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots.

FIction: 4D Future Interaction Prediction from Video Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.850055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.850055Z digest=sha256:5f4e4878392b618dc019afdf2e759149d601a73528f92083168c1cc4eceb967e

Observation ad74e877-1478-4938-bc7f-c3f14648910b · outbound

This paper cites Vins-mono: A ro- bust and versatile monocular visual-inertial state estimator.

FIction: 4D Future Interaction Prediction from Video Vins-mono: A ro- bust and versatile monocular visual-inertial state estimator

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.644501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.853864Z digest=sha256:2a033c58bdbfd1fdb9d8a760c5f8fbfb6704bc8ba733981d68d37e90cb3d9464

Observation a9b6bcdf-c44d-42a6-ba3e-1230e330fa99 · outbound

This paper cites State-only imitation learning for dexterous manipulation.

FIction: 4D Future Interaction Prediction from Video State-only imitation learning for dexterous manipulation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.632033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.856756Z digest=sha256:bfd0f25ac8cee2e456888a57eb67f1dc271422366f049285be20b1e82ba70e56

Observation 4e0c5a99-7e26-42ad-8eb6-e26325e83235 · outbound

This paper cites Poni: Potential functions for objectgoal navigation with interaction-free learning.

FIction: 4D Future Interaction Prediction from Video Poni: Potential functions for objectgoal navigation with interaction-free learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.620790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.860281Z digest=sha256:66c1ab82af1061d024c46025766ff418ceacfcd17fe01b4a6055252dfe98847b

Observation 9c418893-8def-4287-aae8-0409267c61d5 · outbound

This paper cites Humor: 3d human motion model for robust pose estimation.

FIction: 4D Future Interaction Prediction from Video Humor: 3d human motion model for robust pose estimation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.609620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.863253Z digest=sha256:66f8b0e457f764dfd351c23c8bf430f04e3aa277b9d0cd03df1bf8cf8db1181d

Observation ffafcde5-3f30-4535-9a77-390b2dcf1093 · outbound

This paper cites First-person activity forecasting with online inverse reinforcement learning.

FIction: 4D Future Interaction Prediction from Video First-person activity forecasting with online inverse reinforcement learning

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.598098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.866660Z digest=sha256:6eb5bfdb924c422a8e2a5a888c35302974b61c45012f3288722b93a7f17f96c0

Observation 34224ca4-994f-4a09-88e9-2cc83b93bf22 · outbound

This paper cites Structure-from-motion revisited.

FIction: 4D Future Interaction Prediction from Video Structure-from-motion revisited

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.586404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.869991Z digest=sha256:61ab6693b6d0cbb3b8725f89e7b0583ae8b0f6e81910452790e8adc0728953d4

Observation 32b5e6bd-0ae5-450e-8e3d-e2d922e48ff1 · outbound

This paper cites Wham: Reconstructing world-grounded humans with accurate 3d motion.

FIction: 4D Future Interaction Prediction from Video Wham: Reconstructing world-grounded humans with accurate 3d motion

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.575160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.874484Z digest=sha256:d5720e9e4f8a9951aaf841056516ffd4b3d75e10a40402497638ad6f2bcd4617

Observation 9f894892-d0c4-42d8-b1fd-2c988dcaadc4 · outbound

This paper cites Learn- ing structured output representation using deep conditional generative models.

FIction: 4D Future Interaction Prediction from Video Learn- ing structured output representation using deep conditional generative models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.563868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.878602Z digest=sha256:603b6e8cec15ccfa00e5a37a9ff14b977317ed96c1382af757a2e7e63a5647b0

Observation 38969f8c-5b29-4748-a4f8-5b10810fdf06 · outbound

This paper cites Generating notifications for missing actions: Don’t forget to turn the lights off! In Proceedings of the IEEE International Con- ference on Computer Vision, pages 4669–4677, 2015.

FIction: 4D Future Interaction Prediction from Video Generating notifications for missing actions: Don’t forget to turn the lights off! In Proceedings of the IEEE International Con- ference on Computer Vision, pages 4669–4677, 2015

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.552001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.882948Z digest=sha256:52ac872d3f6b332fcdbe580fca08bf9386608d6f6e4e44eb339cd81e4ceb199d

Observation 38789a9a-c1e2-41a3-ad9e-ddbac541ddd3 · outbound

This paper cites Segcloud: Semantic segmen- tation of 3d point clouds.

FIction: 4D Future Interaction Prediction from Video Segcloud: Semantic segmen- tation of 3d point clouds

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.537792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.886883Z digest=sha256:c698579ca623ee1502b1753b6655dcaf66104dd246f97113b89c149a1876c843

Observation 22414549-19af-4e40-96c5-263c29df85cb · outbound

This paper cites Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras.

FIction: 4D Future Interaction Prediction from Video Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.891500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.891500Z digest=sha256:684cf5a73dacaea0558fee533553fd802076c3de4842b74f5975c014352a4d3d

Observation 5e2e33b1-5232-4ec4-8eb3-dbeb01c899f7 · outbound

This paper cites EPIC Fields: Marrying 3D Geometry and Video Understanding.

FIction: 4D Future Interaction Prediction from Video EPIC Fields: Marrying 3D Geometry and Video Understanding

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.518624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.895419Z digest=sha256:1a57557a0f1e4145419f5addc32ad9efc1487e1ce0a26dfeeb8d0da68064ebda

Observation 5667d731-4af4-4431-9556-0195e76ddcf1 · outbound

This paper cites Transformation-Based Models of Video Sequences.

FIction: 4D Future Interaction Prediction from Video Transformation-Based Models of Video Sequences

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.080431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.899330Z digest=sha256:72a791ab282b471b962481867545172050e389777df38ebf91bdeb05d665f343

Observation 2ee9e01c-18a7-4644-8c0d-a8e20bfbfad9 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:56:00.506311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.903006Z digest=sha256:f92deed11d33fd9923d34fab3fc85f09aee380307a0126858bc30343f6b3c82d

Observation cedbaa9f-ae03-4c41-802e-89cd7425b035 · outbound

This paper cites Synthesizing long-term 3d human mo- tion and interaction in 3d scenes.

FIction: 4D Future Interaction Prediction from Video Synthesizing long-term 3d human mo- tion and interaction in 3d scenes

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.493415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.907142Z digest=sha256:8b9b40c9bbf24a2248c406e0ac6a041b6c408f3978fe758df4cb8b7838685290

Observation 3d6d5e86-d9da-4c0d-8d7a-3552690f5732 · outbound

This paper cites Adaafford: Learn- ing to adapt manipulation affordance for 3d articulated ob- jects via few-shot interactions.

FIction: 4D Future Interaction Prediction from Video Adaafford: Learn- ing to adapt manipulation affordance for 3d articulated ob- jects via few-shot interactions

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.480249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.910653Z digest=sha256:34fb4b20390c905cb0b1f12bb3e12def8965f9802db0f634fdea9a42490acb38

Observation 9747d2c8-d121-46a7-865e-8ccc3c9072c6 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

FIction: 4D Future Interaction Prediction from Video Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.467780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:55:59.914437Z digest=sha256:4469db9b6181ca6b38d18cbe4cde09f2c042ccf544ea317a48cf76d88581948c

Pith citing papers

No inbound Pith citation observations are available.