Pith. sign in

Paper Citation Record · LEDGER

FIction: 4D Future Interaction Prediction from Video

As of 14 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 0 inbound Pith citation observations for arXiv:2412.00932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00932 v2

Coverage vector

measured 100 of 119 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:55:59.914437Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 119 outbound references displayed

  • verified exact4
  • verified fuzzy38
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c95a774e-e9a0-4766-a269-52b754da3f11 · outbound

This paper cites When will you do what?-anticipating temporal occurrences of activities.

FIction: 4D Future Interaction Prediction from Video When will you do what?-anticipating temporal occurrences of activities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.397059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.397059Z digest=sha256:ceb66198f683f32d1297d8340a26c7728d857aac1b0e447109d4494e96ccb1f0

Observation 022228d3-b30f-4870-826d-c08cd4edc939 · outbound

This paper cites A spatio-temporal transformer for 3d human mo- tion prediction.

FIction: 4D Future Interaction Prediction from Video A spatio-temporal transformer for 3d human mo- tion prediction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.401821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.401821Z digest=sha256:26ded947aa6579be7d42d75d7614f7f95583f508ed743410098d1973740a6710

Observation b58bab80-8db7-48bc-9ab1-8c9ea6846ff9 · outbound

This paper cites Zero experience required: Plug & play modu- lar transfer learning for semantic visual navigation.

FIction: 4D Future Interaction Prediction from Video Zero experience required: Plug & play modu- lar transfer learning for semantic visual navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.406523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.406523Z digest=sha256:a0b10b11a5c87a6a97f6fe9fae00a7ddef815bf8aa8bd95ae6a783b89253da0b

Observation 7e2b7f73-dc70-429b-83f1-2e990df1e54e · outbound

This paper cites On Evaluation of Embodied Navigation Agents.

FIction: 4D Future Interaction Prediction from Video On Evaluation of Embodied Navigation Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.412778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.412778Z digest=sha256:680511925fdcc5c97700c10da634c30038446568972d9bd3149e181d3e5cb350

Observation dc20e4b7-d67d-4887-9c40-3dd2704841e0 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

FIction: 4D Future Interaction Prediction from Video Hiervl: Learning hierarchical video- language embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.418626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.418626Z digest=sha256:a61ebe1e606570f8f71f61be79a1b99a71217f7ce3a81ae5085fe9cb3f9e19e5

Observation 2542b0b9-ef66-40d1-ae14-91d0690dcb4c · outbound

This paper cites ExpertAF: Expert Actionable Feedback from Video.

FIction: 4D Future Interaction Prediction from Video ExpertAF: Expert Actionable Feedback from Video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.425152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.425152Z digest=sha256:2b82cc88d90f1b41ef627f906b4c4b73a2522ddfbb45ccb266bf4b86346cced8

Observation 1a821230-2642-46a5-84fa-8a3633942ca3 · outbound

This paper cites Video-mined task graphs for keystep recognition in instructional videos.

FIction: 4D Future Interaction Prediction from Video Video-mined task graphs for keystep recognition in instructional videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.429721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.429721Z digest=sha256:7b5288da42b17c7113cf9a6be96b43f76f050ee55d3f3c2a5aea48bc8fbfaf83

Observation f90a7c01-627c-42ef-a83f-4f677f48f99b · outbound

This paper cites Affordances from human videos as a versatile representation for robotics.

FIction: 4D Future Interaction Prediction from Video Affordances from human videos as a versatile representation for robotics

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.434263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.434263Z digest=sha256:f3c1da9e7469ce3ef013050955de0402c2862d4696d6770e0fd6db3f6f9c0218

Observation 95c4d2c7-7c85-46bb-8864-5b7294c1375c · outbound

This paper cites ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects.

FIction: 4D Future Interaction Prediction from Video ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.438939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.438939Z digest=sha256:6cf33cae90947eee1af4383ffbcb240f8712ff54265b1c824386c844573b6875

Observation 2297cef5-8674-44dc-b5a2-c97368b3ec24 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

FIction: 4D Future Interaction Prediction from Video Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.443996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.443996Z digest=sha256:4648b0ddbcc809d853b8ee9c8444707c56fd9bb92588be6d042bd1efbec26007

Observation f8935a8d-f960-4b47-913e-5e629a3e7a4c · outbound

This paper cites Procedure planning in instructional videos via contextual modeling and model- based policy learning.

FIction: 4D Future Interaction Prediction from Video Procedure planning in instructional videos via contextual modeling and model- based policy learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.450889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.450889Z digest=sha256:aabc60f0a5ca5f0efa2e11e79ec17c84cf9d02f7e27cdae90a9391899cb8b59a

Observation f740a5e7-ad07-4b37-9228-1038b8bc3012 · outbound

This paper cites Long-term human mo- tion prediction with scene context.

FIction: 4D Future Interaction Prediction from Video Long-term human mo- tion prediction with scene context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.454637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.454637Z digest=sha256:01746b8ae469164f1601d7e83457015bcd2f0a0d196732a381af386a758365f5

Observation b9a306b7-83f1-4529-a98d-6b87fb4af34f · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

FIction: 4D Future Interaction Prediction from Video Quo vadis, action recognition? a new model and the kinetics dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.457842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.457842Z digest=sha256:69324cf598d095dc557fc440212946afd4c87e564a7d076710a42f4dde1e70a0

Observation 4c1da243-0ef4-44e7-8edd-fb9aab1adbc5 · outbound

This paper cites Procedure planning in instructional videos.

FIction: 4D Future Interaction Prediction from Video Procedure planning in instructional videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.461437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.461437Z digest=sha256:01f7a7b36afb244d609ee380083b26a1b785deee68ad3246a703b8d336eb4a61

Observation 5eb360d5-b888-47e2-9637-920da724e3ac · outbound

This paper cites Expressive Whole-Body Control for Humanoid Robots.

FIction: 4D Future Interaction Prediction from Video Expressive Whole-Body Control for Humanoid Robots

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.464865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.464865Z digest=sha256:4164b48c5a1ca87d058d8ee413011f95a75b6c50a64d3ea2926bfd8ccc0f98f5

Observation 58b08415-f29b-4bf7-9c4e-ef45c24e6c0f · outbound

This paper cites Context-aware human motion prediction.

FIction: 4D Future Interaction Prediction from Video Context-aware human motion prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.468904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.468904Z digest=sha256:5f7da151fb7a26e191be0296cfed3dde5a94374724d6fdcaebcdb7eff665319e

Observation 065ad2b4-b807-45b1-8c15-3fdb8eddb8c1 · outbound

This paper cites Enrichme: Per- ception and interaction of an assistive robot for the elderly at home.

FIction: 4D Future Interaction Prediction from Video Enrichme: Per- ception and interaction of an assistive robot for the elderly at home

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.473018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.473018Z digest=sha256:3f3b23e360173d1a88b42fa7262e6045abe515b87be1cc21969b0268d3d87ba4

Observation 4ac30229-5b29-4828-b673-33e479576cbd · outbound

This paper cites Rescaling egocentric vision: collection, pipeline and chal- lenges for epic-kitchens-100.

FIction: 4D Future Interaction Prediction from Video Rescaling egocentric vision: collection, pipeline and chal- lenges for epic-kitchens-100

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.477471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.477471Z digest=sha256:391e01d9bf7da318b1e08b9dc95837bc584d604579db056bb559542bebe9ea73

Observation ceac0663-40ab-4778-b4be-0bc8c98e4ea9 · outbound

This paper cites 3d affordancenet: A benchmark for visual object affordance understanding.

FIction: 4D Future Interaction Prediction from Video 3d affordancenet: A benchmark for visual object affordance understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.481221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.481221Z digest=sha256:59017603dd3d3625a9711d7dd46b640e91d63a9325f76a5bcd14373f8efd3792

Observation 6dcfee24-9873-4266-ae4a-81bc264f9f9d · outbound

This paper cites Cg-hoi: Contact-guided 3d human-object interaction generation.

FIction: 4D Future Interaction Prediction from Video Cg-hoi: Contact-guided 3d human-object interaction generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.485891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.485891Z digest=sha256:ea5df0a4c24d352d9c12be6f4c3560763a68bd102a0ce83bbaccfe7be09d1d67

Observation 265f3001-992f-4a17-ba35-d57b79014bc1 · outbound

This paper cites The Llama 3 Herd of Models.

FIction: 4D Future Interaction Prediction from Video The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.489732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.489732Z digest=sha256:db9e89c772d1428eebe7705b158d326457f27c30db66594da6ef79f79e1c05b0

Observation 96efb0e0-9d9a-46de-89a9-b2aff306908b · outbound

This paper cites Simultaneous local- ization and mapping: part i.

FIction: 4D Future Interaction Prediction from Video Simultaneous local- ization and mapping: part i

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.494237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.494237Z digest=sha256:0b9e5aca45be1ac804534ea6088792b09c36682e43b4e1b4b2ed6ac869d19f17

Observation 32520ba7-dcf5-46d1-b2d5-95079b5c9a9c · outbound

This paper cites Flow graph to video grounding for weakly-supervised multi-step localization.

FIction: 4D Future Interaction Prediction from Video Flow graph to video grounding for weakly-supervised multi-step localization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.497741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.497741Z digest=sha256:7c0a98a8edad2ff199377f9cb42df87a84f532e6e0bb37f4f8d6716c1dcef19d

Observation 2dc1a19b-e9a0-4a68-8f4c-96490736776f · outbound

This paper cites Tokenhmr: Advancing human mesh re- 9 covery with a tokenized pose representation.

FIction: 4D Future Interaction Prediction from Video Tokenhmr: Advancing human mesh re- 9 covery with a tokenized pose representation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.501481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.501481Z digest=sha256:f9b353492b4511df155d993690ca041dd39e52bf1c3061c71a4365ad2e98df7e

Observation 85643271-4cea-4c54-9799-21e1a05540db · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

FIction: 4D Future Interaction Prediction from Video Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.505259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.505259Z digest=sha256:cb5931b2c17d5f3cc5796562cc250c7c5d7e01559f3c5aef16c6a3050d6d532b

Observation 2299be9d-93e7-4fdd-90b7-400eaaf71e2c · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

FIction: 4D Future Interaction Prediction from Video A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.509543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.509543Z digest=sha256:80c92d635d15bae5a8d516965837e7d07d8433b5d6da7d0c2a37b94896676d82

Observation 9650bfbe-83f2-4201-9ec7-c30b3fd4ba0c · outbound

This paper cites Slowfast networks for video recognition.

FIction: 4D Future Interaction Prediction from Video Slowfast networks for video recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.513886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.513886Z digest=sha256:9904fa2c2e780a4c848a5a0e9373c5c84f0a76bb411e44e0bbca9fdea7aceb6e

Observation 2ca6f636-f66d-4bf6-ae28-7b0be065eb78 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.517545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.517545Z digest=sha256:d697b734bd28a97a738aa788268cdbdb1b533d0fae876038513c420b692628e0

Observation 3f76046c-4380-4c1c-9f2c-70af36c75545 · outbound

This paper cites Rolling- unrolling lstms for action anticipation from first-person video.

FIction: 4D Future Interaction Prediction from Video Rolling- unrolling lstms for action anticipation from first-person video

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.521528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.521528Z digest=sha256:2086ba6f5eb1929d10f3aed5ba28517e544c3d96ff3be95265bd99c049da8204

Observation a70b580e-32ca-4e3e-beba-618376031e57 · outbound

This paper cites Next-active-object predic- tion from egocentric videos.

FIction: 4D Future Interaction Prediction from Video Next-active-object predic- tion from egocentric videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.526581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.526581Z digest=sha256:ecd5dae32762de84c33674d2ef1e7536ddd28104b6995f5dec459de32b30a310

Observation 01abb968-9e33-48bd-b23d-55d9f11abb22 · outbound

This paper cites RED: Reinforced Encoder-Decoder Networks for Action Anticipation.

FIction: 4D Future Interaction Prediction from Video RED: Reinforced Encoder-Decoder Networks for Action Anticipation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.535082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.535082Z digest=sha256:c32bf6b90965366c5a2dc1d057286ba5c9966fe67f6948283efb9ed2d0ec47b4

Observation a6063b88-8054-4b1a-8174-13b6c88ae93a · outbound

This paper cites Anticipative video transformer.

FIction: 4D Future Interaction Prediction from Video Anticipative video transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.538490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.538490Z digest=sha256:bab02ddd507e28ce2ff14402d94f8261582cecd1ee9be536ee902219e833c87a

Observation 004f3e4f-8e7e-4127-88d3-bf299f685a08 · outbound

This paper cites Omni- vore: A single model for many visual modalities.

FIction: 4D Future Interaction Prediction from Video Omni- vore: A single model for many visual modalities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.542273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.542273Z digest=sha256:af3316c1d1d7c5ace23b3c067a8f00d91d1e1240daaceda4e28cad9db493f00c

Observation bf196040-dcb1-4cfb-af48-5c104814bb4f · outbound

This paper cites Humans in 4d: Re- constructing and tracking humans with transformers.

FIction: 4D Future Interaction Prediction from Video Humans in 4d: Re- constructing and tracking humans with transformers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.545672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.545672Z digest=sha256:149e61b86d9c5e7c626599b97c18bcf804462f3b534db6465ce27ecb99131fa8

Observation 47955e9a-7cf5-4cc3-8590-d870b02ffe80 · outbound

This paper cites Con- tactopt: Optimizing contact to improve grasps.

FIction: 4D Future Interaction Prediction from Video Con- tactopt: Optimizing contact to improve grasps

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.549484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.549484Z digest=sha256:c62e365e9b2ab5d74a2bb60db02a916245393e5afa9eda7fb1495c7eca2f2dbb

Observation cbe0b724-33af-4318-b876-7ed2bbe59832 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

FIction: 4D Future Interaction Prediction from Video Ego4d: Around the world in 3,000 hours of egocentric video

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.553880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.553880Z digest=sha256:a7b56df200405093c1f42476ea3b76b57675e983b9cf476fc522c97ba7a471eb

Observation 818be1cb-7ff7-494a-9c1b-967c34453325 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

FIction: 4D Future Interaction Prediction from Video Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.558067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.558067Z digest=sha256:a1df63c5ebf998bef84b5dbef10b72ddfbfc1ad794a5f5063ee6ec3d0fab2f56

Observation 649f042e-568f-453c-9609-b41fc3830ebc · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

FIction: 4D Future Interaction Prediction from Video Lvis: A dataset for large vocabulary instance segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.562199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.562199Z digest=sha256:9a3b3630f31cfaf9a078ed8ab65ff536dee1881b4ebc6501a106669289413148

Observation 896641e9-732a-4cae-a3d5-7b1c5182b8a1 · outbound

This paper cites Resolving 3d human pose ambiguities with 3d scene constraints.

FIction: 4D Future Interaction Prediction from Video Resolving 3d human pose ambiguities with 3d scene constraints

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.566700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.566700Z digest=sha256:c59c6356a91fac947dd0b5fb785d9db9a556208794c6efa371c4d97ead1aa57d

Observation ed044fe1-a82f-448f-b79b-0247bdcd712b · outbound

This paper cites Stochas- tic scene-aware motion prediction.

FIction: 4D Future Interaction Prediction from Video Stochas- tic scene-aware motion prediction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.571042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.571042Z digest=sha256:f74aad683466a44c01fa06348802b1479e6e0b97ed62042e9544ad31f6ab32a4

Observation 14dd2e39-6a5c-4c28-a458-dc4e16fee3a4 · outbound

This paper cites Synthesizing phys- ical character-scene interactions.

FIction: 4D Future Interaction Prediction from Video Synthesizing phys- ical character-scene interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.575363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.575363Z digest=sha256:a930604e31bb29b6139ec56ac007dc0559962c3ad63bc36eb33cb528de1c45e3

Observation ccdee9e4-0959-4787-b098-6152f839eea3 · outbound

This paper cites Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation.

FIction: 4D Future Interaction Prediction from Video Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.579824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.579824Z digest=sha256:3525eed370cd68c6d29645d12e7127d59d6408d2a953cf59afef1a9ed64cd2b3

Observation 90595afb-d08a-4055-9a06-4a6c2fe5c670 · outbound

This paper cites Diffusion- based generation, optimization, and planning in 3d scenes.

FIction: 4D Future Interaction Prediction from Video Diffusion- based generation, optimization, and planning in 3d scenes

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.584042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.584042Z digest=sha256:e9c2066dc3facab575197779a326332f1c0a4a85855ef6c52e2077aaa74d3a81

Observation 7ce5930b-6c19-4f41-ad7e-6a8ee7bda2da · outbound

This paper cites Human3.6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments.

FIction: 4D Future Interaction Prediction from Video Human3.6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.587680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.587680Z digest=sha256:7a06c2f28ee7e020f54e6faabe8e66ee9f1ebb66ef2ac92619848cdd1ce051e1

Observation 9f203ce7-0efa-4553-b54a-9303e78e8174 · outbound

This paper cites Hand-object contact consistency reasoning for hu- man grasps generation.

FIction: 4D Future Interaction Prediction from Video Hand-object contact consistency reasoning for hu- man grasps generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.591402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.591402Z digest=sha256:1d4da0d18b21a119b0b5b59a0271b6d6bc29d69f1d8699341481ddc9785250da

Observation f0a1d69a-e07a-47c5-8728-ec5dedaf5cb7 · outbound

This paper cites Sym- phonize 3d semantic scene completion with contextual in- stance queries, 2023.

FIction: 4D Future Interaction Prediction from Video Sym- phonize 3d semantic scene completion with contextual in- stance queries, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.595135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.595135Z digest=sha256:2e11cc9b5d6ac9b1ca72e8e934e756c184bec6658b0caee208ce773d43ed5694

Observation 07767f29-b574-4145-aaba-ea6861a5e29d · outbound

This paper cites R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection.

FIction: 4D Future Interaction Prediction from Video R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.181188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.598689Z digest=sha256:a54b9ce9365d23502fd202c2bbd9b834447692ecf8a904881607aa28e31a3caa

Observation e35a7de4-ea32-4cb1-9a81-21fade0c5aa1 · outbound

This paper cites Segment anything.

FIction: 4D Future Interaction Prediction from Video Segment anything

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.602804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.602804Z digest=sha256:27e0644547f7655ab1150c449e78520095eb6a9b0e673c01130a304a51f75daa

Observation b05ffc6d-7d7f-4961-9cd1-1a35dc569ff4 · outbound

This paper cites Interactive object segmentation in 3d point clouds.

FIction: 4D Future Interaction Prediction from Video Interactive object segmentation in 3d point clouds

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.969264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.606482Z digest=sha256:9d4aaa2229904cb595c33a10fabda21d2af24865b8c949e92bfbeb2187c772db

Observation 63eef815-7046-4c14-94d7-2ce13dd3dd57 · outbound

This paper cites A hi- erarchical representation for future action prediction.

FIction: 4D Future Interaction Prediction from Video A hi- erarchical representation for future action prediction

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.956226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.609535Z digest=sha256:76a94d0d5ba9f88ace3eb7f85e6b655a27d04c0a3be4b40a8400e9f79799921f

Observation 2d32b3ba-df97-467e-927c-13a462eb3b29 · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

FIction: 4D Future Interaction Prediction from Video LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.612463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.612463Z digest=sha256:c6451ad7e8d7b54df249184df2fe3c19efb61eac72079f6deb6c36c8cf5ae9e8

Observation 1e9820cb-45c1-4420-a34d-f00c8bad3a4b · outbound

This paper cites UniFormer: Unifying Convolution and Self-attention for Visual Recognition.

FIction: 4D Future Interaction Prediction from Video UniFormer: Unifying Convolution and Self-attention for Visual Recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.615541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.615541Z digest=sha256:7d418a176e8bab3f2a7800c4e68b935474767baf7020a5c9d3f48348f6bbd0b3

Observation c0cb4845-502b-4372-a53e-19f406dbed80 · outbound

This paper cites Mvitv2: Improved multiscale vision transform- ers for classification and detection.

FIction: 4D Future Interaction Prediction from Video Mvitv2: Improved multiscale vision transform- ers for classification and detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.943815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.618770Z digest=sha256:27f0b092c5b0c2a0ab097cbc77db80777b568af15cc849ab3bcb3f8bd4938f94

Observation a0c3ed5b-aad4-414a-99b3-572528567b07 · outbound

This paper cites Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar.

FIction: 4D Future Interaction Prediction from Video Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.930511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.622214Z digest=sha256:92932daccde741b1cf987befaeb46e87cb2158a7ceb1eafc2e20fcc14c7322db

Observation 0fb9c6e7-6b18-4534-be8f-4499916cc2e4 · outbound

This paper cites Egocen- tric video-language pretraining.

FIction: 4D Future Interaction Prediction from Video Egocen- tric video-language pretraining

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.915968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.625968Z digest=sha256:84734b3f86ac017cf64f45755d16c90cab6b7ceb7c51f8e7d1d37fc0f6a0a74e

Observation 3c376914-a25f-4b3a-835a-282c30660cd0 · outbound

This paper cites Microsoft coco: Common objects in context.

FIction: 4D Future Interaction Prediction from Video Microsoft coco: Common objects in context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.629458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.629458Z digest=sha256:f30f01e30f246ddad87fbb9c4060d447fb5cb3c698dc19e1023cd3c1921e47e3

Observation d1798fd1-2fc4-40ee-bb42-e11d60d02a8e · outbound

This paper cites Learn- ing to recognize procedural activities with distant supervi- sion.

FIction: 4D Future Interaction Prediction from Video Learn- ing to recognize procedural activities with distant supervi- sion

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.893705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.633145Z digest=sha256:ca6b137e18229bb63b852c97e0828f4f0bae53627848ec0737f5a2f422dd22d3

Observation 04161796-2915-456a-aabf-ab8359d75e6d · outbound

This paper cites Deep patch visual slam.

FIction: 4D Future Interaction Prediction from Video Deep patch visual slam

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.881706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.637095Z digest=sha256:20fdf4dbbfa2b79b96a82aec8e406804f65d9496549bc2907bddc7c87799f48b

Observation 87fed159-0559-4ad8-8d4b-9908c2664220 · outbound

This paper cites Visual Instruction Tuning.

FIction: 4D Future Interaction Prediction from Video Visual Instruction Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.641299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.641299Z digest=sha256:d1122cb4a5006e616360a33b59bd3b8e5d5c476fbbe3a3cc74d0a6de7b1ced9b

Observation cf3b1fc8-af49-4d67-9408-ca464e6ada64 · outbound

This paper cites Fore- casting human-object interaction: joint prediction of motor attention and actions in first person video.

FIction: 4D Future Interaction Prediction from Video Fore- casting human-object interaction: joint prediction of motor attention and actions in first person video

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.870748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.645564Z digest=sha256:a2d64f801cc42a2f135ab4592fead852c885c692917ebdffd69bea6497331e1f

Observation 8c1cdd44-0956-481e-82dd-e18914533a92 · outbound

This paper cites Joint hand motion and interaction hotspots prediction from egocentric videos.

FIction: 4D Future Interaction Prediction from Video Joint hand motion and interaction hotspots prediction from egocentric videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.859489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.649079Z digest=sha256:f86f097872b99e63ab822bbe800b0f095081a3834a2c67b59a35bcf29ae25995

Observation bf640659-1a84-4232-b088-3e464f2819e8 · outbound

This paper cites Joint hand motion and interaction hotspots prediction from egocentric videos.

FIction: 4D Future Interaction Prediction from Video Joint hand motion and interaction hotspots prediction from egocentric videos

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.847867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.652471Z digest=sha256:d0d34be7de2d25108023a3ac5994fe6a114cdba4e2431e4728164f99c5042076

Observation 2860e921-7892-4b06-9a23-8b62fe354f34 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:56:00.837304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.655916Z digest=sha256:29c417f51cccbb4db53ed2e07229647eb19d9cc8c6ead7dcd36c2c67308c2107

Observation 9cbe04d5-4a68-4396-b4fd-0ad035ca3fb5 · outbound

This paper cites Multimodal sense-informed forecasting of 3d human motions.

FIction: 4D Future Interaction Prediction from Video Multimodal sense-informed forecasting of 3d human motions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.826379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.659781Z digest=sha256:5e2a809936dd0e0d855e2e67be9498a5a4da3307ab27c59f80a8d011dab950a5

Observation 5b87f850-56d8-4358-958d-e7f36ea74cf0 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

FIction: 4D Future Interaction Prediction from Video Amass: Archive of motion capture as surface shapes

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.814816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.663603Z digest=sha256:8f47a2866564459ce1df0611e75797fcd032b739a34cc4d08c0dcb84b637957e

Observation fee3d281-a9ba-4e46-96fd-245e975f7d4c · outbound

This paper cites Contact-aware human motion forecasting.

FIction: 4D Future Interaction Prediction from Video Contact-aware human motion forecasting

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.803216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.667609Z digest=sha256:8a427a3b3df3a4e45e429e12f7a1866143696e5c7cf46b530c58b5ea45629481

Observation f8dfe2d4-90ed-4baf-b6a4-7b73c28a4120 · outbound

This paper cites On human motion prediction using recurrent neural networks.

FIction: 4D Future Interaction Prediction from Video On human motion prediction using recurrent neural networks

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.791592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.671250Z digest=sha256:f39676cb0eb15128a299279d67fbf40357cfe48f6e28a3318395c1df74e48be2

Observation 00f2291d-45aa-49fe-9888-a8acafb95bfb · outbound

This paper cites Intention-Conditioned Long-Term Human Egocentric Action Forecasting.

FIction: 4D Future Interaction Prediction from Video Intention-Conditioned Long-Term Human Egocentric Action Forecasting

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.133452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.675387Z digest=sha256:9214c6f73d1d950ab1278c61910bbf20d9519053884b771e62f64d757d8a2671

Observation 4986b394-7980-4259-8baa-bb8b72142364 · outbound

This paper cites Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips.

FIction: 4D Future Interaction Prediction from Video Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.778937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.679252Z digest=sha256:e5235f312c29a424c92d57bc39067ce7013aa334e1658f0b269af620501abb8f

Observation 75393f18-c82b-49eb-aa56-cd909c9a897f · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

FIction: 4D Future Interaction Prediction from Video End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.767520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.682922Z digest=sha256:29ca225c891aa53083ff41566bfb5111986bae35331a757374352809364408be

Observation 877213ac-c196-423a-8cb6-5e895cf645af · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- 11 thesis.

FIction: 4D Future Interaction Prediction from Video Nerf: Representing scenes as neural radiance fields for view syn- 11 thesis

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.753669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.686550Z digest=sha256:d8ba3a9cc5850b1353696ea30e8d6c707400dc0646ce31320723bba78d7b976b

Observation 9ec69155-3b27-4297-94da-1192e797f937 · outbound

This paper cites Graspit! a versatile simulator for robotic grasping.

FIction: 4D Future Interaction Prediction from Video Graspit! a versatile simulator for robotic grasping

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.741183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.690285Z digest=sha256:adf8332565b3ffd9a3bcdcd5282ad59e87c043c80a57860f79f2a3bff8f5b8d9

Observation 44846b5c-4c04-4323-b8fe-dcdc281ff0bb · outbound

This paper cites Where2act: From pixels to actions for articulated 3d objects.

FIction: 4D Future Interaction Prediction from Video Where2act: From pixels to actions for articulated 3d objects

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.729177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.694406Z digest=sha256:13a1063de9c0cec241844a76c7ff85e66e675f7fb5d333ae749369379dd4e3a4

Observation 0b75ac41-22cb-45bc-9a2f-9a0068bcdd1a · outbound

This paper cites Orb-slam: a versatile and accurate monocular slam system.

FIction: 4D Future Interaction Prediction from Video Orb-slam: a versatile and accurate monocular slam system

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.697773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.697773Z digest=sha256:66e798d5eb6cc264aa2ac1808a4cc590aefd9703ebc414c337eefd4ad9259f61

Observation c7b528b3-e084-4f00-a9ea-934d83e910ef · outbound

This paper cites Multi-label affordance mapping from egocentric vision.

FIction: 4D Future Interaction Prediction from Video Multi-label affordance mapping from egocentric vision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.710600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.704470Z digest=sha256:c1130748e37b22539fe8ac6e5d1d26f73f7f7256ab5b0b511cb328ec7299e324

Observation e41b3955-af65-441b-aa9e-703b140f0ec7 · outbound

This paper cites AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation.

FIction: 4D Future Interaction Prediction from Video AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.707512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.707512Z digest=sha256:957c3a25c388f44f334940e5883418430054833c0e09d007b50e626a68214576

Observation 0d4db6a3-d95d-45a6-bebc-b9feb5891460 · outbound

This paper cites Nagarajan and K.

FIction: 4D Future Interaction Prediction from Video Nagarajan and K

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.697433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.710676Z digest=sha256:30293a7ad6e4ac888877f5e6e6b7da1dd81a3a0186f577a0ce7cf005a8971368

Observation 7be45799-0727-413b-8734-fb27ef1dad26 · outbound

This paper cites Grounded human-object interaction hotspots from video.

FIction: 4D Future Interaction Prediction from Video Grounded human-object interaction hotspots from video

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.685731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.713776Z digest=sha256:885482712bcc38332bbba628874d16384575deea3a08cc38585dfa96c037fe5f

Observation 26792741-cff7-4591-b15b-46d282ad7b76 · outbound

This paper cites Future event prediction: If and when.

FIction: 4D Future Interaction Prediction from Video Future event prediction: If and when

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.674007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.717287Z digest=sha256:170af84fc4506514290a3d4d2923aca39996f637a294b0362864e84e157ad7d1

Observation 2472ca63-6afd-48f3-bcc4-77ef094f42c3 · outbound

This paper cites Using Geometry to Detect Grasps in 3D Point Clouds.

FIction: 4D Future Interaction Prediction from Video Using Geometry to Detect Grasps in 3D Point Clouds

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.107862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.837998Z digest=sha256:037efc77fd6c82c62d6c458ea048499d74b4bf02555128cd9279e145a4235211

Observation 1306d568-892c-4691-b533-161300a63894 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.842388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.842388Z digest=sha256:8dcce2a23e50893b678149d8baa37b8dd783115e87276941f2fcb209c02d26a3

Observation 7414532d-4fee-4cef-96e3-b58cb3017e6d · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

FIction: 4D Future Interaction Prediction from Video Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.655407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.846244Z digest=sha256:738580a177e92ebdaaccda373e79bd6afcd442d2f388d6fa5922cc08c3b36851

Observation 4e6dd8f6-41fa-4a1d-add9-743bed487856 · outbound

This paper cites Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots.

FIction: 4D Future Interaction Prediction from Video Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.850055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.850055Z digest=sha256:538da5bd015c3bb0d5742debb18595ed7a950d22d97e9bbab9e811fc58bb8f79

Observation ad74e877-1478-4938-bc7f-c3f14648910b · outbound

This paper cites Vins-mono: A ro- bust and versatile monocular visual-inertial state estimator.

FIction: 4D Future Interaction Prediction from Video Vins-mono: A ro- bust and versatile monocular visual-inertial state estimator

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.644501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.853864Z digest=sha256:b74c8bce72db8d8b98f7bdca6d41168e118c25390e0f723fff7127b08975b64b

Observation a9b6bcdf-c44d-42a6-ba3e-1230e330fa99 · outbound

This paper cites State-only imitation learning for dexterous manipulation.

FIction: 4D Future Interaction Prediction from Video State-only imitation learning for dexterous manipulation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.632033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.856756Z digest=sha256:f0d8227eb9f9b853623ce1a8020214a63a45434646e9770c9399575a81a7be23

Observation 4e0c5a99-7e26-42ad-8eb6-e26325e83235 · outbound

This paper cites Poni: Potential functions for objectgoal navigation with interaction-free learning.

FIction: 4D Future Interaction Prediction from Video Poni: Potential functions for objectgoal navigation with interaction-free learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.620790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.860281Z digest=sha256:55bbbf9040e4407e4dbb7d4d4e2fada71335171f922e44f5639b5e308a161907

Observation 9c418893-8def-4287-aae8-0409267c61d5 · outbound

This paper cites Humor: 3d human motion model for robust pose estimation.

FIction: 4D Future Interaction Prediction from Video Humor: 3d human motion model for robust pose estimation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.609620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.863253Z digest=sha256:7363da90fb105348eebdee30f4d47c8ae33dd6b600b6a3b1cfd8a18ae7124b03

Observation ffafcde5-3f30-4535-9a77-390b2dcf1093 · outbound

This paper cites First-person activity forecasting with online inverse reinforcement learning.

FIction: 4D Future Interaction Prediction from Video First-person activity forecasting with online inverse reinforcement learning

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.598098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.866660Z digest=sha256:2fa8ea254205565eb0250f217d6f2b33ad2566fcd5745b9bfa609d0045c0fed7

Observation 34224ca4-994f-4a09-88e9-2cc83b93bf22 · outbound

This paper cites Structure-from-motion revisited.

FIction: 4D Future Interaction Prediction from Video Structure-from-motion revisited

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.586404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.869991Z digest=sha256:49605f3629e180a22f4b2e5fd9805c1658e76d4848a54142331486d8c1c21b76

Observation 32b5e6bd-0ae5-450e-8e3d-e2d922e48ff1 · outbound

This paper cites Wham: Reconstructing world-grounded humans with accurate 3d motion.

FIction: 4D Future Interaction Prediction from Video Wham: Reconstructing world-grounded humans with accurate 3d motion

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.575160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.874484Z digest=sha256:bc184bc8ba1fde5f65000b676e1ee8afd595c73bd1365f86c0525fa70cadc7e0

Observation 9f894892-d0c4-42d8-b1fd-2c988dcaadc4 · outbound

This paper cites Learn- ing structured output representation using deep conditional generative models.

FIction: 4D Future Interaction Prediction from Video Learn- ing structured output representation using deep conditional generative models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.563868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.878602Z digest=sha256:8c73936dc3955e2e703c15061c84041dff1bdf7e82996741f0201033e57f71bc

Observation 38969f8c-5b29-4748-a4f8-5b10810fdf06 · outbound

This paper cites Generating notifications for missing actions: Don’t forget to turn the lights off! In Proceedings of the IEEE International Con- ference on Computer Vision, pages 4669–4677, 2015.

FIction: 4D Future Interaction Prediction from Video Generating notifications for missing actions: Don’t forget to turn the lights off! In Proceedings of the IEEE International Con- ference on Computer Vision, pages 4669–4677, 2015

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.552001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.882948Z digest=sha256:53696f97e509766ca8db0a7c07c11e596ca7aa3427911a07fbeac8db8c0dc66a

Observation 38789a9a-c1e2-41a3-ad9e-ddbac541ddd3 · outbound

This paper cites Segcloud: Semantic segmen- tation of 3d point clouds.

FIction: 4D Future Interaction Prediction from Video Segcloud: Semantic segmen- tation of 3d point clouds

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.537792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.886883Z digest=sha256:3079e879827eb299bc87a94c61e8c2f3427c62a3830e593d2361eee0f574d785

Observation 22414549-19af-4e40-96c5-263c29df85cb · outbound

This paper cites Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras.

FIction: 4D Future Interaction Prediction from Video Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.891500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.891500Z digest=sha256:901c21f97caf431d437ede2bca4d56393b098a055af3a4830b69ef8be60871b3

Observation 5e2e33b1-5232-4ec4-8eb3-dbeb01c899f7 · outbound

This paper cites EPIC Fields: Marrying 3D Geometry and Video Understanding.

FIction: 4D Future Interaction Prediction from Video EPIC Fields: Marrying 3D Geometry and Video Understanding

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.518624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.895419Z digest=sha256:d6efcae9d9731a3030189785a2a448a81f426c061a4b18682762b99912ba6dd7

Observation 5667d731-4af4-4431-9556-0195e76ddcf1 · outbound

This paper cites Transformation-Based Models of Video Sequences.

FIction: 4D Future Interaction Prediction from Video Transformation-Based Models of Video Sequences

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:56:00.080431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.899330Z digest=sha256:d9b81587b35451d3320b1fb6c837da8937eff6e4b5e73c3621797eefe64d8db5

Observation 2ee9e01c-18a7-4644-8c0d-a8e20bfbfad9 · outbound

This paper cites an unresolved cited work.

FIction: 4D Future Interaction Prediction from Video Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:56:00.506311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.903006Z digest=sha256:e4a1df899ca867786ba148c56fdcce77d0cc1326d9f011a8a4e4e86dc8399a04

Observation cedbaa9f-ae03-4c41-802e-89cd7425b035 · outbound

This paper cites Synthesizing long-term 3d human mo- tion and interaction in 3d scenes.

FIction: 4D Future Interaction Prediction from Video Synthesizing long-term 3d human mo- tion and interaction in 3d scenes

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.493415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.907142Z digest=sha256:21b65e57ea4dd467d09e0db9631db62a29416bd87afd2a132bd9efa9c2bdeb95

Observation 3d6d5e86-d9da-4c0d-8d7a-3552690f5732 · outbound

This paper cites Adaafford: Learn- ing to adapt manipulation affordance for 3d articulated ob- jects via few-shot interactions.

FIction: 4D Future Interaction Prediction from Video Adaafford: Learn- ing to adapt manipulation affordance for 3d articulated ob- jects via few-shot interactions

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.480249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.910653Z digest=sha256:a3099601c30b59354995f9e9366bbff984a15845f58491ba28eecb9029fe0d3b

Observation 9747d2c8-d121-46a7-865e-8ccc3c9072c6 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

FIction: 4D Future Interaction Prediction from Video Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:56:00.467780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:55:59.914437Z digest=sha256:6dc117bc2e1b831f3d42a45d583e610ae0fc3be4bd4d50f363195691452f0a04

Pith citing papers

No inbound Pith citation observations are available.