Pith. sign in

Paper Citation Record · LEDGER

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models

As of 18 August 2026, this Paper Citation Record lists 100 of 135 outbound references and 3 inbound Pith citation observations for arXiv:2506.01608.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01608 v2

Coverage vector

measured 100 of 135 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:35.244016Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T23:20:12.870365Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T23:27:38.821483Z

Reference resolution

100 of 135 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b2be9cd-26d7-410a-9b37-350877afa0c9 · outbound

This paper cites An overview of augmented reality technology.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models An overview of augmented reality technology

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.940539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.940539Z digest=sha256:527d3217dc6b31db4881a256ceb3ffc198333add47003dccd11c7a44708daa24

Observation 6ede1e71-6ca6-41f5-a862-7b472a70ba2b · outbound

This paper cites A survey of em- bodied ai: From simulators to research tasks.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models A survey of em- bodied ai: From simulators to research tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.962237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.962237Z digest=sha256:50f8a97f9a005947e1f8cd32db52b76f30fb0ffb37a6b9ecb163a5ca19eff815

Observation a8dbc1c7-1d27-4c81-bf5f-dd8305866810 · outbound

This paper cites Beyond simple laboratory studies: Developing sophisticated models to study rich behavior.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Beyond simple laboratory studies: Developing sophisticated models to study rich behavior

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.984291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.984291Z digest=sha256:86d62001125a2fd8cdf5e9c51ee46e678d0d9ec44a99adf54bfac265c47a8eb7

Observation 205e4756-c320-4c59-8805-01da11559ccf · outbound

This paper cites Decoding the brain: From neural representations to mechanistic models.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Decoding the brain: From neural representations to mechanistic models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.006140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.006140Z digest=sha256:a0cee0eae446ddc40ec61b80fbf4695e6773d8b3917fb31f04bbadf74e8381e2

Observation d55af09b-dcdb-403a-a930-bdacde5b84c8 · outbound

This paper cites Advanced neurotechnologies for the restoration of motor function.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Advanced neurotechnologies for the restoration of motor function

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.027283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.027283Z digest=sha256:c23e8570b9538c7879478393f75863ba32325833f4348c5c4dbea0ec7553125c

Observation 9ef947b8-7e09-444e-adba-7b5d9cab2b17 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Videomae v2: Scaling video masked autoencoders with dual masking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.055654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.055654Z digest=sha256:698dd325e8a662812028e3980ba19013203414f3a98953fc52a9da5fe2adfed0

Observation ca05c593-1206-4eb1-bd08-5359f852b500 · outbound

This paper cites UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.097643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.097643Z digest=sha256:e0da41ef1e2d19c0d54d8ce9fffba4013b16b837cb6e7464dcdc43f7f319134b

Observation 92a463c6-2bb1-4220-8c8d-a16f313471b2 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.118641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.118641Z digest=sha256:9f9bc9c2d37274fe38fc9fc69d75be5a6227520d85e8c5d91d6ba211d0b630ee

Observation 2c768c4c-e4f2-419a-a06b-7590869e9ed1 · outbound

This paper cites Revisiting skeleton-based action recognition.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Revisiting skeleton-based action recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.140451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.140451Z digest=sha256:04c8e5134f2c5edfbc54c2256dcb090099fc2ed28023664d54fe8ce0cde7b2c3

Observation 7bddfe26-4496-43c7-a8a2-5ca01e5aa12a · outbound

This paper cites Aspnet: Action segmentation with shared-private representation of multiple data sources.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Aspnet: Action segmentation with shared-private representation of multiple data sources

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.162298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.162298Z digest=sha256:c2bd77a481f57d856863d4da13b1eada24c7c78da0274e9a7b76691894894c87

Observation aeb2d922-fd60-4f9d-b131-cd48a67cc018 · outbound

This paper cites Diffusion action segmentation.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Diffusion action segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.197193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.197193Z digest=sha256:ed863f28e8092ebe3a41d35c767d90405840181d6287f7e1c77e04cc10f47b5b

Observation ef1abd6d-5f2d-4370-862b-ae274a9b831d · outbound

This paper cites Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.221391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.221391Z digest=sha256:9c1db1335fdb1c996346bbf8d938c8ad3cdd72342b4311b958b0e7d8b9393757

Observation 96ea514d-bc9d-42a2-ada4-756e47e7cac2 · outbound

This paper cites Elucidating the hierarchical nature of behavior with masked autoencoders.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Elucidating the hierarchical nature of behavior with masked autoencoders

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.243064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.243064Z digest=sha256:030f099c1862450772935fa6c0faf584025f04a99ab51caaa93f66a0d9275326

Observation 5ccab023-f9b4-409a-9d35-9175fe10e0a5 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Motionclip: Exposing human motion generation to clip space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.264417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.264417Z digest=sha256:78d5f34d9dacca2def2cdd4b7a8a3f6b48204338b3966829fce12f3d42304715

Observation 2a583056-30cb-4619-a81b-78ab64ff08a6 · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Generating human motion from textual descriptions with discrete representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.289168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.289168Z digest=sha256:6145b399b15204a7ba928ba70b3080731ca1b536b90ee6d4825b26b594e91968

Observation 1a8210b3-071d-4e5d-b84e-83f0b56e84ae · outbound

This paper cites Momask: Generative masked modeling of 3d human motions.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Momask: Generative masked modeling of 3d human motions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.328815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.328815Z digest=sha256:a9833154da10d211dbe7838e4063e9224de1a7acae95aada6003696b46dd7b71

Observation 3f3d1eff-cc47-462c-b435-7c373f3fc13c · outbound

This paper cites Eye movements in natural behavior.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Eye movements in natural behavior

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.364467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.364467Z digest=sha256:70b59feade0c09f7551954154b4882783b65adfd14146a359859b9795ee1c19b

Observation cab15c84-cab8-498d-8e80-a633512d06af · outbound

This paper cites The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.391216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.391216Z digest=sha256:4396b87bd04978e10475afcde41fcb4bbcee222f52392ce4aa4adab5c110acdd

Observation b1051384-a598-431a-8f7d-a72e443addb2 · outbound

This paper cites The ikea asm dataset: Understanding people assem- bling furniture through actions, objects and pose.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models The ikea asm dataset: Understanding people assem- bling furniture through actions, objects and pose

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.412477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.412477Z digest=sha256:3915834eb2d5d6afd5911620b63c9a3598c0f9051ae38b48dae8e36d5a45ec65

Observation e2c1dcca-9d6f-49cd-9414-81a42f0e229a · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.443947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.443947Z digest=sha256:f6bca3480fd60a752935764016f4104b1497396cc9eac1be616fc585e14a1001

Observation 11549858-064c-4b40-b13d-6ee6cf126e2f · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.465229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.465229Z digest=sha256:dfed98834cf03b2bb493e47be0ca88d7a01db656d1c6e9661181600303853388

Observation 04c666d4-3c45-4ba6-b923-f36d9db3c51c · outbound

This paper cites Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.487014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.487014Z digest=sha256:f05a7f43df985e72c8fdf55dba2a9788a4bdce313b224ba94a02949379570ae1

Observation 9e670a51-d1ee-46ef-a6ca-01e98fe4d32a · outbound

This paper cites Egoexo-fitness: Towards egocentric and exocentric full-body action understanding.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Egoexo-fitness: Towards egocentric and exocentric full-body action understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.508570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.508570Z digest=sha256:ff0f4c461433003fdb8e7feba05d667192c4ea51983f23707e99a631d5c2bdff

Observation da83ee42-8d44-4bbc-a4ec-ca9250b24836 · outbound

This paper cites Troje, Gerard Pons-Moll, and Michael J.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Troje, Gerard Pons-Moll, and Michael J

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.537804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.537804Z digest=sha256:ce4166ee0ccaa5b102a3edef3c3d3368ec56a2ff0c1a4ff5b5e7bd79d7a287eb

Observation 031bfbe2-f352-4044-8160-a6fcf4c89067 · outbound

This paper cites Babel: Bodies, action and behavior with english labels.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Babel: Bodies, action and behavior with english labels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.559086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.559086Z digest=sha256:88a3c44443ea20e5e4f5be67cc73ccef198f62dd88dcfc375e7ef75d9446cb5f

Observation 903b9e00-8f85-4d33-ba82-e3d84adc6ef5 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Generating diverse and natural 3d human motions from text

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.588603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.588603Z digest=sha256:bfee341ae3070769c3c6d58388e3fcc06a4dad83790847eb6878292fe876f43f

Observation f506e1b9-1b78-48da-b95e-fbe85fedd997 · outbound

This paper cites Action2motion: Conditioned generation of 3d human motions.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Action2motion: Conditioned generation of 3d human motions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.624410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.624410Z digest=sha256:0c045f1aa7cd7dd3eb89192672dd257c56053bfc2dbc1c72450ad33a58016669

Observation c4481867-cce2-48c5-aa10-277bcd7e963c · outbound

This paper cites Assembly101: A large-scale multi-view video dataset for understanding procedural activities.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Assembly101: A large-scale multi-view video dataset for understanding procedural activities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.657659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.657659Z digest=sha256:0650a28798dc31c2e481fb7b15c59a118d5af33b2efcab521b75c843d11145b9

Observation 8e1319d1-f53e-4513-b3cc-8caf1e64fa79 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models H2o: Two hands manipulating objects for first person interaction recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.681208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.681208Z digest=sha256:d07c2a98bf30ebd123a0c33d502fb4abb38928b3fd104fff18f3d0b8d2c8455c

Observation 3f769998-0d71-4d1f-81a2-a7197d35fdbc · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Motion-x: A large-scale 3d expressive whole-body human motion dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.703072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.703072Z digest=sha256:aa313479b20fdcf4e5b62111fba212954cad5788099b8852d1ab9fcaa9caa93f

Observation 2b0fc9ac-b2a3-404a-92fb-9ad0559f4c33 · outbound

This paper cites Nymeria: A massive collection of multimodal egocentric daily motion in the wild.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Nymeria: A massive collection of multimodal egocentric daily motion in the wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.724618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.724618Z digest=sha256:3c20bd7c4df2de2b3d5cbc9cfa23dddb28ec272d5fe4b13f4c3d019c9353bb40

Observation 53c20bca-3d0d-4319-85b4-d3a20b6846ef · outbound

This paper cites Ntu rgb+ d: A large scale dataset for 3d human activity analysis.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Ntu rgb+ d: A large scale dataset for 3d human activity analysis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.746391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.746391Z digest=sha256:dd8f3186b79661cceac9894b64c75a5ecb66e84da513ffaad0a09883ade0c135

Observation d923b017-9273-4cf5-98a2-b357845f63cf · outbound

This paper cites Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.768482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.768482Z digest=sha256:81a55229730d16a00b509a89ffd2bf888d0faa7a684b4148fbcc5ab27d099506

Observation e5fdd136-4f14-4a08-be0a-467e5a317c68 · outbound

This paper cites Flag3d: A 3d fitness activity dataset with language instruction.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Flag3d: A 3d fitness activity dataset with language instruction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.791474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.791474Z digest=sha256:4a642683a1edcf0e713e7fc8d75c10d4c0e1098c07c2ef2f0ac3b935f4107a56

Observation 7ab12692-76d1-4e9a-8bc0-1a3a62271a83 · outbound

This paper cites Unsupervised learning from narrated instruction videos.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Unsupervised learning from narrated instruction videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.812117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.812117Z digest=sha256:b8eeccb2e039f586e57a42486bf36ef5bd65b38e20df6a8c99ee468e30f90165

Observation 7178622f-f37d-44ef-b819-dc5199dbc1c6 · outbound

This paper cites Humans in kitchens: a dataset for multi-person human motion forecasting with scene context.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Humans in kitchens: a dataset for multi-person human motion forecasting with scene context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.833133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.833133Z digest=sha256:fd2e35fc11acb33530ac396cc27ad0eb1a4f375f23fe41e874974de9cc5f39cd

Observation 33257411-ba05-441c-bf78-d8b343c19f87 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.854918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.854918Z digest=sha256:8ce055d409f2633491b7d65b88ad9aef2c4f22e89cc8581de6b8c12f5d3c9955

Observation 8d7d89ad-b21b-4cde-b782-88936c002a59 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.893021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.893021Z digest=sha256:2f9649359bbaea104defb76f1f9f09838256847c5ad918a929e530b4349c5472

Observation 8d6d4156-c5c8-4e27-82cd-14f1a6c2e556 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.916143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.916143Z digest=sha256:44f060c2c9d28dd3fbc5904cd3823950dd0c31b0d89bbd89686797c077a062c1

Observation bc9bb55a-417a-46de-885c-16b6320c065c · outbound

This paper cites Egoschema: A diagnos- tic benchmark for very long-form video language understanding.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Egoschema: A diagnos- tic benchmark for very long-form video language understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.942537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.942537Z digest=sha256:476e22248000586c7da5c72720d0fd28d5480efecfabc3d1c5ce39e42ef2d60a

Observation b12b6a4b-d8c8-4ebf-8023-186a0b0a5815 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Egotaskqa: Understanding human tasks in egocentric videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.962811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.962811Z digest=sha256:6b670575a4b7b4b7bf69c7c898644f80f9e2f664f8fe95ef6f5756a2cc62cbe2

Observation 5036638e-e3eb-4725-b7d4-60094562dc2b · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.984557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.984557Z digest=sha256:d773dd4ae447abd4ea1abe807599cf78b9b784f9039816dbd91b9e125ac1346e

Observation 64ae0596-6111-4a47-98c8-9a285cbce2c4 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Next-qa: Next phase of question- answering to explaining temporal actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.006811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.006811Z digest=sha256:c28a770b92c40ee8fce96f5ab1ed0b6e393cbe9b9ed7b48c09a878ab8e1758bf

Observation 0773ce32-d25b-4c9a-a11e-7d9cc1e2eb03 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.028134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.028134Z digest=sha256:6d430cfcb6f9a95b9af947125f0cd3b2300ede71f44f03af7de29e4f42548970

Observation dc42cb8f-a263-4a4a-b30c-8f972028b1a5 · outbound

This paper cites Multi- view action recognition using contrastive learning.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Multi- view action recognition using contrastive learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.054212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.054212Z digest=sha256:bdbe7c470930dfd50f5bd750d1bc0c190038c97ff25de246bd81af7b64de12ac

Observation 88de4ef9-58ee-4263-b140-61f023c9875b · outbound

This paper cites Generative multi-view human action recognition.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Generative multi-view human action recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.075666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.075666Z digest=sha256:52f35b53351f0ba32331a0dc547a76dca2f16feb9dc8c11b929f20ae276af6bb

Observation f16ebc21-63f5-4b97-bc72-d69bfbd8bb77 · outbound

This paper cites Learning video representations from large language models.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Learning video representations from large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.123869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.123869Z digest=sha256:604e94c147a9c515b1fb5e9e67b954bae9f1e5bd279297abbf366601a6ad6a08

Observation 0b30383a-89cc-4c0c-baa4-d010977ffd9c · outbound

This paper cites TIM: A Time Interval Machine for Audio-Visual Action Recognition.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models TIM: A Time Interval Machine for Audio-Visual Action Recognition

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:42:35.724537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.161147Z digest=sha256:9a54e0e9d61054ce9506f4937226034cb2cb8f10f6731b22144728864b45fc03

Observation 89b87024-bfb2-4ced-b430-40c7b94df63b · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Audiovisual SlowFast Networks for Video Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.185270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.185270Z digest=sha256:ca89e602c3ea754c3161588e5945dd827a3749bac9b3558de4ccf481eed43b4d

Observation 515a5db1-8a7c-4e15-84d7-f8cf8cd01815 · outbound

This paper cites On the Utility of 3D Hand Poses for Action Recognition.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models On the Utility of 3D Hand Poses for Action Recognition

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:42:35.685612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.206599Z digest=sha256:54fec4d07424862d14a76c9719f83158da34441f7d1fedf80685d78d25f90287

Observation e094a8bc-8e8d-4558-a503-0e8c9fe37d96 · outbound

This paper cites Executing your commands via motion diffusion in latent space.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Executing your commands via motion diffusion in latent space

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.228558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.228558Z digest=sha256:880bc75ac0caf2bca202fdb50806adcee1a08e327f209eee7407011c735dc4c5

Observation d2fc2dfd-4cc9-47ef-bf20-fefe6bfcb9c5 · outbound

This paper cites Human motion diffusion model.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Human motion diffusion model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.250145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.250145Z digest=sha256:0c76cf4dc2ff2f95683e6a636cad586f572825838a9239f43ba93013e7dec5f3

Observation 07051925-73eb-4e12-b7c9-c1db4b27660d · outbound

This paper cites Mmm: Generative masked motion model.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Mmm: Generative masked motion model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.274650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.274650Z digest=sha256:ea6d9502ba6240950c31d18b92c346b9ccccb97ae964a736b572874c1063b63f

Observation d44785eb-f1c9-4a7b-9eb7-a790a139f5ae · outbound

This paper cites Remodiffuse: Retrieval-augmented motion diffusion model.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Remodiffuse: Retrieval-augmented motion diffusion model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.299445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.299445Z digest=sha256:7d2b893b595a424f9d632a4a7d8918fc30aef4ead06ba40da1b81d6ea84525d2

Observation 6e3aff71-8e4d-4dba-85c1-874ce8939635 · outbound

This paper cites The kit motion-language dataset.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models The kit motion-language dataset

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.323234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.323234Z digest=sha256:7e9b78a6fb1b8b79c91e15475bf355c22e54270fb9cc0679d5936ab5a7a493d0

Observation 636181aa-38b4-414f-8f4b-8b3ca3913c9b · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.352593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.352593Z digest=sha256:1bdbeca65a6dd8234aa473f9fe20a17789b7a603babf243ab967fc1d81af17eb

Observation 467b9160-d4ec-4bd0-badf-c2d1c1d8f121 · outbound

This paper cites Qwen2.5-VL Technical Report.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Qwen2.5-VL Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.382040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.382040Z digest=sha256:c5e8aec8497fbea23150e9e07742d31dc3adb8031eb5c9dabd326a666c56e347

Observation 6cd2554c-7425-4f18-bc56-520cf1f8c0bf · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.402950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.402950Z digest=sha256:8bcd4cfa033fa2686532ddd6ef00bac74078f793332b2af25bfe61b16e796249

Observation 12337c45-d00d-4ba3-838e-031430cadd7d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models LLaVA-OneVision: Easy Visual Task Transfer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.426590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.426590Z digest=sha256:1a93d703c2f79a03227d3b6fcfecf9bac6dd0321bbbed656b9be34e220ee3354

Observation 382cb532-05cc-4f13-85df-7c346f2d3915 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.447772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.447772Z digest=sha256:062b279c40c50b75541670af942879408777123d0f85761e8adb8a488a9c02cc

Observation 80bf06b9-a341-4af5-bf9a-dde5ec0e8f6d · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Llama-vid: An image is worth 2 tokens in large language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.468933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.468933Z digest=sha256:e6b84bcf9a8c1e78d4b476b7e129ba1b058b538fb19e56eb5b9e96279b138ab8

Observation 4f98908d-3376-4030-b86e-991b66b6e86d · outbound

This paper cites Vila: On pre-training for visual language models.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Vila: On pre-training for visual language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.490449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.490449Z digest=sha256:024fa0e8db3ecd952e733ed141022623dfcecc12b55bddc157c636fa557f2136

Observation 7ced5595-6bca-413e-848d-b883a54ec928 · outbound

This paper cites Long context transfer from language to vision, 2024.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Long context transfer from language to vision, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.513936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.513936Z digest=sha256:feec312baced6868636184abd4b0170bf5de99fad1bd53fa29dd5d47fb8b4119

Observation 466d48a1-b1eb-4457-8796-16a063f26ebb · outbound

This paper cites Longvlm: Efficient long video understanding via large language models.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Longvlm: Efficient long video understanding via large language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.538689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.538689Z digest=sha256:5ff0d0d190dfc5d8df1e967cbb94bf62a5cc7e99fac1cf60fe3868572a05044b

Observation 388f79e1-50ab-4210-b808-69698edb3de7 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.565742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.565742Z digest=sha256:576540328f135385fe44eeac04e9169c1dfd06257bc53a2d9ba9c3ed39cfcbf0

Observation e0c450ec-1717-41dd-b0b7-ab3213950d53 · outbound

This paper cites World model on million-length video and language with blockwise ringattention, 2025.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models World model on million-length video and language with blockwise ringattention, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.614402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.614402Z digest=sha256:766c83654479bd5e9d8424bbee2f80ce7fc8ae706f503e127a4c55c84456c09d

Observation a1675ecc-30e0-4151-9828-6e4f85cc131a · outbound

This paper cites Azure kinect dk documentation.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Azure kinect dk documentation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.641766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.641766Z digest=sha256:f4e1cb4c407aad12cbfb94697b41097b1140f9da2c63b82c8594a4e4daa75087

Observation 700f82f1-a3ee-4844-b57d-da923debd695 · outbound

This paper cites HoloLens 2 Research Mode as a Tool for Computer Vision Research.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models HoloLens 2 Research Mode as a Tool for Computer Vision Research

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.667717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.667717Z digest=sha256:2acf28599afc0236cda1ba6833394883ec7b4e23d4cce23b93ebd6967ef84993

Observation 2120ad00-7744-45d7-96e3-37f3e7428164 · outbound

This paper cites RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.689266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.689266Z digest=sha256:1aeaf2abf2b76f1e09117a74c00f1bb1f6ccd1e8c7577805daffb0fc257adea2

Observation 64f451a6-85ac-4da0-92e7-a00962199158 · outbound

This paper cites Deeplabcut: markerless pose estimation of user-defined body parts with deep learning.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Deeplabcut: markerless pose estimation of user-defined body parts with deep learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.711205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.711205Z digest=sha256:f92473b4a8047733def1083ddc46e947590e96fc3a9705e0de4eac629db659e6

Observation 888dfe29-671c-446a-b849-9491011ff0a6 · outbound

This paper cites Azure kinect body tracking sdf.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Azure kinect body tracking sdf

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.731798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.731798Z digest=sha256:b4ceabe10c3a8f8705c9b61de1b68a040f6a5db7a248bed3ec98403c9560c8c9

Observation 59712c2b-2a84-494e-947c-88cf96843867 · outbound

This paper cites Hand tracking: Hololens mixed reality toolkit.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Hand tracking: Hololens mixed reality toolkit

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.359104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.755488Z digest=sha256:dddd060b03cbd500fed4dabd5da57b1acdfa1891f436179de35cf9d61e26fad0

Observation 680d7736-dc6c-4bd8-9861-f6d1ce24e474 · outbound

This paper cites an unresolved cited work.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.785544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.785544Z digest=sha256:d1f14c2222b9b02e9f1d7e30398df1c660fec5bf71331fc8aded4d943cea5aa0

Observation 125d4d28-5435-4c97-a4c6-1f1707b35624 · outbound

This paper cites an unresolved cited work.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.807991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.807991Z digest=sha256:ea4c167ddaf58c16ae84a4fc0ca302c359f3cf08501e5ccce1766a6421fe1504

Observation 5455a2fc-6e49-4821-941e-b0468254825e · outbound

This paper cites Analysis of the synergies underlying complex hand manipulation.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Analysis of the synergies underlying complex hand manipulation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.331052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.838736Z digest=sha256:9c01f1da53a7d24a7192f1d8f85f91bf354ba59eb05ba545ed536f653b03f1e3

Observation af4b35b0-2d1e-426a-b2d6-832b16a26492 · outbound

This paper cites Postural hand synergies for tool use.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Postural hand synergies for tool use

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.317517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.866713Z digest=sha256:45b9461eb5b3bec33818b3462fdff73f4feda3827ef58a98a824c53b2de7d458

Observation 2144d9ef-0bdc-4bc4-bd6c-b0bfee2f57e7 · outbound

This paper cites Acquiring musculoskeletal skills with curriculum-based reinforcement learning.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Acquiring musculoskeletal skills with curriculum-based reinforcement learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.304452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.897911Z digest=sha256:59ab78576061c30a043fe9dd0f148e90bddadc06d7587ef67a0119574fd5140e

Observation bcfd542d-36e5-436f-b60b-80e626029944 · outbound

This paper cites Toward a science of computational ethology.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Toward a science of computational ethology

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.291378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.919806Z digest=sha256:5f16ef91e817254958989208012c8c3b15747adfe1c0068e2045e8a021510988

Observation 1c39d3c3-020b-41b9-9f61-2a80690ab4a2 · outbound

This paper cites MammAlps: A multi-view video behavior monitoring dataset of wild mammals in the Swiss Alps.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models MammAlps: A multi-view video behavior monitoring dataset of wild mammals in the Swiss Alps

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.940728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.940728Z digest=sha256:e2b18325ddfca8d4d1401b30c752c392fa5b173ceb1da546504e9f40445d9fd6

Observation 377f150a-c92c-472e-8526-74f35f3bbdce · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.962462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.962462Z digest=sha256:01759c561f47b0c7e68a6d29613988d87e8406c87377396988a17a456abdabc7

Observation a45239bf-2396-4784-a3b2-d1bde95f9ec2 · outbound

This paper cites Lmms-eval: Reality check on the evaluation of large multimodal models, 2024.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Lmms-eval: Reality check on the evaluation of large multimodal models, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.279504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:34.999609Z digest=sha256:0df6808993f624a90684579df9fdf2524f6458a33eb22f8f066a3037e72bcedd

Observation 300e4bd1-d8a0-47c0-83f6-bff07c8fbe4a · outbound

This paper cites Llavaction: evaluat- ing and training multi-modal large language models for action recognition.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Llavaction: evaluat- ing and training multi-modal large language models for action recognition

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:35.021183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:35.021183Z digest=sha256:53e2fdd8b4441b0b9a9beaed9277b0570d1c47ad1a8a87055d1ca05b9a5082ab

Observation 251481f0-354d-48d7-a8c1-9c781350e6d6 · outbound

This paper cites Favor-bench: A comprehensive benchmark for fine-grained video motion understanding, 2025.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Favor-bench: A comprehensive benchmark for fine-grained video motion understanding, 2025

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.267974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.042341Z digest=sha256:a8a9922857b58a76bd0f645bcf707c8293f882c27bc2b185255bdddee942c8b1

Observation 70836335-fa58-477f-b72b-0a9e53017727 · outbound

This paper cites Introducing Gemini 2.0: Our new AI model for the agentic era, December 2024.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Introducing Gemini 2.0: Our new AI model for the agentic era, December 2024

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.255203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.074791Z digest=sha256:f22f2bf3d4e8d7beae5e00db02ad8b13a9252b829dbff9405fbee8ce9fffaaca

Observation abb8efd7-4a01-4324-94e3-b76cdcffce03 · outbound

This paper cites Posegpt: Quantization-based 3d human motion generation and forecasting.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Posegpt: Quantization-based 3d human motion generation and forecasting

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.241714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.116961Z digest=sha256:245fee4d7859144ae63a074a24adadbeac6e9d89b4dd32cddff3258d04aba648

Observation 16c58b1d-d306-42dd-abe3-585c2166456c · outbound

This paper cites Chatpose: Chatting about 3d human pose.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Chatpose: Chatting about 3d human pose

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.228351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.150139Z digest=sha256:7fdc08c3f3f1646f839e8b69e23127206e002aef6e8a87b61b97786bf73dd771

Observation c52bbcfb-dd6a-4428-a307-b01dbb87ef5a · outbound

This paper cites The language of actions: Recovering the syntax and semantics of goal-directed human activities.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models The language of actions: Recovering the syntax and semantics of goal-directed human activities

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.214576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.171141Z digest=sha256:d07862b9261a47490c558395a6e0d8a600da96663ac6556a4ab8dc21d40a3a33

Observation e43a664a-75a1-42ed-9acb-04c8cfcb8222 · outbound

This paper cites Combining embedded accelerometers with computer vision for recognizing food preparation activities.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Combining embedded accelerometers with computer vision for recognizing food preparation activities

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.200865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.188375Z digest=sha256:30a9ba71714590952ab85bffdbdc0b696c19d468451abd90bf9a44ee7c08ff66

Observation 73059ca1-7b34-4031-ae64-cf5360e6bcb0 · outbound

This paper cites Ms-tcn++: Multi- stage temporal convolutional network for action segmentation.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Ms-tcn++: Multi- stage temporal convolutional network for action segmentation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.188014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.193512Z digest=sha256:1f8d4f0469e1867ec8ef92eeb5161447327353bea8cee0a75782e571251471d3

Observation b2ffffaa-df10-4027-8c01-ce49092e083b · outbound

This paper cites Flynn, Rene Vidal, Austin Reiter, and Gregory D.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Flynn, Rene Vidal, Austin Reiter, and Gregory D

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.174858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.197386Z digest=sha256:314f32474dede63ab86fbe74b914f2e9d5d19a25d75f332f0e6c3582d9ee22ee

Observation f5ba2dfe-03ce-4e1c-9f8e-3193c011187a · outbound

This paper cites Iterative contrast-classify for semi- supervised temporal action segmentation.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Iterative contrast-classify for semi- supervised temporal action segmentation

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.161016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.202345Z digest=sha256:85d59c4ad5b4eb31e0de6d8310913298c2b716bf8181b00ee7a6c9c9f887f7c2

Observation 3d17b358-df35-4c42-979a-42c00d4891a2 · outbound

This paper cites Learning transferable visual models from natural language supervision.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Learning transferable visual models from natural language supervision

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:35.206622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:35.206622Z digest=sha256:6afbc1ee8fc0f7da6807eb193ed411c14288c0f9c863db8afaf5ae692bd71aca

Observation 031c0866-474f-48f0-a915-433cf3211cd8 · outbound

This paper cites Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.140026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.211291Z digest=sha256:741b6f2ef49b9e59b4d5c179a2f5595d9388b40f077fc1f60bfa355b40fd1f7d

Observation 3f7f4a61-3dd6-407c-b1cb-2d66c27f1977 · outbound

This paper cites 4M: Massively multimodal masked modeling.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models 4M: Massively multimodal masked modeling

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.125604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.215052Z digest=sha256:ff0dc4ebabb0e4475d022ae62d4995445be42eff2db2ac83c270eab63259d344

Observation c0e2f556-b8f2-4823-a6fc-b6815cff9d6a · outbound

This paper cites Github, 2021.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Github, 2021

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.112527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.219410Z digest=sha256:7c61e66132b4e5637eabeaf487835ba7ab66893c830eb534f6f5c8a5b8732caa

Observation f2f72c66-deb8-4943-ae04-954cca9e966e · outbound

This paper cites Microsoft coco: Common objects in context.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Microsoft coco: Common objects in context

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:35.223237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:35.223237Z digest=sha256:ee792743fdd4b813a1b0f0af7563adbd8ab51f80bf0671bc2e89c5eda3111e61

Observation 7a023a43-733f-42d4-8978-895db140fefc · outbound

This paper cites The montreal cognitive assessment, moca: a brief screening tool for mild cognitive impairment.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models The montreal cognitive assessment, moca: a brief screening tool for mild cognitive impairment

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.090198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.228587Z digest=sha256:9d6cb5a29c766e7124bb9b8c55ce24c4a796a2d197e3fff335fbe515fc46d799

Observation fd5407f8-0a2f-4b7e-9cd5-c90569f64a49 · outbound

This paper cites MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:35.233869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:35.233869Z digest=sha256:5b44b005990b88b568b2213515f9b4fcc8d818bf33c3f10e5b19de5ac139589e

Observation de23679b-d42f-40b7-9874-f7aa71241d79 · outbound

This paper cites Omnivore: A single model for many visual modalities.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Omnivore: A single model for many visual modalities

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:51.075459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:42:35.239273Z digest=sha256:a7eebf6325d48f00b7e6ab83ba93b8c1e1f4ba15fcbb4ca584aecbd2a68b5d4f

Observation 2785d080-d669-4d40-a479-088c7aa28f2f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:35.244016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:35.244016Z digest=sha256:c1574348b4476b12d05acd097016fbfb1cd770a62b5a108206da336e0d985edc

Pith citing papers

Observation 351ce3f8-4f59-421d-b968-78714f8f3100 · inbound

Towards Human Motion World Models via Executable Behaviour Representations cites this paper.

Towards Human Motion World Models via Executable Behaviour Representations EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:10:09.774872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T04:54:29.565609Z digest=sha256:b846b61d560bc7a7931a3dcadf425f994dad6163d9fb3d367f7b0d5cd35af28b

Observation e5600034-30eb-4550-b38c-61bb99302148 · inbound

REACH: Hand Pose Estimation from Room Corners cites this paper.

REACH: Hand Pose Estimation from Room Corners EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:34:42.615844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T07:34:11.107131Z digest=sha256:60716a12abcda02304e743e03bba1254494f639b9510fe6c806b7e1e7a408736

Observation 5a1432c9-2836-4989-a452-f3418bebee9c · inbound

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views cites this paper.

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T23:27:38.837005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-10T23:20:12.870365Z digest=sha256:f3e7bf875cde11a2e07190d620c83b85d63f7479a449090ead63a4ad24d61340