Pith. sign in

Paper Citation Record · LEDGER

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

As of 14 August 2026, this Paper Citation Record lists 100 of 132 outbound references and 2 inbound Pith citation observations for arXiv:2411.08753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08753 v4

Coverage vector

measured 100 of 132 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:27:16.884770Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:45:38.441181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T06:00:58.843605Z

Reference resolution

100 of 132 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb025d11-7abc-45fa-941e-be55fa65712e · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Deep Learning using Rectified Linear Units (ReLU)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.239636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.239636Z digest=sha256:a5ac7ea01f61df1da134db392d01bd1aefdf4067cf8dc5da27e1b0c298bc3ffe

Observation bbbd5838-3012-4a00-bed1-68e5f1a12eb2 · outbound

This paper cites McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.246965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.246965Z digest=sha256:c4889203d8f2cfaf9b4de09bfa749972c0d75c09a0c39953362ac7f2919fd7f6

Observation 5f7e63c4-d784-42c4-9cd6-80c062bfff39 · outbound

This paper cites A dataset for develop- ing and benchmarking active vision.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos A dataset for develop- ing and benchmarking active vision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.253155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.253155Z digest=sha256:93f726a31e43cb6edfa53212a8bb104935ecab73b72ab33c4fd4ee615b252801

Observation 5f35b769-c0b0-401e-b12b-d158305d7e1d · outbound

This paper cites Automatic editing of footage from multi- ple social cameras.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatic editing of footage from multi- ple social cameras

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.258665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.258665Z digest=sha256:af2cea394c9428b9e076c7a4902b0da8c943b5a99fc5979a9ed74b470e1e3bcc

Observation 2b809dd6-454e-4425-a264-bae399839d18 · outbound

This paper cites an unresolved cited work.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.266710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.266710Z digest=sha256:625beec397a7899370d0d94c0884a55d752af6621126ba7a8ac8e8c8619a120c

Observation cc8180e4-325e-4e9e-8930-f4a1fede591c · outbound

This paper cites WeaQA: Weak Supervision via Captions for Visual Question Answering.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos WeaQA: Weak Supervision via Captions for Visual Question Answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.273850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.273850Z digest=sha256:4111a0953b7dab713715a17154ce2d2b41e4ca5e8b36079ac204fffb1ebb934a

Observation a3137ed3-695d-457c-b997-5a27ba7b661f · outbound

This paper cites METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.281097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.281097Z digest=sha256:ba1c13409fb7f625a2b8007ac9a499deafcf2815035f636fc6223641c780a1d9

Observation 1fd79043-4e34-4a2c-892b-dd4ab3137967 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.287219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.287219Z digest=sha256:d46af5f323db0e70dc125b0de498891c426434dac1760dbcefe8254ad1c7ef40

Observation 71bc575b-f19d-4b46-8a5c-939063e2b0b4 · outbound

This paper cites High- lightme: Detecting highlights from human-centric videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos High- lightme: Detecting highlights from human-centric videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.296787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.296787Z digest=sha256:c222505ce1895bf2a873477bcbdd6a5d593f67dccaf2152af2c7b481cd7265f8

Observation 1970d812-711b-4c45-a19f-73c8b38c4a28 · outbound

This paper cites Extreme rotation estimation using dense correlation volumes.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Extreme rotation estimation using dense correlation volumes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.304503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.304503Z digest=sha256:4685664ef4cb8cee0189ffe8fadd60a7fde9625046a05d30bfda11d49307ec67

Observation 0a0c9b5d-fb6a-4dba-a770-a30c788473ae · outbound

This paper cites Davis, and Lei Zhang.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Davis, and Lei Zhang

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.312059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.312059Z digest=sha256:76ca4988a27b124f9fa39dedaea182cec65afccc6bf6e9c9b29d2368e09d3cc6

Observation e8215ca4-bb01-45c4-917f-b677ab0198c0 · outbound

This paper cites Enhanced interactive 360° viewing via automatic guidance.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Enhanced interactive 360° viewing via automatic guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.318489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.318489Z digest=sha256:746dee1efe9ae1e52003aca87bd9d85c4937c432cc61df65b19dd6e2eda8a4b6

Observation 20ccd794-4833-45b6-95bf-6553e0024242 · outbound

This paper cites Learn- ing sports camera selection from internet videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Learn- ing sports camera selection from internet videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.324741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.324741Z digest=sha256:8719b722845df5fa942ca8e0ad77ce6fb9f7f199ceeb21854aca47d41ec9be3f

Observation 9f637a90-1a36-4165-9a78-1202d46acbd4 · outbound

This paper cites Wide- baseline relative camera pose estimation with directional learning.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Wide- baseline relative camera pose estimation with directional learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.331782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.331782Z digest=sha256:c44c58127f516b0985261dcacef7b196da365c1b04698dbe4c017e3886498b69

Observation 3899a749-9f6c-48a5-bf29-582af09d9ade · outbound

This paper cites Geometry-aware recurrent neural networks for active visual recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Geometry-aware recurrent neural networks for active visual recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.337905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.337905Z digest=sha256:36f064cf95dc7307b83dc11a55c4b579595948160722f2df9918ab19b554d1f8

Observation 128e8bf9-6538-4aff-8651-901cd8535971 · outbound

This paper cites Towards a richer 2d understanding of hands at scale.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Towards a richer 2d understanding of hands at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.351292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.351292Z digest=sha256:d6503966ff2f86421278601f719086546986329a0101338c3e779fe6d69b6edd

Observation f512be94-a5f9-47ad-9c48-3fe5ebc5064b · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Gonzalez, Ion Stoica, and Eric P

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.357941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.357941Z digest=sha256:a16e3b1959008f65e3645cb2fa5b89e9260b5b41e594e8d039716efd6c8c8838

Observation 85501816-c5e0-4811-a324-9150f4030214 · outbound

This paper cites Self-view Grounding Given a Narrated 360{\deg} Video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Self-view Grounding Given a Narrated 360{\deg} Video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.365365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.365365Z digest=sha256:55ca3475897c480005e6028038caa490604e3ea341666646b0e42b3414c7bb1f

Observation 0b7dded2-0479-4e31-860c-2989caae0cf8 · outbound

This paper cites Video co-summarization: Video summarization by visual co- occurrence.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video co-summarization: Video summarization by visual co- occurrence

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.374896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.374896Z digest=sha256:0c0bb74c1b16ac770a8d9cc281d8f0713950afe10ef6a88498e54326b89c6544

Observation 7ab7f5a7-5e64-401d-b3e0-1b9d56e9b9a2 · outbound

This paper cites elochoice.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos elochoice

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.381981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.381981Z digest=sha256:09246301e6a7e5a7fe9eac8636182d70dc5d46a998090370e8ce8fb1df309439

Observation 15cdcf51-af90-48d1-9276-bccd1ddac3ad · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Scaling egocentric vision: The epic- kitchens dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.387951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.387951Z digest=sha256:3b6c09243053e1d50f0c5489a314815e5c82b7306152f68faf6de09cb962ee1f

Observation 13f61efd-951b-41cd-a55b-661009310c20 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.393588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.393588Z digest=sha256:472c6abb78334e526262ec1c0a7d5e2be93fe99a127ba9253cffde58ece4a3c7

Observation 58020f30-05f1-4628-8f05-3ada0da5fc75 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.399483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.399483Z digest=sha256:e6507cc070eeaf98bb9b92ba8bfa190383bd8fc22997a6456de83bc6eb7a7b5d

Observation 9cb4724f-b0dd-49f3-825f-fe2f18162725 · outbound

This paper cites Velastin.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Velastin

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.404733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.404733Z digest=sha256:d876275bf47ba94f27a02ee1b07246c712a8a3061fdb910fa3b43b83998acf96

Observation d0be0b1f-e948-4cb3-b0af-4b86378f439e · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Virtex: Learning visual representations from textual annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.410436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.410436Z digest=sha256:c0906390f6736cc808a04998e8d7ade2ce6ee7da5babdc671b49e05b5ddbebd0

Observation acde62f6-cf94-455a-bc65-539a3d3f779d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.416234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.416234Z digest=sha256:c815b73d0adef44ec2807a83182abad09315c61b7423d9dfcb1b59e13f508f13

Observation 786245bb-34cc-4782-b6da-f8e960b5843f · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Dense and aligned captions (dac) promote compositional reasoning in vl models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.423479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.423479Z digest=sha256:fd3b7fe25e140fa7c63c797f804c4be76ee768606fbd908ce35bce3486f39665

Observation face0574-4d45-4a9a-b547-52f30f5a2cb4 · outbound

This paper cites Multi-view active fine- grained visual recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view active fine- grained visual recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.429622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.429622Z digest=sha256:71cfd38fd096ca83dc486ed2951d51a3869e3af15efc7b10bbf752ea0450e52b

Observation ae7829db-5a61-4529-b7b2-53d081487132 · outbound

This paper cites Multi- stream dynamic video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi- stream dynamic video summarization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.436093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.436093Z digest=sha256:0fa1f60ff50aeee498d2d5092461bbe2d0eb598b297d2b0db47261c1c2408f16

Observation 6c41b363-bd5c-48fc-8940-827deb8740d4 · outbound

This paper cites Elson and Mark O.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Elson and Mark O

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.441602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.441602Z digest=sha256:9300c5f88be604ee736d287949439e4b52e9639c67d5681183239a8c772bae60

Observation 20a10721-ce03-4f24-8f5d-68767cb25750 · outbound

This paper cites Foote and D.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Foote and D

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.450299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.450299Z digest=sha256:eb67ab0d9a526529e1b571d20e0b661ac40b73c35a54055a875aea6357875302

Observation 0904ea70-ee80-4a76-9aac-ec3227472dea · outbound

This paper cites Multi-view video summa- rization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view video summa- rization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.457120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.457120Z digest=sha256:9bc714ee2370b114e8ba8cc14d1e5ef7110ea561a1ef77078fe3fd6ba899dbe1

Observation aea03a6d-8bf6-41c2-823e-ae6aaa5262d3 · outbound

This paper cites Gleicher, Rachel M.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Gleicher, Rachel M

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.463492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.463492Z digest=sha256:31ef3e323a94f0e0f6e5aa62f95e88ae926e2c04ebb5ea8456d6ae91808ac9c4

Observation 69cd9eef-ad25-4bde-be6c-0f56da3ca3fa · outbound

This paper cites PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.469572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.469572Z digest=sha256:f14bd8452c980dc94105a203228fa0484e1ff6dc1cf28bfeedd6e0d1fc62b32a

Observation b399e38c-c960-45a5-8172-81faef191d2f · outbound

This paper cites Diverse sequential subset selection for supervised video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Diverse sequential subset selection for supervised video summarization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.479571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.479571Z digest=sha256:5ba7f92e4d4c5b6742e684e546054833fb08f7215664a88487fdc9acbf281798

Observation 1fb9a401-425c-4436-beaa-101ae1ca7ace · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.485303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.485303Z digest=sha256:054e7d2129483c6e89e52af49b83529f6a59d4433421d8ade22d536a9f37341d

Observation c0b2dc9a-8d8c-43ed-9738-c519d6a0c022 · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.490447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.490447Z digest=sha256:bd6026b94c54661847fb60d0843d6819ecd1722ec2c35937ad10d7ef9e1ef2dd

Observation 36ada9b3-36b4-4a4b-bdaf-73478b66a597 · outbound

This paper cites Temporal Difference Variational Auto-Encoder.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Temporal Difference Variational Auto-Encoder

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.496223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.496223Z digest=sha256:7e4e5c8896ba30ba1880745232d43532fb43a815e2f2601bf20f7b6f244a6f34

Observation 82cf2f02-60ce-48a7-b2b7-45c7745d63b4 · outbound

This paper cites From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.503459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.503459Z digest=sha256:2e07b27ba7823a774f7366e245df877f444549ec54afc5eff62d74ba44947650

Observation bb7f5658-868b-4da4-8003-6fea52ca632f · outbound

This paper cites Using closed captions as supervision for video activity recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Using closed captions as supervision for video activity recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.510211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.510211Z digest=sha256:26f1b9259a5f6fdd80c963b7846933f454bab9f53407eb3959d51f5cffc7cd5b

Observation ee3b7321-6456-4c50-a586-d310a09ced60 · outbound

This paper cites Creating summaries from user videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Creating summaries from user videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.520424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.520424Z digest=sha256:d5dcaaf6a4f0f8b4c1c672f4fa24e4d67b91b1d26d0509ee724c4fc81a572f8d

Observation adf155bf-20d6-4c5c-b6cc-30895c8ea6ba · outbound

This paper cites Video summarization by learning submodular mixtures of objec- tives.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization by learning submodular mixtures of objec- tives

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.533989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.533989Z digest=sha256:ad2c21ab0d3b189d91aec4d8c34cd991fc036503f4c53570d9430d2856e2927d

Observation 51c0c39d-eb07-4531-b1e0-32883fb4dd9b · outbound

This paper cites Align and attend: Multimodal summarization with dual contrastive losses.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Align and attend: Multimodal summarization with dual contrastive losses

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.539969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.539969Z digest=sha256:65a51782c753a22db85015e1d45403758c157417d56f51e935e311c47c07c751

Observation ccfe8115-fc24-4cb6-8957-db3605f10956 · outbound

This paper cites Cohen, and David H.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Cohen, and David H

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.545269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.545269Z digest=sha256:21712ba25eb51eb8361c4a2c82fcad6614965a6f3f9a895ce42402b3af54f14a

Observation 4771a9c2-2ac3-4484-be01-c8bc7857d19c · outbound

This paper cites Cohen, and David H.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Cohen, and David H

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.551100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.551100Z digest=sha256:2f25c9274095826567275e788416826f4ab2c0f5d7a49590a6e5437d3b7d8413

Observation 772d4bde-00fa-4203-a507-cb95825b6041 · outbound

This paper cites Vir- tual videography.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Vir- tual videography

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.556351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.556351Z digest=sha256:d279f8e6e4d586331825467d2ff3fd1efead3d89547010078313e6bd0a0cb150

Observation 00f95f17-6312-4f34-91fb-bd507296c3bd · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos LoRA: Low-Rank Adaptation of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.561315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.561315Z digest=sha256:8185fdd137410835add490f18218cdb3217b1b8e3fe5effa9872ad0d3551ac3e

Observation d38383b6-f858-4707-a017-1dede39c461d · outbound

This paper cites Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.566908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.566908Z digest=sha256:394f7219ffa1eae3cce789b608eef53e8e329d32217a9fbd6c8db3dd28ff6c3c

Observation eaaccbea-74bd-4aa4-a9a1-0ed54ce64a55 · outbound

This paper cites EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.572054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.572054Z digest=sha256:5f6b45bcfa313f87b12317a068fc373c08d608444a0ed68e98111d3aac17e464

Observation d3df167d-f2ce-44a1-97d9-fcefb64f1699 · outbound

This paper cites Batch normalization: accelerating deep network training by reducing internal co- variate shift.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Batch normalization: accelerating deep network training by reducing internal co- variate shift

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.577856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.577856Z digest=sha256:2806d5ad66989c43bad3e810ab542d84ed2ebf43d1390c0ca21a20d1147918d5

Observation c0397a64-a4a6-40dc-9e1d-f306b0a6255a · outbound

This paper cites Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.584139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.584139Z digest=sha256:046593fad8b9dfecfb57a9e11dfeb498f4cd9f9339291de399a1813985a7a27c

Observation 8ea4e95b-7769-44eb-b8c4-b79cb4233be3 · outbound

This paper cites Learning to look around: Intelligently exploring unseen environments for unknown tasks.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Learning to look around: Intelligently exploring unseen environments for unknown tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.589659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.589659Z digest=sha256:34e4bbcecb257b60845b5923787f0a432d5dc0c222ca25ba86a2da0c7971694a

Observation 512bb144-e0c5-4bf9-a5f8-977456cfdf53 · outbound

This paper cites End-to-end policy learning for active visual categorization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos End-to-end policy learning for active visual categorization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.595297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.595297Z digest=sha256:7a47b0328a232ea769ad94b3af377e5066b411178d06ac45b5cb3fda372d3bdd

Observation 2835bd81-9f36-41e4-a29d-cbf613bc3efc · outbound

This paper cites Time-Agnostic Prediction: Predicting Predictable Video Frames.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Time-Agnostic Prediction: Predicting Predictable Video Frames

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.600840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.600840Z digest=sha256:d7909430e455d6af88a8f4495bafd38287fa831be26a5cd11529c7c58c2891c3

Observation 5c8021b6-ddd4-4df0-a69b-51c0e2d0c7ce · outbound

This paper cites Simglim: Simplifying glimpse based active visual reconstruction.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Simglim: Simplifying glimpse based active visual reconstruction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.607025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.607025Z digest=sha256:2c38a850a907a679760853c70366c6f4c048c77047c863a8ef840918837994c7

Observation d3adfba1-3679-4606-90ce-c0a8591e3d1d · outbound

This paper cites Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.612958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.612958Z digest=sha256:d86733e8afa7a3a6b6b93d7f818e5ee117ee96b7b2d0234f79a2e45dd8256a00

Observation 75304319-c3b5-4630-854f-269cd0168e43 · outbound

This paper cites RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.619160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.619160Z digest=sha256:056836e73c0edb30ab73f8286500e9ae34086d04287573ea38cf7dcf5e2280b1

Observation cece5df7-2619-49f6-be83-94a62c0fa7f0 · outbound

This paper cites Large-scale video summarization using web-image priors.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Large-scale video summarization using web-image priors

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.624405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.624405Z digest=sha256:e2ed776c785afa4f1a768d35137ad8514ebdc3645dba18f132e2071505cb7afc

Observation 4afe155d-955c-48ed-bbcc-329d56768cd9 · outbound

This paper cites an unresolved cited work.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.629375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.629375Z digest=sha256:2605223876ba3f4f89ef1f0aed1bdcdb90fc24dfde16acdaf40a2692c676a5bd

Observation cb661ab1-ee16-4e98-81ab-c87070f9aa7b · outbound

This paper cites Segment any- thing.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Segment any- thing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.635697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.635697Z digest=sha256:71a5254dfd4acea9df923b9efa331e2db884ebf207b1336491656178cca93161

Observation 883c89c2-d078-4ec9-9dee-f367fe283eba · outbound

This paper cites Hyperbolic Learning with Synthetic Captions for Open-World Detection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Hyperbolic Learning with Synthetic Captions for Open-World Detection

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-12T21:27:17.290381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.641426Z digest=sha256:a4ee457a8ac1a7abfabdf940e36ec3183d1c5cb5acad17a0420a27ca327d1a00

Observation 41dfb3fc-a295-4cc4-895d-81fc3c3afa56 · outbound

This paper cites A memory network approach for story-based temporal summarization of 360° videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos A memory network approach for story-based temporal summarization of 360° videos

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.650230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.650230Z digest=sha256:11da35eff595815986a60f99c3eb157f2366db645f983b6e438262fad08afa10

Observation f6e45f0f-680a-4afb-9c6e-ea8b2a8a707a · outbound

This paper cites Predicting important objects for egocentric video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Predicting important objects for egocentric video summarization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.657826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.657826Z digest=sha256:c12f5d758667f6c5cebbb7907135727150b6ee8b0e009bd820b77d9f0ffed27b

Observation ebfefeba-e6cc-4e98-ad04-5e03f150fd7a · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.663653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.663653Z digest=sha256:a5ea4b51bba89790ddceac789580500b57cd740df0aa6fbe9c960980a37c2b3c

Observation e15fd5d7-482f-47e8-9912-453c1c504e1c · outbound

This paper cites Grounded language-image pre-training.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Grounded language-image pre-training

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.669888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.669888Z digest=sha256:f22ddbc7cd8560dc577c12592185b32448bf29340554f0639833b00acc07b085

Observation cb28cbc6-1078-4f06-883f-5b54aacfb4cb · outbound

This paper cites How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.675226Z digest=sha256:b95a40d23f0bfbc2b3efd1cb400d3f20251a180a0f0f62151edddc729be58b8d

Observation 756366c4-920a-4fa0-adff-8122bda308df · outbound

This paper cites Egocentric video-language pretraining.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egocentric video-language pretraining

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.569353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.680256Z digest=sha256:bf2286e95733f2724dc7a23f9e6e55516e5ae956a922e96cedda4c0a9b97cbdf

Observation cadd1a59-85f1-4e16-aa5a-d1d422355b18 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.685434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.685434Z digest=sha256:f6347551a7b19358bd1cce4d0b6d4b776872ac04d2ba9eb4811d10f1f272a151

Observation d116433c-7261-411f-861d-40dd7605c0ca · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.691686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.691686Z digest=sha256:a1183040a4c0a1830f48246f46430bbe39f62756e0b8d56d8e619352a93f9117

Observation 30c3a2d4-49dd-4644-9a7e-f73abf1a3b68 · outbound

This paper cites Decoupled Weight Decay Regularization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Decoupled Weight Decay Regularization

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.697244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.697244Z digest=sha256:224247ae6af9ea31e320588611d9fde132d1bd9aad8501087d7ed7c5a2c1f56a

Observation 5fefd853-04f0-49ae-bd8c-3265e7192f38 · outbound

This paper cites Story-driven summariza- tion for egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Story-driven summariza- tion for egocentric video

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.549057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.702606Z digest=sha256:535a4a9df68b64c2b78a2f5306d3841c0662aa953ebf2fad64d8cb9d78b91b6e

Observation a9ae5bc4-44f2-43e1-b13f-e77415b31e02 · outbound

This paper cites Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.709927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.709927Z digest=sha256:1b86c5afc4b4b1d7505a1f36b5eb3b9939cbdc87b97797552ff14090e5890d89

Observation 04784fba-8090-4375-bf5d-e45497f5a31e · outbound

This paper cites Video summarization via multi- view representative selection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization via multi- view representative selection

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.525516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.715943Z digest=sha256:de959bf2ca2dc4ff38abb989d49776385b8c5c38871e1b7f991f3e24c0ce19f4

Observation fe0807fc-36ba-4265-b1fc-4fa9b5e49f6a · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.496074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.721421Z digest=sha256:23ac24a73dc3db713b71a234c36c0c30b7c0fbf1a82fff46f95576b34cf3edf2

Observation 8106cd4f-a3de-45d2-9b5d-29a59d8cef9b · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Srinivasan, Matthew Tancik, Jonathan T

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.726553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.726553Z digest=sha256:954f17338e4e6a9cbc364954c67b0b69911e714adb063423a574160ce014f476

Observation 03ae2a1e-6682-4d9f-a1ec-1b14efc1a1e2 · outbound

This paper cites Automatized summarization of multi- player games.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatized summarization of multi- player games

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.454733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.731426Z digest=sha256:a8ccff22b5a80c7db93d9a8fe7ae84cd4c3da666da9268b5b503a48f711a9796

Observation 61da99fb-d31c-4453-92b7-c0c58c127885 · outbound

This paper cites Egoenv: Human- centric environment representations from egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egoenv: Human- centric environment representations from egocentric video

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.410732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.745576Z digest=sha256:b489a345db2f9dc9ef54a77c9956f1dd82eab24a23ed16cae1d6db6faf652545

Observation 67c3248e-d510-4503-9440-8e7f5cc3a3bb · outbound

This paper cites Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.385894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.750940Z digest=sha256:63488738526613d80008c8923a7ea006e0955eac3b8ab9b27f600bb63ed7ab51

Observation 82e5f804-aefd-4974-b8dc-ef24f38a6c01 · outbound

This paper cites Adaptive skip intervals: Temporal abstraction for recurrent dynamical models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Adaptive skip intervals: Temporal abstraction for recurrent dynamical models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.357873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.756958Z digest=sha256:dd24698bff89b5740dab88fa95d40ac5035d5d85550e08704aaf0cd710bcd923

Observation 3e63b266-ec5b-4f6c-abe9-3c919f19deda · outbound

This paper cites Au- tomatic video summarization by graph modeling.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Au- tomatic video summarization by graph modeling

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.336833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.763234Z digest=sha256:ba219f2aa456947e6cb0694c8b3ea7e24cac8af0605a7fa984c1a844458fcfe4

Observation ad9bab32-aae4-4415-8174-2a602a225541 · outbound

This paper cites Collabora- tive summarization of topic-related videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Collabora- tive summarization of topic-related videos

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.316696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.771983Z digest=sha256:1a31dae75c921d3e554efdc9452ed41c8ede41a6f93a6889810e716cdfc60716

Observation 1deecac3-7788-4b32-ab77-5d47a72a0932 · outbound

This paper cites Multi-view surveillance video summarization via joint embedding and sparse optimization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view surveillance video summarization via joint embedding and sparse optimization

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.295299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.778053Z digest=sha256:a92ee250b06fd101853e39ee7dbc499b9ff75acf8298cf65d1a08188a53e6a1a

Observation 8f1a830f-752c-4357-a063-ac2da8b6188e · outbound

This paper cites Roy-Chowdhury.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Roy-Chowdhury

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.274325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.783830Z digest=sha256:6c455fe1978be7307c572036969314f10a6915a72b6c7120f70d8d3748e844cb

Observation c21a3630-06c2-4b52-af5b-4584cb1ac0f1 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Bleu: a method for automatic evaluation of machine translation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.257179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.789486Z digest=sha256:d92a079efdf7f45ae6204163bef3e9858377e2cf45b6d2919f243f0af2bd6137

Observation 5a1a688b-0e61-4405-8da4-40835fb472d8 · outbound

This paper cites Sumgraph: Video summarization via recursive graph modeling.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Sumgraph: Video summarization via recursive graph modeling

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.238943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.795037Z digest=sha256:5a0fead504137fe746fd0970b09b9f1fab38b76232a8babab0a51efaf53105d2

Observation 7d7ac8df-72f4-4c62-9db1-cc869eb1d79d · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.215657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.799995Z digest=sha256:ec29a0ea27df3dadbe02b2e0e38998db37d570c3d6a68b31ac668e0aec63cd23

Observation 378c0f34-b5b3-4334-b4f1-51e9efeac2e5 · outbound

This paper cites Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.193781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.805015Z digest=sha256:2784e87e40b45b23dc6df41d38139b4deca9f4524b54f136f3692ab4d18e1718

Observation a4a08e10-1781-4d09-a398-42aed2923331 · outbound

This paper cites Sidekick policy learning for active visual exploration.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Sidekick policy learning for active visual exploration

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.176412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.810155Z digest=sha256:a53e77bd93e49d8b642019f24ac05c69846339aa5d1acd0676f116b291d1ca2f

Observation af605da0-9852-40f0-bac3-870705489a30 · outbound

This paper cites Emergence of exploratory look-around behaviors through active observation completion.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Emergence of exploratory look-around behaviors through active observation completion

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.160511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.814977Z digest=sha256:a834e195bc503d44da3e69a8f23cb53eb35e77e092503c67d8f8a152a4c8b7af

Observation e496eb70-24af-4f73-94f6-c2f539eb4e2a · outbound

This paper cites Naq: Leveraging narrations as queries to super- vise episodic memory.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Naq: Leveraging narrations as queries to super- vise episodic memory

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.143654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.820501Z digest=sha256:f0736283db3c10ca4b7581483bc8bb70b6ad098b128f402a1983c68f41017d44

Observation d19f8cfa-e3cc-4253-ad30-7ea6ae3fa489 · outbound

This paper cites Video summarization by learning from unpaired data.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization by learning from unpaired data

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.126653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.829184Z digest=sha256:2499e86acb8594bc52366244411c4130bed9535d6c1e4706361cae4ed47d22d2

Observation 0d4725b1-c25f-457f-9870-5f756309db84 · outbound

This paper cites Adaptive video highlight detection by learning from user history.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Adaptive video highlight detection by learning from user history

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.835818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.835818Z digest=sha256:605031bea5ca7f588dfd930ba06406a127a8f083d83e75f8d7e8085997f70379

Observation 2fc92f75-2740-4f13-a654-84a748667606 · outbound

This paper cites Chowdhury.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Chowdhury

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.097889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.841151Z digest=sha256:b5bbe042088be0ceb2737ed086eeb0349704719d670f402d94935390227d166d

Observation c5a9690f-59df-487e-b0b1-bb2c99d7e9e3 · outbound

This paper cites Attend and segment: Attention guided active semantic segmentation.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Attend and segment: Attention guided active semantic segmentation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.846399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.846399Z digest=sha256:d2072b2f32f4d9d62effa97c1e4021e56aa1c8833545b81f85c989e7e22c801b

Observation 121c8f1a-836b-4abb-a24b-3acd089351ec · outbound

This paper cites Glimpse- attend-and-explore: Self-attention for active visual explo- ration.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Glimpse- attend-and-explore: Self-attention for active visual explo- ration

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.070161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.851964Z digest=sha256:ef5918902104dab6c37099af70de8807915877dac2c516f7ddeefc1d8ad3c21c

Observation d6b9d89a-40b3-4049-8dfd-16e538200d9a · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Actor and observer: Joint modeling of first and third-person videos

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.052780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.858613Z digest=sha256:9b4729b6ba8c0e0d21a9f6ea3143217eb3381414e57aaa084cce7bb69585ce41

Observation 8481d850-ddf3-403a-8c8c-56f49059cce9 · outbound

This paper cites Tvsum: Summarizing web videos using titles.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Tvsum: Summarizing web videos using titles

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.034387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.863934Z digest=sha256:2a4b157d07ccd6747b43747a1d0da5217564fda6d599d7e1bc8a6306baa27ef0

Observation 6266f243-3dcb-4190-a5e2-0dda93bef23a · outbound

This paper cites Making 360 ° video watchable in 2d: Learning videography for click free view- ing.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Making 360 ° video watchable in 2d: Learning videography for click free view- ing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.015861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.872398Z digest=sha256:9dbebbb485b11b643b84f256c2a97937d2c27df2aa43c1e01add8807cff322b1

Observation 67a9f5c6-e0a6-41cf-9c32-e40ba2a166ae · outbound

This paper cites Pano2vid: Automatic cinematography for watching 360 videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Pano2vid: Automatic cinematography for watching 360 videos

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:17.994590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.878297Z digest=sha256:d9bbd81466710a15c7ab349ef670d0e9ed1e3dd78b516a51792e25c7df27a424

Observation fbe32558-5b46-4a57-9a06-394c4ba8a462 · outbound

This paper cites Automatic con- cept discovery from parallel text and visual corpora.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatic con- cept discovery from parallel text and visual corpora

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:17.975142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.884770Z digest=sha256:aa7a12b49320eea3cffe86056ccb26c7d7eeb7fb143b249bbac9e648ec6e24da

Pith citing papers

Observation 6cd707ae-c6d6-4b32-b3a1-eec071956336 · inbound

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos cites this paper.

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:38.441181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:38.441181Z digest=sha256:f0213939fc110ec8050620f7f9fbc389f62942b7e0f75bc15e500f690c0d708a

Observation ebe94ece-1e15-49a0-bcc6-e24f2cfb1f29 · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

Reference 187

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:00:58.847019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T06:00:58.555825Z digest=sha256:434640818de0c776988fd9ef0ee86348d97fd6e337590c5824eec5cf9d84fc99