Pith. sign in

Paper Citation Record · LEDGER

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

As of 14 August 2026, this Paper Citation Record lists 100 of 132 outbound references and 2 inbound Pith citation observations for arXiv:2411.08753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08753 v4

Coverage vector

measured 100 of 132 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:27:16.884770Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:45:38.441181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T06:00:58.843605Z

Reference resolution

100 of 132 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb025d11-7abc-45fa-941e-be55fa65712e · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Deep Learning using Rectified Linear Units (ReLU)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.239636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.239636Z digest=sha256:3a46220e34570c75ebe29bd4e644c3886162bfe911b816b1cc8679d2553bf229

Observation bbbd5838-3012-4a00-bed1-68e5f1a12eb2 · outbound

This paper cites McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.246965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.246965Z digest=sha256:89313189c9ed9bfea206c9920b383c242c5e76bb14569dadd1b338afd876e3ff

Observation 5f7e63c4-d784-42c4-9cd6-80c062bfff39 · outbound

This paper cites A dataset for develop- ing and benchmarking active vision.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos A dataset for develop- ing and benchmarking active vision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.253155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.253155Z digest=sha256:33213fb86b1cc291c5fb39d58e415da1699ff9b9e4472b572a1a8c69505487b8

Observation 5f35b769-c0b0-401e-b12b-d158305d7e1d · outbound

This paper cites Automatic editing of footage from multi- ple social cameras.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatic editing of footage from multi- ple social cameras

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.258665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.258665Z digest=sha256:9082c2260ff4d9dc380baf95b7c47a13ac1cf63bb482193b8a0c915723de5ff1

Observation 2b809dd6-454e-4425-a264-bae399839d18 · outbound

This paper cites an unresolved cited work.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.266710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.266710Z digest=sha256:63a13581bcac91eb40008eeeac04c3ba671502334a1636695a6787e39b658a23

Observation cc8180e4-325e-4e9e-8930-f4a1fede591c · outbound

This paper cites WeaQA: Weak Supervision via Captions for Visual Question Answering.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos WeaQA: Weak Supervision via Captions for Visual Question Answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.273850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.273850Z digest=sha256:573de93743d75c3060aa6ba4d5e0f849f794b5ec445d4e031998a8783cb354a1

Observation a3137ed3-695d-457c-b997-5a27ba7b661f · outbound

This paper cites METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.281097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.281097Z digest=sha256:5c1ae5285983579ebbbaac025919a85533fbabefb06e06bdc61b8f1d084f8025

Observation 1fd79043-4e34-4a2c-892b-dd4ab3137967 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.287219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.287219Z digest=sha256:04a52d274544ff6e651b5681f7d6fceaa3613630d5349b6d3b90ba157c924ed1

Observation 71bc575b-f19d-4b46-8a5c-939063e2b0b4 · outbound

This paper cites High- lightme: Detecting highlights from human-centric videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos High- lightme: Detecting highlights from human-centric videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.296787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.296787Z digest=sha256:0580fa9a6d74948fe9b7cf094cf578bde2ffe3d21e502a90269cb6cbab8c6e18

Observation 1970d812-711b-4c45-a19f-73c8b38c4a28 · outbound

This paper cites Extreme rotation estimation using dense correlation volumes.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Extreme rotation estimation using dense correlation volumes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.304503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.304503Z digest=sha256:a0e819cb9897fbdc4bef4ac3e1b60b845665991e5902386744ac8bd6cfbd1b4a

Observation 0a0c9b5d-fb6a-4dba-a770-a30c788473ae · outbound

This paper cites Davis, and Lei Zhang.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Davis, and Lei Zhang

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.312059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.312059Z digest=sha256:750f2810cd3bf0493ea17233dd49a3cac7a174fd689c267b77197a0a88e751f1

Observation e8215ca4-bb01-45c4-917f-b677ab0198c0 · outbound

This paper cites Enhanced interactive 360° viewing via automatic guidance.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Enhanced interactive 360° viewing via automatic guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.318489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.318489Z digest=sha256:743df688973e01432d53b914519dffecb6d51cf002e8879a6f190390465a21ef

Observation 20ccd794-4833-45b6-95bf-6553e0024242 · outbound

This paper cites Learn- ing sports camera selection from internet videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Learn- ing sports camera selection from internet videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.324741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.324741Z digest=sha256:fbde3b5e09a2c2cea6125448f9ea3d557bfa7d5c0406655bbeac1c4f0a186a89

Observation 9f637a90-1a36-4165-9a78-1202d46acbd4 · outbound

This paper cites Wide- baseline relative camera pose estimation with directional learning.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Wide- baseline relative camera pose estimation with directional learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.331782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.331782Z digest=sha256:a490c2a93f5bdd9f38320f9f05b335d112ebfb7eeb94bf3d1a8c20c14b65f081

Observation 3899a749-9f6c-48a5-bf29-582af09d9ade · outbound

This paper cites Geometry-aware recurrent neural networks for active visual recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Geometry-aware recurrent neural networks for active visual recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.337905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.337905Z digest=sha256:eb5e5fbbe4a3644d339ccf41835a68b95e847919643866265f66cc2c40a26929

Observation 128e8bf9-6538-4aff-8651-901cd8535971 · outbound

This paper cites Towards a richer 2d understanding of hands at scale.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Towards a richer 2d understanding of hands at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.351292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.351292Z digest=sha256:103b7302d739283b00515943bf38c8376e08cb0c32ebdb578550da087c24c323

Observation f512be94-a5f9-47ad-9c48-3fe5ebc5064b · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Gonzalez, Ion Stoica, and Eric P

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.357941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.357941Z digest=sha256:aae06d3bfd69f99019078b7931da5735b7da3711d55f80d8f8b2f145d9d0a690

Observation 85501816-c5e0-4811-a324-9150f4030214 · outbound

This paper cites Self-view Grounding Given a Narrated 360{\deg} Video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Self-view Grounding Given a Narrated 360{\deg} Video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.365365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.365365Z digest=sha256:33e5d0e020b066394f67c0181a3eba4ae9db062f1b314291f66c8177c309e6cb

Observation 0b7dded2-0479-4e31-860c-2989caae0cf8 · outbound

This paper cites Video co-summarization: Video summarization by visual co- occurrence.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video co-summarization: Video summarization by visual co- occurrence

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.374896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.374896Z digest=sha256:79fc0c777b074d7c1947ce72017f34413987c76aeb552c4343d5d19f6bf56db3

Observation 7ab7f5a7-5e64-401d-b3e0-1b9d56e9b9a2 · outbound

This paper cites elochoice.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos elochoice

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.381981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.381981Z digest=sha256:8f402bf52bb92835fe0f88048d0eb0e36887abfdf79333e53a7447b60e5a2433

Observation 15cdcf51-af90-48d1-9276-bccd1ddac3ad · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Scaling egocentric vision: The epic- kitchens dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.387951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.387951Z digest=sha256:d4a0c64b615cee64c79175abc306407ec826f7d9d707563619228529e7d1ab2f

Observation 13f61efd-951b-41cd-a55b-661009310c20 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.393588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.393588Z digest=sha256:0b102a99205b69b794cee7c7f3f491b9d61619f678a794fc47fdd2707608dc8a

Observation 58020f30-05f1-4628-8f05-3ada0da5fc75 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.399483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.399483Z digest=sha256:bc99ef5ce0c5b1ac559238e1073da45ee7327854e5b31cc6a771d9c3f2ada833

Observation 9cb4724f-b0dd-49f3-825f-fe2f18162725 · outbound

This paper cites Velastin.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Velastin

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.404733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.404733Z digest=sha256:6016dbeb962e291391b80dd9053e486917065d86400b5daae441c6641340e209

Observation d0be0b1f-e948-4cb3-b0af-4b86378f439e · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Virtex: Learning visual representations from textual annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.410436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.410436Z digest=sha256:45c850476e53d1f5608b129aa5496f7c97f17615686c901e201b07a159e9c7bf

Observation acde62f6-cf94-455a-bc65-539a3d3f779d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.416234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.416234Z digest=sha256:0467d0637381af2e3885a2e22c469e12318a9259d401cafcd7f9945b5c07949c

Observation 786245bb-34cc-4782-b6da-f8e960b5843f · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Dense and aligned captions (dac) promote compositional reasoning in vl models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.423479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.423479Z digest=sha256:0abfd733f1a085e46f7d39197504c7ebf297708d34fe40b6ed59baf27f3151d0

Observation face0574-4d45-4a9a-b547-52f30f5a2cb4 · outbound

This paper cites Multi-view active fine- grained visual recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view active fine- grained visual recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.429622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.429622Z digest=sha256:ed56242d98b30738109d7dcf1e1c481e41221a0c692c7d8552d92742f70b47d0

Observation ae7829db-5a61-4529-b7b2-53d081487132 · outbound

This paper cites Multi- stream dynamic video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi- stream dynamic video summarization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.436093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.436093Z digest=sha256:317771f7bd9fb31e5fb5e16b99c9e8ffbf2ac8c7b2155ad4f2e4528d79dcbe5f

Observation 6c41b363-bd5c-48fc-8940-827deb8740d4 · outbound

This paper cites Elson and Mark O.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Elson and Mark O

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.441602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.441602Z digest=sha256:1eaa30cff752d225be9577da51c87e312436ab1ac2b1a2855d6c671e90865a2c

Observation 20a10721-ce03-4f24-8f5d-68767cb25750 · outbound

This paper cites Foote and D.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Foote and D

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.450299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.450299Z digest=sha256:0a4d9d0e14abf2a0ed39a6eb917d23b5bf22655a931478abe0b7bfa20852e06f

Observation 0904ea70-ee80-4a76-9aac-ec3227472dea · outbound

This paper cites Multi-view video summa- rization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view video summa- rization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.457120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.457120Z digest=sha256:3091195e366652f84ee83d6ec71c22ec0a50a14c0ffe6453338973ca06369b22

Observation aea03a6d-8bf6-41c2-823e-ae6aaa5262d3 · outbound

This paper cites Gleicher, Rachel M.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Gleicher, Rachel M

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.463492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.463492Z digest=sha256:4af44867cd4ca604699bfaee9b28922319b30c0c614ccd252e1a502550de8a9b

Observation 69cd9eef-ad25-4bde-be6c-0f56da3ca3fa · outbound

This paper cites PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.469572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.469572Z digest=sha256:efe84c454bc10303c825159816eae78da2f6bf20a9e4696d97def49a34507a56

Observation b399e38c-c960-45a5-8172-81faef191d2f · outbound

This paper cites Diverse sequential subset selection for supervised video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Diverse sequential subset selection for supervised video summarization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.479571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.479571Z digest=sha256:ff2ed3cf619959b0b012175588212c42b742ac2c95da9b7ce186d2c69954b8c0

Observation 1fb9a401-425c-4436-beaa-101ae1ca7ace · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.485303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.485303Z digest=sha256:aab6d274001233a4c752ab7058c5b854e765b9e42a4eb7b5a8d38183477e0970

Observation c0b2dc9a-8d8c-43ed-9738-c519d6a0c022 · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.490447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.490447Z digest=sha256:a0fb9728fe89fc2ae3f42e6dc29bc76cf4a8076077846aa3d71c7060d90fdd37

Observation 36ada9b3-36b4-4a4b-bdaf-73478b66a597 · outbound

This paper cites Temporal Difference Variational Auto-Encoder.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Temporal Difference Variational Auto-Encoder

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.496223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.496223Z digest=sha256:e977d37616fc7c82b9bc43ca2606ffa503f4a38f3d24ae3a7c55ebf24719e6d4

Observation 82cf2f02-60ce-48a7-b2b7-45c7745d63b4 · outbound

This paper cites From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.503459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.503459Z digest=sha256:edd4c819bacf7e30306e5cb54f424779bfb1cca3ca961c403eed6066bc2bb535

Observation bb7f5658-868b-4da4-8003-6fea52ca632f · outbound

This paper cites Using closed captions as supervision for video activity recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Using closed captions as supervision for video activity recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.510211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.510211Z digest=sha256:12196bdea35b8c6d2113d51abec4d5e0e95fa10cdf02bc0738d31e741eaf5ff6

Observation ee3b7321-6456-4c50-a586-d310a09ced60 · outbound

This paper cites Creating summaries from user videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Creating summaries from user videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.520424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.520424Z digest=sha256:8e62423b19d440ff44c0c060d9d6181cc0bc026fd39b9b2d5cd6a57c6c632818

Observation adf155bf-20d6-4c5c-b6cc-30895c8ea6ba · outbound

This paper cites Video summarization by learning submodular mixtures of objec- tives.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization by learning submodular mixtures of objec- tives

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.533989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.533989Z digest=sha256:08a56c350ed952091dbc807e263aa39b85148aceac91de8776d79d71cefec6ca

Observation 51c0c39d-eb07-4531-b1e0-32883fb4dd9b · outbound

This paper cites Align and attend: Multimodal summarization with dual contrastive losses.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Align and attend: Multimodal summarization with dual contrastive losses

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.539969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.539969Z digest=sha256:05a1c767366cfffd75566a18bf59cb9dcb7f785d77776ab01c4b923676799aa0

Observation ccfe8115-fc24-4cb6-8957-db3605f10956 · outbound

This paper cites Cohen, and David H.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Cohen, and David H

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.545269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.545269Z digest=sha256:33772ed4504ddcc8ae8ea3288afc17ebd1f69caa4aa3981aa4cef2d808f29617

Observation 4771a9c2-2ac3-4484-be01-c8bc7857d19c · outbound

This paper cites Cohen, and David H.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Cohen, and David H

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.551100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.551100Z digest=sha256:3f17dec6c0a37b0f2f6f894bb06a92acaa09b4ff9bf901c09074888ffdaa1037

Observation 772d4bde-00fa-4203-a507-cb95825b6041 · outbound

This paper cites Vir- tual videography.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Vir- tual videography

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.556351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.556351Z digest=sha256:863d861754d78efadabfeb670f9e59e10e248e6fe0a3366d62e4f572f39a9482

Observation 00f95f17-6312-4f34-91fb-bd507296c3bd · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos LoRA: Low-Rank Adaptation of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.561315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.561315Z digest=sha256:61b7e55858f16958bfe69fa13706548776576a473a3620986757a6cf0b0e978a

Observation d38383b6-f858-4707-a017-1dede39c461d · outbound

This paper cites Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.566908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.566908Z digest=sha256:67be56a9e836936bb0663f787462e86d415ad10eaa4bc2be0a00c1bf61d782a5

Observation eaaccbea-74bd-4aa4-a9a1-0ed54ce64a55 · outbound

This paper cites EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.572054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.572054Z digest=sha256:1f6e10383071e4cc8bbbd95451138f955847dcd247cf077f26e9b4fdd5dc3a25

Observation d3df167d-f2ce-44a1-97d9-fcefb64f1699 · outbound

This paper cites Batch normalization: accelerating deep network training by reducing internal co- variate shift.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Batch normalization: accelerating deep network training by reducing internal co- variate shift

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.577856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.577856Z digest=sha256:cc889cdfb43f3df47035b92dfe7b0b876f4a13abe9b621fb851ee736b5fe5dd8

Observation c0397a64-a4a6-40dc-9e1d-f306b0a6255a · outbound

This paper cites Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.584139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.584139Z digest=sha256:ff852bcf24178dac3fd5dae83082ea3530692c1e1060bd02e580a5aa9436f841

Observation 8ea4e95b-7769-44eb-b8c4-b79cb4233be3 · outbound

This paper cites Learning to look around: Intelligently exploring unseen environments for unknown tasks.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Learning to look around: Intelligently exploring unseen environments for unknown tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.589659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.589659Z digest=sha256:3a858d14acdf8450a350cfb52516fdc812377c5b6f3ee561b7629d507f929631

Observation 512bb144-e0c5-4bf9-a5f8-977456cfdf53 · outbound

This paper cites End-to-end policy learning for active visual categorization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos End-to-end policy learning for active visual categorization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.595297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.595297Z digest=sha256:95d29ee5976c138e3a7b9516418dab377751efa1f16c028a146f3137d73076fe

Observation 2835bd81-9f36-41e4-a29d-cbf613bc3efc · outbound

This paper cites Time-Agnostic Prediction: Predicting Predictable Video Frames.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Time-Agnostic Prediction: Predicting Predictable Video Frames

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.600840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.600840Z digest=sha256:349d8e2177e858bc3b331372da472266b516a7231e5dc573d8a29153ca412881

Observation 5c8021b6-ddd4-4df0-a69b-51c0e2d0c7ce · outbound

This paper cites Simglim: Simplifying glimpse based active visual reconstruction.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Simglim: Simplifying glimpse based active visual reconstruction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.607025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.607025Z digest=sha256:968751f5ccdeea0e180ab7fb3fa027e8abd6d6ccdfc4470ff2b585859268b1c9

Observation d3adfba1-3679-4606-90ce-c0a8591e3d1d · outbound

This paper cites Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.612958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.612958Z digest=sha256:0fbf0a5d4f5afa87de7ee506d0f7ebb44e01c3a1ea2fed11d1d9ff85363c3d1b

Observation 75304319-c3b5-4630-854f-269cd0168e43 · outbound

This paper cites RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.619160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.619160Z digest=sha256:fba267c136284788e10cfdff54d08c3a6f78fde077b5915cde2ed5cf2c11d374

Observation cece5df7-2619-49f6-be83-94a62c0fa7f0 · outbound

This paper cites Large-scale video summarization using web-image priors.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Large-scale video summarization using web-image priors

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.624405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.624405Z digest=sha256:ea1a4eda26266d422e7c12ce0caef27163276fffe49f5097e3afb7564d3c2c82

Observation 4afe155d-955c-48ed-bbcc-329d56768cd9 · outbound

This paper cites an unresolved cited work.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.629375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.629375Z digest=sha256:b24f9e976534ee9b5ca552b58546f6f20bd2dbbe5654a22220111db4c3c3fc1f

Observation cb661ab1-ee16-4e98-81ab-c87070f9aa7b · outbound

This paper cites Segment any- thing.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Segment any- thing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.635697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.635697Z digest=sha256:9ded63a5c23f8703a93d16b6f9425fdf206591d0c8d12b04e4d13c228c5946be

Observation 883c89c2-d078-4ec9-9dee-f367fe283eba · outbound

This paper cites Hyperbolic Learning with Synthetic Captions for Open-World Detection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Hyperbolic Learning with Synthetic Captions for Open-World Detection

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-12T21:27:17.290381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.641426Z digest=sha256:328af4268e15a3bf0e5ab6896dc5bcd5058ed742b502eb4b4c411a0411453961

Observation 41dfb3fc-a295-4cc4-895d-81fc3c3afa56 · outbound

This paper cites A memory network approach for story-based temporal summarization of 360° videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos A memory network approach for story-based temporal summarization of 360° videos

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.650230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.650230Z digest=sha256:1bce3d24348f18a4a5a6f5a48f3adbcd4f2c57a57df1d4cb8fb2d18af2a74a08

Observation f6e45f0f-680a-4afb-9c6e-ea8b2a8a707a · outbound

This paper cites Predicting important objects for egocentric video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Predicting important objects for egocentric video summarization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.657826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.657826Z digest=sha256:8975c92561513a104cf62fdfecc66c6e6f5d6a0e0c84190ae12c98e8c677892e

Observation ebfefeba-e6cc-4e98-ad04-5e03f150fd7a · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.663653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.663653Z digest=sha256:49e864988a6c85c21b5495e82927fc039da04b658a4d11811647c1eefe1c4831

Observation e15fd5d7-482f-47e8-9912-453c1c504e1c · outbound

This paper cites Grounded language-image pre-training.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Grounded language-image pre-training

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.669888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.669888Z digest=sha256:ed4109587d539d4c6ea45cc8bb636a2b1356b02e6897bcd64448a819e21facf0

Observation cb28cbc6-1078-4f06-883f-5b54aacfb4cb · outbound

This paper cites How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.675226Z digest=sha256:6483535d82a356b3918402757a0b6fff12d95bb143c9dd90ddbb316dc6cbc2ca

Observation 756366c4-920a-4fa0-adff-8122bda308df · outbound

This paper cites Egocentric video-language pretraining.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egocentric video-language pretraining

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.569353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.680256Z digest=sha256:d2bf75887d0a45a6dac6e1be1606af237ad2cab03d45eb675d9aad1f46991681

Observation cadd1a59-85f1-4e16-aa5a-d1d422355b18 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.685434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.685434Z digest=sha256:d6dfc805ffcf7470d69710741d792583bbd1c7953e19a13e028c68f30daa08c7

Observation d116433c-7261-411f-861d-40dd7605c0ca · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.691686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.691686Z digest=sha256:6c8eb37265da0fcb103abcaa8e552b4add0d2d82679428b449a26a965c92b5ea

Observation 30c3a2d4-49dd-4644-9a7e-f73abf1a3b68 · outbound

This paper cites Decoupled Weight Decay Regularization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Decoupled Weight Decay Regularization

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.697244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.697244Z digest=sha256:e174451af26b9d28b120a7566e9078e8ebee5933f40b2e14aa6a14b59a0c7183

Observation 5fefd853-04f0-49ae-bd8c-3265e7192f38 · outbound

This paper cites Story-driven summariza- tion for egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Story-driven summariza- tion for egocentric video

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.549057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.702606Z digest=sha256:dcb45a3c5bbbbcc3321fd05d168d7c0c0f495924edab8c140767b3574122a5cc

Observation a9ae5bc4-44f2-43e1-b13f-e77415b31e02 · outbound

This paper cites Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.709927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.709927Z digest=sha256:185dadbc067b4ee97c31e00d4d33f5d467a25684081eaba9c14aa54e89f61bbc

Observation 04784fba-8090-4375-bf5d-e45497f5a31e · outbound

This paper cites Video summarization via multi- view representative selection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization via multi- view representative selection

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.525516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.715943Z digest=sha256:02a044a0e4038cab1ab1c69162ba1e7616a49ffc39d87d812701003d4dcf95c5

Observation fe0807fc-36ba-4265-b1fc-4fa9b5e49f6a · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.496074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.721421Z digest=sha256:cfe272cf9567e283983d8a755d24189e310910d9915c2cb35d7299e20f6912ee

Observation 8106cd4f-a3de-45d2-9b5d-29a59d8cef9b · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Srinivasan, Matthew Tancik, Jonathan T

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.726553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.726553Z digest=sha256:f2aac42a75e6810ae3136378d77bc4ebc4503fc9d43d8164c79d5e33b2c2de68

Observation 03ae2a1e-6682-4d9f-a1ec-1b14efc1a1e2 · outbound

This paper cites Automatized summarization of multi- player games.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatized summarization of multi- player games

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.454733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.731426Z digest=sha256:230e972cee4bd54196993b6cec418235631336547c58c77dff8bf1350fcfdfbb

Observation 61da99fb-d31c-4453-92b7-c0c58c127885 · outbound

This paper cites Egoenv: Human- centric environment representations from egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egoenv: Human- centric environment representations from egocentric video

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.410732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.745576Z digest=sha256:5bd685afe068324e101254dfec9c0e0258d8529dc48682c85fbea01fba47976b

Observation 67c3248e-d510-4503-9440-8e7f5cc3a3bb · outbound

This paper cites Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.385894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.750940Z digest=sha256:6b4f0e43559059e50f6e9aa347c5eff9b3e5c27a74329f01236dc4185b4b2e87

Observation 82e5f804-aefd-4974-b8dc-ef24f38a6c01 · outbound

This paper cites Adaptive skip intervals: Temporal abstraction for recurrent dynamical models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Adaptive skip intervals: Temporal abstraction for recurrent dynamical models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.357873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.756958Z digest=sha256:196ef9e6b54c1b3ac5a0a35b39efc0de89662f9d8b9f2f26fead16a97964b219

Observation 3e63b266-ec5b-4f6c-abe9-3c919f19deda · outbound

This paper cites Au- tomatic video summarization by graph modeling.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Au- tomatic video summarization by graph modeling

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.336833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.763234Z digest=sha256:8a3f30e48b7bbb18337099e68ee4beda92f19aa76a163483cb950b59d89ef096

Observation ad9bab32-aae4-4415-8174-2a602a225541 · outbound

This paper cites Collabora- tive summarization of topic-related videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Collabora- tive summarization of topic-related videos

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.316696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.771983Z digest=sha256:5fa152fbb7ae78da1629c4531c1cd8584835f68a73fe823f52856efee734f4b3

Observation 1deecac3-7788-4b32-ab77-5d47a72a0932 · outbound

This paper cites Multi-view surveillance video summarization via joint embedding and sparse optimization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view surveillance video summarization via joint embedding and sparse optimization

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.295299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.778053Z digest=sha256:eedcadf404943467023e5c2593661d290e403137e77731c2194ec3b91bbc9158

Observation 8f1a830f-752c-4357-a063-ac2da8b6188e · outbound

This paper cites Roy-Chowdhury.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Roy-Chowdhury

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.274325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.783830Z digest=sha256:146d6f23366645fcffbb371c7432028d1f3054552b93ae00fdd21721b11bbb8d

Observation c21a3630-06c2-4b52-af5b-4584cb1ac0f1 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Bleu: a method for automatic evaluation of machine translation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.257179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.789486Z digest=sha256:523ed63808848f00c67f32300a74248ad9e94993c583d98d5c6a5b915e72b873

Observation 5a1a688b-0e61-4405-8da4-40835fb472d8 · outbound

This paper cites Sumgraph: Video summarization via recursive graph modeling.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Sumgraph: Video summarization via recursive graph modeling

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.238943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.795037Z digest=sha256:acc8158faeba8efa172cd2c711940e19005adc16e9094387654016ddfe3352f9

Observation 7d7ac8df-72f4-4c62-9db1-cc869eb1d79d · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.215657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.799995Z digest=sha256:3893de533fbb8eec0b25ccf79a437ee9ae595fface1cd72cdf24a44ed0159d7a

Observation 378c0f34-b5b3-4334-b4f1-51e9efeac2e5 · outbound

This paper cites Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.193781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.805015Z digest=sha256:f36b99ead3c219c586800f324a07db2ef7ad70dc13777ec3bf035a1bfc517d59

Observation a4a08e10-1781-4d09-a398-42aed2923331 · outbound

This paper cites Sidekick policy learning for active visual exploration.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Sidekick policy learning for active visual exploration

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.176412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.810155Z digest=sha256:1dc2e90a59e608101ad75bf6d37b1a230a961b9b8bf7e587f8b0ca23466e8d7b

Observation af605da0-9852-40f0-bac3-870705489a30 · outbound

This paper cites Emergence of exploratory look-around behaviors through active observation completion.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Emergence of exploratory look-around behaviors through active observation completion

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.160511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.814977Z digest=sha256:c850b55bbe89d46ec9ed55af7d79ee353553383442983c07c358da8dfaddc3fd

Observation e496eb70-24af-4f73-94f6-c2f539eb4e2a · outbound

This paper cites Naq: Leveraging narrations as queries to super- vise episodic memory.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Naq: Leveraging narrations as queries to super- vise episodic memory

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.143654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.820501Z digest=sha256:7a1da76d1b7e3b125da3f2c178070e0932e11b43fa34f8dac0adf5fe4c210a89

Observation d19f8cfa-e3cc-4253-ad30-7ea6ae3fa489 · outbound

This paper cites Video summarization by learning from unpaired data.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization by learning from unpaired data

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.126653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.829184Z digest=sha256:ba3a3e5cf6ad66968e85ca89526d3c2a37687c3f151d0e971e0451c7c5606a79

Observation 0d4725b1-c25f-457f-9870-5f756309db84 · outbound

This paper cites Adaptive video highlight detection by learning from user history.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Adaptive video highlight detection by learning from user history

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.835818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.835818Z digest=sha256:904d86f1c079e10439d1685416266621d14200c3df98f975fc2bb50296a9f148

Observation 2fc92f75-2740-4f13-a654-84a748667606 · outbound

This paper cites Chowdhury.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Chowdhury

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.097889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.841151Z digest=sha256:5c913eaf4009c9ac8bfc977115926ded952d3b9b60f2a66813d528d77f3b8bb8

Observation c5a9690f-59df-487e-b0b1-bb2c99d7e9e3 · outbound

This paper cites Attend and segment: Attention guided active semantic segmentation.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Attend and segment: Attention guided active semantic segmentation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.846399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.846399Z digest=sha256:5057963a20f62b15eee7b75b6672ea89bf41ae41dc1cde24f839054a875a6032

Observation 121c8f1a-836b-4abb-a24b-3acd089351ec · outbound

This paper cites Glimpse- attend-and-explore: Self-attention for active visual explo- ration.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Glimpse- attend-and-explore: Self-attention for active visual explo- ration

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.070161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.851964Z digest=sha256:2810db60ff598253eef4740d76de92d23f77c27d9780fc809a5b7ebb0137443d

Observation d6b9d89a-40b3-4049-8dfd-16e538200d9a · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Actor and observer: Joint modeling of first and third-person videos

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.052780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.858613Z digest=sha256:923942c2271706273a5d0ceb6e5819e25f07ffd819d5f8bc1542ac9486e2db41

Observation 8481d850-ddf3-403a-8c8c-56f49059cce9 · outbound

This paper cites Tvsum: Summarizing web videos using titles.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Tvsum: Summarizing web videos using titles

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.034387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.863934Z digest=sha256:809fb9cb8e258bf0bd128f385827c20d5b2337c999c42dbe0f5abc71222e05dd

Observation 6266f243-3dcb-4190-a5e2-0dda93bef23a · outbound

This paper cites Making 360 ° video watchable in 2d: Learning videography for click free view- ing.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Making 360 ° video watchable in 2d: Learning videography for click free view- ing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.015861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.872398Z digest=sha256:cb07040994406d992c690946fc64411c5e305c02cad2af3e113c70f465c82897

Observation 67a9f5c6-e0a6-41cf-9c32-e40ba2a166ae · outbound

This paper cites Pano2vid: Automatic cinematography for watching 360 videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Pano2vid: Automatic cinematography for watching 360 videos

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:17.994590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.878297Z digest=sha256:fde7e9d0fc481db8d6441211b3ae3a5c1104b3c2923fecc7f60cebeb5dddaa4e

Observation fbe32558-5b46-4a57-9a06-394c4ba8a462 · outbound

This paper cites Automatic con- cept discovery from parallel text and visual corpora.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatic con- cept discovery from parallel text and visual corpora

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:17.975142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.884770Z digest=sha256:b5c5dc6b9db685652caa83c4186df8fabb44b5df188121f7652d25eb530ff064

Pith citing papers

Observation 6cd707ae-c6d6-4b32-b3a1-eec071956336 · inbound

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos cites this paper.

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:38.441181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:38.441181Z digest=sha256:98420a05001f27b9abc4d4cd09251b83894218d0a45b059c9f2c475eb2ae3601

Observation ebe94ece-1e15-49a0-bcc6-e24f2cfb1f29 · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

Reference 187

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:00:58.847019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T06:00:58.555825Z digest=sha256:1328d2c2736c00901bb1fd7be866358e8959e5b5246654079d26acbb6cefaa81