Pith. sign in

Paper Citation Record · LEDGER

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations

As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2508.03410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03410 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:58.855512Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:31:30.702800Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T04:31:30.785644Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 853abf69-ed9e-4ba1-9e1c-aac611e0a97d · outbound

This paper cites Self- supervised object-centric learning for videos.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Self- supervised object-centric learning for videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:05.285548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.315227Z digest=sha256:ccc64af3ea3e8a5c797a1b8fa271e51b88d8dde1bd414165af54afb2061ac743

Observation afa7eafe-1c6b-495e-aa81-aad8e3b9446f · outbound

This paper cites Invariant slot attention: object discovery with slot- centric reference frames.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Invariant slot attention: object discovery with slot- centric reference frames

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:05.128437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.385986Z digest=sha256:4901041c215b6bd8464c9d48e1536c798110350c78caf0c56d982ad37264b2f1

Observation 788d4df3-a31c-4f4f-98b0-77e2e490960f · outbound

This paper cites MONet: Unsupervised Scene Decomposition and Representation.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MONet: Unsupervised Scene Decomposition and Representation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.451793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.451793Z digest=sha256:986578c8f004354bda9b11c6ed7ae7f084cb39a254b53b53d77d677eb2184309

Observation e0b0dcf3-e677-4c6c-bdf3-e74073ec048f · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Emerg- ing properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.939414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.539206Z digest=sha256:d3ddb4ed31c1d3d51a1d33916be21edac19ac5dc094ee333956b7d67d74d87d7

Observation a835b49b-c214-476c-becb-a1db42252975 · outbound

This paper cites Sobolev training for neural networks.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Sobolev training for neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.752038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.608318Z digest=sha256:970c1eea10ac9318373b192554f72579a8d8505d195add85e8f4bb2aa688780c

Observation 7ea62f96-40c5-4b5c-a6ae-7891b212fdbe · outbound

This paper cites CTRL-O: Language-Controllable Object-Centric Visual Representation Learning.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations CTRL-O: Language-Controllable Object-Centric Visual Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.664516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.664516Z digest=sha256:2fc7b9e181655f28cdc51ff3c940d37079a8499e80fc74114fc160031dbb5bbe

Observation f930be02-3da2-4fcc-92d6-29c250c9037a · outbound

This paper cites SA Vi++: Towards end-to-end object-centric learning from real-world videos.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations SA Vi++: Towards end-to-end object-centric learning from real-world videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.725716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.725716Z digest=sha256:7602d9205c46370b3874077663e7740437673dcba045207485a302b8a2d88c81

Observation 13164b92-a76f-4537-9434-634e112842e6 · outbound

This paper cites Adap- tive slot attention: Object discovery with dynamic slot num- ber.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Adap- tive slot attention: Object discovery with dynamic slot num- ber

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.774375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.774375Z digest=sha256:b656daf673ec0b49b63a996fe62c0ceafe279e7be5bc84558932e6432ce6fe55

Observation d236eda9-cf8d-4093-ab34-4d7c9f54fc61 · outbound

This paper cites MoVi: A large multi-purpose human motion and video dataset.PLoS One, 16(6):e0253157, 2021.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MoVi: A large multi-purpose human motion and video dataset.PLoS One, 16(6):e0253157, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.582904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.839087Z digest=sha256:bffeb2729a5477418ad8bff62df96f669b66cf9f3d758a32245f29ae83166c50

Observation 9bebab07-a7b3-4fab-b6e9-c7711aca7fd3 · outbound

This paper cites Tagger: Deep un- supervised perceptual grouping.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Tagger: Deep un- supervised perceptual grouping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.421492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.896320Z digest=sha256:ae6bfbf0c6c075351b9612e2d859145dddaa3ec013707212b41650a61e7d0869

Observation 3dbd392c-dc5c-4cfe-87f1-d1009247cf78 · outbound

This paper cites Neural expectation maximization.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Neural expectation maximization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.211712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.954368Z digest=sha256:46d50249eab55ded1a65c5100a05a47db26eaa429937029ae893029090828fb0

Observation 5d1fec2e-4287-4024-bcb6-510d26a8e419 · outbound

This paper cites Multi-object representation learning with iterative variational inference.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Multi-object representation learning with iterative variational inference

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.031103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.010076Z digest=sha256:a3cd5c8f4789c631aabf0eae8f35f6a506fe62b85c50925e4d23c5314173d1d5

Observation f8ef2d4b-b7b5-40ef-922d-7bb24ddd1f00 · outbound

This paper cites MiniLLM: Knowledge Distillation of Large Language Mod- els.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MiniLLM: Knowledge Distillation of Large Language Mod- els

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.799045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.073725Z digest=sha256:d03b0aaab61bb83d3412d57e9878c47dcb0fb53fc73f7c7177702720af12b667

Observation 5782148d-ac6e-4093-9c78-edd1cd0cfef9 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Masked autoencoders are scal- able vision learners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.587311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.122200Z digest=sha256:a9d0206d03d32a02623f851c590916dd2137c2b70d075c01b4ddc0da364a5573

Observation 00520d13-c187-4762-9a1a-6a8358dab990 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Distilling the Knowledge in a Neural Network

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:57.206596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:57.206596Z digest=sha256:2c7d11478fee5a2a462ddde95e3d1d1f45837eb3d3637965180014d13b0a47cd

Observation bc62d036-bb5f-4dc8-bd36-387ca04eca9e · outbound

This paper cites Multi-level feature distillation of joint teachers trained on distinct image datasets.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Multi-level feature distillation of joint teachers trained on distinct image datasets

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.365189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.274758Z digest=sha256:00a8c6b5786dfa09a61577c062ff79cecc596c020b68c78ee9f9543e3c77d46c

Observation d71c245f-57bd-4c87-a2cc-7f44af9ecddb · outbound

This paper cites Improving object- centric learning with query optimization.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Improving object- centric learning with query optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:57.347323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:57.347323Z digest=sha256:da03cf8d26cf1216d09520a2acb7f19b9c4c2fb279e7eaa1c130b4a3a4e1e7ce

Observation 00622ac3-50ac-4062-885f-51679a8509d7 · outbound

This paper cites SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.157503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.389490Z digest=sha256:0a8acd34c88b697b68ece5278da7330666dbe8d0708d03b1db37ba372306d491

Observation 4b2b0fd7-dc79-405d-a25a-498f0a23ad8a · outbound

This paper cites DIOD: Self-Distillation Meets Ob- ject Discovery.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations DIOD: Self-Distillation Meets Ob- ject Discovery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.866586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.480626Z digest=sha256:1a76b2dd47fb4e6c736a896eb8155de97bccea7e2ba81c1acbd7afcbd2092bdb

Observation 62c6b32d-722d-4918-a9fb-e5f478849866 · outbound

This paper cites Elsayed, Aravindh Mahen- dran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jon- schkowski, Alexey Dosovitskiy, and Klaus Greff.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Elsayed, Aravindh Mahen- dran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jon- schkowski, Alexey Dosovitskiy, and Klaus Greff

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.637383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.547467Z digest=sha256:2b04076cf647a6e2e1149e2b61d4f08f8e9ed974a1e627a2a20617d6c01a35e1

Observation 8c313c9c-04d6-4ee1-82ae-ddb00b39d19e · outbound

This paper cites Object-centric cross- modal feature distillation for event-based object detection.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object-centric cross- modal feature distillation for event-based object detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.437533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.618336Z digest=sha256:e74d8f89a0b8160319f031c5df324802f8c7a89c52351d7f17cb0e3480c700fc

Observation 21f4a567-791d-4453-ab7a-96d1ec00b5b7 · outbound

This paper cites Hashimoto.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Hashimoto

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.138735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.690861Z digest=sha256:864829b714c04a5eddc627eb76963b90078d847078d53db72bdd9cbab14f1926

Observation d0f3c9ce-1e28-420f-822d-f92e0ddfed7a · outbound

This paper cites Microsoft COCO: Common Objects in Context.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Microsoft COCO: Common Objects in Context

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.961664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.778202Z digest=sha256:ac1d6cbc090c071a9fcaba0555d94e6fcf8d5efe81fb6aab18c673e006c4002f

Observation 492b0791-372a-44d9-a45a-58deda931865 · outbound

This paper cites Object- centric learning with slot attention.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object- centric learning with slot attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.761165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.874877Z digest=sha256:e9388da8b8c2fa905f6967ad10b6c8588347fe2083467eddb9d6a821f9de3367

Observation d13e6d06-ee1d-4dfd-866b-a4b6a6ad3d0e · outbound

This paper cites Temporally consistent object-centric learning by contrasting slots.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Temporally consistent object-centric learning by contrasting slots

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.576513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.969711Z digest=sha256:548d43fcff3b450ec8bcaacfaa889e9563a5662db100e5db6c96f2bcb68948ec

Observation 01100cb7-ef8f-4908-ad72-21d653f1ccaf · outbound

This paper cites DINOv2: Learning robust visual features without supervi- sion.Transactions on Machine Learning Research, 2024.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations DINOv2: Learning robust visual features without supervi- sion.Transactions on Machine Learning Research, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.311964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.071804Z digest=sha256:7a670ec5cc7dc68162c4f31d1c946d5bf6b1be07a06e49edfaa864b21f2c39f9

Observation 998a1e18-719b-4f53-b3f6-dbb637f54a42 · outbound

This paper cites Perazzi, J.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Perazzi, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.016609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.150727Z digest=sha256:0ea8006d69e514e28851f75be898c24167c2d3855ef088f46750bd18a194a480

Observation f3dca698-d562-411f-996e-b2d9e1717786 · outbound

This paper cites Torr, and Song Bai.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Torr, and Song Bai

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:00.692495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.223978Z digest=sha256:5c50b18a76363f92c561074d1ceaef86cd6eaf78bc27e3e8a0fe97ddf0921e3f

Observation f2d0397d-f3bc-4e86-8f64-26d5d111339e · outbound

This paper cites FitNets: Hints for Thin Deep Nets.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations FitNets: Hints for Thin Deep Nets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:58.248095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:58.248095Z digest=sha256:112cc59fd71dbef3660d4ff651cb8cda7aba38ca453b76127d942397026f0c35

Observation c0adba32-53f7-49ed-83e2-d08b174b2882 · outbound

This paper cites Bridging the gap to real-world object-centric learning.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Bridging the gap to real-world object-centric learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:00.372198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.306675Z digest=sha256:3b54d7c973c31fc1b9ac7c0ae342aef5f42edf5f8dfd1b71d3edb3cf6c84bd11

Observation 623dff23-c2ad-4960-95f6-3d55c66f6bb5 · outbound

This paper cites Simple unsu- pervised object-centric learning for complex and naturalis- tic videos.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Simple unsu- pervised object-centric learning for complex and naturalis- tic videos

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:00.032058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.384400Z digest=sha256:12eb15fe23889575960948de057d4b156133339a39fd6a850a88731b6645590f

Observation 4109caf7-3d09-4470-9f27-fc1938d7b9df · outbound

This paper cites Self-supervised video object segmentation by motion grouping.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Self-supervised video object segmentation by motion grouping

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.789129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.454779Z digest=sha256:afe72f980927c9e15bef56b0064fb71bf0a2aef04802ee9e8f9e0240e6146f26

Observation 60a88745-0076-4b62-9070-000ae8c311de · outbound

This paper cites The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.606311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.585604Z digest=sha256:f24175be82d283d1adcc08b32316a3fdfaa9bc04f870f165e7c98caab1fbf736

Observation 92a05be0-e7c7-42ad-ac55-0297dd531122 · outbound

This paper cites Object-centric learning for real-world videos by predict- ing temporal feature similarities.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object-centric learning for real-world videos by predict- ing temporal feature similarities

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.442514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.699996Z digest=sha256:936ce06637c0b1a6f6c252f2ed57cb72f5c9c11b46e4135eb9bfe3011990bc8b

Observation 8855759e-aea4-430e-ab96-00dc24dfd321 · outbound

This paper cites The SLOTMATCHstudent based on DINOv2 is compared with an equivalent architecture without distillation (no KD), as well as its corresponding teacher model.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations The SLOTMATCHstudent based on DINOv2 is compared with an equivalent architecture without distillation (no KD), as well as its corresponding teacher model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.223002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.777597Z digest=sha256:ae1eda95119ee0d020aa7bd00a0da380a0580b29e76bcbfda99ea316b537662f

Observation fe00adde-0429-40ae-8600-88453e16f351 · outbound

This paper cites In preliminary experiments, we found the standard deviation for FG-ARI, and mBO across seeds to be within±0.06 and±0.29 on YTVIS, indicating stable convergence behavior.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations In preliminary experiments, we found the standard deviation for FG-ARI, and mBO across seeds to be within±0.06 and±0.29 on YTVIS, indicating stable convergence behavior

Reference 2048

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.064396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.855512Z digest=sha256:4e0578dd430f3e022f3accf61dde1f7d094ffb29ab75bf1c9c71f0cd434416cd

Pith citing papers

Observation bb072915-1027-4b07-bcb8-ca1f9d69bb33 · inbound

Cornelis Easton:The Milky Way as a spiral galaxy cites this paper.

Cornelis Easton:The Milky Way as a spiral galaxy VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:31:30.867218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:31:30.702800Z digest=sha256:5f5c34b89d87a825dd66e89f083c27f39ab9c9afaf73d060da92265811729e22