Pith. sign in

Paper Citation Record · LEDGER

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

As of 11 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2507.07744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07744 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:40:05.200664Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c171e9f0-1219-48da-b4e1-b594575d6f54 · outbound

This paper cites Joint visual and audio learning for video highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Joint visual and audio learning for video highlight detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.354245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:00.652223Z digest=sha256:ccdb3d1dcb5744705a0eaa1c943470579eb3f90c28a836f533734fdd80cdaebe

Observation 9dcd0f8f-f622-4c8e-b43a-e26cea825954 · outbound

This paper cites FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:00.715128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:00.715128Z digest=sha256:f9cdded37afe8b191b15ba239bb97eaf6a3d49afd750741d01d328071d552541

Observation 8c7eba0d-06a4-4bbc-842e-ddab6282ee95 · outbound

This paper cites End-to- end object detection with transformers.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding End-to- end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.193724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:00.779071Z digest=sha256:09cd1b08618874040b8cdc94952fd9c6a37fdb0cd6478aa13b59bc5d230cceaa

Observation 654734bc-97cd-4a42-8fa6-90ae36976868 · outbound

This paper cites Slowfast networks for video recognition.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Slowfast networks for video recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:00.848713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:00.848713Z digest=sha256:479309a0462d6b66c5435cf30b46595270f80c4fdf5f9ac213cb4b9c0d5bdde4

Observation 75d26b24-e8b2-4486-8b32-79c9a7cc3419 · outbound

This paper cites The use of ranks to avoid the assumption of normality implicit in the analysis of variance.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding The use of ranks to avoid the assumption of normality implicit in the analysis of variance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.007086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:00.913841Z digest=sha256:18df9a44c8b035c9b5f82c737a4cd6156122ecb25f7eca003a1808975790a6ba

Observation cdf82518-2ae2-4220-849a-0bf54a004368 · outbound

This paper cites Tall: Temporal activity localization via language query.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.839389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:00.995731Z digest=sha256:0bb06f7b0f8e3694c0197cf930de44d04396faa74b1708a026560d936ed0a8a9

Observation 9e3b8fdc-7227-40c6-b3f7-6591f537a59f · outbound

This paper cites Clip-adapter: Better vision-language models with fea- ture adapters.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Clip-adapter: Better vision-language models with fea- ture adapters

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.672575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.076972Z digest=sha256:351da3d50e8182c7ea77632b45cd6035e69a4a3dc36dde1ade5bb6807de765c6

Observation 7058ff4b-37a9-4dfe-a399-c044ba2fbe00 · outbound

This paper cites Saliency-guided detr for mo- ment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Saliency-guided detr for mo- ment retrieval and highlight detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.160740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.160740Z digest=sha256:0b117ed55dc1d2de6c3d8bd8dd61006519fa328b56d5aea6995e2fa0f94cc4fa

Observation a7a763af-351e-423e-b119-84f00fc9f336 · outbound

This paper cites LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:40:06.018099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.243689Z digest=sha256:fa453f73428ebc98908658e1b156954c8958f2f9b5e775fd672141d3862b4dca

Observation be569355-1b5e-4e5e-9ee8-d15ddfbaf142 · outbound

This paper cites Unleash the Potential of CLIP for Video Highlight Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Unleash the Potential of CLIP for Video Highlight Detection

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:40:05.757945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.315277Z digest=sha256:0eb8f445710f132044d0ddd9e5cded6cd77dfc6ecc50bbf34718a04ed14504df

Observation c87d8e20-ff76-4ffa-8e6b-08460fe55dd3 · outbound

This paper cites Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.379564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.379564Z digest=sha256:c72fc3d7f79ac41f90b3e51062d5a2eb12ff907ad246f8a8b28a54c41c4eb26a

Observation a5466ecd-48ef-446e-a5ba-6f557beac4a3 · outbound

This paper cites Mini-net: Multiple instance ranking network for video highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Mini-net: Multiple instance ranking network for video highlight detection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.490094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.461875Z digest=sha256:b41c0f4461a750c02bbdbb6b7fc22728a5c2c2d28a2e4477c8d1c8b0a7370662

Observation c99d1de7-495b-47b5-9e05-6b70969c5966 · outbound

This paper cites Fifth berkeley symposium on mathematical statistics and probability.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Fifth berkeley symposium on mathematical statistics and probability

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.152005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.530384Z digest=sha256:d2d8bba8fd6f1637f161b598d179f8894687a52bba8757d7faef9cbbcf1c92d9

Observation 4bb2cb10-9f50-4a4e-b88b-35c86fc67807 · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Knowing where to focus: Event-aware transformer for video grounding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.779308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.644826Z digest=sha256:ed362bc59d9bea43047f1bf58972c6241aaa784aa22671eeca57c1e2ee3f092e

Observation d6455af2-90d2-4edc-8066-b87cf2bbe5d1 · outbound

This paper cites Vi- sual prompt tuning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Vi- sual prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.378488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.720406Z digest=sha256:211ee8fe98dc5b8880ef4a4f77cc4d92022cb17ef62219ce5fd19b957c22566e

Observation 027ddf54-2842-4fb2-af89-78088e6f6da9 · outbound

This paper cites FractalNet: Ultra-Deep Neural Networks without Residuals.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding FractalNet: Ultra-Deep Neural Networks without Residuals

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.791278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.791278Z digest=sha256:25262e8329d05678443ddb514a961ba6d4eb6cde978527a5427ce97a41af9a74

Observation 1ade79a0-108e-4edd-aaba-28d11e4eb1c9 · outbound

This paper cites Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.187552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:01.879168Z digest=sha256:15cf469cbe158f2428e666f5d07457bbe2f0e5bc2980562d1b31c0ebd290d71e

Observation 26a573f4-6ffc-41d4-82b9-e287e889ca43 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Detecting moments and highlights in videos via natural language queries

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.085279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:02.017528Z digest=sha256:fce8d90b46eb52815f088b38a78b005b110328b74eb3e81cddd3fefa9430f2b5

Observation ac65712d-020c-41bf-9d4c-13f2eb98229f · outbound

This paper cites Dn-detr: Accelerate detr training by intro- ducing query denoising.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Dn-detr: Accelerate detr training by intro- ducing query denoising

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.967704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:02.079804Z digest=sha256:bb57e714bf2e7a8ea3eed05984e0bf5fa43c3cb783f51c3c95aa72b86fb4f448

Observation b3a7ea83-9220-4053-984c-0aa808346661 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Univtg: Towards unified video- language temporal grounding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.857558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:02.199351Z digest=sha256:55a7c2ebd6ca1d1967ac91c394322b42f8cafcf142609357fb8bf537821af256

Observation beb6e54a-05e5-4ed0-927f-41a94c3afc0e · outbound

This paper cites Focal loss for dense object detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Focal loss for dense object detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.316676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.316676Z digest=sha256:9ae3bf4b2bda4ef1a6a148cfc520273657de1f4756ea9e720cab64efbede9b0a

Observation 9d306071-2c72-4600-a409-34c126147cb0 · outbound

This paper cites Frozen clip models are efficient video learners.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Frozen clip models are efficient video learners

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.737363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:02.409539Z digest=sha256:8edc3750523d33ff5c550b235f26a8bc08c653050504fe5d3c9c250d6c2f761e

Observation f82a852d-52f8-4d0b-b24c-e873451843be · outbound

This paper cites DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.500974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.500974Z digest=sha256:fb3643f94b620be30e3c7e8b1c6e5aff8488aa8cfa7c10ae0688dd0c6c80a6a5

Observation a4d84933-1429-4601-952d-6e011314d4f8 · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.576985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:02.567378Z digest=sha256:eed31d0dde67814470048fb3f3f5e26ca1e009001a367c925d56230bec9a009e

Observation 962d6db7-09bb-4a60-9218-3d4e5fe0ff3f · outbound

This paper cites $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.643521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.643521Z digest=sha256:a670f2db9f38c8e1d0915d7a01f6e65daa8b0105262bc0e3c810b551f346f4cb

Observation 4bb7e029-ef6d-4f59-977d-318d65752d03 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.450671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:02.714083Z digest=sha256:9259268b165538630fa8e02382c06c371e125cc0791d55ff78d8e5b7e4dc387a

Observation 34b3b1a3-7981-425a-a102-7663f633fb81 · outbound

This paper cites LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.771931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.771931Z digest=sha256:bd276e678e86896d4a24737412b926604319aa53e37c1cf4c34515b81c3f05fc

Observation d8b4e273-c1b5-45ef-b906-8cbe788b8875 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding The surprising effectiveness of multimodal large language models for video moment retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.837075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.837075Z digest=sha256:8f76e7fd98eeac672633eb42a93542c7d3fac9612a7f7775eb18c5c1283a8e11

Observation f40d9eda-caae-4d09-a7d6-9c140d7b66e5 · outbound

This paper cites Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR, 2023.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.354168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:02.921358Z digest=sha256:c6a778d60fd76d27ac7b95f355c0677f007e0723a319a041d80a83ea14ba533d

Observation 96ce5b25-054a-4166-a25e-886ed9a96a9b · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.205916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:03.042994Z digest=sha256:8f9a1ef56bb6bf8084855d296e212a814d19b825f4b5839083f93c7828d3dd68

Observation dd6a6245-faf9-4110-ad95-d857ae3466e3 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Representation Learning with Contrastive Predictive Coding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.121427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.121427Z digest=sha256:e69f6b2573e220e11185fb52c9cab28901628c0e23756c8414a218009b40c57d

Observation da0e6ad1-6cb4-4050-8c67-39d70e34a7a4 · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding St-adapter: Parameter-efficient image-to-video transfer learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.051372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:03.202428Z digest=sha256:94385bf1bc61d7b8f052f0b28bc9bea6eade88d6f60588a9f1e537728ab81ea9

Observation 0019f4b3-10f4-4eb8-81ba-ed5f4fb2d95c · outbound

This paper cites Sada: Semantic adversarial unsupervised domain adaptation for temporal action localization.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Sada: Semantic adversarial unsupervised domain adaptation for temporal action localization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.867577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:03.289501Z digest=sha256:41d403151c975b80d67a40c695ba888eba422468d2c28a78b756e6dc13deed55

Observation d16eecaf-28d0-4610-b94b-afdcbbaf92c1 · outbound

This paper cites Disentangling spatial and temporal learning for efficient image-to-video transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Disentangling spatial and temporal learning for efficient image-to-video transfer learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.643095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:03.370771Z digest=sha256:2fcc2d9238460a1f2cc27042bcc5614d22a993efab5abf488fe70f8977417ada

Observation ed20371f-0403-4410-b007-1bdccf5441a7 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.510044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.510044Z digest=sha256:952b74b56cc8d05d7b60220d7413c0d5e7a92a7a5c0f4504c0d630666e943190

Observation cc3e3217-270b-42e1-b80f-c8f76e0fb56f · outbound

This paper cites Coherent multi-sentence video description with variable level of detail.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Coherent multi-sentence video description with variable level of detail

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.444433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:03.632073Z digest=sha256:f7f9e8fc006406496a0308e583bea6826df5d3bfd5e3474caab80505acade15b

Observation 22c1ed98-f3ce-4cd9-ae74-f64b4bd54f25 · outbound

This paper cites Ranking domain- specific highlights by analyzing edited videos.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Ranking domain- specific highlights by analyzing edited videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.271012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:03.734332Z digest=sha256:e0238326ae1ad7d6f955a35bce7a9c8948773b71e014579b336e521fee3d01c0

Observation 55b6973d-c485-45ac-9314-207f4d83df5b · outbound

This paper cites Lst: Lad- der side-tuning for parameter and memory efficient transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Lst: Lad- der side-tuning for parameter and memory efficient transfer learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.112200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:03.833480Z digest=sha256:4644157b844061c65c0e2de5af4af3293cdff35d9ff52c6f87acc0d02b562d79

Observation 0afbd36f-a5ef-489d-9838-1bf659250beb · outbound

This paper cites Attention is all you need.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.914238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.914238Z digest=sha256:80a965199b5ae6674c6309d43a7393f6ffcff90371ae9a9a166b6478a3b15bad

Observation dc814c50-2861-428e-a020-1ced0515292a · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.944574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:04.067540Z digest=sha256:c4357c45e1c8959119c8257691325cad90e3908952d80bfd9932993366da96cc

Observation dae05960-f941-4840-985e-c37e44273954 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:04.192757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:04.192757Z digest=sha256:f60096398ff55f68f88595681f0ff1cf31a09211bbde2ab5c5bb4bb7c4b2d1f8

Observation fcaac928-c242-44ff-ba16-b035f1054970 · outbound

This paper cites Vision transformer with deformable attention.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Vision transformer with deformable attention

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.710516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:04.293577Z digest=sha256:fdfd8ec9c1bbabb86b5fbe1df8f1577bfd96b04d07645c929f54f8ea9ec1cac3

Observation 6e3c69d9-d6f6-42f5-b21b-21bce7ccd6d1 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.511209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:04.371035Z digest=sha256:d228320f4a2e2a2f1d1e4e814be360eb020c180a1ebaa54397ac0ecf71246e05

Observation 542b7c04-9546-4022-bcc1-47839712a2bf · outbound

This paper cites Cross-category video high- light detection via set-based learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Cross-category video high- light detection via set-based learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.318203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:04.445706Z digest=sha256:b20f71020ea89fcaeafad12eb438535be9f88cdaaa5ed66ff624511efe8a3559

Observation f6177304-f351-4ff1-9927-97870806d7c9 · outbound

This paper cites Unloc: A unified framework for video localiza- tion tasks.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Unloc: A unified framework for video localiza- tion tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.110265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:04.563858Z digest=sha256:a80740e83379a3f347e8562fc71e44409c61052520ff6196a141f272cd79ff7e

Observation 15576df0-3d5b-4402-aa8c-5705814ab4a4 · outbound

This paper cites Understanding negative sampling in graph representation learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Understanding negative sampling in graph representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.893073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:04.698446Z digest=sha256:a0c914780685eec73944415955d558bb42f04baec1b8686410c52337ae4d3e4c

Observation a89e8865-aa01-4c01-bda5-86a15edc9dd0 · outbound

This paper cites Parameter-efficient is not sufficient: Exploring parame- ter, memory, and time efficient adapter tuning for dense pre- dictions.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Parameter-efficient is not sufficient: Exploring parame- ter, memory, and time efficient adapter tuning for dense pre- dictions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.739159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:04.818146Z digest=sha256:47ec14515df6e56df2ace00ef850445d151cfc3a7e2a3c0d2e2c83b62a140b32

Observation d05ac6e5-583e-49f8-a42f-a901e059b357 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:04.919669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:04.919669Z digest=sha256:8c2675e49e1bae9b777cd40b5fa77661c3c1679831c33d7a1264759e02a2461f

Observation f196f109-36e6-44b2-a4b1-68282aa34209 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Conditional prompt learning for vision-language mod- els

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.521597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:05.020736Z digest=sha256:eb51274fe20497f11bcc968bc457d3f0686cc483ec8d93d46cd64c5fe68aef5d

Observation 93b92d5b-bd99-4e09-8c12-71c2c64db6fa · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:05.125881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:05.125881Z digest=sha256:ce4a3a1fa4dfc4e5d91871fe4c37e63021e8ea4937e94666fd9897b105cebb89

Observation 60b4fea0-4029-4292-816c-9e4963f77bea · outbound

This paper cites Notably, these losses are applied to all the intermediate layers inde- pendently.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Notably, these losses are applied to all the intermediate layers inde- pendently

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:40:06.366128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:40:05.200664Z digest=sha256:f799a9904df3e7b54e4491a07e33dcb5553181e5d6e997e0dc90b77268fc7911

Pith citing papers

No inbound Pith citation observations are available.