Pith. sign in

Paper Citation Record · LEDGER

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.10932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10932 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:06:16.767341Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 638df983-383b-43a4-a9a7-634935fe8208 · outbound

This paper cites Layer Normalization.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.718672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.718672Z digest=sha256:7ee0d84c32b6364a44251529c9caeb8673df3fa63704dc26813941b6943b3be2

Observation 7ef75b17-3d7f-45bb-a374-8f04efbfe663 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation LLaVA-OneVision: Easy Visual Task Transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.737321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.737321Z digest=sha256:7d060c2bd650650dab405481f9a77a5abec58b392b9b070b399c4fab80a0ccfd

Observation ef6adfc5-8d40-421e-8d1c-141d5467603e · outbound

This paper cites Can video generation replace cinematographers? Research on the cinematic language of generated video.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Can video generation replace cinematographers? Research on the cinematic language of generated video

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:06:17.112505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:06:16.740457Z digest=sha256:a3934770e972e70ce8646f18fdbce42f1d2462f89de35905b5be4b7a122ce063

Observation 8ba3fc51-3215-4174-8864-3070c349db14 · outbound

This paper cites OpenAI GPT-5 System Card.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation OpenAI GPT-5 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.743214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.743214Z digest=sha256:eddd7d7d7cbb02de6d5c20eda6e0051a4d67d5be619ee604d93be0ff718c1e46

Observation 0e44e870-1d8d-47c5-bb89-dbbee7901342 · outbound

This paper cites Per-class difficulty and the long tail.Figure 9 reports frame-level F1 for all 20 direction-aware labels.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Per-class difficulty and the long tail.Figure 9 reports frame-level F1 for all 20 direction-aware labels

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:06:17.693248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:06:16.767341Z digest=sha256:a2d25a43e3f99c555bfdc6e68c66d63303cc237078a0839e737deec102b3b962

Observation 13730a63-c831-49d2-8bd9-5f1a2051b3b3 · outbound

This paper cites Cambrian-P: Pose-Grounded Video Understanding.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Cambrian-P: Pose-Grounded Video Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:06:17.020445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:06:16.754066Z digest=sha256:f379386e964d3f7d3f8f1e2e93e9d2aa89d132e7ea9d22d4e0bc48391ed2d612

Observation eacf236a-1c70-4939-b6bd-afc919ad34b1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.759467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.759467Z digest=sha256:67327faa71dcd9943937f96acbc1476b1565f83d5097e42d779dee9c4eb8bcb9

Observation 932e66e0-37f0-4116-ae26-8141d93de4a3 · outbound

This paper cites These pipelines can be slow and brittle under low texture or pure rotation.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation These pipelines can be slow and brittle under low texture or pure rotation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:06:17.709719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:06:16.762050Z digest=sha256:99ad591b1021a14b2160359bd11a39f812a6b2615d757cdc74d40556785d687e

Observation 9fcab7f1-5cb8-4dd2-af0f-7d28cb9472af · outbound

This paper cites Large annotated resources such as SpatialVID (Wang et al., 2025a) further support progress in video geometry.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Large annotated resources such as SpatialVID (Wang et al., 2025a) further support progress in video geometry

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:06:17.701653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:06:16.764609Z digest=sha256:8557ecee19fc6604def81db24d58fbc0146b1d023032b231de9c66d9afe14cec

Observation 919fec3a-63a6-4dba-b4a8-1a2f85c64d0d · outbound

This paper cites Geometry-guided camera motion under- standing in videollms.arXiv preprint arXiv:2603.13119,.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Geometry-guided camera motion under- standing in videollms.arXiv preprint arXiv:2603.13119,

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.727896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.727896Z digest=sha256:b652fa3e3691e654475e856ccc971d235adae1e22d847390117658ee752d74e3

Observation 3c7af0bf-4133-4724-b9ea-e5fa8756da28 · outbound

This paper cites ViPE: Video Pose Engine for 3D Geometric Perception.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation ViPE: Video Pose Engine for 3D Geometric Perception

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.734244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.734244Z digest=sha256:9ccd975046496f7297e26506e84d5db0777cab2686e8f31e2398a6ae197788e4

Observation 4f0defbf-9fb8-4006-bcc5-bd6e43443222 · outbound

This paper cites Qwen3-VL Technical Report.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Qwen3-VL Technical Report

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.722169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.722169Z digest=sha256:b60403b68b8b578d54551de89cb19de7a81ad43a76140bd3bfd96706274664af

Observation d506abec-4dcc-43c0-aa0d-b756d3525552 · outbound

This paper cites Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.748622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.748622Z digest=sha256:18efce52a113c81f26b2492e0a3c1664efa79905651f2769c2961de28eaaf5c5

Observation 96cb87b2-6162-49a9-8064-6cc4d9067153 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.745790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.745790Z digest=sha256:daefdbe9a947869c26f801cf330759f0e2d6ea865551ccc38b75889fe858bc6a

Observation 73db782a-147a-489f-82ec-03ac4a884bdc · outbound

This paper cites On the generalization capacities of mllms for spatial intelligence.arXiv preprint arXiv:2603.06704,.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation On the generalization capacities of mllms for spatial intelligence.arXiv preprint arXiv:2603.06704,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.756804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.756804Z digest=sha256:4a8234be09d87c213c7ba9e3ecd2384a97fc0dbb25d0d62dff7341c666da5b1c

Observation b520f7e9-dbf6-4e28-b679-a1633d6d9d15 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.751373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.751373Z digest=sha256:4e05093caceac38c97cb261a7870e014fc0bd9afca631ca00bed9c8c61313c8e

Observation 643ce842-46c7-48c9-8ce9-c661f39bd7e2 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.725005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.725005Z digest=sha256:3013530f67e8699fea3949c9bdc50557898fc1a80084ab8a7f23799b2800f97d

Observation 6ec89010-04af-4130-aa4b-7eaade060fe8 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Distilling the Knowledge in a Neural Network

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T14:06:16.730579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:06:16.730579Z digest=sha256:7c15bddcb5a43ffe0406d3b25f4d8c88156319e3ad50a447806fe068c42585fc

Pith citing papers

No inbound Pith citation observations are available.