Pith. sign in

Paper Citation Record · LEDGER

Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.06189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06189 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T06:02:34.065885Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.790278Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 950c5de4-7ac2-446d-abc0-7c45252c75f1 · inbound

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training cites this paper.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.065885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.065885Z digest=sha256:c11cb2ffa83a68955a49e92af76ccafc2e52a044a8986edeba46b983e3b793d5

Observation 7bbb85ce-bd01-4a79-94a0-db5e973566a4 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.831110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.831110Z digest=sha256:4407eacaa3a27a39df5b1b716dbb90cb6efe7a23417bc9930a41309c5b562a42

Observation 3c592c7c-ab04-44c7-8713-b9e59e2c950a · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.330222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:fdfa51bcefe48f88a671c9e6ae00cb394cba58eb2648e3e789c8a12250b7e159

Observation cbd74da5-b010-4ae6-9fc1-70794fdbb2eb · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.335661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.335661Z digest=sha256:c91b9153e8c7747ee1be23fe77c35babbd1d0cfb1fee3dc051e7a3b8a75437b7

Observation e637b7cf-0310-4c32-940f-16401aff3cc1 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.741475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:bc8c022d4f27d5256d5c0699c6759fcdd6eded737a2647a6d7e231e9e28840a0

Observation de93c1eb-8727-46ec-a8d0-58a2198a708d · inbound

See, Think, Learn: A Self-Taught Multimodal Reasoner cites this paper.

See, Think, Learn: A Self-Taught Multimodal Reasoner Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-08-03T19:04:29.148776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:04:29.148776Z digest=sha256:9d47e5fea64806b62270d8c92a6791891f5a3d663315d6e5363d52d7d0c12bd3

Observation 27b97f05-089c-498c-b445-df9a8fa132ef · inbound

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding cites this paper.

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.690880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T11:52:26.828633Z digest=sha256:070cd4dc528c20d686a7d52f30d29c73de8504c32e3183a24ad60df32a1fd291

Observation 082d1ef7-ff9d-4be5-88c6-a5803e6dd4ec · inbound

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding cites this paper.

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:45.455092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T07:00:21.192082Z digest=sha256:12ae5c6ae5653bc4aed1d2e5bb84bbd9680bbf70ee9abcbed892b4170112903b

Observation 8f2e2cf0-ce10-4037-a833-ae03c78aef61 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:40.791756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T12:14:35.109298Z digest=sha256:d1b0985260d11ce615c21a5e18bcd5b1820af4236b0ba8d40208ca028ad3c811

Observation ffa5e41b-1a77-46a7-8fe8-57b3a5625bdb · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:39.637541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T06:29:51.635039Z digest=sha256:a287595804595459e164577d0fbcd6c516ea6a6afb6eb64e1bdd76675b4b28d2

Observation 0e6c1eb6-3d0a-4e8f-8554-fd73f8a1d567 · inbound

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models cites this paper.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:4cf6ace8fb891c7c830f3f5ccbe228389051c2fa79b55b4627cd18a2cd3d27a1