Pith. sign in

Paper Citation Record · LEDGER

Self-Chained Image-Language Model for Video Localization and Question Answering

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2305.06988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.06988 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:52.748814Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.467463Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 01864bbb-665c-4f1f-bc48-dbb15c4be93d · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.509676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:573908d8146f4066a41f7c09680eda2661017b544b8b794f3fe09c685889ed7c

Observation 5b099f7e-35c1-4349-a697-ba25452e77cf · inbound

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning cites this paper.

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:52.748814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:52.748814Z digest=sha256:05e0ff6117c4f551ce5842bf9919e715921db6430ac7cbf09f9cabc0b5169342

Observation 0a16f30d-b100-4137-849c-bc3eb44ebdfa · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.185509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.185509Z digest=sha256:cb3b0418e7a5016c00b88b5aee8c23a3b46ac9b0609c95d87005886c075a8aee

Observation 1d09069a-466c-42fb-bde3-01b85e1e5ce5 · inbound

Rethinking Video-Language Model from the Language Input Perspective cites this paper.

Rethinking Video-Language Model from the Language Input Perspective Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.468853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T13:24:46.360149Z digest=sha256:5c87fe4814eb567a9be724ef38ab09b940a62f860d1982e6c7080ee9ea2e961d

Observation cfa677f9-9df0-448e-8108-7469cf311057 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:0b3812910e300d2857dc51ab82eb81c9d16f2463e00fb8eb812ee648f316e437