Pith. sign in

Paper Citation Record · LEDGER

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2508.03039.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03039 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:49:42.909084Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0fc53e2-d681-4262-9f49-feb7351126b9 · outbound

This paper cites The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:44.836759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:37.193595Z digest=sha256:a2954c09cea5ad5787d7f1853832f70f0fe8e27b3ce368fcaa43bc05cc81dbf2

Observation 8916b5d6-5696-4ab6-85bb-0c4570017a2e · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.262227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.262227Z digest=sha256:77294a6cb3ea5e6f12505f497d239b8acf38cb9e3258d7afe90a4d0274154f45

Observation dbbd0beb-e5af-453c-862f-f1b19e196efe · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering A Short Note on the Kinetics-700 Human Action Dataset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.375757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.375757Z digest=sha256:f0f3d91f95c5d450116aeca002ce216e58617506324749fe5738b9e04e6a69c5

Observation 457862a8-afb3-48f7-85d2-efa803b70330 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.464967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.464967Z digest=sha256:df2ab790082cb6b7fd61b146749c9af203285cd2eed6e688e00fe3a8f8029114

Observation 4ed2ed85-0e9d-4dc1-a83a-a32078302cf0 · outbound

This paper cites Enhancing Long Video Understanding via Hierarchical Event-Based Memory.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Enhancing Long Video Understanding via Hierarchical Event-Based Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.562864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.562864Z digest=sha256:6faa9e2f6559d2be4af4f4437e64413f2b51df288fabe2e4c89efaac7bf69012

Observation e691ea34-fc65-4152-ac4c-de2035cca43e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.643512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.643512Z digest=sha256:7ad1637c03d8c8b98a7c8bedf56994f80f4d5f13e04068f187538e923c814078

Observation a74b6568-0c14-4b40-bcfd-65d75c08ab60 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.982958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:37.738710Z digest=sha256:6c5bede41b8e94765c6e3026bf3ccf2d5dabc193d19df082ea7726c9386af2c4

Observation 59dd60d2-be20-446d-ac9c-a41ce768a327 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 8

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T04:49:44.688355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:37.840755Z digest=sha256:a33a68d61ed4de467b8de5d53d10fa2d0b46ec910ffed4ad7ddb64894dfa2ad4

Observation d8855dac-cc08-4c80-8e64-ec425fe59203 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.951565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.951565Z digest=sha256:75768d0d40afcf5b206936675471b3f1a6fa305e0c682b64c94b59f925c8fca7

Observation 39f00fc1-9517-4c66-b420-af0ad379cf9d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.032079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.032079Z digest=sha256:73f916a74dcd2534ccad79512c6db593c8daf422c7ffd65831b01c4fdea6d2c1

Observation 4c12f38d-35bd-4389-bdcc-01ab01d93e16 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.149590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.149590Z digest=sha256:f7eda0e8e2a22d60f2558373a7cd617c6cc31e46e84bbda997a735425bdc5c25

Observation c69bb0d7-baac-4ee6-80ee-9b06a3803f67 · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.388595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.388595Z digest=sha256:57e9ddffb44a0202a24bc7470acef5a7bccae0c23d3a5862c35799def593b916

Observation ada452eb-d963-4838-8da9-cf70e5235c73 · outbound

This paper cites BIMBA: Selective-Scan Compression for Long-Range Video Question Answering.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.487988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.487988Z digest=sha256:31fc2a014b944ed719fa8c526e5da780aa6cf942bb6b44a6d8d28e49166c1b0d

Observation f4bcf83e-1057-4518-8e37-70fa5b0f5f89 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.734096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:38.592793Z digest=sha256:ad22bac1d4c87d911d6ef8b1807380774e698052331ac2dd798f58f6a8f4d003

Observation 3abd9c83-6526-4618-8dd2-b7e698badfed · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.715526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.715526Z digest=sha256:e0e651c19a564e843b90f406353264f5fc469ff1fba1286737e7b2bf7871432f

Observation bb82cd92-7085-45ee-9194-c19c7c338fa8 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.819208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.819208Z digest=sha256:2e5d6e7874d7164a5f81cfcda6f729efcd69a90d50b8509bcfa07519a9e04c2e

Observation a6aafc0f-31fb-4274-a25b-06f213b47ac4 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.925919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.925919Z digest=sha256:29534270900781684e659d681ac7230f732d184be8d2ee5572af160a79eaa582

Observation 215cecf1-67bb-41ca-83fc-6a30a3038aa5 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.008817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.008817Z digest=sha256:644a5f631ee545b436ba8da3b80abf516e6d3841123cd061dfcd8206de1670e5

Observation 880ab0e9-887c-47b5-bbbd-b542464beaca · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.656764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.090836Z digest=sha256:e86588c657199806d20786a7f2e23b305e8a1f445cb90978337d7654203371e3

Observation 4be1116f-7d98-4a42-97ed-97c8dac97951 · outbound

This paper cites Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:44.268102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.256408Z digest=sha256:7156b7bf4086fa8d3990fb9ec9bd43d4e2f77a8e764a76b76c6e19a22d155219

Observation 1fe38cbe-5629-477a-b57f-8a8f54bea77d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.355423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.355423Z digest=sha256:a9344c2a3de0b10d7f93bb4dfa4fb7c1e46ba6d08d9df1eba1090af2766ae857

Observation 43064c23-18b9-4948-b091-01583c16cbb1 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 22

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T04:49:43.187724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.449214Z digest=sha256:499cd357136666466c4774ffcd64588d100f3d9340e89d394573a254915b2cbc

Observation d55e937b-66a2-44c1-97c4-9820b5cc216c · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.517323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.548025Z digest=sha256:ba5cc78e2458ec2b2da72b21208988128720bd99a8b9ee4c50dd279a6abe88b4

Observation f6d41882-8c87-4c39-8e1f-71d88ac17d0f · outbound

This paper cites FineAction: A Fine-Grained Video Dataset for Temporal Action Localization.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering FineAction: A Fine-Grained Video Dataset for Temporal Action Localization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:43.956206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.607583Z digest=sha256:cda683637bad288292c281fa07a17b164e1f4f1dfd377e90165b42d0253b603d

Observation c9ffc015-b9e8-4de1-b45f-f0574edc4c8a · outbound

This paper cites DrVideo: Document Retrieval Based Long Video Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering DrVideo: Document Retrieval Based Long Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.709238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.709238Z digest=sha256:6e283fa44d0f93f652fb1d7ef397a2e951d366936abb5ed5db132ba99ee644e5

Observation 4f84d07b-d150-4a22-a3e6-d5af65ef3d1d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.834064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.834064Z digest=sha256:f4e6d80d0bd93db46f4aac87709c3484d3a086c1d979aca01ad33f7c230fd124

Observation 085a55b7-2b21-43a7-8a36-80c70466fbb1 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Foundation Models for Video Understanding: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.068720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.068720Z digest=sha256:bf1a1535edeb0932c5fa1c2e61c3e0829554dfa11bb4f1d0d19fb6b468847c86

Observation a6dd47bf-e0be-4ea9-9140-bfcfb4aac0f3 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.186187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.186187Z digest=sha256:f171c6f3453f09cb9f390e8a7965f0442ae95f9f97b508f098eab3ea2a13379e

Observation 58c488f5-236a-4915-8b1e-458f1a1cce65 · outbound

This paper cites InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024).

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.952238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.952238Z digest=sha256:78b216ab2955d132b506839b3a08840bdbcf2646c3f9c9de19e072c62d98bceb

Observation e049e15f-6f1d-43c2-b1ff-b5fc0c38de89 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.413518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.413518Z digest=sha256:19de5283fda93619451c293988e5a91917e2eb967832dd7eabbc1ccbfc8ed719

Observation 48bf7a61-c99d-485a-a7e5-114d49c266a4 · outbound

This paper cites Qasim, R.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Qasim, R

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:49:45.411683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:40.562692Z digest=sha256:8927503de83428bd4395160887248072df86f1a5fddbb196499d1fc6893a23f7

Observation bfe8e4e3-9465-4421-9f4b-1e4d1c9e580b · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.297857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.297857Z digest=sha256:4cf8eb1105ed641c9053b2241b213abc38fe7dd6edc98a637b23043f04fc12df

Observation 0404bfd9-8da6-4e13-a968-9f43d86969a0 · outbound

This paper cites TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:43.602596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:40.870463Z digest=sha256:8395bd2ba7e60617e247d9a677ea46153f29f1ca2f62a9584fbca7bd9347fd5f

Observation 384eec76-6808-4fd5-8fea-1a8de3ccfbc5 · outbound

This paper cites Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.021647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.021647Z digest=sha256:bde6ec759362df3deaa31f68b64d55f7083db74ba6e4f2efe8ba397365bf0e2d

Observation 4d5c72c7-0faa-494f-b3f4-6453f13a14c2 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.708535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.708535Z digest=sha256:836cc2acbf593f97cd7e5930252cfca559ab9cd37f79e1ba029eaad44d507305

Observation eb9e33e5-fe94-4d8d-824e-90f52617c15d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.322353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:41.240798Z digest=sha256:2b837ae53e09102bec0473d3ea0cfdbfe501522c8c197884f0fbff48b14dfa39

Observation cdb02b9b-67d9-4830-be24-be0b0b3b54f8 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.380659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.380659Z digest=sha256:9e1acd97a485daaeafa68ba874e53ce96a5f2cb4c51e946220fde558169fe668

Observation 0dd44e4c-e53b-44f3-be17-484d8fd9ff9b · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.124471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.124471Z digest=sha256:584a23b847d1acb2e7df324dcd420d242273e37b57911ed2e7dcef396d5366e6

Observation 72be860a-0026-4722-bb6d-14242836a9c7 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.166685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:41.636346Z digest=sha256:3c0235938c1a51e7b9227a1398e61d95e5fc23ac8ee31d8f20978ab84da0eabb

Observation 28fd947c-bfe8-4d0b-8954-e19b2f9e491f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.776592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.776592Z digest=sha256:0671609354043c0c8af9d996f91124c59ff47b9502659dceb8ef16092f9c12e4

Observation 2ac3d27f-7c44-4e0a-be78-6d8ba40ee92e · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.523576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.523576Z digest=sha256:885e3d161bf0b34866542cd6b6f3cd3e8cc25ec9a7a719ce28bb7dbd39bc033c

Observation 20751638-ba9e-4f60-84aa-533399259549 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.003797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.003797Z digest=sha256:d75d077b92be88a16bde4e1dc08b99c30ef05d2450b4c446ec19ed2c1fc9d312

Observation 7633f68b-5142-46fb-be42-0dcf7ba2cfcf · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.097293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.097293Z digest=sha256:079c751bee2aa21b20e5d072bec248907b30dfdd3939d9f55d747cab5c925c50

Observation 0f9d426d-bf5f-4d1d-a70a-ea7ea25b8dd6 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.086855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:41.911771Z digest=sha256:c7cfca1b1833ee0cc6cdab356a8ba503d544f7951466e3594a10545a56c64418

Observation 34e11361-29a3-4833-90e7-58ddcf13c67e · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering A Simple LLM Framework for Long-Range Video Question-Answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.312571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.312571Z digest=sha256:b2249c8ede5481f87d6e20bb5f08fc19d365782943d4b4d6859534cb8177450f

Observation e28f971e-4ef0-4741-b281-96adc07bbee8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.423364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.423364Z digest=sha256:5b4960412306f208955d80a8eb7e8941cf43c94d4b776ad6a33776e8c17a5123

Observation 0a16f30d-b100-4137-849c-bc3eb44ebdfa · outbound

This paper cites Self-Chained Image-Language Model for Video Localization and Question Answering.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.185509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.185509Z digest=sha256:e7893b2c1ef7668307b57716def0cd462ae9e9a29b6ab302744e284c2b9b1d5c

Observation 9efbab19-1769-48cb-9f6a-b522e115e48d · outbound

This paper cites HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.667270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.667270Z digest=sha256:c8bdcf503c7fb90dbf54dd0674c8675883f0b2f159bf08c410a8a7f6ad222c48

Observation 22477e2d-58ab-4c4a-bd9a-878b1ce5b9da · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-06T04:49:42.770981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.770981Z digest=sha256:e43cd6e271907897a476d88ec66bc13dd5435652f26dad6db54e8871fa1bc51c

Observation 274ddc68-04d3-4e29-b61a-c2eefc215d8d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.523814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.523814Z digest=sha256:ac8c7bcb8e7d24fad09bd414782d7c00521f91e2681091a836627d1df8dbb815

Observation bb1fe118-e6d6-4eb4-a643-551b0be153ba · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:44.985780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:42.909084Z digest=sha256:96acd02e256e9d732414e16768305ca5a49f281e5fa098138a46667d2d865a4c

Observation ea88ecab-50cd-4258-909f-1cbeec8dc364 · outbound

This paper cites https://api.semanticscholar.org/CorpusID:1710722.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering https://api.semanticscholar.org/CorpusID:1710722

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:49:45.879225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:38.295335Z digest=sha256:33d896077437384a7a66bdbd765340df1f794da46501aa40ff6a0d3381f3a009

Observation 7d166b95-e89a-4822-9a73-8309863d4e2a · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.159644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.159644Z digest=sha256:ecd4ab9bf535262dc7f68121e31472e10363eae53b63fdc60d7185c45955f807

Pith citing papers

No inbound Pith citation observations are available.