Pith. sign in

Paper Citation Record · LEDGER

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding

As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2509.09263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09263 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:28:35.241017Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T21:37:55.887477Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6cea1c6b-09c3-4af2-9218-279d935743ce · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Flamingo: a visual language model for few-shot learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:30.931046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:30.931046Z digest=sha256:e76b537266a592a4d24cf7ee382d45d39c90ebb159b2dbf10cecf29c9f688933

Observation 6c75c933-e1ba-4588-86d6-462aacd90fcb · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.043095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.043095Z digest=sha256:398d8a504edc6058ec1f76ce4472ed794a44b16faf82debdce847b29ba9ce634

Observation 5dd4e9d0-4a99-4d0f-98db-59072db3e3dd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.153924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.153924Z digest=sha256:ef18af9d3d6f03199b29de021503a9e78df7583bd91315b48689f152780406ca

Observation e166203b-70e5-4cfd-8f51-2b98926b70e7 · outbound

This paper cites Qwen2.5-VL Technical Report.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.349484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.349484Z digest=sha256:257ff74dbba53aad6e69658154b5aec9cf9414e5c6f80d6c21bdfbdf6933a825

Observation e97afe11-87d4-490f-8de7-50293f2847aa · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Roformer: Enhanced transformer with rotary position embedding,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.468049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.468049Z digest=sha256:1e29a2e5266f9ad38a0c852b74d2d4b17a31a87d4500c4d43ed1770e1d214afc

Observation cf1d15db-9e10-4c41-918f-dccbe1e2c010 · outbound

This paper cites Adaptive Keyframe Sampling for Long Video Understanding.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Adaptive Keyframe Sampling for Long Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.595682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.595682Z digest=sha256:0ad7481905da0a4ecf80a62d463a1ac8e42bb0129d5eee3ad36aecef5207332e

Observation 29e0a2e4-663f-4f7d-943e-8ffade73c217 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Learning transferable visual models from natural language supervision,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.720004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.720004Z digest=sha256:24f4c77127ae7fb243b31827b3882fcc84e9895a941bba584556e7aace5734dc

Observation 6634ecc5-1a9a-4aef-b5a1-b43d5f433feb · outbound

This paper cites GPT-4 Technical Report.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.798706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.798706Z digest=sha256:d8b8082f1bedced106640d12f4578f0667d06ecfa15b6039beb8536b00eaf788

Observation cc6ebdbc-ffea-4b5d-89ec-aad8bc74d4bc · outbound

This paper cites Language models are few-shot learners,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Language models are few-shot learners,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.899722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.899722Z digest=sha256:40e585664f39a009bdc0c011f59cd9a4f10877cfceb9e80e8c01c762fa3338d5

Observation b0bfc4e1-4f82-4885-950b-01b2402ad1bb · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:31.957700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:31.957700Z digest=sha256:690dc47658e6971838e9f07b195777ddb00c1c96eddc51e3b2fdd142182c3020

Observation ca93cbda-637c-41bd-8fae-f84d434f64dc · outbound

This paper cites Palm: Scaling language modeling with pathways,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Palm: Scaling language modeling with pathways,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.155504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.155504Z digest=sha256:ae0bd53550d2a65da770b0531c5b321ec21c0c827edc2881cd7580774644115f

Observation 0e27cd1c-cb58-48de-9798-7d54c9a12564 · outbound

This paper cites Scaling instruction-finetuned language models,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Scaling instruction-finetuned language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.255963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.255963Z digest=sha256:a1f3463966c475cac479d8bbcdcfce3ae8c95665398f2557cf783c4547d014a1

Observation 7c3db572-5161-4eea-9d24-ce54b66efbae · outbound

This paper cites The Llama 3 Herd of Models.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.332129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.332129Z digest=sha256:1e1ac6f97d619ca84d1a2a5d1eb9c8538375312b7a794a08ff70128977c15ecc

Observation cce955e1-8ccc-49e4-9d1a-fafbafd35fc2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.409193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.409193Z digest=sha256:c222ae8543278737e3a507ef2e25b6ab5f8d7c8ba0f79fd0cbee81f7ab3259cb

Observation 5ce0c97a-7680-43e3-a505-1d8127a1d2fd · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.482894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.482894Z digest=sha256:d3628d3a23dd9584044ff96dcc45dd99dd9492b9cea009a3fc4b635070fefb5c

Observation 46c9fea0-ba0c-4c96-9e4f-5f447ab6a67b · outbound

This paper cites Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.552972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.552972Z digest=sha256:18ae32d163ef1c2cae65790da9115a5fdabe340e40ce70476878672650588f4d

Observation 857f07fd-9954-45a6-ad11-5ab223d512e0 · outbound

This paper cites Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.618929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.618929Z digest=sha256:39905bf5d82368cc61468c415ffd0343e743e2e692e25f55e1c5e6e1a9b94e6f

Observation 55a0459e-e30a-47eb-93d2-2ea7c978ded2 · outbound

This paper cites Lisa: Reasoning segmentation via large language model,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Lisa: Reasoning segmentation via large language model,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.687197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.687197Z digest=sha256:40142b7173c49e5b7c5fd7ab0ffbb7f82f5f123d89873403e1b07e14150293f3

Observation 5bfac942-3429-4c47-9b39-dfd7e4d5055a · outbound

This paper cites Visual instruction tuning,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Visual instruction tuning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.742111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.742111Z digest=sha256:7270b1b3cd37aee3ad090737b61951b7043037fd2dfc8b8baa6d5db821384f24

Observation 4c588222-a42d-4351-b04c-d24c9f3e7c07 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.849254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.849254Z digest=sha256:21e1dcda4a2871f193effcc76c93f2361aa5af0b7f828480fc86b473a970e291

Observation 93b70808-69bd-43bb-b880-3955f547ac61 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Sharegpt4video: Improving video understanding and generation with better captions,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:32.973420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:32.973420Z digest=sha256:281a7499cd9f71c63b11542b67d90fe329321fb503d1d5efba8ba117add38ea1

Observation 147cd8f3-db37-4d87-995b-fb11bb041f78 · outbound

This paper cites Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.032028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.032028Z digest=sha256:0d5b48e78dcb5f6c1935bddc50d60e04f9a936ea31c2b3869d32fb4e71ad518f

Observation 476d495c-d00b-44d6-919c-42593243e6ca · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Morevqa: Exploring modular reasoning models for video question answering,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.121208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.121208Z digest=sha256:e637d6c369a8919678b5712d3159f660f80f539803dfdccdea19fee4c02408ab

Observation 7efcc816-3e85-4a3e-9cc1-8dabf50b7804 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.205521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.205521Z digest=sha256:44e8f1b82ea8c84392bcef76f8413456ab179b0f3624dd27d17011a411f89c49

Observation bb5557a8-9227-45ae-abb6-bc505cfc1a82 · outbound

This paper cites Negative sample matters: A renaissance of metric learning for temporal grounding,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Negative sample matters: A renaissance of metric learning for temporal grounding,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.305062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.305062Z digest=sha256:e8d87b4ded6515a28fcf29084b546830bbe7787182a3284fe57540050baff0a7

Observation b59b4d2e-d85d-439c-89f8-bb58049600ca · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.385051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.385051Z digest=sha256:ad2415d676ee9bfeadbeaa63e82983772d5c3d2716ac837c21e948f18ff2ab45

Observation 2a6b7a38-761c-4dc1-80c4-79e67f6e1854 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.485739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.485739Z digest=sha256:2a2cea0081c02cd6586ba3a7afdbd4e2f35ee91ac3f5ba7e25df6b65697d34fb

Observation fdb33a5f-58f8-4185-8734-36fbeeb8005a · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.593857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.593857Z digest=sha256:88674eceb2ab6921ba4de59ce0436c4c1722e69851915df6f2752c1be4743b29

Observation 98bc26e1-b95b-4169-a7de-4e94406e3e86 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.688874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.688874Z digest=sha256:cb3a166d8e36c3f0e1edb5c0f213d8e6e768eefea63fcc6eeac4270c9ad35061

Observation b36bf2d3-60fd-49bb-82f2-07bc32b5146b · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.774070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.774070Z digest=sha256:1cc54ed3aed191450654335328aeb7d97e31f0300501f341083dc9a37355188e

Observation 6b128814-feca-4e5a-bb51-019d44bc8eb8 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video- language understanding,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Longvideobench: A benchmark for long-context interleaved video- language understanding,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.879367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.879367Z digest=sha256:684bdd1a77194cdb39a1dbb1b1d8bd0187e6a67967d404f9c718ad73a3ef797e

Observation 69d74be3-8a17-49e7-b34e-6e81746dc78f · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.960072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.960072Z digest=sha256:2d46d49733996f60321b708c5fcd90e9b86809a7408d64da7115e8c96504dd20

Observation feb1c0be-5028-459b-94a2-88cc60b01236 · outbound

This paper cites Interpolating Video-LLMs: Toward Longer-sequence LMMs in a Training-free Manner.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Interpolating Video-LLMs: Toward Longer-sequence LMMs in a Training-free Manner

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.040878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.040878Z digest=sha256:2369f1260114619fd62368cd355aa898d2e49ffd9835aec75eba3e0730f0d1e4

Observation a75667df-189b-40d5-a2e7-ab369304ba27 · outbound

This paper cites Long Context Transfer from Language to Vision.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Long Context Transfer from Language to Vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.150112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.150112Z digest=sha256:05aeadd84efda622986ddbc0b50cda502ab78e42ff760694b645703c7bb6a302

Observation 22e1696a-3703-4675-9506-ef2c6cae2c87 · outbound

This paper cites Visual Context Window Extension: A New Perspective for Long Video Understanding.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Visual Context Window Extension: A New Perspective for Long Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.247158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.247158Z digest=sha256:110a6a7eeea4889228c9a96308413a3e0fbcdd9aaf9379c9b21d8986d856a23a

Observation efb52306-5104-48a7-aef7-4735da755979 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.352539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.352539Z digest=sha256:d33f112c5c9ca9abce5bbad61995f27fee734fd6263e970963d6e2a5339a69c8

Observation c2c22972-8ae0-4bfd-8dab-8065d9bb208b · outbound

This paper cites AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.417628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.417628Z digest=sha256:9e8f5ba7a3641632b3b6d1663a3e519354f26c1892c2a3289c8d6b96e939cd9a

Observation bdb74b2a-7170-4308-8f2e-da0a092a7ddb · outbound

This paper cites Enhancing Long Video Understanding via Hierarchical Event-Based Memory.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Enhancing Long Video Understanding via Hierarchical Event-Based Memory

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.513278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.513278Z digest=sha256:e1b59672e0449d1c7ed4fc6003e04872febed1b6878841bae4a0f676f883aac8

Observation a2c89837-b407-4cab-898d-d6e0d3c2c885 · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.591425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.591425Z digest=sha256:4c9a8042b47780a739d77529b43df766eecca5f1d413b57ed0d01496ffe5c7ea

Observation c73f5009-f744-476c-ad1e-6f9c6992c300 · outbound

This paper cites Ma-lmm: Memory- augmented large multimodal model for long-term video understanding,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Ma-lmm: Memory- augmented large multimodal model for long-term video understanding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.696017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.696017Z digest=sha256:5e04377170931aadd8d964aa217d6bd1425f5a859baac4289bf92538165b2fbe

Observation 3b556654-4211-4f6f-b6dd-fc0ec69fb8fd · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.762522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.762522Z digest=sha256:f1abd542729f9999ec1f6ca43ac446b72337c8233f00b2aaba9ce80a1bc491c1

Observation fd14bf21-e9ae-4f60-9646-65fb995f1861 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.870488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.870488Z digest=sha256:d850878c31b8c153022d7fa099c389f57937273d01855b34f0216e63d31247a8

Observation 0ec38891-e23f-4f47-a80a-d702265396f9 · outbound

This paper cites Qwen2.5-Omni Technical Report.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Qwen2.5-Omni Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.963515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.963515Z digest=sha256:731ed37076dd3c40b498a423f7b4c47a71d218ad05ba152f1ef788d040f374fa

Observation 9800e012-7f9c-4229-b3cf-744921626fca · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:35.059185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:35.059185Z digest=sha256:ba3c6f61ed09a508bb3a984fa7f2cadeec0ea43fca12682bd2abbbb963fc85f8

Observation 65c7535d-2b92-4fda-8351-eed921ad8a2a · outbound

This paper cites BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:35.159055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:35.159055Z digest=sha256:68bd1b2a41fe8ebdd399c9980060568d7b04c78bbc3ecf0e87e8581b666ef050

Observation 04b4fa8c-bc72-4680-8df4-4e81a2d28cd1 · outbound

This paper cites DeepSeek-V3 Technical Report.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding DeepSeek-V3 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:35.241017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:35.241017Z digest=sha256:48f4c482437a05e467627e4675e8641298c584771cf9fb31a0ac366a0420b5f9

Pith citing papers

Observation 5c04fc4a-164b-495d-8250-15706bd20a54 · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval DATE: Dynamic Absolute Time Enhancement for Long Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:88c48cf4d4a6bfa2230125574bc6d5fadcb9902b7f6bbd760774f871309cead7