Pith. sign in

Paper Citation Record · LEDGER

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 3 inbound Pith citation observations for arXiv:2505.15447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15447 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:59.938701Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:28:56.047745Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T05:56:07.964144Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf12f22b-1d92-4016-afcf-f6ebf19bc729 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.425325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.425325Z digest=sha256:56f3ad49c9d520ba8936079ad3783ac7dbf7f07e8e5ec5bcd63ceef7fd6cc671

Observation 27ac58e4-903c-438f-8d0d-c130fca152aa · outbound

This paper cites Qwen2.5-VL Technical Report.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.475604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.475604Z digest=sha256:137adc2af8b88ed5b598c97b2565e042596dd40cc00f55492760b93f482ef9a3

Observation 342666bf-c2fd-4629-ad13-108355f1159c · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.552758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.552758Z digest=sha256:3ece1c88c9fe329024cf1bd713332fa3464cfe967f05b28b24dd8382f79cc813

Observation 6bd6c3fb-fa54-444c-8940-686612390029 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.612041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.612041Z digest=sha256:5a56cf7679caa33da042d8cd83738ca775b5727c80874c367a80f79136702c85

Observation 7df3f253-5ae7-484a-8a85-caa50585f1e0 · outbound

This paper cites Supervising strong learners by amplifying weak experts.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Supervising strong learners by amplifying weak experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.700519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.700519Z digest=sha256:e1436e86b97d515c3069c0f75bead4ab7244386f691948ca3c870594834e7e0f

Observation 4c265e10-9dcf-4610-9ba4-825985f4bb4d · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.775593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.775593Z digest=sha256:ff4760772cf43b339eb787d3b5705c55cfc5c1f70b09943e32db34b3d1e01613

Observation 71f78dde-a786-407f-a9c0-4cac870c7a37 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.873824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.873824Z digest=sha256:d104e9530bd264e54cb87838dbd5de22faec77fcd78d173eb13fc99b3c01a7b8

Observation dd2ede75-5c5b-4b78-bb51-b4595d376349 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.957053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.957053Z digest=sha256:b18eeb91285ec164be7633cbd01f0a26e8cc6642fe0f6eee230906159f6c9787

Observation f85709bc-2a0e-4bd8-bbe6-daaf4d5c108b · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.028711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.028711Z digest=sha256:03fce586a6c19cf4db4bcbdf265b025229b1fda93f09aaceb7430f8b0d96c6ce

Observation db2ab293-089b-4639-b534-f9246a05f5a8 · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.091933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.091933Z digest=sha256:d9507253f26bed6b5b9672e51c3a350bbd3f537a9b7dd06e94ec56ddb77fd006

Observation 0e53acd5-2aaa-4d8f-9a02-327391ec0747 · outbound

This paper cites M-LLM Based Video Frame Selection for Efficient Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning M-LLM Based Video Frame Selection for Efficient Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.174134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.174134Z digest=sha256:e6e4a5cb4db98f8f9ab642257782e4426495d511871e778144558057cd98234a

Observation 113e3bc5-eadf-4c37-92bb-29c545e23c57 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.242364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.242364Z digest=sha256:df058697e6297eed33c202c16618b250017e141f14dea955ddd43c3848bc29ae

Observation 0cf1332b-7be7-4177-b5e0-de0665d5f8ad · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.315079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.315079Z digest=sha256:b2edff57f64cc9aab4562abe86bb6c8a5e993a3b7093d25a4d5709f3e8e759a7

Observation 2e2e35f4-2831-4741-985a-97dad4e1fcc7 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.409103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.409103Z digest=sha256:ec5a1726d5239803f499bd685b67f6847847da1c3d2a12cc41a1b4fa8a2ea0f0

Observation da9d7c4a-a09d-46f1-aee2-ea29148cb717 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.485440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.485440Z digest=sha256:3b88581acd417c88591baadd2668e67bf70ff928aef974de4769a71512ace69e

Observation 720ecb11-afd2-4c2e-a1bb-ad1569845d9f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.562364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.562364Z digest=sha256:0b6da663dbd096cee5d4633de98ba6f9ed31496eab199917b0c1171311ea2b98

Observation 399c84bd-a020-40a0-a430-6bdb27af0921 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.634142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.634142Z digest=sha256:5aecad4257523fe829e3a9fa993517cba3ffc3f8589bd813e691378c889f6797

Observation 8335557e-cb75-467b-b018-6d8712a4a5b0 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.694251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.694251Z digest=sha256:84756d38f7115f1fa94ea7a8ec1e949ba4669388293f4112bb1bd64ffdf02616

Observation 3a699b2b-43ba-4426-a63a-3c27efa0be05 · outbound

This paper cites Hello gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Hello gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.763628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.763628Z digest=sha256:368c41c856fe5438fd68cd70cfeae14815d35d0a81a22d87d63e6845388485d5

Observation c66a802e-73d2-4f05-8583-4eb95c48834c · outbound

This paper cites Ai 2027.https://ai-2027.com/, 2025.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Ai 2027.https://ai-2027.com/, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:00.704398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:20:57.848572Z digest=sha256:d561178d684625687373cf7bc5e900ca2587f4fb5edf5f6967290229f4f41297

Observation 74c24c73-18f8-4c77-a611-7f3807c6e079 · outbound

This paper cites Introducing openai o3 and o4-mini.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Introducing openai o3 and o4-mini

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.917067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.917067Z digest=sha256:48c99ccd6ce33ae4fda51fbdc5b1cd33f5da3ab19aa471784fd6654c57663273

Observation fbfa3356-28f4-4ebe-b8a7-afdc691d7104 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Learning transferable visual models from natural language supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.009144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.009144Z digest=sha256:d9ddac10621069166e315002052703b83d8a65980533409d319615e0b30f02cf

Observation 436ddb25-59ae-4f29-aed5-494b4d635f43 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.095416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.095416Z digest=sha256:93cc029ff80f0185a995cbe3d7979e5abee2262944d40476e2b07f5e90b1193a

Observation d1e8c877-96e5-4be1-834f-8bdbff7da2c3 · outbound

This paper cites Trust region policy optimization.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Trust region policy optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.248304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.248304Z digest=sha256:12d67fa1bc802878da8199942c5977f492c32f5988d17c64f4acf19348cae835

Observation 9115f05f-a2f0-4b8b-9f4e-6fe819ccecad · outbound

This paper cites Proximal Policy Optimization Algorithms.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.362338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.362338Z digest=sha256:17957f8bfdb34c8c1eba55b7ab0de8d8276a1babd10c149068acbc197c4d9936

Observation 87adb3c3-3276-4d74-b086-2120528e9b70 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.456192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.456192Z digest=sha256:c0245ea9c1cf03dc9e46c3e32cb0c5554c923b45d66ec426a853d1a963c00b55

Observation 8d9c289e-2d5a-4082-a8ec-fcc72a4a1192 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.584507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.584507Z digest=sha256:bd31dc22658ad2679d56bc3c50f48702459513db26b6f16c75089507ca62c6ed

Observation dbe22634-eae3-4339-8748-7cdf90877cd4 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Moviechat: From dense token to sparse memory for long video understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.676634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.676634Z digest=sha256:644151cac4474bc5d5dbacfec7f7e89f0a83b41dbe182d8d0f6bb241a4ac5ebe

Observation a5fc3f40-2601-4cb8-a4ee-395a76b206cf · outbound

This paper cites Adaptive Keyframe Sampling for Long Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Adaptive Keyframe Sampling for Long Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.751103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.751103Z digest=sha256:220bff64120a43e5a397938e2819eccb97b36e4ea8380b71e22d900b4a588e1a

Observation 65c61772-c52c-4bbc-9c30-2dd5b0872121 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.826451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.826451Z digest=sha256:427eeac3eabf571e982f0bb7cbc3bd0ba63991592708c0475818b9d00425bf9a

Observation be03e616-7c36-4067-ba95-416a5bc8a359 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LVBench: An Extreme Long Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.975786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.975786Z digest=sha256:00d6d35f1016035158365b4fea7fe63ca48eecc62463c2c7bd83c657f4d6e21c

Observation 8ffb3ee0-50c0-4ea5-affe-48f2ee8e725d · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.072257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.072257Z digest=sha256:9c211b84860eed0aaa777f0f88fe7ff1faa80b6acaf246a5c9810684c441dc6f

Observation 1445af28-cabc-47dc-8c25-d4d3be3d296f · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.162625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.162625Z digest=sha256:b7a8d3398ef2cdf2ccb6ede13f5822466ddd88689e2109179d0258e203695d79

Observation 4e3c51c5-4ff5-45b9-bf7d-a5636c50b639 · outbound

This paper cites Number it: Temporal Grounding Videos like Flipping Manga.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Number it: Temporal Grounding Videos like Flipping Manga

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.262348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.262348Z digest=sha256:5bfa7ecd734a11affbc1778f854a0182269f87119321e01acded3f285cd8e691

Observation 7a6ad6ee-18f4-468e-92b1-152590b8d05f · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.359347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.359347Z digest=sha256:6104e98d7fd8d502c44c74eaef5de7e5836eb550efc94b2442e14e4b489ca67f

Observation 8961aedc-0025-46cf-a0aa-5d0cca75b54d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.427397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.427397Z digest=sha256:1f9b1e117537a29f9902779e9116c2972af14c9fd31e805ea8b1a3390866e3bf

Observation 686d382c-f69d-4325-9780-7c48d5248f4d · outbound

This paper cites Long Context Transfer from Language to Vision.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Long Context Transfer from Language to Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.552615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.552615Z digest=sha256:f46040c50e779e43a16e1286bee96322ae478a7c5f62e36f22bee754af5bbd01

Observation 75ba5317-b5ea-4f21-b12a-517e7916add4 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.633738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.633738Z digest=sha256:c2d48d0018fd2920ee373227137ffc2c41361551d5eab6b33aff69a9698ed2a9

Observation 2855bbcd-de4e-45a1-b7e1-4e80a065abd8 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning MLVU: Benchmarking Multi-task Long Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.712696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.712696Z digest=sha256:902ee2d851a7145612a46358a5dd62c7fc2825efc6b5c306a4331b0dcd919eae

Observation 28a12429-787f-48e2-805e-b199c7f94349 · outbound

This paper cites - Check if the occurrence time is mentioned.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning - Check if the occurrence time is mentioned

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:00.510642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:20:59.765758Z digest=sha256:2871d308e3c04d8288e7eec3cbfcb0ee2e47cc36655b72381e530b30f61aef74

Observation 4940706e-4e41-4fd1-92a3-f8c3efbb7516 · outbound

This paper cites an unresolved cited work.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:21:00.388794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:20:59.845090Z digest=sha256:6282cbc78cabf04f4a8e5cef565895564074e08f2c98d0ed6ba35a327d7585d0

Observation 3571a2f4-bc9c-4d10-a20a-9d55c9d3ba98 · outbound

This paper cites mRNA" and.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning mRNA" and

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:00.261706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:20:59.938701Z digest=sha256:adab9b2d988a77f5682a14b5b873931f3be27b6be49ad62bce3d035890bc91c2

Pith citing papers

Observation 8527e4c0-d110-4e6d-818e-d6e4fa602931 · inbound

Phi-Ground Tech Report: Advancing Perception in GUI Grounding cites this paper.

Phi-Ground Tech Report: Advancing Perception in GUI Grounding ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T10:28:56.047745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:28:56.047745Z digest=sha256:92c4934f06ce6da264fa3b86c49c53e88a25f378c0a2391414d304f1fa3db1da

Observation 125702e3-0b83-442c-b5ec-d8ac8b39a4d9 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.967892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:7d104c1b79ef02f36807d564d2847361fe33434d464623f152a73a3204580362

Observation 3638ec97-6a7f-47c4-81fd-d04f151aeebf · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.510039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.510039Z digest=sha256:16f899624e261e2cdebee7fc20e718069f71fb3b5bf158177ba9a6521a22fad9