Pith. sign in

Paper Citation Record · LEDGER

AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2410.03051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03051 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.853647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.928668Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2c41478-032f-491a-9901-03bd4df7a710 · inbound

SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory cites this paper.

SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:30.236866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:30.236866Z digest=sha256:205f5fc079df05506c306bb223b32789c10b622fe92d85c450c41e705c5dc0e7

Observation 7ea5e044-5a17-44a0-8c83-4e88c9eb8af6 · inbound

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts cites this paper.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.189481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.189481Z digest=sha256:78900dbf695394aeaa5140087069813f220ca85c72489aa2d3a39b33dd8d9e9c

Observation cd12837a-9450-408b-9fd3-ad712256895b · inbound

Efficient Multi-modal Large Language Models via Visual Token Grouping cites this paper.

Efficient Multi-modal Large Language Models via Visual Token Grouping AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.321557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.321557Z digest=sha256:97c097a8fbd31a162cc3f2e9e43fb70cfa881312765f10a40fe39dd7c60aa9b9

Observation f8955a32-afb1-41a5-a212-f21c16161038 · inbound

Progress-Aware Video Frame Captioning cites this paper.

Progress-Aware Video Frame Captioning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:58.284406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:58.284406Z digest=sha256:3d61ee4c1507c3de17c445e8c2c06dbd65442ae253096be8ea860356402f03c4

Observation 3ef28860-7006-4c2b-aeb7-fd5b7f854b48 · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:31:44.310669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:31:44.310669Z digest=sha256:50ab4c0a9a5b8bcc2ef79531dad1b4765a3ba7405639fcedf697320f2acada56

Observation 43f0032f-7b25-411f-88c1-c5f3d1eace6c · inbound

Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration cites this paper.

Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:02.058887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:28:02.058887Z digest=sha256:05a27439007fd38e43de9a26a063433077c13b1dfeb0f2b42bbc0ac01c468b0e

Observation 831294d4-84b3-4832-bc3a-b50709e32390 · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:05:29.239902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:7ce84876495f5904005db10d660c972dc1dd0225431e062eb936af29f79ec9c9

Observation e4b62b3e-d556-415d-aa59-f5a387c059c3 · inbound

MVTamperBench: Evaluating Robustness of Vision-Language Models cites this paper.

MVTamperBench: Evaluating Robustness of Vision-Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:55:17.536674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:55:17.536674Z digest=sha256:cdd3840a3665e2209fad11089939d85c676179434f7aeb54f373c7a0d28518bd

Observation 09174efa-e7cc-4fe4-973c-ca7ef49863e6 · inbound

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval cites this paper.

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:00.725465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:00.725465Z digest=sha256:a23c30416733204f3d0600b617b8d2212718f5d259d681d1ab90123071a09225

Observation 28ef629a-2f2a-476d-a16e-7da6b4b5d589 · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:24.963803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:24.963803Z digest=sha256:8b9875d78b972bf22df027a2c475de65c1e34782b1256c69adc8cc5b05ce45fb

Observation 68652cac-f1a8-4700-8de1-5f779442999b · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.152534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:789749d55b7db439ab33751e44fd919402f99699e1b070cefd76e1418c67ebe5

Observation b5837546-afe1-4b76-bd4d-509e4deaf5f0 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.088592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.088592Z digest=sha256:30c1e746c351c00a3f77a6fe6c687db91ef9defe37a294b3f438f937c7a19a49

Observation 2537b212-0d3b-463a-bec1-b35f1d15dcc5 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.878492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:3a424776b1eb5b342022adde5f18c7488d33636e434210abdbd4f4d147ceb4ba

Observation 24faaf20-db2e-40d4-8566-e87b52a3e902 · inbound

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding cites this paper.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.853647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.853647Z digest=sha256:f39b83f5e38d69cca8184a1741f05a90d354bde240ed47f94afc423ed16a2ea4

Observation f3c2bf1c-8de4-446f-86e8-eadc1e6c7509 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.396187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.396187Z digest=sha256:1d0416eb7ec1dd2e987ce16c121d9b3ce16808c8a66dbbe483941c181210c578

Observation accaba5b-d68a-491b-9d1f-baec2974b67c · inbound

Towards Understanding Camera Motions in Any Video cites this paper.

Towards Understanding Camera Motions in Any Video AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.274698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.274698Z digest=sha256:7349da51fdea9b7b095c60ef8c038bda5e23e462226e5dbd2f008dc36ae59c56

Observation 647c0f4c-cd02-47b7-9f82-f91f92bb0ec4 · inbound

MR. Video: "MapReduce" is the Principle for Long Video Understanding cites this paper.

MR. Video: "MapReduce" is the Principle for Long Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.593130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.593130Z digest=sha256:295bcdcb4286163ed94a8d4dec1ebccdc2b8ecd7122d3e1c1d6c2d6f044cbfb3

Observation 0e8ddad6-119a-4e62-af48-654c8cfc74a0 · inbound

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action cites this paper.

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:40.071985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:40.071985Z digest=sha256:02c328d0cff0ab6fa97ad6c31ae489954ac5782d4a3ababa7db1f1ddb4e9ed7a

Observation caca77dd-0c7f-436f-a4a7-f8dbfe66e7db · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.686789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.686789Z digest=sha256:610263d19dba0581f0f84d4940b0fb09005c097ef76072d57ee8b033f768de4a

Observation 286408a4-2bf5-439d-b8c2-c1f170a644df · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.661583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:02:59.661583Z digest=sha256:6afa6255ad6adad32a4be9bfa105c75b67b5037df1b6e52fd2ac32bb03f843d0

Observation 14d775da-00a6-4a49-9a83-208061a1d3d8 · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:00.159271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:00.159271Z digest=sha256:b4b06a549f0ca42ce90c0459a9433ecd5a9e299782a9491ead3dcc37ddea689e

Observation ea527a28-a167-4066-bada-0ccd3e18f566 · inbound

ARGUS: Hallucination and Omission Evaluation in Video-LLMs cites this paper.

ARGUS: Hallucination and Omission Evaluation in Video-LLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.591343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.591343Z digest=sha256:5b70b1d478a9492f25d5564c650d0262a5b0f8a9a98d9f79c4cf41d497581227

Observation 3bf9bb16-7765-40f5-b839-9587f6f7bfcb · inbound

ToSA: Token Merging with Spatial Awareness cites this paper.

ToSA: Token Merging with Spatial Awareness AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.947783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.947783Z digest=sha256:2197c5fad16d66afe937ed431220779b772495a2f5d5321b0280a92539b486ef

Observation 2e94a8c9-1420-4ad4-97ce-400db8aba0f9 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.764060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.764060Z digest=sha256:b4e417a87a94369ec61d4dd1ed5db2f3abba441782181460c21dda8f363bcc09

Observation 7b21e852-8f3b-4f2d-9372-9a4dce0f8077 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.227081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.227081Z digest=sha256:7693998c0650468feb7f02fbeaf2f8981529565de857c4ded4d400f84e75dbc5

Observation 50b612f2-6f4e-4f4d-b273-6ace6548816b · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:26.518534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:26.518534Z digest=sha256:8b4b6e5b7993bbef35893d2253a8455894017ff3aa86a1a18ee3aaf8e742cb3f

Observation 307b969c-1c54-4663-83d5-2209560690f9 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:27.013397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:8adc3e6f0eabe08acfff988c32bba39bb9edbb8375b7d817c1c057578ac0e7da

Observation cdf41a37-fada-4535-901d-9ea0ad180cda · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.635849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:838bf8523f9074da1233475cd0cdcdd275c8ae59fef131e579053fc5e81aace1

Observation 9fdf966d-545e-41f9-a686-acde1a8d4700 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.576239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:ed865973f3094faeb51932ecb9cb086e1de64db98f13702c105cb999a92ae5a0

Observation 1f19ee89-936f-4124-b487-3852d4baad26 · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.544729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:e107cd51df1da88902f9dff21a4206e31a9cfcd33a27b65da1c70d1c67c25239

Observation 5715f666-a5d1-451f-928e-f5e25c060ee4 · inbound

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation cites this paper.

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.313840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T15:47:33.435653Z digest=sha256:933f0413427a1b67c325047be1b0f15f75c7ca4a0588b7af932f0feec2630b90

Observation 9c1e0d0d-2b0d-4e9c-a200-53bc6e838f59 · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.406296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:9124d5993c0dba055525f33c6d8dd50c165b6b66dac19f041ef8bb4cc5f5fb49

Observation 6098b20b-c083-409d-ae59-491209b36e95 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.382456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:4d96f3420aafb07da8015c5782ce988ffe55438f71cb5112ea28bb318d3a0e18

Observation 00fc8319-181c-4837-ad2f-97b1f9e2570d · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.575211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c2bb8d223e68fb60280cb8434b8b4f65925160db7344537e206f059c2da1ab6c

Observation 7aff1fca-7e12-45ca-a8f0-9ba6eb33cb73 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.272203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:bdba1cdf355af49dec9554e08a4e433c35ffc24021679845aa0d229102ef3a55

Observation 48128038-df63-4ab8-b3bf-897687c62909 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.930565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:744be7ed28eae6ad49a50a49578bd239612f498bdb92776d22f6d3d6021abe7d

Observation 4a2ab20c-32cd-4207-afa5-4298ace80d12 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.141536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.141536Z digest=sha256:2516bd6b07b3be78856982742592a2d8e04d7725e8d862812afda0d1e6f12d2e

Observation 2f5f4e4f-4e40-44e3-8325-75d037f4915d · inbound

Visual Token Compression Enhances Robustness of MLLMs cites this paper.

Visual Token Compression Enhances Robustness of MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.202730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.202730Z digest=sha256:d4ada94a2ad7ceb05ab9829d606f82addddf5605e7e2719f1e0e732f76a39d52

Observation ec9bfde8-4dd7-4aca-8880-33a90c8adda8 · inbound

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning cites this paper.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.515974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.515974Z digest=sha256:c92e76b6940673a7e5f9596fbb02a33ae41d337aef055160d6fd6f800c726942