Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T14:19:34.622854Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 6 inbound Pith citation observations for arXiv:2505.16416.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T14:19:34.622854Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:17:12.549851Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-01T13:05:45.066160Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation be3155ce-62be-48fa-b442-2733e97c150a · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 99a3544c-aa83-4909-8c4e-2ecf90561e91 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 296a26b6-e263-46aa-a12c-60b1eb267054 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 050d2c3b-ca16-4e0d-857b-5540d678d2de · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2028d01-687a-41f2-a5b2-7f76926d2d2b · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d21366e8-78ce-4120-92db-4a02f47b2928 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b532ea4-62d2-495c-80ce-b2bdeed8c3b1 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models On Path to Multimodal Generalist: General-Level and General-Bench
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 703e8ac5-584c-48aa-bd45-2f71d209e61b · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 622429d7-7447-466a-bc82-9128e192904b · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models A diagram is worth a dozen images
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97348b64-4114-4889-9a2a-b871dc941f75 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5ca492b-1455-4894-9c1a-f14b39c3a1ca · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Transformer-based visual segmentation: A survey
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 823798cf-2643-4c61-8cfa-936ac8ae3ada · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Baichuan-omni-1.5 technical report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 74310671-fa47-4a38-a913-840134086350 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Visual Instruction Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 42de5efa-0cb5-431a-83ba-d78fd3927afc · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bedb7fab-876e-4f26-929d-834c6f0dda3a · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76420bee-e81f-491f-86c1-0dd85a6789b6 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27d21c87-cfb0-4cf0-9ce8-3a1f3b7b226f · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models InfographicVQA
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a5a7488-3a23-49d7-8382-328a111d4aa5 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4698198-98de-4154-8165-dc4cbc530d78 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 960d2ccc-21ba-43c8-8c42-92b1a788810a · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73719d88-6bd1-47d4-87cc-2276caa773d2 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Emu3: Next-Token Prediction is All You Need
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a78a7883-6d7d-4472-902f-582dae61d412 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7a351f35-fb6c-44bb-b844-ce62efad7530 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e53f9576-cd13-4120-8856-d0129d3a0d11 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1866e258-2d55-448d-b56c-dbcb34d455ad · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0668538a-cf7b-40e9-a8cd-75764d2ee628 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9a31ce4f-9beb-4920-89ad-5d56ac396aff · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82e1864c-7d36-43f0-a806-d9d4663d0c17 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc562104-9b98-43ea-937c-f9a34300c4a6 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38a3333f-6f02-4b51-b56e-9b78381f26e0 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Pixel-sail: Single transformer for pixel-grounded understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d43770c4-9261-42b9-901c-4092cd57a8e4 · outbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fa80d22-adfc-48f5-88a0-06a3fa545159 · inbound
RAP: KV-Cache Compression via RoPE-Aligned Pruning Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17295a5-fcef-4cf8-8c07-f730491a2d67 · inbound
MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8adb6880-4177-47d0-9622-fe1d16d75d9d · inbound
Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9cf1f5d4-1cb6-432f-bc33-1fe085cbacad · inbound
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbc14466-2750-401b-83d6-019f212d1863 · inbound
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b29f0c5d-5e88-4d88-a178-75f94fcd5ca9 · inbound
DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.