Pith. sign in

Paper Citation Record · LEDGER

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

As of 6 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 6 inbound Pith citation observations for arXiv:2505.16416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16416 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T14:19:34.622854Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:17:12.549851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T13:05:45.066160Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact25
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be3155ce-62be-48fa-b442-2733e97c150a · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.716042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:b88bac515ffd019cce2274766dc0439b22980401bc7ca06117a5912ecf5d9faa

Observation 99a3544c-aa83-4909-8c4e-2ecf90561e91 · outbound

This paper cites Qwen2.5-VL Technical Report.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.720635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:436e873d70b3c0ed8af532a6ce3a33d0ee0ba90b4963cd6b483edf86ae3e43ed

Observation 296a26b6-e263-46aa-a12c-60b1eb267054 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.692249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:0cf225d5a946620285b178cc589ed13cf198d089a0be2f99479c101f8fbd4e17

Observation 050d2c3b-ca16-4e0d-857b-5540d678d2de · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.737877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:ff7509717fe61607d04a62e62b262220864c9bc1f895bc32391f459535ca1cc4

Observation b2028d01-687a-41f2-a5b2-7f76926d2d2b · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.726967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:f9fc8a6d28cabe8c1db30d1fa2deda4cf8f0e74cb2f4fe284f0af8b7df67025f

Observation d21366e8-78ce-4120-92db-4a02f47b2928 · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:36.652705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:1758347734ec90785382f425c5d64ff5f421e3a4e152119c181be65cf112a65d

Observation 2b532ea4-62d2-495c-80ce-b2bdeed8c3b1 · outbound

This paper cites On Path to Multimodal Generalist: General-Level and General-Bench.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models On Path to Multimodal Generalist: General-Level and General-Bench

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.794474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:678ba4de654b06f103fddcd5bd6677512c739b26edd87c4771449a1db6eec781

Observation 703e8ac5-584c-48aa-bd45-2f71d209e61b · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.771573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:bd10a2664ad4ed03c74a802f678a1821b8e54d3f5032f9bc77bcb16f9d813036

Observation 622429d7-7447-466a-bc82-9128e192904b · outbound

This paper cites A diagram is worth a dozen images.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models A diagram is worth a dozen images

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:21:40.161567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:26ce6bccb9cd06d48d29d70e9b10edb3b6ffea6ec6570fdfe3466b88989df502

Observation 97348b64-4114-4889-9a2a-b871dc941f75 · outbound

This paper cites The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.804522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:b6a7b07ae0be5e4c944ff34a905d0091938d17ef9ef51b9ff6944192d2aa363d

Observation d5ca492b-1455-4894-9c1a-f14b39c3a1ca · outbound

This paper cites Transformer-based visual segmentation: A survey.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Transformer-based visual segmentation: A survey

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:21:40.158297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:de03b71ef9c51d630d6a28e25bf7d5057d54935c597d09c4ce032a9a31494505

Observation 823798cf-2643-4c61-8cfa-936ac8ae3ada · outbound

This paper cites Baichuan-omni-1.5 technical report.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Baichuan-omni-1.5 technical report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.704589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:2d09f03677147caa32b856746e83eaae2e5bcd9e0363e914c6f7f260ae208bdb

Observation 74310671-fa47-4a38-a913-840134086350 · outbound

This paper cites Visual Instruction Tuning.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Visual Instruction Tuning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.687079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:cdbf2809568ad6229001772501558395cb9d096282b46fe356e4eda225a222b3

Observation 42de5efa-0cb5-431a-83ba-d78fd3927afc · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.732552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:22f1f561e49b18513ccebf5c3a0dc396531eb413af849053922d90e6f80bc9c2

Observation bedb7fab-876e-4f26-929d-834c6f0dda3a · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.829730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:c992cc1b4311144074d080b686230a8d40982b631729a16619f6a415c26bc4e2

Observation 76420bee-e81f-491f-86c1-0dd85a6789b6 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.788591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:bde05582394442cb8cbcaccb9f0a978a864f5bf4b0b85f70a50353d50d501f17

Observation 27d21c87-cfb0-4cf0-9ce8-3a1f3b7b226f · outbound

This paper cites InfographicVQA.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models InfographicVQA

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.835811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:b1d503bcc7134f1132086f059f4820fd3bdbd4bfbb151eaeb01ad78897476efd

Observation 4a5a7488-3a23-49d7-8382-328a111d4aa5 · outbound

This paper cites Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.710173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:86cd827b56857e80e8940ca047ba9d369fd5fa5a914e8ca7f73795cfe7d3a17a

Observation a4698198-98de-4154-8165-dc4cbc530d78 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:21:40.164892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:d792655bf23b73a8b5807d8e2741ba82901184d7e557634e0f3170cd4b61275e

Observation 960d2ccc-21ba-43c8-8c42-92b1a788810a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.743572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:ae6afdf65dd07369c96c00159ec22d778d00ebc082b9ccf1819a13127b385755

Observation 73719d88-6bd1-47d4-87cc-2276caa773d2 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Emu3: Next-Token Prediction is All You Need

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.819559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:cea11d618c42d77d4f8f96e1c2bc3dfad085d0f2ee2297afddc68f22ec32e24c

Observation a78a7883-6d7d-4472-902f-582dae61d412 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.810267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:d72ef36c7d9e1daf24d5b2d04807b70f715ad8ba16d58c9fb54722f89b3c0eff

Observation 7a351f35-fb6c-44bb-b844-ce62efad7530 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.799625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:4349bf878c0572592954627f4da65eb0f9d8e145da5a52506a3f2cc19cbacda4

Observation e53f9576-cd13-4120-8856-d0129d3a0d11 · outbound

This paper cites Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:21:40.172125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:870ef0299889241fef4bef3414ce023166bf30c8bc641d90af77db50655ad75b

Observation 1866e258-2d55-448d-b56c-dbcb34d455ad · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.814756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:9a35e4a6c53312ac981517e2a1f2411d7c4fe56dff9abd0e93bd2c0c361e8142

Observation 0668538a-cf7b-40e9-a8cd-75764d2ee628 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.824392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:9710d7d98ef19f29b07599d9287771f4b5453418c5a0ffe250ec2aec5a9835b1

Observation 9a31ce4f-9beb-4920-89ad-5d56ac396aff · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.754537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:20217db897fe1a3d811ae5c7b913fd379390a2f1dd3840aa6d1dddc71eb1d4c5

Observation 82e1864c-7d36-43f0-a806-d9d4663d0c17 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:21:40.168553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:f956d6657ac899c2fbe40fcf21eda714ee114badf3f6a82c9ae4f54e438c9568

Observation fc562104-9b98-43ea-937c-f9a34300c4a6 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.782161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:089db67d18317fa8da1f0e99ffd9768626a14c70b1a0958bb7f18bda06d16a0e

Observation 38a3333f-6f02-4b51-b56e-9b78381f26e0 · outbound

This paper cites Pixel-sail: Single transformer for pixel-grounded understanding.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Pixel-sail: Single transformer for pixel-grounded understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:21:40.176227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:ba586282a17ffe4be4b263a5395d6c43bde76fc1a9f69b19928ec3af763631d7

Observation d43770c4-9261-42b9-901c-4092cd57a8e4 · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.749504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:cdbad4e66921ee491a9cb3d783ca4a5b4d74ca37c6ece307a784c99a1aa384e3

Pith citing papers

Observation 3fa80d22-adfc-48f5-88a0-06a3fa545159 · inbound

RAP: KV-Cache Compression via RoPE-Aligned Pruning cites this paper.

RAP: KV-Cache Compression via RoPE-Aligned Pruning Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T06:17:12.549851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:17:12.549851Z digest=sha256:6e7d85451a71715dcf7f56bd6806f4aef1932b4e1040778b03537580a9b4ffed

Observation a17295a5-fcef-4cf8-8c07-f730491a2d67 · inbound

MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models cites this paper.

MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.153387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:02:34.007217Z digest=sha256:87f5cd2fe44ba0af2a8ea7ef8aba5619bdc36a84958f762854a2524b80381ce5

Observation 8adb6880-4177-47d0-9622-fe1d16d75d9d · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.067720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T01:10:48.111387Z digest=sha256:05661f02dd9d88d60d06ea6d7a639dc0efd7f33c0555ff670713f708b7179704

Observation 9cf1f5d4-1cb6-432f-bc33-1fe085cbacad · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.153387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T02:14:59.487486Z digest=sha256:26839bf5bc17e28fcda6e8ca1bec25248e5ed365082f29a86a691e476512353a

Observation cbc14466-2750-401b-83d6-019f212d1863 · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.153387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T21:47:32.112057Z digest=sha256:ee52d71c635cbd076ca34a26175a9fa1d9ba2d5ac5ea44abd48895b8c20b481d

Observation b29f0c5d-5e88-4d88-a178-75f94fcd5ca9 · inbound

DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers cites this paper.

DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:15:44.837433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T05:39:21.642584Z digest=sha256:be04b51435f4d2d889823b1e40d82b90b359fd40d1f0eb3367d8ae7c8cceb4a8