Pith. sign in

Paper Citation Record · LEDGER

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

As of 22 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 4 inbound Pith citation observations for arXiv:2501.07783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07783 v1

Coverage vector

measured 100 of 119 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:39:35.669077Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:46:59.953404Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:50:09.964594Z

Reference resolution

100 of 119 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7e8f682-f319-42d2-95b5-c4564048b26b · outbound

This paper cites Parameter-inverted image pyramid networks,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Parameter-inverted image pyramid networks,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.228489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.228489Z digest=sha256:9859410b8dab73854ec4d97b13fd102502f6dbaa6d064bada27c3660a830fccb

Observation 029c8cdf-2149-4570-924d-c41da79d3ae9 · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Training data-efficient image transformers & distillation through attention,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.233258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.233258Z digest=sha256:1ea6c11170575142517c5d03eb0fad1aedd225dc80156bf000f1058beae07a86

Observation e78e1773-0c99-429b-8ab5-2bd92524d524 · outbound

This paper cites Deit iii: Revenge of the vit,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Deit iii: Revenge of the vit,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.237534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.237534Z digest=sha256:295d800aaeb463db2bd0d303c58046dffe190b394edbf5da3bdf07c541ee291d

Observation e5243b14-c4b2-4cf9-929b-7b1f841bd6d8 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.243205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.243205Z digest=sha256:bdb113b60114f7c278cc634b66ba7c5e96057261634fd14bb28d228e373fc0f7

Observation c475c6dd-c813-4c94-9dac-e74c39c3f33d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Learning transferable visual models from natural language supervision,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.248072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.248072Z digest=sha256:11f85c4e7bf72ceb02af7cabf599d346db327caea23b24f819747757fbed424a

Observation 256ba91d-8ad1-461a-be28-07649b1cf307 · outbound

This paper cites Cascade r-cnn: Delving into high quality object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Cascade r-cnn: Delving into high quality object detection,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.252463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.252463Z digest=sha256:6bb1d6c84f6fe104627e48c6261e214c4ee3530f7a3cde7db247eba26a7d3834

Observation 1ca56dbb-42ee-4cb8-ad19-19922f0a2a5e · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Deformable detr: Deformable transformers for end-to-end object detection,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.257063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.257063Z digest=sha256:837206e26a6c85cd668619af7cb9ebcd712abf50e05c2565945629913ee5b523

Observation c5df4f9e-0507-49fe-b22a-b0ffeed51125 · outbound

This paper cites Scrdet: Towards more robust detection for small, cluttered and rotated objects,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Scrdet: Towards more robust detection for small, cluttered and rotated objects,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.260922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.260922Z digest=sha256:d5dd7271cf370ecf3cb3fe6b87b05435cb7b6b5b1508c9d36e9ad88c0eda159b

Observation 940c0986-541e-477e-af35-7b8f080cc418 · outbound

This paper cites R3det: Refined single-stage detector with feature refinement for rotating object,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding R3det: Refined single-stage detector with feature refinement for rotating object,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.264912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.264912Z digest=sha256:7a9e441c9ec2faafba6a317bc51017b1bda16cee514f4ad802c6ddb982d1c450

Observation cfe9cb20-e890-4986-90d2-b6208b5e34d3 · outbound

This paper cites Mask r-cnn,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mask r-cnn,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.268910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.268910Z digest=sha256:240bd81197e109b915ac3b4f74f8964cdee4df7575015f5ce36eac639371ec27

Observation 320083c9-d497-415a-907f-3d4b6129fc59 · outbound

This paper cites Unified perceptual parsing for scene understanding,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Unified perceptual parsing for scene understanding,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.273429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.273429Z digest=sha256:25b27cc9061364af671ceb535da5e711969eca8efccc16dd2fb6c845b6e4fa56

Observation ef282b32-b5fb-41ed-aa6b-3184d4faccff · outbound

This paper cites Patchdct: Patch refinement for high quality instance segmentation,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Patchdct: Patch refinement for high quality instance segmentation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.278197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.278197Z digest=sha256:6525a0b9f188a4a2d53269dc19c06ae625f107c8f087e8e399d410c93d9d56f8

Observation f12696d0-e72b-4505-8a32-76d3c7d017c1 · outbound

This paper cites Sniper: Efficient multi-scale training,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Sniper: Efficient multi-scale training,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.283453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.283453Z digest=sha256:907d9b503cc46b8edf7e194d96461ff96f4a8226ba57e68149cc960437dbb001

Observation 81c7d120-2dbf-4aa1-ad1e-27f405b90cfd · outbound

This paper cites Autofocus: Efficient multi-scale inference,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Autofocus: Efficient multi-scale inference,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.287408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.287408Z digest=sha256:015c1385ee921cec88eb8c32ccd412b70e587e5193c5dddfa4cea27aad5075dd

Observation 5cf6fed6-8eeb-4277-8896-1da576c87ae4 · outbound

This paper cites Feature pyramid networks for object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Feature pyramid networks for object detection,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.291185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.291185Z digest=sha256:ecf8a9028545aaa7c3d0e2e63bb3cc9e485f6fe274c5b844f094a6f5b966957c

Observation 6f67304e-2c60-4c9b-9eb8-b275c878c472 · outbound

This paper cites Nas-fpn: Learning scalable feature pyramid architecture for object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Nas-fpn: Learning scalable feature pyramid architecture for object detection,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.294672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.294672Z digest=sha256:51af32d4f444aba75ece9016fdb133aab88161fe5f31f2d92c09534800aa430e

Observation b2980713-7bec-4f24-8a78-af591550901d · outbound

This paper cites Efficientdet: Scalable and efficient object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Efficientdet: Scalable and efficient object detection,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.298374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.298374Z digest=sha256:1252f9d3f0e93085027eab59e7b4c72c2fc639987791323351a778924c0dee86

Observation e47daed9-0496-4245-889c-0072d5d332c6 · outbound

This paper cites Internimage: Exploring large-scale vision foundation models with deformable convolutions,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Internimage: Exploring large-scale vision foundation models with deformable convolutions,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.302468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.302468Z digest=sha256:89f6bc9bed068aaf5f4f43afbeb30febc30897caa60f1730ef586fb1c1e08b28

Observation 169c9456-e0ce-45bd-a9a7-7ec2d60d0ae9 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Eva: Exploring the limits of masked visual representation learning at scale,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.306659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.306659Z digest=sha256:d6e13fe1eb6d7df6516e12a575b6a7e54ed8339bb4c74fd43d443f87b36f3f82

Observation 23830b55-6131-4881-bb02-921b41f0f346 · outbound

This paper cites Detrs with collaborative hybrid assignments training,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Detrs with collaborative hybrid assignments training,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.311183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.311183Z digest=sha256:e9ecd42de5c8c1aec0ad608e28e7fe82cd4267ad8cc19b35af06214561bcb01b

Observation 0da3636d-2b38-4d2b-952e-6a1d7c820a27 · outbound

This paper cites Vision transformer adapter for dense predictions,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Vision transformer adapter for dense predictions,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.315441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.315441Z digest=sha256:2dd6d66e0030a80027451bccbcc7f5daaa1cbeb8bf49ec979cab12c8e633d219

Observation 16a8cf13-fedf-4606-9f05-4532c4c2ce0f · outbound

This paper cites Microsoft coco: Common objects in context,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Microsoft coco: Common objects in context,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.319289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.319289Z digest=sha256:6a57a30821ee0530acb93cb0971ef02a9cae55ab8c55fb65b6df4cf6fccd726b

Observation 3a0ce5a6-1d29-496a-b2b9-39a12dc01dd9 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.324007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.324007Z digest=sha256:db1beaba95c28e10f40e2a3752443ab5141abbe1b06c2794fb0ef9802182595d

Observation c7068d50-913f-4c8e-b238-28308cd59340 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Improved baselines with visual instruction tuning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.328800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.328800Z digest=sha256:b85073346dc732c472580a4a39694823565cff2f2f0ebeae2f218d43e2f0a0ed

Observation f9a996a6-a0ac-4fc7-b081-ac19495c1656 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.333202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.333202Z digest=sha256:fb211c2e991067784bbf62b25efa399cc82870ce4cb7b94d37b6bfbeb02a1665

Observation f56fa86e-db31-4618-a2a3-65e0feffd7b9 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.338553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.338553Z digest=sha256:29a9d7783070e144ab25dc24566f9817a8ab1a05fcc1f353d897b935739c1e03

Observation a419daf5-a1ba-43ff-81be-775195a155d6 · outbound

This paper cites Sparkle: Mastering basic spatial capabilities in vision language models elicits generalization to composite spatial reasoning,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Sparkle: Mastering basic spatial capabilities in vision language models elicits generalization to composite spatial reasoning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.343641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.343641Z digest=sha256:68a782ce185b019e196c31d1bfe2516cdeddd5d9181f80994fde8affbfcef3bb

Observation 3ca48a23-2f3c-4c8d-8017-3912c28b8f2c · outbound

This paper cites $\gamma-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding $\gamma-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.348099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.348099Z digest=sha256:53c6c531ece20fc83903d9adfa11e4321a4bca4945d78061c6d8b91555c5f588

Observation 60684f37-4c40-42f8-9d86-0485d009adff · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.352932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.352932Z digest=sha256:5982969c2886d43ce50d500f9887943f47e7309091c5ec216f21b15e9cc12c9f

Observation 888e426a-573f-4d9e-8849-bf9d040fcaf9 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.357961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.357961Z digest=sha256:cce3b1eea5cde80e5ba7701ffd8da6bfae94119a85c1102dca3245abe16fc4d8

Observation 7d708c53-5c60-483c-9dd3-5262ce2f297f · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.362552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.362552Z digest=sha256:89450112ebf31e0cdb3766d3ce6c3f2ba12a1f1ee03d15a86c851769887fadf2

Observation 5a3c5f21-5511-499c-9425-12092afbb2cf · outbound

This paper cites MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.367058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.367058Z digest=sha256:39b381277dfe8fa0ff0e383b6ee79d085f8b5893289a7a11a8fa9fafe858568a

Observation e4a339a9-e3fe-4dd7-b175-4ba91d3260c9 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Llava-next: Stronger llms supercharge multimodal capabilities in the wild,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.371383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.371383Z digest=sha256:852db960097171b2a3f77cd44ceb5799466e494ffbf0e321b8316550465835db

Observation d3175493-4f5e-4f12-9cb0-331661bf21b3 · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.375935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.375935Z digest=sha256:1ae85030f7123b132504ac73ae7829240b16ac9296176b5cfa9b9d193c302e6b

Observation 73a88199-3b83-4783-b056-416ef5d6628f · outbound

This paper cites LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.380728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.380728Z digest=sha256:0109a10cce54875cf97bd003c5c006b26445c07545ded6a41daa5b1b21a9744f

Observation fc2bafc9-bc71-45f4-8c11-87fd6d268878 · outbound

This paper cites Unveiling encoder-free vision-language models,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Unveiling encoder-free vision-language models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.385117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.385117Z digest=sha256:b96536dbc30232cc74ff2c10a40b2956b97806742ac7712927223c2dcf1bc9bf

Observation e5a29765-d179-4b64-885b-ec795bec9a16 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.388753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.388753Z digest=sha256:49b40b9c55c797ad3bb8fd670388bfd4a4b64ce8b82bbaaf79026fa63e74609e

Observation 6bb2dcce-70dc-440c-a329-573ea56d2d4e · outbound

This paper cites A convnet for the 2020s,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding A convnet for the 2020s,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.393271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.393271Z digest=sha256:bfa784fd24f9f32702d84bdfd02b67195851c68b737728592f04e1221ffae52a

Observation 959e9ed6-9e06-4c90-a2e7-3d7afb3cfb00 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.397230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.397230Z digest=sha256:c7a05949310fc6c6dd45dd5622d4b58259a6d3eb1733287790bba1ad7d5d182a

Observation c017aaf0-c275-4c60-bb61-1b9dd1e856d6 · outbound

This paper cites Scene parsing through ade20k dataset,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Scene parsing through ade20k dataset,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.401533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.401533Z digest=sha256:ebf2eeaaf0f8ba719fa0a1fa2550ecb58f1a01e1f61b8550fa01305110b3d509

Observation 826b7bb1-9468-4d03-bbf4-c5dc32583574 · outbound

This paper cites An analysis of scale invariance in object detection snip,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding An analysis of scale invariance in object detection snip,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.406760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.406760Z digest=sha256:2824f0adc935a6838ff334b88ccec3107ece21592a36c7b3b3ab9a5cc06e6d3c

Observation 8bbf7be4-aff6-4ba1-9e76-99d9d2a1a29f · outbound

This paper cites Crossvit: Cross-attention multi- scale vision transformer for image classification,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Crossvit: Cross-attention multi- scale vision transformer for image classification,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.411176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.411176Z digest=sha256:1c0fa576c1701aa25c32604b9974d240b27140036e44de99536df660fbb50da1

Observation 74c8b773-2756-4cfd-a7b2-7e99de065c03 · outbound

This paper cites Deep high-resolution representation learning for visual recognition,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Deep high-resolution representation learning for visual recognition,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.415887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.415887Z digest=sha256:f988110560f54028d72d841a8cf3e65738f5ef0abb2c47dd7eddf707a4e90899

Observation 8395ea4c-bfce-42c3-bd42-e4756d2a5f91 · outbound

This paper cites Cbnet: A composite backbone network architecture for object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Cbnet: A composite backbone network architecture for object detection,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.419674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.419674Z digest=sha256:53398e3170513c5a8c9c1c17c5acbae1de477b78465d4b5168cbfd131fb2c337

Observation 6bfa0c4e-0b7a-435f-8df9-794293e3a6d1 · outbound

This paper cites Vit-comer: Vision transformer with convolutional multi-scale feature interaction for dense predictions,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Vit-comer: Vision transformer with convolutional multi-scale feature interaction for dense predictions,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.424603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.424603Z digest=sha256:4c94a86cf54bae0b8df4bcb96cc2a7e0289efa4f3cb08170fc3a73700903b11a

Observation adc38dbd-3e19-48b9-9c27-73c1cbeb0bf1 · outbound

This paper cites Hrformer: High-resolution vision transformer for dense prediction,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Hrformer: High-resolution vision transformer for dense prediction,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.429079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.429079Z digest=sha256:e5030b135dabb52a0d431ce400576cfed345149613844b07249256d80fa665f3

Observation f2048d8b-e39d-4b8f-9b0f-755d4c258c4e · outbound

This paper cites Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:39:36.312005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.433082Z digest=sha256:d0cc0bac4f567e7b4bbd669431be568cc414ebf1eba21c547e22ade9eea64a03

Observation 03960af4-c56f-416d-8e62-9fa561179df5 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding CogAgent: A Visual Language Model for GUI Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.437102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.437102Z digest=sha256:37edcefc02a9a99e8fc93a30d9c22ff0ead9d73de75eae18b29eba887b8475be

Observation 3691f387-e568-4d16-8164-ccad926f3728 · outbound

This paper cites GPT-4 Technical Report.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.442096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.442096Z digest=sha256:6e67de38636a32a5f076c8ba15fbd4e58259f28ffceb7fa247a9c1b50eab847f

Observation dd59da37-4cad-43da-afca-ecd9f9d6c118 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding The claude 3 model family: Opus, sonnet, haiku,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.446999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.446999Z digest=sha256:c1450a4c829de50831044cd2e35de51ae3c301fa70a8577d5fa8bc9f39ed6f84

Observation 31e42b01-9b38-4de0-9665-43dc738913ed · outbound

This paper cites InternLM2 Technical Report.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding InternLM2 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.451538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.451538Z digest=sha256:bede6ba82b9d69cd83d86933befca326ee408bc0d25d3b8f5f91efbc1461cbb7

Observation 0a0dc1a9-d7d6-4be0-8ce6-4515995133ee · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.456542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.456542Z digest=sha256:784672a89e923b276323abc7e0ac51e319053ecc929d163666dd06b62993fe5a

Observation 162c2ddd-af01-4734-92d3-f22ad3b34f0c · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Emu3: Next-Token Prediction is All You Need

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.461397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.461397Z digest=sha256:3c866a34b773aa75968e9dd61ddf76e6ba84d1f2205d60bffdfc7131c311c45c

Observation 853fab78-30c7-40f5-af36-9aec3406c865 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.466176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.466176Z digest=sha256:2a7a1e460e22d1ed08e66b3f2ffebd19116d61beddf0acf2e6e660c8ccc3c92a

Observation e7f2eaf5-f397-45dd-884e-7365d0cdd142 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.470582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.470582Z digest=sha256:71a36f6a6535e75572ae4262b3596ae514901d337a7971eee7589531d2e483f8

Observation 066c389f-2024-47b9-9c5d-29ad544e20e9 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.474197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.474197Z digest=sha256:1f3d7c1575cd6812779f80b552a28e68ef98df3e626caa851f59d42acb9ce61e

Observation e600f1ea-3796-4743-af71-30635ca88f4f · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.478614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.478614Z digest=sha256:a043a5e6be33ba4ff0e449c744f53969b35d71e4da62865f498f44cb064f94cd

Observation 3e6176a0-382c-4d2b-bac5-d6f9e47c5455 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.483390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.483390Z digest=sha256:3d32d1553d05d7e5a592da09f1fcbd6076b8417d14cc3220775386c84ad134b7

Observation 5d2659a4-f25a-436b-a142-ff1f94ad2f2b · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Dynamicvit: Efficient vision transformers with dynamic token sparsification,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.487519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.487519Z digest=sha256:39bda79edf07a372fda54ffc54ae5e24ac1ab9338980c589451c5edb7a789ccc

Observation 8f51329b-0577-4bb7-8585-bc5fbf01e61d · outbound

This paper cites Adavit: Adaptive vision transformers for efficient image recognition,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Adavit: Adaptive vision transformers for efficient image recognition,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.492638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.492638Z digest=sha256:a84b936085f75f6d961276de1c499e53b39d3d2560ff1e53763890fcb3ef9f93

Observation 09d1fe35-7876-4bca-9900-7157d577401c · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganizations,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Not all patches are what you need: Expediting vision transformers via token reorganizations,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.496759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.496759Z digest=sha256:2ca7c286ff43c42c90a0436aa3255de1772c68028c684f211dd6fbad1a9fdfd7

Observation 5dd8a7bc-8f91-4813-9938-14695b57bfcf · outbound

This paper cites Evo-vit: Slow-fast token evolution for dynamic vision transformer,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Evo-vit: Slow-fast token evolution for dynamic vision transformer,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.190403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.500397Z digest=sha256:8657ac9fe90e6c0d1819b6eeadf408299535cacef32449a898b234fc5fc86b57

Observation ff3aefb7-882c-4ed3-a197-218ab5076def · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Linformer: Self-Attention with Linear Complexity

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.504513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.504513Z digest=sha256:270e8ec36ae700f8f6516e11d201c25143304a54a5b2fc6b7383b40e023bbb3c

Observation f7804fe2-caa6-4418-8670-02027e83e692 · outbound

This paper cites Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.176392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.508609Z digest=sha256:d5bb744bbbf464486aeca791a14bc280238bf6c20ac10fdce9a57a607a921e60

Observation 8c2fef07-2c6f-48b6-b5cd-8b48f81cec9a · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.163579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.512731Z digest=sha256:7155ecb8dccf6cf403c53b2faacdab29eacda4581494b4d2c0862d8b505871ea

Observation 736759cb-f5e7-478a-8124-84fa5a0fd60f · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.149677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.517386Z digest=sha256:0998169444f87cc4383a129c906ce82950a539e78156c570d67bd8f6189ad650

Observation 7d2283ec-96e3-40be-a04f-cc09ba18c725 · outbound

This paper cites Layer Normalization.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Layer Normalization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.521445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.521445Z digest=sha256:84884fcc3430c7c02b0698e0b6e8ee5c1133b9e4e8e8a1c9944c273a5ceda431

Observation fa03723a-893b-4f90-a7c9-97150cea905c · outbound

This paper cites Exploring plain vision transformer backbones for object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Exploring plain vision transformer backbones for object detection,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.135798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.525372Z digest=sha256:1bb7d072c603a371d8fd0a131e0728107d3104e3e72f0161797b935f3c253c15

Observation dde65495-e56a-4c9d-a18d-2de038f7ee3e · outbound

This paper cites Group normalization,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Group normalization,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.117546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.529346Z digest=sha256:1d84635f0a60ebc82a5758297d59523f823265e12e3622c28f9c6b9fbce985c8

Observation 3755918b-1af8-4c06-8d6c-fca6d6d054e2 · outbound

This paper cites Visual instruction tuning,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Visual instruction tuning,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.102082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.533159Z digest=sha256:4038590edbb2c0311e89d67ca987812bfd02fbb216672e582cf879d1ec72f79c

Observation b94bd380-387a-40b5-a2b9-53cf1c373a7b · outbound

This paper cites AutoAugment: Learning Augmentation Policies from Data.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding AutoAugment: Learning Augmentation Policies from Data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.537486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.537486Z digest=sha256:a9f819194bbc65c58021839568caa1429d41aac11b615d9b37cbc59c17625d4c

Observation aec7fc9d-356d-447b-897a-d1fd448a8163 · outbound

This paper cites Benchmarking Detection Transfer Learning with Vision Transformers.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Benchmarking Detection Transfer Learning with Vision Transformers

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.542024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.542024Z digest=sha256:ae047064ca8e593963c36c6649d0cedb900739ba7edc1f7e23625e2809205394

Observation c41814c5-6437-4706-8f4a-59d572ad3960 · outbound

This paper cites Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.546574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.546574Z digest=sha256:a7a0f14c5d052762e22076504cc33b05f2b0f08fe56e74ce3cfd42b94d97a969

Observation 14c89420-fca5-416a-99ad-6a802aea0d96 · outbound

This paper cites Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.088762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.551904Z digest=sha256:687f0b0975b7384cc020672149e87efefba6dfc15efefb4da7a702050b38ec3d

Observation 9b3df187-a2b4-45db-9cdb-184d048791b0 · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Imagenet: A large-scale hierarchical image database,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.065572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.555816Z digest=sha256:ee00904c9228a416c7a25382ec20655f3be8e7c940d4b381f5550c6ab034ca71

Observation 27f28aab-b875-470c-8215-ab72ebb7a3f3 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Masked autoencoders are scalable vision learners,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.047194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.560310Z digest=sha256:b0537bf99dba0a8eeba188811e18a47daf740e2446bd2d0de3709ac4ac687781

Observation 3ff4d424-951c-4b1e-912f-52158cbb93ff · outbound

This paper cites Decoupled weight decay regularization,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Decoupled weight decay regularization,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.030384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.564148Z digest=sha256:adb61386623a00368318efe30eca80012f3ab85822f31bb877021a0e8e157d2a

Observation 72a95629-4c8e-43ca-8ed5-6d0b59a3b28f · outbound

This paper cites Beit: Bert pre-training of image transformers,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Beit: Bert pre-training of image transformers,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:37.015591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.568552Z digest=sha256:3382ee54195c8033083c4f5307ca4f9b7fb663531ca109641f439534bbee25d5

Observation 851a0a9e-90e8-420b-8abb-94dd9213e39a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.999968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.573046Z digest=sha256:bba086ccdfaaed79957f6d0402a5cbe439de177c836d6c63548382a3f2a0ec3f

Observation a636aa6d-774e-4008-9ac7-de84b5462f71 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.577461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.577461Z digest=sha256:8cf5f2112791a8d31dab627a6eab794732ba650aaa009d7af914d94ba6c93e9d

Observation 6b712daf-9a19-4e73-8628-2b8f624e7e83 · outbound

This paper cites Uni- perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Uni- perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.983731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.581816Z digest=sha256:d1c7b61932da722c6834bade9d0e96042bc9851553751e34737f6322596a369a

Observation 9dbf7394-ff04-4446-9e5a-24ae2f02268d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.586115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.586115Z digest=sha256:15b911ae2df779abcb0847395b5e4df2118629a09880347f82bc1491d9748638

Observation b2ab7337-a888-4a7f-b038-083a8b1d7003 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.591284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.591284Z digest=sha256:c1cf862cf86f729734a49b285d3ad8769c093c31f9834a7265e642b186c9dc43

Observation addd6608-da4a-41e7-a8d2-d48f537aabb5 · outbound

This paper cites More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.961542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.595548Z digest=sha256:413acbfec15ad83e759725d7a49365b3638cc028d8e160ae93f830517c7bb98b

Observation 13d8b986-7f58-4761-8e22-d3d542002b64 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object detection,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Dino: Detr with improved denoising anchor boxes for end-to-end object detection,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.942855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.600252Z digest=sha256:8083102745dd03ecfa74b22f3267039e420ea7f9e21487389b86fa2ceaafa8df

Observation d9cc708b-af9e-459d-8eb7-fa196434a50c · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.927844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.604605Z digest=sha256:9c16f615fbfcdee409d1da77bf7e9eabb725aa84468f0e059dba2bb678a527aa

Observation 60f2bb9d-162d-4615-a534-936e022368d9 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.912236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.609071Z digest=sha256:352b335af724aa365ce527e658ea08b638675e7edf2b320ec62d1f0eae88ec83

Observation dc8c22e9-d5b5-4603-8f06-c712c2941f6f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.613164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.613164Z digest=sha256:03c45bb93c1bb746c8d857ffdb6f1ac9eda9ff2d247aa1360410607d26f4faaa

Observation b517905a-ce01-43ed-ac12-fbc4ebf740d5 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.898017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.617758Z digest=sha256:425a7332c5d07b77a55acb4fd52c9a51d807873fb11623344433dca00a9684c4

Observation 2cbadcd9-caee-4e34-9511-f6bea7f32226 · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.622849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.622849Z digest=sha256:6ecf4f190a9a2ed79084309934b05e5a024cf8a0901693263d551d520e9a057b

Observation 06e0b997-010a-4be1-bb5a-16b4b227a6d5 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.628119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.628119Z digest=sha256:253cdb60c44286f414166c44d8050c8ee99eae2d92065cdbd951e2b45e970a62

Observation 8281498a-dce9-4af3-8a88-137af24558f8 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.632670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.632670Z digest=sha256:ca0ca3000ecc2ee6852529a358301953085f7fd9dd96012ad5b3d6ca1fbc4427

Observation bc08064d-b88b-4533-9d7e-c73d409d769b · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mm1: methods, analysis and insights from multimodal llm pre-training,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.884692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.637482Z digest=sha256:7d4d2e6bff179bae1e157c82af3693494f91e6a8041c56360abb09a1b7483fbd

Observation 0fdd2be2-99ab-491a-ad97-bcdd3b92cbdb · outbound

This paper cites MMSegmentation: Openmmlab semantic seg- mentation toolbox and benchmark,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MMSegmentation: Openmmlab semantic seg- mentation toolbox and benchmark,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.642191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.642191Z digest=sha256:a3b8009c290d9eeb5d78f70de86b95661e447733e971e5be6215d88c50b89d8c

Observation 5be28d97-aa78-46cb-8742-a1db127e3ec2 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.858551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.646949Z digest=sha256:a1822ea0ff45156c923f9f2884936fcc23499c002a71fc2806b77db1209cfa2d

Observation 4a779076-0fac-4efd-b4d1-dca22efd6942 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Sharegpt4v: Improving large multi-modal models with better captions,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.842153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.651268Z digest=sha256:e438ac962d12a519dfd4ea61e6977eb7adb3d559d212fc8c13a0d11ec1b76639

Observation 24677da9-dfd0-4eeb-89f3-cba5a1b268de · outbound

This paper cites Laion-gpt-4v,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Laion-gpt-4v,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.827460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.655136Z digest=sha256:57f5931ceaedfe19abe3f68d8cf3faa4896e626790a9ebcaa4e9e4cd84309fd6

Observation c6799925-9d93-44ef-b7fa-94346482bdf2 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.659703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.659703Z digest=sha256:3ba2c1742c11a78dad286e6a08bdc022f82736812ae8dce3ef69885323b38141

Observation 31014970-ffc0-45f5-8e80-b5be2a30ed3a · outbound

This paper cites Lima: Less is more for alignment,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Lima: Less is more for alignment,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.812661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.665075Z digest=sha256:a88c722fa0e5d951b82dbae80bee2dfd633d73f52303e6a46a332b40e8721d10

Observation fd6363c0-240e-4139-be3e-a9cb40c143f1 · outbound

This paper cites Openassistant conversations-democratizing large language model align- ment,.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Openassistant conversations-democratizing large language model align- ment,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:36.798170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:39:35.669077Z digest=sha256:ee21ec93ab54c29caf977eddaa2726665eee0bb43929115b36931239a24e1cf0

Pith citing papers

Observation 1e9e7224-d411-4b1a-953c-d69e251da7cd · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.953404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.953404Z digest=sha256:de36f64893a074c77514696874944d624d1c1047eb5fb035d33f5ef405be7c61

Observation fbd66841-47d1-4c4f-9ab3-d5ed719acaa5 · inbound

From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models cites this paper.

From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:36.348501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:31:36.348501Z digest=sha256:6efb4c1909f005df310c6f469dfcf5c47b0e99ac4194eed0d3cdc1144251ede2

Observation 3b412e13-b0c9-482f-9f09-a2dae75db265 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.338434Z digest=sha256:f8fc9f1a454ea5c2201267bbbf9e00d241bbc3423f9d4e1b26ee92f0bed8af0b

Observation cc75399e-c724-46e9-8338-dbd706c97401 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.051759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T16:49:57.683627Z digest=sha256:4ab547fc695bb5af76454673160adb7444764ef356f80c714123ba19e6bb9184