Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:39:35.669077Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 4 inbound Pith citation observations for arXiv:2501.07783.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:39:35.669077Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:46:59.953404Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T16:50:09.964594Z
100 of 119 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c7e8f682-f319-42d2-95b5-c4564048b26b · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Parameter-inverted image pyramid networks,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029c8cdf-2149-4570-924d-c41da79d3ae9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Training data-efficient image transformers & distillation through attention,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e78e1773-0c99-429b-8ab5-2bd92524d524 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Deit iii: Revenge of the vit,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5243b14-c4b2-4cf9-929b-7b1f841bd6d8 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c475c6dd-c813-4c94-9dac-e74c39c3f33d · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Learning transferable visual models from natural language supervision,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 256ba91d-8ad1-461a-be28-07649b1cf307 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Cascade r-cnn: Delving into high quality object detection,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca56dbb-42ee-4cb8-ad19-19922f0a2a5e · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Deformable detr: Deformable transformers for end-to-end object detection,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5df4f9e-0507-49fe-b22a-b0ffeed51125 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Scrdet: Towards more robust detection for small, cluttered and rotated objects,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 940c0986-541e-477e-af35-7b8f080cc418 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding R3det: Refined single-stage detector with feature refinement for rotating object,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe9cb20-e890-4986-90d2-b6208b5e34d3 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mask r-cnn,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320083c9-d497-415a-907f-3d4b6129fc59 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Unified perceptual parsing for scene understanding,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef282b32-b5fb-41ed-aa6b-3184d4faccff · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Patchdct: Patch refinement for high quality instance segmentation,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12696d0-e72b-4505-8a32-76d3c7d017c1 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Sniper: Efficient multi-scale training,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c7d120-2dbf-4aa1-ad1e-27f405b90cfd · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Autofocus: Efficient multi-scale inference,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf6fed6-8eeb-4277-8896-1da576c87ae4 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Feature pyramid networks for object detection,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f67304e-2c60-4c9b-9eb8-b275c878c472 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Nas-fpn: Learning scalable feature pyramid architecture for object detection,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2980713-7bec-4f24-8a78-af591550901d · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Efficientdet: Scalable and efficient object detection,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e47daed9-0496-4245-889c-0072d5d332c6 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Internimage: Exploring large-scale vision foundation models with deformable convolutions,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 169c9456-e0ce-45bd-a9a7-7ec2d60d0ae9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Eva: Exploring the limits of masked visual representation learning at scale,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23830b55-6131-4881-bb02-921b41f0f346 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Detrs with collaborative hybrid assignments training,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da3636d-2b38-4d2b-952e-6a1d7c820a27 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Vision transformer adapter for dense predictions,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a8cf13-fedf-4606-9f05-4532c4c2ce0f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Microsoft coco: Common objects in context,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a0ce5a6-1d29-496a-b2b9-39a12dc01dd9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7068d50-913f-4c8e-b238-28308cd59340 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Improved baselines with visual instruction tuning,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9a996a6-a0ac-4fc7-b081-ac19495c1656 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56fa86e-db31-4618-a2a3-65e0feffd7b9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a419daf5-a1ba-43ff-81be-775195a155d6 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Sparkle: Mastering basic spatial capabilities in vision language models elicits generalization to composite spatial reasoning,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ca48a23-2f3c-4c8d-8017-3912c28b8f2c · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding $\gamma-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60684f37-4c40-42f8-9d86-0485d009adff · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888e426a-573f-4d9e-8849-bf9d040fcaf9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d708c53-5c60-483c-9dd3-5262ce2f297f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a3c5f21-5511-499c-9425-12092afbb2cf · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a339a9-e3fe-4dd7-b175-4ba91d3260c9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Llava-next: Stronger llms supercharge multimodal capabilities in the wild,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3175493-4f5e-4f12-9cb0-331661bf21b3 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a88199-3b83-4783-b056-416ef5d6628f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2bafc9-bc71-45f4-8c11-87fd6d268878 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Unveiling encoder-free vision-language models,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a29765-d179-4b64-885b-ec795bec9a16 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb2dcce-70dc-440c-a329-573ea56d2d4e · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding A convnet for the 2020s,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959e9ed6-9e06-4c90-a2e7-3d7afb3cfb00 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c017aaf0-c275-4c60-bb61-1b9dd1e856d6 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Scene parsing through ade20k dataset,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826b7bb1-9468-4d03-bbf4-c5dc32583574 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding An analysis of scale invariance in object detection snip,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bbf7be4-aff6-4ba1-9e76-99d9d2a1a29f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Crossvit: Cross-attention multi- scale vision transformer for image classification,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c8b773-2756-4cfd-a7b2-7e99de065c03 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Deep high-resolution representation learning for visual recognition,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8395ea4c-bfce-42c3-bd42-e4756d2a5f91 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Cbnet: A composite backbone network architecture for object detection,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfa0c4e-0b7a-435f-8df9-794293e3a6d1 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Vit-comer: Vision transformer with convolutional multi-scale feature interaction for dense predictions,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc38dbd-3e19-48b9-9c27-73c1cbeb0bf1 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Hrformer: High-resolution vision transformer for dense prediction,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2048d8b-e39d-4b8f-9b0f-755d4c258c4e · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 03960af4-c56f-416d-8e62-9fa561179df5 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding CogAgent: A Visual Language Model for GUI Agents
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3691f387-e568-4d16-8164-ccad926f3728 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding GPT-4 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd59da37-4cad-43da-afca-ecd9f9d6c118 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding The claude 3 model family: Opus, sonnet, haiku,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e42b01-9b38-4de0-9665-43dc738913ed · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding InternLM2 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0dc1a9-d7d6-4be0-8ce6-4515995133ee · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162c2ddd-af01-4734-92d3-f22ad3b34f0c · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Emu3: Next-Token Prediction is All You Need
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853fab78-30c7-40f5-af36-9aec3406c865 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f2eaf5-f397-45dd-884e-7365d0cdd142 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066c389f-2024-47b9-9c5d-29ad544e20e9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e600f1ea-3796-4743-af71-30635ca88f4f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e6176a0-382c-4d2b-bac5-d6f9e47c5455 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2659a4-f25a-436b-a142-ff1f94ad2f2b · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Dynamicvit: Efficient vision transformers with dynamic token sparsification,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f51329b-0577-4bb7-8585-bc5fbf01e61d · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Adavit: Adaptive vision transformers for efficient image recognition,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d1fe35-7876-4bca-9900-7157d577401c · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Not all patches are what you need: Expediting vision transformers via token reorganizations,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd8a7bc-8f91-4813-9938-14695b57bfcf · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Evo-vit: Slow-fast token evolution for dynamic vision transformer,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ff3aefb7-882c-4ed3-a197-218ab5076def · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Linformer: Self-Attention with Linear Complexity
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7804fe2-caa6-4418-8670-02027e83e692 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8c2fef07-2c6f-48b6-b5cd-8b48f81cec9a · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Swin transformer: Hierarchical vision transformer using shifted windows,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 736759cb-f5e7-478a-8124-84fa5a0fd60f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d2283ec-96e3-40be-a04f-cc09ba18c725 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Layer Normalization
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa03723a-893b-4f90-a7c9-97150cea905c · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Exploring plain vision transformer backbones for object detection,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dde65495-e56a-4c9d-a18d-2de038f7ee3e · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Group normalization,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3755918b-1af8-4c06-8d6c-fca6d6d054e2 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Visual instruction tuning,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b94bd380-387a-40b5-a2b9-53cf1c373a7b · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding AutoAugment: Learning Augmentation Policies from Data
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec7fc9d-356d-447b-897a-d1fd448a8163 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Benchmarking Detection Transfer Learning with Vision Transformers
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c41814c5-6437-4706-8f4a-59d572ad3960 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14c89420-fca5-416a-99ad-6a802aea0d96 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9b3df187-a2b4-45db-9cdb-184d048791b0 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Imagenet: A large-scale hierarchical image database,
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 27f28aab-b875-470c-8215-ab72ebb7a3f3 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Masked autoencoders are scalable vision learners,
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ff4d424-951c-4b1e-912f-52158cbb93ff · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Decoupled weight decay regularization,
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 72a95629-4c8e-43ca-8ed5-6d0b59a3b28f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Beit: Bert pre-training of image transformers,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 851a0a9e-90e8-420b-8abb-94dd9213e39a · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Laion-5b: An open large-scale dataset for training next generation image-text models,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a636aa6d-774e-4008-9ac7-de84b5462f71 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MMDetection: Open MMLab Detection Toolbox and Benchmark
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b712daf-9a19-4e73-8628-2b8f624e7e83 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Uni- perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks,
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9dbf7394-ff04-4446-9e5a-24ae2f02268d · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding DINOv2: Learning Robust Visual Features without Supervision
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2ab7337-a888-4a7f-b038-083a8b1d7003 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation addd6608-da4a-41e7-a8d2-d48f537aabb5 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 13d8b986-7f58-4761-8e22-d3d542002b64 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Dino: Detr with improved denoising anchor boxes for end-to-end object detection,
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d9cc708b-af9e-459d-8eb7-fa196434a50c · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning,
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 60f2bb9d-162d-4615-a534-936e022368d9 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dc8c22e9-d5b5-4603-8f06-c712c2941f6f · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b517905a-ce01-43ed-ac12-fbc4ebf740d5 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2cbadcd9-caee-4e34-9511-f6bea7f32226 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e0b997-010a-4be1-bb5a-16b4b227a6d5 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8281498a-dce9-4af3-8a88-137af24558f8 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc08064d-b88b-4533-9d7e-c73d409d769b · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Mm1: methods, analysis and insights from multimodal llm pre-training,
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0fdd2be2-99ab-491a-ad97-bcdd3b92cbdb · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding MMSegmentation: Openmmlab semantic seg- mentation toolbox and benchmark,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be28d97-aa78-46cb-8742-a1db127e3ec2 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4a779076-0fac-4efd-b4d1-dca22efd6942 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Sharegpt4v: Improving large multi-modal models with better captions,
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 24677da9-dfd0-4eeb-89f3-cba5a1b268de · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Laion-gpt-4v,
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c6799925-9d93-44ef-b7fa-94346482bdf2 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31014970-ffc0-45f5-8e80-b5be2a30ed3a · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Lima: Less is more for alignment,
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fd6363c0-240e-4139-be3e-a9cb40c143f1 · outbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Openassistant conversations-democratizing large language model align- ment,
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1e9e7224-d411-4b1a-953c-d69e251da7cd · inbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd66841-47d1-4c4f-9ab3-d5ed719acaa5 · inbound
From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b412e13-b0c9-482f-9f09-a2dae75db265 · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc75399e-c724-46e9-8338-dbd706c97401 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.