Pith. sign in

Paper Citation Record · LEDGER

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

As of 16 August 2026, this Paper Citation Record lists 100 of 287 outbound references and 0 inbound Pith citation observations for arXiv:2608.12748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12748 v1

Coverage vector

measured 100 of 287 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:09:18.495392Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 287 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5ada3e7-5e40-4954-926e-02eec0f49532 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.131394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.131394Z digest=sha256:366eea8b75b9bab719e31145002032eeda23724c1642fb72114bbdb809026fd6

Observation a81588f9-4d07-4de5-8d0e-2e7ba330ccfe · outbound

This paper cites Applied Sciences , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Applied Sciences , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.135472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.135472Z digest=sha256:36e73e9e65bd121552cbfd9ad34a09c5b57ccf10d45c6b47464de9e70d1b5bec

Observation 7a1532a6-cb00-4695-8031-9a8beedf3cce · outbound

This paper cites Signal and Data Processing of Small Targets 1993 , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Signal and Data Processing of Small Targets 1993 , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.138865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.138865Z digest=sha256:85901bb16d5a4afa7b41e4410d0fb128df5fa9dcbb21e55e8cea6343466db7ff

Observation 3febe7c0-0ae0-40c8-80fe-e7aed9825ef5 · outbound

This paper cites Sensors , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Sensors , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.143478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.143478Z digest=sha256:dcdd6365320682af5b9d7cdf60d4bff2ce49bf2a2330fe42e680f3dfb01ea1be

Observation f5a86670-90e2-4c0a-aba4-449af04cb6c5 · outbound

This paper cites Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.147495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.147495Z digest=sha256:fc9bd7d77389eb76302592989da8e3331f44830c9c4115395748ca6ffbedd61f

Observation cb5ff312-aa1b-4059-9149-a9c07845df58 · outbound

This paper cites Sensors , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Sensors , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.150786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.150786Z digest=sha256:19273a1119816df361aeaad4c59e99afeda5768979b4e40f98bf4aa7116eee06

Observation 6cac7bdf-ef22-4c44-96bf-66fdb7166eae · outbound

This paper cites 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.153828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.153828Z digest=sha256:dce48c0f8f4d4c364718c2f6d4cc9fb6424ebd4f0c54b094ab8bdd4cc6dad59d

Observation c637ee4e-6a73-4ed3-acfd-8baccf73b7a7 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.158526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.158526Z digest=sha256:5984b4f013cb6d71f5ffb90ebb8190c6616ca1457f4a582d95aa966106ac221f

Observation ee327802-dd29-482f-9599-43871f64980d · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.161713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.161713Z digest=sha256:3ad884a6f4e95eb7e011de2dd6eccb939040253f6b165b0e28d26c029cf60ba9

Observation 6713d00b-68a6-407a-ac83-7f0e7df22efd · outbound

This paper cites International Journal of Computer Vision , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Journal of Computer Vision , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.166091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.166091Z digest=sha256:ebd89efdb18a9a51eba344b179da4d4116ee4606390721b8c9d041f2d9946ac4

Observation 198ced30-915a-4898-b3d7-56b6f1096078 · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.169235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.169235Z digest=sha256:48c8584bb770cb5a6ceeb304cac221c83468c6853a5a46e34ba6e6fa1993d6a0

Observation 1ef4a677-338f-4948-8ef1-88dd4e14e538 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.173688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.173688Z digest=sha256:b906d2ee97dd5481add04b524172dde55ccfb1ff700cd4c9e44eb61ebe898a55

Observation f40c3e1e-bc8a-4f09-ae70-5f1976c6468d · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.176903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.176903Z digest=sha256:1d66dca380eccbdd071b1a5542fd281d91dcd799bdc848d3e448a9eb3f79e02a

Observation a44725a2-b514-444d-b8f9-0221777dcbf7 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.180236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.180236Z digest=sha256:33702a742801f5ec077c30ec769b0ef5cb09e5d09713c06840432a785e985713

Observation 2384866c-174a-4bfc-b306-6b373233115a · outbound

This paper cites International Journal of Computer Vision , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Journal of Computer Vision , volume=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.183249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.183249Z digest=sha256:9db607f523800dd8106326d11386aeb37095c796ce9a86934220efb14d47834a

Observation 3150fbd9-0146-4dc5-9f1e-09ffd12aff8c · outbound

This paper cites Infrared Physics & Technology , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Infrared Physics & Technology , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.186095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.186095Z digest=sha256:8d6dc67358c21d06cb248267ba5a33ecaba3c4632c50ce9a10fe8e35a0107e42

Observation a41111da-8d77-4612-a8ce-e501f6f99bfb · outbound

This paper cites Kaur et al.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Kaur et al

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.189270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.189270Z digest=sha256:8472777ef13f2b3e831fdd749e5973fd8447ee9f90c2352e8a75d83ba4debcbf

Observation ea7d1d10-5292-4302-826b-a1e991a320aa · outbound

This paper cites , author=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding , author=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.192080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.192080Z digest=sha256:6655bf66ae71d5c8421173171f6f2e4124a54d05ed4002a8f84a9b63093c9b6b

Observation 738f01d4-a2f8-47b3-8b23-7b6b60045d01 · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.195382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.195382Z digest=sha256:f82cd8a6b3db6f31632b689df5a016cbc6bd56559a66b031a42cd203ab3e3a6f

Observation 169abf66-44fd-43ea-a19e-a84eb2cf17fb · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.199801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.199801Z digest=sha256:b564b309a242cf151f4dd336a2f5263adc9f0786297c6159d54b1ec8de212e95

Observation fb978c6a-5f8e-4476-9cb5-8f2be0c79f39 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.203629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.203629Z digest=sha256:cdbbc997392220e6943242c6123c2c99be57c0f6225248cc9b8ea25ace4d3e0f

Observation 6a3d78cb-6c88-43a7-8995-2200ec1fd3a1 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.206600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.206600Z digest=sha256:f9683f36c116b674adcd78affcf9fe8c9d52045f5b5eb98abfe93f061ee31279

Observation 9e8b2e84-e73c-4461-91c4-9bf00688c6d6 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.210068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.210068Z digest=sha256:b63272639fe468c5577aff879b3a7a4cc42a813e6bd4a4a497aec992492e2d86

Observation 198ab2ef-01d2-41d7-923f-fb195d791173 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.214296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.214296Z digest=sha256:390157e69da85b6818c25c342a85aba482776a63360c6d2621f7f91a94e7dad9

Observation 790836ec-2a49-46b1-80d6-4ad9eb6be0a5 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.217495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.217495Z digest=sha256:04c586d81e6791551309af7bd4a79ee56cce9ab14b4add443ef52ef5d9b50d75

Observation e5373747-d872-4110-aa85-a96428c7291d · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.221276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.221276Z digest=sha256:8f97ded317aecb42c6d261a7483b0d51faef78dfbaa5a5699ee8ece7211582e7

Observation 0af7ddf2-e113-46d1-9b70-d2ab4e6bab68 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.225163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.225163Z digest=sha256:9da0832ae13fe6cebd87a861f8735cf3673171ba0ae28d6bdc1f271cb2fa165e

Observation a9b04ec1-e611-4fc4-9a3b-085206799bd3 · outbound

This paper cites International Conference on Machine Learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Conference on Machine Learning , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.228795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.228795Z digest=sha256:deb6f4dfae435425bd104559758e66dabd9df014337039b2d0138b88d3cbeb14

Observation 7c113b82-daa2-47f6-aabb-a6533ce590bc · outbound

This paper cites 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.232127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.232127Z digest=sha256:3842932a9ead5ba07cf8595f8ceeddd83e20840276f5182a9907c4895187a959

Observation 166969f3-0ec7-4d50-8c18-99cf64e4e238 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.235430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.235430Z digest=sha256:6a3c7382551f1b754c0a368aa97c717bd7685621391211a0e65671565887f5e5

Observation 631494c5-6dca-4ca1-a8b1-6655b275f04d · outbound

This paper cites Visual Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Visual Intelligence , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.238598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.238598Z digest=sha256:ab5f56bdc463d0a0146db4944a1ce2120d0cc8d683607a6090829d7d2e2d00fc

Observation dd53a003-ce60-4e95-b197-78914fbff2e7 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.241956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.241956Z digest=sha256:cba8c548d999a17b9335e7e72a3f3d62eda69914608b25e93779eba9f08507e1

Observation a14fb8f5-6243-4fdc-85da-f6b0fde0383a · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.245336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.245336Z digest=sha256:1031f7c6a92c5629b1b6221578204005be6170d8701c474c6b5a5f97cd4fd633

Observation 24ddecf9-0999-43a8-b6dc-1a18525ae5f2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.248499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.248499Z digest=sha256:307cce9c514800db59bb9e8a6cd7fa30b81e81d985e8f4a47c29ea161953974b

Observation 38806458-f364-4829-b4fb-fc5b09da3680 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.252010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.252010Z digest=sha256:cdd9f59638fdc7161a1e7a09f59e68968c67b2b2424d0d076e766dda361533ef

Observation 47feca59-be9f-4838-aecf-0d40e502dc86 · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.255311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.255311Z digest=sha256:8a195b3f291929ba7be321218f58a9cd914ed75cfd1703c08850a81ef42980cb

Observation d0ff5c32-6006-4d1b-9c83-d82dc1017147 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.259063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.259063Z digest=sha256:bd484480ecdd014039e91e083e2c2ecc8aa944e29690935dd7278086db41fc8a

Observation bbd3066d-cf2a-47a7-a5eb-dfda3484a6bf · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.266450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.266450Z digest=sha256:0502de3f03c8332f698e5a65a384965193acbf1f918ea492b95160543f94f295

Observation ff44c456-9f22-4a7d-8893-38d52e44efd7 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.269408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.269408Z digest=sha256:5ee5ca2ad7325d2715ac27bc4450eaced7c38c2e8c6a8de5312f5b543d97557a

Observation b413ac14-a37f-46c9-8167-fba1bd896f1f · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.273874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.273874Z digest=sha256:bdb42a08f0183e305c4ed50124e2d8a96c6afffa41d55d5be1fb38f6a974000b

Observation b8c98ffb-b1ff-4c18-a79d-eebcbcaeb8b9 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.277272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.277272Z digest=sha256:4a7c4c722dc30ad053e69f80a721063bdcafe15bb4f8b4e6011e56c959b11ca0

Observation f1431835-d342-4cc3-88c2-5a7d44c75e60 · outbound

This paper cites International Colloquium on Automata, Languages, and Programming , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Colloquium on Automata, Languages, and Programming , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.281345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.281345Z digest=sha256:7749ac3dc492efd5ce2ea6248af37467a2ceef371263c3416791f48a56000fa5

Observation ba2da6cd-5520-4f45-987b-2d330b2ec766 · outbound

This paper cites Proceedings of the IEEE international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE international conference on computer vision , pages=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.285481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.285481Z digest=sha256:000b2b5c1a1b6de38ae250d1d97ff336efaa639d9c04c9a3c4f56fb29fe3f9dc

Observation 3033de1b-51b1-4341-b512-1827915139ba · outbound

This paper cites Tensor Fusion Network for Multimodal Sentiment Analysis.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Tensor Fusion Network for Multimodal Sentiment Analysis

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.288993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.288993Z digest=sha256:ef57dfa40d4df0d04a58b43431c3153ab2a166ff4cc501c7f4cfc2300cb879d8

Observation 93cf72aa-8317-4099-b5b8-1fd6a6d3c3ac · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.292484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.292484Z digest=sha256:0bdb4517f1dbb98be51e8d0e1ffcd332455d56e785acb0350b9c84fb3b780139

Observation 862e57a0-a390-4f67-8b7e-6b008c20bb70 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.295725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.295725Z digest=sha256:2caf212d7b7b1e1b7ea93de5181b4dee1994a5c1e1ff2f4b6e835ebbbe3f1f5b

Observation d13c8183-9b69-418d-9ef0-3948308fe9d1 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Fast Transformer Decoding: One Write-Head is All You Need

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.299513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.299513Z digest=sha256:419f9749e6f3ccc56d675db3ccef6503f347ce21ec18acacc58f499f3e834c35

Observation b342f152-1b0d-45bf-8ac5-deeb7ca75eeb · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.303079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.303079Z digest=sha256:dd18dabc42a5edf8721101480d8cde6565ec15c04a89cc245dcb826d5aa724b0

Observation aeb73f93-bd3a-49a8-88a1-8b5313b2fb04 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.306273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.306273Z digest=sha256:8a7836ef0947230c1e398a86f37bca87c8db989e4868ddef0f5feaf0b1b5f5fe

Observation 58aa6810-536d-42ef-8104-bda0302ae315 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.309435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.309435Z digest=sha256:d06ee44223a7b9a52c3fffb5ef95de91b9c662a1d326c4046559ecd44f17e6ac

Observation 20c76a08-459a-4705-bf22-b9973526bb4c · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.313282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.313282Z digest=sha256:f96c940b4f202e558bcec1cf6217e5aaa84164cdf42a7c3298d0fcec1eca8cb3

Observation 5c05fb9c-087a-4a51-a5f0-11564571a76e · outbound

This paper cites Scientific reports , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Scientific reports , volume=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.316532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.316532Z digest=sha256:410a62e350b3124723d5d0f138eae37c894a8b1dd0a651806af08ad487759e78

Observation 434536f2-3d05-492c-9ff9-c499cacfbf0f · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.320625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.320625Z digest=sha256:333791073bb9607710450add8d8448f94844fcbab11492ba93a156168839122b

Observation 20aeb9fc-85c8-4089-b7b9-e16e7c68811d · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.324123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.324123Z digest=sha256:6dad6a60e11a336e310bc86fd1ace36948cc58d0808eccfab019ddc7456ee379

Observation 4f1fd045-68ac-451c-992e-7c6c08929aa8 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.327291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.327291Z digest=sha256:c5d4ab491290df19bba1b10b78c08523a3d7fed81a777bdf44e5b8d2b0122af6

Observation fe591112-dc85-4e93-b70f-bc762eb20c72 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Retentive Network: A Successor to Transformer for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.330658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.330658Z digest=sha256:434b9b857d74015bf898a563cf18a99b897069914c8de2151affd934c908c2ea

Observation 834b29cc-6fe3-49c8-a74f-c1a1d8b1ec3d · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding RWKV: Reinventing RNNs for the Transformer Era

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.334068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.334068Z digest=sha256:51ed4391c6e2748e238c14144bd317b919e526785e0dee88f8b7f28fc8ae693e

Observation 037c8e0f-fe59-4ff0-89cc-69a691be0dea · outbound

This paper cites A Systematic Analysis of Hybrid Linear Attention.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding A Systematic Analysis of Hybrid Linear Attention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.338225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.338225Z digest=sha256:dd2a1a94499d0c503494656bf85c0e02c0fbd95c7dbbf82a34d605847dc19b86

Observation b631e196-28cc-4f9c-b96d-fe9e23d8d46b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.342324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.342324Z digest=sha256:616ced6bbb7d9948516bb202649a220afdc1b52b0c57497881f838c5f56aa807

Observation 0be8a52b-9bda-4cc7-bec2-dc520ed3ac70 · outbound

This paper cites IEEE Transactions on Knowledge and Data Engineering , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding IEEE Transactions on Knowledge and Data Engineering , year=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.345488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.345488Z digest=sha256:8484701c18cfe6b729b3ece57e70f9df8ffa4a66ee76aa7457831eacf1a277d2

Observation ae418849-8255-46b7-b87d-7f14f0bec818 · outbound

This paper cites Learning deep representations by mutual information estimation and maximization.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Learning deep representations by mutual information estimation and maximization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.348668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.348668Z digest=sha256:ec8c5602e801dfac0c2c234d4cc21a34757fae317356d204146928e154a0a7db

Observation 173da82d-718b-401c-ba00-c9b11063f096 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.352825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.352825Z digest=sha256:6b40edba63678c692a8d0b2263be2d1c71b8efa1f58e7786f152e7194ebdfa26

Observation c8d034c2-0382-4b57-95fc-bebd575fd027 · outbound

This paper cites Towards Achieving Perfect Multimodal Alignment.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Towards Achieving Perfect Multimodal Alignment

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:09:20.152184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T00:09:18.356102Z digest=sha256:57f5bbf82a1396c440f6e8fe94bc4aa66ada9dad22d97806387e0e788651d2c6

Observation 94bcdb0f-c482-477f-bbbd-02555ecee220 · outbound

This paper cites Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.359522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.359522Z digest=sha256:724db563e8713ac9341afba0a287e48e2b4d1f1d8251c7a9728b2b84a6a20bb2

Observation 6cddadf8-7db4-4e33-b6a6-4492a9084913 · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding The Thirteenth International Conference on Learning Representations , year=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.362682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.362682Z digest=sha256:eaceade2c6cbb78eda3895d71377b07b6974085313149eb137eb7f3b2f5da0dd

Observation d0c07504-d2c5-4e3e-8e0b-66e28e7418a5 · outbound

This paper cites IEEE Transactions on Visualization and Computer Graphics , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding IEEE Transactions on Visualization and Computer Graphics , year=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.366327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.366327Z digest=sha256:fa3264e2f3584d42e9df5d826948b44c02991d9afe695d3c8dca01014bd6afd9

Observation 294af071-c99c-47f8-9fba-cfcf71802ec3 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.370409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.370409Z digest=sha256:7489632b9756bce18f1c6affa774ce40b265a6bd4989fcea1cf72e31a40ede92

Observation 76608d80-5443-43ae-a212-acd4d5b81c2b · outbound

This paper cites Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:09:20.132947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T00:09:18.375123Z digest=sha256:70e19d2c03500cea0191cf111afa34c716321b94ac59e1270e4c5bce50ade0bb

Observation 3663de83-66cc-40f6-80e1-6fc00da5e1fa · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.378889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.378889Z digest=sha256:8191c19eb646e09c6d3552f103ffa0eadce68392b9a6f482918f1d19c29efd82

Observation 9a27dc3a-387a-424d-80b6-d1842df26b8f · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.382502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.382502Z digest=sha256:6f95356602e47c636879120fd12bbbc936d1552d00415621e92b9b27a2b059d6

Observation 582bd7ee-ebb1-42f8-9fe3-c54da476fd65 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Representation Learning with Contrastive Predictive Coding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.386507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.386507Z digest=sha256:9faf7ef08be491fe367509cbc2286709a07b20741d6dfdf0e197d14118fae7eb

Observation 3f693ce0-b55d-4375-ace4-aff1f1285ad5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.390945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.390945Z digest=sha256:be8ebc2e593750b52446d27a3f8163c30971f40bc3362d3cdf39ba8d24f4aed2

Observation 807fff39-ebf6-415e-8287-d348b5d9300c · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.394896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.394896Z digest=sha256:17b701c91fae7c8df1c2353b66d5ffb04f29f827c9e4734565ee7ea62e331e0f

Observation 84c5ce49-402f-46c9-a776-285a972c0e64 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.398783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.398783Z digest=sha256:43b44e5f9baf876ab35d565325b5d195758c7dac27f4fbe0831b0fd8c9b95fb4

Observation c6fa70a8-fc79-469b-963b-e9ce45a69b42 · outbound

This paper cites DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.403553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.403553Z digest=sha256:efb005e96bb04c89504103cc9018dda702dabe0c18a68a459723e5c0ce6a516a

Observation 61759e2c-5b3e-42ff-a09e-69345ff19fa6 · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.407130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.407130Z digest=sha256:0aacd3cc6439ebd0284a9ba5860bbaef8afa518c3e8deb04a67ce1b7e7cb4689

Observation f235c6f8-8702-4b7f-a42a-dd3fdcfcf2b1 · outbound

This paper cites arXiv preprint arXiv:2503.07465 , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding arXiv preprint arXiv:2503.07465 , year=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.410782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.410782Z digest=sha256:9b1b3dd185c6c30ba4fa7b66d7afee55a3d83592323e6ac76178d98b8efc5178

Observation 757268e6-1c12-4301-ad2b-0de76dddf4d6 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.414921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.414921Z digest=sha256:7bd8ea3961324daed23a547901f260edf292fd5b19ecf24de1e1cb58cb765c3a

Observation f442bf30-05d6-41d0-ab8f-0d770b481288 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.418075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.418075Z digest=sha256:ce730c2954b069de5eef4c3f457d8fa3dc0a3521378490e21aa458a3b0b1ce3a

Observation 4a822874-dff6-4f9a-9e56-ec55b0e4d893 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.421565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.421565Z digest=sha256:c343cf2a85de9087bee8027c4600975fb5223a1af0deba714edf0d24b80432f5

Observation e908ccd2-a150-485d-bceb-743d546d6521 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.424416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.424416Z digest=sha256:5dca1703b2a8b5a187408c6aadaed1c8c35554f3a2991c8c0fd10de575025448

Observation 94186c65-14e1-4ef6-9934-bd93689fb394 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.427521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.427521Z digest=sha256:3eff4b519ff3d408cdb3f9ee8274b638a64da3b6c83f04b0060518a6c3749786

Observation b545fd16-3ec7-416e-9ce4-00c77ae285a1 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.431661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.431661Z digest=sha256:9e4d5f493d173c3a4d407988321b7c9a8bce67f2da27fd1e750ac90ceef86d02

Observation 3e262bb6-28bc-4f79-872e-2f0e029ba1dd · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.435447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.435447Z digest=sha256:595e860994f9feb9d35a0f009d516a8474cf88d9e75b74d8856bf68b34d751bd

Observation aa95a6b9-2d9d-4569-9f10-9f499d54cb9f · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.439386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.439386Z digest=sha256:2a149939abf1764a93369915c27c1dae67a72d330296d8f8d590c6fd3eaedcd4

Observation 16e0af0a-7474-46f1-97c5-a6ce637d7988 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.443344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.443344Z digest=sha256:1e46950344ea42a0b4ff1ca0308600c8bb181b745171966bba8c2c8f461625f6

Observation 58e84b30-1e5e-4bfc-82ff-4a46363cabb2 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.447397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.447397Z digest=sha256:021942dab2cca35b04f083b6f560896a665acd5ac4dc947fabc1753a6a25431a

Observation fa101d85-636c-47f0-9dc6-9f4878968d35 · outbound

This paper cites Reconstruction Alignment Improves Unified Multimodal Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Reconstruction Alignment Improves Unified Multimodal Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.451539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.451539Z digest=sha256:5e959bb0410cddb1a196837271e9341ed74cc9158f87114c99dd512ec9222b58

Observation c4606724-0164-4c2d-921f-7e9932cb4b89 · outbound

This paper cites Reconstructive Visual Instruction Tuning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Reconstructive Visual Instruction Tuning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.454827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.454827Z digest=sha256:7b09448e7951e47801e0e6161a1b74469d306f7437b18991d013860a8b2477e0

Observation b9059908-295c-4d50-99aa-3e19fd988604 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.458692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.458692Z digest=sha256:14cbffe7bad3baaefd81cc2db7e7e171ee627e0fc77438ebdf62d9ed5d417b5b

Observation b044f2a9-f4b1-4c70-bb86-18d3710013ed · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.462406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.462406Z digest=sha256:196919d5736c82931945faa60e2203da7a49f7e797edd098a9693ee402b59be6

Observation 2305c43d-1a7b-475a-bb5b-92fcffeac25e · outbound

This paper cites AutoVP: An Automated Visual Prompting Framework and Benchmark.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding AutoVP: An Automated Visual Prompting Framework and Benchmark

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.465655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.465655Z digest=sha256:3c67507df8e7bcdb3254961c2efdee16cfafa4db835d27fd1d5a0564e3c62293

Observation 2098a7e2-3c1e-42b0-ad3d-8915e61b50df · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.468974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.468974Z digest=sha256:0924e8fefd5c1d6c39178a3a1a03f046d165bdef3c04933d902a1f476c524e42

Observation 5a7ef476-59b4-47a5-a3ba-43babdb96744 · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Exploring Visual Prompts for Adapting Large-Scale Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.472945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.472945Z digest=sha256:e25c4470ba97b29f64e3c24f8d739865e1926b6bacfcad8a1d0463e659ca89b6

Observation 7b18dd3c-38af-4f9b-a832-10a6e057418e · outbound

This paper cites , author=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding , author=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.477311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.477311Z digest=sha256:bf518596245f61e80e5d6703ec2fb073e89c06651761349f368372807a2b80b6

Observation 6b4b153e-f91d-494f-b8b1-5e528e43d0a3 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.480340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.480340Z digest=sha256:050a12ffa1d95c9ad555ccfa77c876b35461d91a368f1f7fd78c985e4419e836

Observation 35dfea57-56f9-4d8f-a039-9ece697abcb4 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.483633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.483633Z digest=sha256:d50760eeb2ab13f189572bdf3576d9bfd0304e7d8c38da78159030e79ab02c97

Observation 431e6702-d31f-4fcf-be6d-7a75c5fa7469 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.487512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.487512Z digest=sha256:00fae3a125187b4ec30cd876b9c1a425a11d7b460a9736e485886f5418e692e4

Observation c7c9fe7d-7655-4784-8ee2-dc942b7961f6 · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.491464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.491464Z digest=sha256:e24a1fbcdfe82e5d17e61f8fa951b5ece66beaea835a511b316bb3ca32d9ba99

Observation e587b3c4-fedd-4f32-9535-011c1aad0ba1 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.495392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.495392Z digest=sha256:e208d5e1641bea56972b470f692bb3750e2a81b09a847d1ec0523f5d8aa35be7

Pith citing papers

No inbound Pith citation observations are available.