Pith. sign in

Paper Citation Record · LEDGER

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

As of 16 August 2026, this Paper Citation Record lists 100 of 287 outbound references and 0 inbound Pith citation observations for arXiv:2608.12748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12748 v1

Coverage vector

measured 100 of 287 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:09:18.495392Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 287 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5ada3e7-5e40-4954-926e-02eec0f49532 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.131394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.131394Z digest=sha256:7724de1d37c92c94da542b99789ed32f24218f7c02201d60fc582e1675c33c70

Observation a81588f9-4d07-4de5-8d0e-2e7ba330ccfe · outbound

This paper cites Applied Sciences , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Applied Sciences , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.135472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.135472Z digest=sha256:f656a6f53c1c41ee4e76f0a990a7a2916480333fe9461d82a9deb019178c4601

Observation 7a1532a6-cb00-4695-8031-9a8beedf3cce · outbound

This paper cites Signal and Data Processing of Small Targets 1993 , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Signal and Data Processing of Small Targets 1993 , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.138865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.138865Z digest=sha256:9db9bf30b3471855bb51d2f052770aa58cc6f44e4a96548a20a23a5ddf9652d2

Observation 3febe7c0-0ae0-40c8-80fe-e7aed9825ef5 · outbound

This paper cites Sensors , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Sensors , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.143478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.143478Z digest=sha256:d5603ac1b31d168b85ea1d6de773d3e3c6ba3c8b1849d9a663064dce4f8c2351

Observation f5a86670-90e2-4c0a-aba4-449af04cb6c5 · outbound

This paper cites Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.147495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.147495Z digest=sha256:8dc4a89b412886be3eb015c5238a4550e935681e86ec54e7ab7559ed1e8d465f

Observation cb5ff312-aa1b-4059-9149-a9c07845df58 · outbound

This paper cites Sensors , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Sensors , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.150786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.150786Z digest=sha256:618ac5342f8ff1a3561e8a86b5d06a99902e8ac1e7ea80562c572e8f1f7b1b6a

Observation 6cac7bdf-ef22-4c44-96bf-66fdb7166eae · outbound

This paper cites 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.153828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.153828Z digest=sha256:69b4d06f8334324f6d74594ef8fd63a7691d9608d973ce955b6eb54e44840c73

Observation c637ee4e-6a73-4ed3-acfd-8baccf73b7a7 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.158526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.158526Z digest=sha256:ae3e6f643eef2e0ef2e36f563960360a282802f31b37fd48f3afa3dd08939c42

Observation ee327802-dd29-482f-9599-43871f64980d · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.161713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.161713Z digest=sha256:753665da716af9685948c51b5f2b09360979b32b2adb1a2638dfabe367e5e6c6

Observation 6713d00b-68a6-407a-ac83-7f0e7df22efd · outbound

This paper cites International Journal of Computer Vision , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Journal of Computer Vision , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.166091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.166091Z digest=sha256:6d70ff70482dd38fe07eac19acaf72bc7968be789f59733c1523cd0d660bd04f

Observation 198ced30-915a-4898-b3d7-56b6f1096078 · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.169235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.169235Z digest=sha256:072e4271f78cb71bc5df804e0aea3da67f2fabe12a18d65258cf8b35f7922c3d

Observation 1ef4a677-338f-4948-8ef1-88dd4e14e538 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.173688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.173688Z digest=sha256:8c1d3f73773d067f3e23e45e8da876a03631d023c90e3a8e42230c25a9ce3d8a

Observation f40c3e1e-bc8a-4f09-ae70-5f1976c6468d · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.176903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.176903Z digest=sha256:9a877190438d95600ab360f904b57e91edc0efeffc032c658933bc2b9ca379a0

Observation a44725a2-b514-444d-b8f9-0221777dcbf7 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.180236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.180236Z digest=sha256:b5ade3746eb5c9d1a91be0df96913a5953ff43922471f228c213114e6b7784a1

Observation 2384866c-174a-4bfc-b306-6b373233115a · outbound

This paper cites International Journal of Computer Vision , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Journal of Computer Vision , volume=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.183249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.183249Z digest=sha256:c12f02b7a124bb4c4d4ac297187db30a4fe74423f631c61dc93209e47bed2597

Observation 3150fbd9-0146-4dc5-9f1e-09ffd12aff8c · outbound

This paper cites Infrared Physics & Technology , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Infrared Physics & Technology , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.186095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.186095Z digest=sha256:4716e02131bcf9446862ffee0cb1329a055b1ff3e4c68c2f4cfee22aa012d0f5

Observation a41111da-8d77-4612-a8ce-e501f6f99bfb · outbound

This paper cites Kaur et al.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Kaur et al

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.189270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.189270Z digest=sha256:b1429778bac25b3471c4fd80ee46d581d099b8d4cae8a70075d7dfdcb1b851e4

Observation ea7d1d10-5292-4302-826b-a1e991a320aa · outbound

This paper cites , author=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding , author=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.192080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.192080Z digest=sha256:48050881efb6d87ebe5f9721ece429d35370f9b0633af65015614fc044d85b47

Observation 738f01d4-a2f8-47b3-8b23-7b6b60045d01 · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.195382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.195382Z digest=sha256:b1e60c5c57224d163cc74b19a7e23a73fb4b9b5979c8a975cf9d32881b089225

Observation 169abf66-44fd-43ea-a19e-a84eb2cf17fb · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.199801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.199801Z digest=sha256:4fa790337a3b36da0da4f55f039884ecacb8c15b8ec6c85c20509dc2ab083441

Observation fb978c6a-5f8e-4476-9cb5-8f2be0c79f39 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.203629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.203629Z digest=sha256:a626d8951d7f255045ec36fc302a441a5bddf3ff5d505df13726eb84e299a25a

Observation 6a3d78cb-6c88-43a7-8995-2200ec1fd3a1 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.206600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.206600Z digest=sha256:9ce022c5e2b94dcd229524bb5c4f90096e664f930586b7362e13cd98840525dd

Observation 9e8b2e84-e73c-4461-91c4-9bf00688c6d6 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.210068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.210068Z digest=sha256:edb14addf79186f90286cb6d4b0e4d603836321c16b112dbe343cb8580b56c29

Observation 198ab2ef-01d2-41d7-923f-fb195d791173 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.214296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.214296Z digest=sha256:d809405694189b9967fb806f6848b593d31447eff78065271d0a8c93c708e194

Observation 790836ec-2a49-46b1-80d6-4ad9eb6be0a5 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.217495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.217495Z digest=sha256:e81501f9d3305a196cbca5cc9f89af849aad60c53172a6554d68fe4622709f79

Observation e5373747-d872-4110-aa85-a96428c7291d · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.221276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.221276Z digest=sha256:6a04502051870299223329d5b85f2a6f78e056e432378f0d5258b72b6296a1c7

Observation 0af7ddf2-e113-46d1-9b70-d2ab4e6bab68 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.225163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.225163Z digest=sha256:7d8dec5417d81edf5df6aaae3f7e650b7f6bd2d77d35c85a9949f82c1ba05712

Observation a9b04ec1-e611-4fc4-9a3b-085206799bd3 · outbound

This paper cites International Conference on Machine Learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Conference on Machine Learning , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.228795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.228795Z digest=sha256:49b936e077b567b4ab8cf60839fdbde50455d36822a525a189ae846ca1bb1171

Observation 7c113b82-daa2-47f6-aabb-a6533ce590bc · outbound

This paper cites 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.232127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.232127Z digest=sha256:47d4044c5ecf2fdcf9d144ff6669bc0753e9c721fdf4e009c6b1e92c720a3cd2

Observation 166969f3-0ec7-4d50-8c18-99cf64e4e238 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.235430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.235430Z digest=sha256:df8629bd688965c5109f4480f86cfc541eddd21acfe5da851ce178f88e57a2a0

Observation 631494c5-6dca-4ca1-a8b1-6655b275f04d · outbound

This paper cites Visual Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Visual Intelligence , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.238598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.238598Z digest=sha256:d1236adc7209923f173f8c4e408e2d5fa626f90752f9a6b14e9a900ff280e5f3

Observation dd53a003-ce60-4e95-b197-78914fbff2e7 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.241956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.241956Z digest=sha256:f682a608364089f3cac53d783226782b0c280834a5ff881e378a1d5a80719dd5

Observation a14fb8f5-6243-4fdc-85da-f6b0fde0383a · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.245336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.245336Z digest=sha256:326bebe799019cd403b8c305d6c039e0fe6ffe188cc9cdf1f13aa827d60756e6

Observation 24ddecf9-0999-43a8-b6dc-1a18525ae5f2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.248499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.248499Z digest=sha256:3694071e0a46ad7bfa53e935a0199d8d2e216878af37e059bb6dcf42424429db

Observation 38806458-f364-4829-b4fb-fc5b09da3680 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.252010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.252010Z digest=sha256:969e1fd01294bc37bbce886547e0f119930567791abf690acce1f965541aa492

Observation 47feca59-be9f-4838-aecf-0d40e502dc86 · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.255311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.255311Z digest=sha256:375aac03bfd127ffa6c29092c0aa64616da57fd3cef428fc8813d3e8c4863f71

Observation d0ff5c32-6006-4d1b-9c83-d82dc1017147 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.259063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.259063Z digest=sha256:ad2afd3a5d75399bffd7a38b0d3a82df357e24b7199c5ec819cf20a471ecb661

Observation bbd3066d-cf2a-47a7-a5eb-dfda3484a6bf · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.266450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.266450Z digest=sha256:9024240bf140f80048b9ab0031f5cbd714eb58c6cea3c475247d76e412c626da

Observation ff44c456-9f22-4a7d-8893-38d52e44efd7 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.269408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.269408Z digest=sha256:f25ecc1e5eb5da2a63568eca6bd6e7cc5a5633254f39e6f82c646b178c1d13c1

Observation b413ac14-a37f-46c9-8167-fba1bd896f1f · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.273874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.273874Z digest=sha256:d48031f98ddca24fd48cc2e0b82882f2a5af47a83bc9f8246b4b9b24dcfd1000

Observation b8c98ffb-b1ff-4c18-a79d-eebcbcaeb8b9 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.277272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.277272Z digest=sha256:fb4ef0705dbd1972e72ff1c5afcab1f84d0f0bcffd1b98b83cb08404ac48e185

Observation f1431835-d342-4cc3-88c2-5a7d44c75e60 · outbound

This paper cites International Colloquium on Automata, Languages, and Programming , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International Colloquium on Automata, Languages, and Programming , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.281345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.281345Z digest=sha256:52640ad0ca9747c441ef6a9fb30e167be1b37d6925405b66fa68d8cf4e6e309b

Observation ba2da6cd-5520-4f45-987b-2d330b2ec766 · outbound

This paper cites Proceedings of the IEEE international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE international conference on computer vision , pages=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.285481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.285481Z digest=sha256:08cdb286d932cd813e881269afb432dcb60ef3abbc250e30136295c2b619731a

Observation 3033de1b-51b1-4341-b512-1827915139ba · outbound

This paper cites Tensor Fusion Network for Multimodal Sentiment Analysis.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Tensor Fusion Network for Multimodal Sentiment Analysis

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.288993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.288993Z digest=sha256:ce5efe3773370d10f1ebd553981e75364faf294189e5ea7a722479d324bf5185

Observation 93cf72aa-8317-4099-b5b8-1fd6a6d3c3ac · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.292484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.292484Z digest=sha256:82fe36709bef5553eade9abe4a33a33f77dec328bbc92275e9c9093cae4da189

Observation 862e57a0-a390-4f67-8b7e-6b008c20bb70 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.295725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.295725Z digest=sha256:77a8dfa9cfc70a2765293ee2b476822e2064e8c24e04e950093baec3ae4c64c0

Observation d13c8183-9b69-418d-9ef0-3948308fe9d1 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Fast Transformer Decoding: One Write-Head is All You Need

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.299513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.299513Z digest=sha256:185c21041b4029baa19af996b9c443b991569210e39c449249d7d5c34b6d0d0f

Observation b342f152-1b0d-45bf-8ac5-deeb7ca75eeb · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.303079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.303079Z digest=sha256:36f821034fcfb9b330f6f7e894eefbe9a8b874123b47b7ad62becd8883e1c8dd

Observation aeb73f93-bd3a-49a8-88a1-8b5313b2fb04 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.306273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.306273Z digest=sha256:e621e2c08aeaba8d8f4ed895f5094ae2837bce0440d5d036c7f3c46541174d9e

Observation 58aa6810-536d-42ef-8104-bda0302ae315 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.309435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.309435Z digest=sha256:95a81523a09f983c00afce6d16fb15ac996f112526a18e74b36b1fe5c221c19d

Observation 20c76a08-459a-4705-bf22-b9973526bb4c · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.313282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.313282Z digest=sha256:5838655c098ba306fffda05110858a756eae272423e024a06e147d5c2fbd0411

Observation 5c05fb9c-087a-4a51-a5f0-11564571a76e · outbound

This paper cites Scientific reports , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Scientific reports , volume=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.316532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.316532Z digest=sha256:f4e5132f66404411003f3f45b45ca2d528cf5d0eaed662607f7ba0e71ac1c969

Observation 434536f2-3d05-492c-9ff9-c499cacfbf0f · outbound

This paper cites Advances in neural information processing systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in neural information processing systems , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.320625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.320625Z digest=sha256:c82771445ad8403ee8e2f1b50d5c906d0855e29b0cb1238bfae7fe6ff9c5f577

Observation 20aeb9fc-85c8-4089-b7b9-e16e7c68811d · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.324123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.324123Z digest=sha256:a37469f248d1416a6340f8215dce12a98e5b186ec17135f606dbe060c3d19fc6

Observation 4f1fd045-68ac-451c-992e-7c6c08929aa8 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.327291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.327291Z digest=sha256:4e4dc824b437a62ea0f381ea1773a0440969f3d6d8730eddb95a87b7e2859591

Observation fe591112-dc85-4e93-b70f-bc762eb20c72 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Retentive Network: A Successor to Transformer for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.330658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.330658Z digest=sha256:d6b6ca91ae4527364485c2fce85f37330f19df8973b1b3c1e16e12a162a7831d

Observation 834b29cc-6fe3-49c8-a74f-c1a1d8b1ec3d · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding RWKV: Reinventing RNNs for the Transformer Era

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.334068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.334068Z digest=sha256:4e891bfd8db536a40894d3eb1376f0ff69240b03aef1a3205aff0c17b6aea2fb

Observation 037c8e0f-fe59-4ff0-89cc-69a691be0dea · outbound

This paper cites A Systematic Analysis of Hybrid Linear Attention.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding A Systematic Analysis of Hybrid Linear Attention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.338225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.338225Z digest=sha256:53793da1190b90103275ffc93d077ca0561f52ad0bb7fb25d6554f156131c060

Observation b631e196-28cc-4f9c-b96d-fe9e23d8d46b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.342324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.342324Z digest=sha256:c5cb35c6001f39e0d1d8ea0a7df6dd9c64cf661dc83bb400ec01c463b351161b

Observation 0be8a52b-9bda-4cc7-bec2-dc520ed3ac70 · outbound

This paper cites IEEE Transactions on Knowledge and Data Engineering , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding IEEE Transactions on Knowledge and Data Engineering , year=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.345488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.345488Z digest=sha256:ba83111cc966a5783cc1f16d328a4d3086e84b41971575b47ef3464f63439bc0

Observation ae418849-8255-46b7-b87d-7f14f0bec818 · outbound

This paper cites Learning deep representations by mutual information estimation and maximization.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Learning deep representations by mutual information estimation and maximization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.348668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.348668Z digest=sha256:72b2ca57380de8ab60b36bc5113af8467ed0acf9f001357949499e2ae1e45dce

Observation 173da82d-718b-401c-ba00-c9b11063f096 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.352825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.352825Z digest=sha256:e8ad20b9f90255a8d9e14af4c73c3eebcefbd82bde656fe98ef5292f883434eb

Observation c8d034c2-0382-4b57-95fc-bebd575fd027 · outbound

This paper cites Towards Achieving Perfect Multimodal Alignment.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Towards Achieving Perfect Multimodal Alignment

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:09:20.152184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T00:09:18.356102Z digest=sha256:0999cde9983746c2e69f30c426c85ee507bcc41a3f1e4b3cd3b0d86058f4892f

Observation 94bcdb0f-c482-477f-bbbd-02555ecee220 · outbound

This paper cites Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.359522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.359522Z digest=sha256:d98d054c4f99a3bcaa20dec6bf89c5681ca48411fae5e93690fef1ca0d1eea87

Observation 6cddadf8-7db4-4e33-b6a6-4492a9084913 · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding The Thirteenth International Conference on Learning Representations , year=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.362682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.362682Z digest=sha256:be9a685c82570810da5426eb3191b82f7d629cac7690795c79ceea627fd76c8c

Observation d0c07504-d2c5-4e3e-8e0b-66e28e7418a5 · outbound

This paper cites IEEE Transactions on Visualization and Computer Graphics , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding IEEE Transactions on Visualization and Computer Graphics , year=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.366327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.366327Z digest=sha256:7b67f5346d2819549581dbfdb15a16afe4b9f7043a8dd726cf8d568a92061ad7

Observation 294af071-c99c-47f8-9fba-cfcf71802ec3 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.370409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.370409Z digest=sha256:7b68af21c18d7f9f0f0bc12057ca76b03324d00812d992f677c0d81dd6e08745

Observation 76608d80-5443-43ae-a212-acd4d5b81c2b · outbound

This paper cites Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:09:20.132947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T00:09:18.375123Z digest=sha256:fc5a278676b2a37436ae9fda0aee5c7987740d9f66fae1a2617e59b6514c97e3

Observation 3663de83-66cc-40f6-80e1-6fc00da5e1fa · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.378889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.378889Z digest=sha256:52da743a8b976343daa86fc58abc8948ac934b7361dcc29b1bd70a24f08ebde4

Observation 9a27dc3a-387a-424d-80b6-d1842df26b8f · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.382502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.382502Z digest=sha256:3f89dfa4406ce12c3604f76d54b95c13fe81354e49134edcec97cbb6f3099c57

Observation 582bd7ee-ebb1-42f8-9fe3-c54da476fd65 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Representation Learning with Contrastive Predictive Coding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.386507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.386507Z digest=sha256:0b39ff7e59186b428527ce05771d5dcc73721974c8237254016f1e0fb53b58f4

Observation 3f693ce0-b55d-4375-ace4-aff1f1285ad5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.390945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.390945Z digest=sha256:3e97c56ef4462a3b4599487bb4011858e4ca675c09ce703ec616a7576121cf2a

Observation 807fff39-ebf6-415e-8287-d348b5d9300c · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.394896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.394896Z digest=sha256:cd1a28abafa7bca6c497cdb2517af3fdf1d2c3480b48740a8e42be5ae1bf0154

Observation 84c5ce49-402f-46c9-a776-285a972c0e64 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.398783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.398783Z digest=sha256:6a50dc63c2bc62a6fdaf39c93543fd6300e753c4c2282bfb1343940941510d78

Observation c6fa70a8-fc79-469b-963b-e9ce45a69b42 · outbound

This paper cites DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.403553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.403553Z digest=sha256:6c5afc5e04df2d6e4f3ebf316d255326d2b484facaf0c533622b243f61de7299

Observation 61759e2c-5b3e-42ff-a09e-69345ff19fa6 · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.407130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.407130Z digest=sha256:e1e2d6b07e03f13c9e75b4a190e1f2ba2856d7d475aa2c23dc304930e89d743c

Observation f235c6f8-8702-4b7f-a42a-dd3fdcfcf2b1 · outbound

This paper cites arXiv preprint arXiv:2503.07465 , year=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding arXiv preprint arXiv:2503.07465 , year=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.410782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.410782Z digest=sha256:050498458128b9122d3a659a5e021c090d42b16d4f8e9c222242a45ac94ee2a3

Observation 757268e6-1c12-4301-ad2b-0de76dddf4d6 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.414921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.414921Z digest=sha256:a2e71404747fb2e6b74c787efd1c3abd9ce693da3de74d8c6d2be51ed96f827d

Observation f442bf30-05d6-41d0-ab8f-0d770b481288 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Advances in Neural Information Processing Systems , volume=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.418075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.418075Z digest=sha256:b0a2059138ef2628b1775eb8edb0719bcab10abc642abbcaa6d6c552d6183b10

Observation 4a822874-dff6-4f9a-9e56-ec55b0e4d893 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.421565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.421565Z digest=sha256:8d8955e9902df97368635a93605bb1f5160d5d225ec4aa22c5099ad674bfa5ac

Observation e908ccd2-a150-485d-bceb-743d546d6521 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.424416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.424416Z digest=sha256:014a11b10a9c26ec06b8536de19cf51f19b47edf057b416d3e4752afd185e6df

Observation 94186c65-14e1-4ef6-9934-bd93689fb394 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.427521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.427521Z digest=sha256:f2cbf8699d7ecacade013353371dab0faf4e62d65c6de8ddb683849d987e3897

Observation b545fd16-3ec7-416e-9ce4-00c77ae285a1 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.431661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.431661Z digest=sha256:6577f5007955423e08ce47caca395dc482c2da7e7048be4c59f3662c5f86f512

Observation 3e262bb6-28bc-4f79-872e-2f0e029ba1dd · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.435447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.435447Z digest=sha256:e51ad4881ad861f7cfe69d1861bd84f331228dbef805e83cc9ac09fb85408e82

Observation aa95a6b9-2d9d-4569-9f10-9f499d54cb9f · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.439386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.439386Z digest=sha256:b2b90337c9542a6e01c34387980b7300d5254011d012528834408c3fd4b84e24

Observation 16e0af0a-7474-46f1-97c5-a6ce637d7988 · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.443344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.443344Z digest=sha256:afd39899ca3038e6d9798453e06dd73b5332e005f7e10deaac68973464215aa0

Observation 58e84b30-1e5e-4bfc-82ff-4a46363cabb2 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.447397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.447397Z digest=sha256:46febbc87b1eb30b62c4f8a5b302a186555bbf5a395ba3a26e7c047a7040ee9c

Observation fa101d85-636c-47f0-9dc6-9f4878968d35 · outbound

This paper cites Reconstruction Alignment Improves Unified Multimodal Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Reconstruction Alignment Improves Unified Multimodal Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.451539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.451539Z digest=sha256:fd2dcfdd765bde364bd61c4d7ed69d3ca2b01aea2564132d5837f6d8054be40e

Observation c4606724-0164-4c2d-921f-7e9932cb4b89 · outbound

This paper cites Reconstructive Visual Instruction Tuning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Reconstructive Visual Instruction Tuning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.454827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.454827Z digest=sha256:9b3b4cac3eddde294eed083d0cab04d9dda3619e0c9e3b18135dd85259abb030

Observation b9059908-295c-4d50-99aa-3e19fd988604 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.458692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.458692Z digest=sha256:7fbaa7feadfabe175dfa34fe99ac1bbf944e4ede79881ad08ec9dc4b4df4f8d4

Observation b044f2a9-f4b1-4c70-bb86-18d3710013ed · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.462406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.462406Z digest=sha256:5da13ecca100a3de561f08417446097fb5a77f6aa52f891412bbf716bf6672e7

Observation 2305c43d-1a7b-475a-bb5b-92fcffeac25e · outbound

This paper cites AutoVP: An Automated Visual Prompting Framework and Benchmark.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding AutoVP: An Automated Visual Prompting Framework and Benchmark

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.465655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.465655Z digest=sha256:fe2b289b36d5f3aa7dbe859a073ddd5ad173f1310ff619db8d7fb47278c5f9f4

Observation 2098a7e2-3c1e-42b0-ad3d-8915e61b50df · outbound

This paper cites European conference on computer vision , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding European conference on computer vision , pages=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.468974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.468974Z digest=sha256:f2e9282c201c995d8f19265274eadae8b62ef9511a42700ed9f68858a14e559b

Observation 5a7ef476-59b4-47a5-a3ba-43babdb96744 · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Exploring Visual Prompts for Adapting Large-Scale Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.472945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.472945Z digest=sha256:053805c110546f063e279547cea89666855630ec49cc3b9c14fc155a78f21060

Observation 7b18dd3c-38af-4f9b-a832-10a6e057418e · outbound

This paper cites , author=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding , author=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.477311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.477311Z digest=sha256:181d3edb162044dc335301b312f1570fc2a37bcbdddfc3d9dcbeae258837a40f

Observation 6b4b153e-f91d-494f-b8b1-5e528e43d0a3 · outbound

This paper cites International conference on machine learning , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding International conference on machine learning , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.480340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.480340Z digest=sha256:28df884f16b92136ac430139d9b5d245dee5d31a1998592519638376f7511e30

Observation 35dfea57-56f9-4d8f-a039-9ece697abcb4 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.483633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.483633Z digest=sha256:96a7cc75fc33e1c680e418c29226502d355d7fcda56ec5b7318882d67b7f7aca

Observation 431e6702-d31f-4fcf-be6d-7a75c5fa7469 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.487512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.487512Z digest=sha256:49bec1f402d76587ae45e50f7733a4312564fe5fe1d3464ec25f0e5de1225516

Observation c7c9fe7d-7655-4784-8ee2-dc942b7961f6 · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.491464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.491464Z digest=sha256:5ac488116bbe5bdc69704482a5e68c6642a5c7b0750ab0a62d3f8fa0235c0e2f

Observation e587b3c4-fedd-4f32-9535-011c1aad0ba1 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.495392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.495392Z digest=sha256:0a32adc6eaf186174a47a23a001ba16081b65b36b09839bd2e88f66193220938

Pith citing papers

No inbound Pith citation observations are available.