Pith. sign in

Paper Citation Record · LEDGER

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination

As of 23 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2509.04833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04833 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:30:35.361785Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy70
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 765229e9-51ec-4c6d-b8ae-bd2c9b503962 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv, 1 (2):3, 2023.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv, 1 (2):3, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.913648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.913648Z digest=sha256:c17597de7d3b69a83432e47f5c0e061d2d39105142119f7970eef30cd21ee61e

Observation 5b1333c4-4e84-4484-81ec-a2e20de91120 · outbound

This paper cites End- to-end object detection with transformers.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination End- to-end object detection with transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.919228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.919228Z digest=sha256:2e92b57017fa6a8416da88e5005309fd3d901e653da969316b7deabdb0838a61

Observation b6c361a7-e809-4865-916f-d958f137ff60 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowledge.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Lion: Empowering multimodal large language model with dual-level visual knowledge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.924724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.924724Z digest=sha256:57250b05ce613dc14e7ee282db32722e47e775258cf8c74ec1200b7973dcc2a4

Observation adcc4ed3-bc77-422b-95c4-82c1df66f4e1 · outbound

This paper cites Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.931379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.931379Z digest=sha256:e9843a1a419e6547e14d2074dfe775cdce17831c1c51f6f88bfe5e92d8ba4024

Observation da0f035f-081b-4310-b3b4-c9f814b6aee6 · outbound

This paper cites An efficient and effective transformer decoder-based framework for multi-task visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination An efficient and effective transformer decoder-based framework for multi-task visual grounding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.936619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.936619Z digest=sha256:2f455c099ade97e7424b8d5afec328c0d6b7bc1da8a6c8009217dc9f19aa322a

Observation 34c5bbfc-ff38-4bec-9bf5-3e97ad4d0094 · outbound

This paper cites Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Sam4mllm: Enhance multi- modal large language model for referring expression seg- mentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.942326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.942326Z digest=sha256:81ce83d2a728cf1799759172a559e92441af1e1f7904034c1d458aaf526bc896

Observation dc3f05ff-7616-4224-a639-5fa0b0d07323 · outbound

This paper cites Parallel vertex diffusion for unified visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Parallel vertex diffusion for unified visual grounding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.948052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.948052Z digest=sha256:3203de75153265de1b26ca2ec0c5cca5dddc4baaea6dde10cd8d1406b1b51596

Observation f709f40d-8c1d-4978-8adb-ea7a99cb3b78 · outbound

This paper cites Mask grounding for referring image seg- mentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Mask grounding for referring image seg- mentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.759092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.953226Z digest=sha256:669055f59897f5ab13140e91edd57f6b08afad2fcce3ba7df0f9fa4f2edd48b5

Observation 364a295b-e8a4-489f-8455-0927806a9e0e · outbound

This paper cites Simvg: A simple framework for visual grounding with decoupled multi-modal fusion.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Simvg: A simple framework for visual grounding with decoupled multi-modal fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:34.958852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:34.958852Z digest=sha256:0ea772e40f35db764e4844eb2b590a8c9081c2fe84a1359aaebe0167c8134f8f

Observation cb26641c-eee3-4e65-9e94-4141070bed72 · outbound

This paper cites Deris: De- coupling perception and cognition for enhanced referring im- age segmentation through loopback synergy.ICCV, 2025.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Deris: De- coupling perception and cognition for enhanced referring im- age segmentation through loopback synergy.ICCV, 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.731285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.963928Z digest=sha256:14aea78ab478859114628640fc6ba46341ebd2de2cb0267609afd4805ac67710

Observation 649f2a40-bbb0-4113-8508-c72bd8727245 · outbound

This paper cites Multi-task visual grounding with coarse- to-fine consistency constraints.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Multi-task visual grounding with coarse- to-fine consistency constraints

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.714323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.969480Z digest=sha256:0b3ea2503837a1c3e346289028f56af1caf288c5329e043a9914047e0984cf79

Observation 90f590c0-ec82-4c7e-825b-c476c97e5767 · outbound

This paper cites Transvg: End-to-end visual ground- ing with transformers.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Transvg: End-to-end visual ground- ing with transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.697483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.974699Z digest=sha256:e201fab573b40dfa7a342d68a660a3919b2a6c088c66787deb680a5af4806409

Observation d7650124-3954-4132-8d5e-306bcd43b77c · outbound

This paper cites Transvg++: End-to-end visual grounding with lan- guage conditioned vision transformer.TPAMI, 2023.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Transvg++: End-to-end visual grounding with lan- guage conditioned vision transformer.TPAMI, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.680382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.980115Z digest=sha256:fdff89dbd14a78ff914ae113fa274b3f6e2943769fb10de0fa3bb81244f6d9ea

Observation 2b3363f5-146e-46c5-8960-c05949fa9b1e · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Vision-language transformer and query generation for refer- ring segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.663659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.984999Z digest=sha256:9e0d808f31eae882b308000950925ff613f8f11f9e75f8cd80e39802514d70c5

Observation 6fe7b854-d9c9-4e88-be9d-61e0abd7db4f · outbound

This paper cites En- coder fusion network with co-attention embedding for refer- ring image segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination En- coder fusion network with co-attention embedding for refer- ring image segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.646875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.989692Z digest=sha256:da9b12399b9597f10629a35aa0a809caa01f6965a77389ea5f266c3f881c40a4

Observation 96f86004-9e27-4a98-998e-c59243da7a63 · outbound

This paper cites Mask r-cnn.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Mask r-cnn

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.629978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.994512Z digest=sha256:b13ffa2d078d9a840f6636aa43f95df1c382438db1a9f1c30dbe3716fb9b3dc3

Observation d68347c4-d6d1-4e73-a07f-0e4ab0e5bb9a · outbound

This paper cites GREC: Generalized referring expression comprehension.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination GREC: Generalized referring expression comprehension

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.614403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:34.999801Z digest=sha256:69416b51bed23a77dbdc2399ac83abf93abc7ce1b210f71154d60b950a97ae59

Observation 134483bc-420b-4c0b-8a28-e73b5e42837a · outbound

This paper cites Learning to compose and reason with lan- guage tree structures for visual grounding.IEEE TPAMI,.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Learning to compose and reason with lan- guage tree structures for visual grounding.IEEE TPAMI,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:35.004958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:35.004958Z digest=sha256:6c78c982b859e313426dcb4ca16daadf3637e765a287b33005f8fcc1d7226d5b

Observation 2c93d68f-8385-411b-9d7e-65663fc4d444 · outbound

This paper cites Seg- mentation from natural language expressions.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Seg- mentation from natural language expressions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.587391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.011137Z digest=sha256:31bc3770c22f01a587a0fde59eb9c97bd2e78eebb93235e4426be15c8e326a02

Observation 6e1debb1-c172-4ae2-91aa-ce9b9d19c3d0 · outbound

This paper cites Natural language object retrieval.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Natural language object retrieval

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.571108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.016131Z digest=sha256:2c21fd43a7384e9662d3b88d0ff9f46ef123132f188e78f93e6c715a0cf3c4ac

Observation b10e7153-7c18-44e9-9049-a1d7da179e70 · outbound

This paper cites Modeling relationships in refer- ential expressions with compositional modular networks.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Modeling relationships in refer- ential expressions with compositional modular networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.554653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.022036Z digest=sha256:b09a69af246ab5572224bd91e442c808bef4ebe5fd9e940deea52311ac025cf6

Observation ad3c0b03-f102-45ad-a059-e8addd1b4471 · outbound

This paper cites Beyond one-to-one: Re- thinking the referring image segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Beyond one-to-one: Re- thinking the referring image segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.538519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.028078Z digest=sha256:8b8be46c11c2297e9faf5d11c87b7a0ab2fd958c3a25d21f4ea218bdf8deaba4

Observation 7250de98-a3cc-41c3-8882-aa6b2d35e113 · outbound

This paper cites Densely connected parameter- efficient tuning for referring image segmentation.AAAI,.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Densely connected parameter- efficient tuning for referring image segmentation.AAAI,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.520959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.033843Z digest=sha256:0d2c5f2cc3d3b8f18cf41861ddbcf0b91d52360217ac3c1d0fe99afdb647e4d3

Observation ce62fc49-8421-46e6-aec3-b3131b7ccaa9 · outbound

This paper cites Referring im- age segmentation via cross-modal progressive comprehen- sion.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Referring im- age segmentation via cross-modal progressive comprehen- sion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.502995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.039292Z digest=sha256:842592a8d65c7975ae50f1494d3e3d56a2494436816d6eaca4a5fb5cddf4a9ad

Observation 30d73a67-d4ca-4fa9-8a36-614b1ba5a24b · outbound

This paper cites Mdetr- modulated detection for end-to-end multi-modal understand- ing.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Mdetr- modulated detection for end-to-end multi-modal understand- ing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.487048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.044178Z digest=sha256:3e6e4adfac79385f182578470e185cb2351ea6c07e08c41a464224e383d5b4ca

Observation ebe22ec4-b737-4e47-beb5-c01dd628759a · outbound

This paper cites Segvg: Transferring object bounding box to segmentation for visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Segvg: Transferring object bounding box to segmentation for visual grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.470972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.049337Z digest=sha256:409be47ddcdcfd234e28621e2cc9a812562d27ab865c63a983a349a089f7040d

Observation 32edafaa-a7e1-4f8b-9468-8cb196493698 · outbound

This paper cites Restr: Convolution-free referring image segmentation using transformers.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Restr: Convolution-free referring image segmentation using transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.454881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.055190Z digest=sha256:08a4febb2f3a19d97b19ee232bbd524d10caf8ee07f8cd019fae599cbe320cb6

Observation 3d5a1919-8a79-4211-b698-1259fafb1328 · outbound

This paper cites Segment any- thing.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Segment any- thing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.438074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.067703Z digest=sha256:0567697b46d4b391d28841b3726469c07e7f6d0afc32f6a3762f1ace6c75086b

Observation dec76a0f-d129-417c-b6bf-3881e0b54bbf · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Lisa: Reasoning segmenta- tion via large language model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:35.072843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:35.072843Z digest=sha256:be2c2a7816b5343831595ed5ccf33e4b9376c4b9cd664a2ed09ea666cc029563

Observation cbf0d661-ab40-43d8-97cc-d74993ddc0e8 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination A Survey on Benchmarks of Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:35.077904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:35.077904Z digest=sha256:4d39a68fc8f81fae3a0afa155bf56bdf71b6731566650770072791e64f970b97

Observation 0f5dec88-7522-479f-96d3-f4f4fa5a2510 · outbound

This paper cites Grounded language-image pre-training.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Grounded language-image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.411211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.083881Z digest=sha256:2d341bb66f1384482b469439d25e7f8a2760e014fd261cdf68bf1ebfc9ece627

Observation 865d57be-32ae-4aaf-ae28-ba14c8be83be · outbound

This paper cites Referring transformer: A one- step approach to multi-task visual grounding.NeurIPS, 34,.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Referring transformer: A one- step approach to multi-task visual grounding.NeurIPS, 34,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.395214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.088975Z digest=sha256:b41c6eebf1f8d12616892cdaa180420ce47334cbb55486b0f200533f4764f6f7

Observation 0f692e05-dad9-4140-a94f-6593473bd908 · outbound

This paper cites Bring adaptive binding prototypes to generalized referring expres- sion segmentation.arXiv, 2024.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Bring adaptive binding prototypes to generalized referring expres- sion segmentation.arXiv, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.379145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.093877Z digest=sha256:6dc9fdcec9be63aacf2e06f87a10faa9ea89282a9765d68c1d58a09161237f8d

Observation 361aacd1-8b4f-4147-8c32-0b0c2f3bba59 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Exploring plain vision transformer backbones for object de- tection

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.361264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.098740Z digest=sha256:2cbaee3b0245b658ff88ab1ef828bbe9c0cdf287071ca0e7ffdee4b9d8000a69

Observation 30bd41d5-1578-43e0-8116-9dae4abb464b · outbound

This paper cites Ground- inggpt: Language enhanced multi-modal grounding model.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Ground- inggpt: Language enhanced multi-modal grounding model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.342064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.104592Z digest=sha256:8c9eae49a5b79732c9de2d6b1c2fa3ef1ce02c6278131cab4e99b33ba91eaea0

Observation 4a9e49fc-6dda-4417-9dd7-d78bae9f9621 · outbound

This paper cites an unresolved cited work.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:30:36.325636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.109564Z digest=sha256:4459620baed05b7f378285f65d7523193738d24d4bf29a8ddb5c5bdf161959b2

Observation fe9027ca-1e02-45b8-a7f3-e6866ac8260d · outbound

This paper cites GRES: gen- eralized referring expression segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination GRES: gen- eralized referring expression segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.309021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.114561Z digest=sha256:210daa2be98ab703fa73837f5f3e181fafdec261e22455e529b54580a37af87f

Observation c02f88b6-1f52-4080-8c35-8035e2d3403a · outbound

This paper cites Multi-modal mutual attention and iterative interaction for re- ferring image segmentation.TPAMI, 32:3054–3065, 2023.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Multi-modal mutual attention and iterative interaction for re- ferring image segmentation.TPAMI, 32:3054–3065, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.291837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.120016Z digest=sha256:0ec53811e4e5c808116c7fd4664964b292663ce24c8d42c6e600f0baa0aa30c2

Observation 80fef60b-bea0-4108-a359-da8f27db7200 · outbound

This paper cites Learning to assemble neural module tree networks for visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Learning to assemble neural module tree networks for visual grounding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.275374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.125087Z digest=sha256:a3d97c2828fa1cb766bc627fe7c854e1d6748ea820fdc3ea4a906df6cf465cd7

Observation 9c17dbfd-224e-4bff-b66e-b0619bbc25dd · outbound

This paper cites Visual instruction tuning.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Visual instruction tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:35.130130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:35.130130Z digest=sha256:e50af9960aa5ea7d475ae4b1bf31b1988bce6a12b41883713e3dba86d272663b

Observation ee286886-bd7e-4acf-b76b-c3831fb03afd · outbound

This paper cites Poly- former: Referring image segmentation as sequential polygon generation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Poly- former: Referring image segmentation as sequential polygon generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.247388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.135981Z digest=sha256:6c2af147ae7443466806a6124ec7f5fa44616b4e905e6afebfccd9df5cf84f48

Observation e1bedebc-0bf8-4d6b-ad2c-5e6f85baf1cf · outbound

This paper cites Dq-detr: Dual query detection transformer for phrase extraction and grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Dq-detr: Dual query detection transformer for phrase extraction and grounding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.231330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.140776Z digest=sha256:6fd92a96653d483575699401127a674bf13e9c946d3089391662260dc1b20de9

Observation e2299073-1046-416a-b21d-6952d82a1dce · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.arXiv, 2023.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Grounding dino: Marrying dino with grounded pre-training for open-set object detection.arXiv, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.214540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.145436Z digest=sha256:b73b3c91295a57451d75876736ea165117229cc41ee0f262181a03a3f2f60a29

Observation ae9c7165-eeaa-4975-b38e-3fa755808fea · outbound

This paper cites CARIS: context-aware re- ferring image segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination CARIS: context-aware re- ferring image segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.198737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.150568Z digest=sha256:fd8b6243757f302ad98f9b552e8b7c96d67cdd52ac01cd70964b5826e6e107bc

Observation 144e30ea-94b4-4726-aad6-eec1fe209284 · outbound

This paper cites Dara: Domain-and relation-aware adapters make parameter- efficient tuning for visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Dara: Domain-and relation-aware adapters make parameter- efficient tuning for visual grounding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.181964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.155698Z digest=sha256:e5083d07a8bdf4e4e8313c41d83b67f7cc8fb13c451ea08c517b895eaf2ee20b

Observation 7248157d-b4d2-4ef1-8578-983a0b618103 · outbound

This paper cites Mapper: Multimodal prior-guided param- eter efficient tuning for referring expression comprehension.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Mapper: Multimodal prior-guided param- eter efficient tuning for referring expression comprehension

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.166058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.160526Z digest=sha256:653b2ef9741c249fdcea87cf5654b8b0b3ce64dda8af63576d7cb75811719b3b

Observation 732cb52f-7317-476f-8bf4-d86d80738868 · outbound

This paper cites Improving referring expression grounding with cross-modal attention-guided erasing.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Improving referring expression grounding with cross-modal attention-guided erasing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.150362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.166059Z digest=sha256:0768d7de1058951763b8ad84e5789e500573735da301fc73248995647b452bdd

Observation a3644e9a-66ce-454b-a448-31476398d77a · outbound

This paper cites Improving referring expression grounding with cross-modal attention-guided erasing.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Improving referring expression grounding with cross-modal attention-guided erasing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.134863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.171179Z digest=sha256:7ab08f6c0a2b55dc89caf3ace793ee1006d0eac69527e02aef00a351be8f5c09

Observation 903331b6-e3fb-4acf-a3ba-8d39a8d25247 · outbound

This paper cites Multi-task collaborative network for joint referring expression comprehension and segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Multi-task collaborative network for joint referring expression comprehension and segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.118887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.176132Z digest=sha256:0e0727740f4d4cf47455716cbd8eb0bfd0731b57259c2cb229dcad3f84c7340c

Observation 666517f0-1600-4300-a6e3-b6a8d37787b2 · outbound

This paper cites Hdc: Hierarchical semantic decoding with counting assistance for generalized referring expression segmentation.arXiv, 2024.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Hdc: Hierarchical semantic decoding with counting assistance for generalized referring expression segmentation.arXiv, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.102962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.181176Z digest=sha256:2d38c3959b0abc4ef4464ed5955ac74c703f048b56a3764ceaca197ca05580fc

Observation fc41aa72-5228-4b98-980c-68e0ece38947 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Generation and comprehension of unambiguous object descriptions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.085925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.186068Z digest=sha256:786199d8224f5f54506d4b1692e2938227b244031c3417c50f1bafad36e637fb

Observation fdfd0865-ecc1-4f62-af00-5e3746739e69 · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Mod- eling context between objects for referring expression under- standing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.069994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.192007Z digest=sha256:308ed0598b25a56ac1ae2467fda53efa577a589b925ac1ecbe48f851f9c021e7

Observation 37b95990-9226-4b39-a79d-8d70e842fedf · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.arXiv, 2023.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Kosmos-2: Grounding multimodal large language models to the world.arXiv, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.052942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.196910Z digest=sha256:4de0c57803a31e5ee0b66fb2f49bdcb49c36e90756ec28ee0081224301808e3e

Observation da0cc831-8f2a-42be-9e9d-86939ac3d49d · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.036560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.201601Z digest=sha256:2be3df701106b6dc00f481282a2565884579f63957c158c78becb332d3c0d763

Observation 8fdda0d5-87e4-4ebd-a963-2387eab9fc62 · outbound

This paper cites Yolov3: An incremental improvement.arXiv, 2018.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Yolov3: An incremental improvement.arXiv, 2018

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.019077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.206218Z digest=sha256:2bb3b49e452985e64e4c6a517f58af5c6ee7fd4fe1cc6954a7f421b6bc77e326

Observation f3a9473f-6cb4-445d-b924-382f092628c4 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.TPAMI, 39(6):1137–1149, 2016.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Faster r-cnn: Towards real-time object detection with region proposal networks.TPAMI, 39(6):1137–1149, 2016

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:36.002864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.211792Z digest=sha256:a18f304d7dde6a13e5fc00719feeb0aa0f494e5e1fd5f10dac5b9bec046d5c75

Observation 7550a75c-f2d8-4416-9123-ead8de55bd27 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination U-net: Convolutional networks for biomedical image segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.985815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.216680Z digest=sha256:e05c33fa01e3364320fc7f63b4caa456862da9757bac78faa089305cfd4f9d22

Observation c6350e41-27ab-4737-8f5c-33dab9fefd91 · outbound

This paper cites Lqm- former: Language-aware query mask transformer for refer- ring image segmentation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Lqm- former: Language-aware query mask transformer for refer- ring image segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.967897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.221492Z digest=sha256:8f419c232ccb6f63395b9e46a9617c066e4270f15978740bec13c08d8fc1b64e

Observation 0cac53e5-3dc6-4e3d-bae4-d242df80734e · outbound

This paper cites Dynamic mdetr: A dynamic multimodal transformer decoder for visual grounding.TPAMI, 2023.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Dynamic mdetr: A dynamic multimodal transformer decoder for visual grounding.TPAMI, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.950112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.226421Z digest=sha256:9a1e16b5c1b0c6d4a4361a80d781de3e15272dac4e233f79208ea774f80795ed

Observation bca06384-1eec-4415-8f6d-89d4c8f4f1d5 · outbound

This paper cites Referring expression comprehension using language adaptive inference.arXiv, 2023.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Referring expression comprehension using language adaptive inference.arXiv, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.931223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.231395Z digest=sha256:9c4cbcf9fe47228d2039c01fc0ee3595453605e05ad02096872edbbc0a450439

Observation 6ce66470-8965-49ab-840f-7d3bd3b0c700 · outbound

This paper cites Language adaptive weight generation for multi-task visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Language adaptive weight generation for multi-task visual grounding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.912841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.235964Z digest=sha256:84d0981482913545e3cbf3758f73d09bc9a22de8aa32f67be5eb4df8f6081927

Observation 19ed48bb-2000-4160-9f9f-37ad610720a0 · outbound

This paper cites Scan- former: Referring expression comprehension by iteratively scanning.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Scan- former: Referring expression comprehension by iteratively scanning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.896261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.240828Z digest=sha256:a4f3e3634ee397094a143a31d8b56ffb70f15e8554e4c38f4a5a42f32eebad4f

Observation 6c8e062f-8fc0-456f-ba6a-f414cb5f7802 · outbound

This paper cites Llama: Open and efficient foundation language models.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Llama: Open and efficient foundation language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.879593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.245537Z digest=sha256:411a4b8b5f5f7c67e115b72a453c512160bef5d8b7f984f7a94aec6ce5cdc936

Observation 1cc41093-fc52-443e-b639-3b48ea7e9af3 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Llama 2: Open foundation and fine-tuned chat models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.862796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.250186Z digest=sha256:f33bec59b7df2568a8e4408776995c6eafad5f727ffdc260cb717ffbf58195ff

Observation 4830c549-6b98-4fbb-b057-c22005df6e8e · outbound

This paper cites Attention is all you need.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:35.255166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:35.255166Z digest=sha256:29cb4402489b5762d252ff7800c20276d32d4d1be16af8df941c0245145789b8

Observation d8da6a70-a676-453f-8945-258cac8e28ed · outbound

This paper cites Image as a foreign language: BEiT pretraining for vision and vision-language tasks.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Image as a foreign language: BEiT pretraining for vision and vision-language tasks

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.834626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.259911Z digest=sha256:b43c564a3234ca578eab37287344ac284dabb9523d1365591a9cdb2f80bdf752

Observation c9c963d0-7533-4e75-8660-1fb99bfce7db · outbound

This paper cites Towards robust referring image seg- mentation.TIP, 2024.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Towards robust referring image seg- mentation.TIP, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.817940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.264665Z digest=sha256:60acb143357a87427fdcd2b6ca6e104441090021e289ba51dcc266a79aebb3ff

Observation 0ecee307-80d2-40cc-95cf-6e46e31466bf · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Gsva: Generalized segmentation via multimodal large language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.801329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.269731Z digest=sha256:97b963a5f33b2110a4e2df031be5b1b2df7eaa234079774068b7e85506e0dde5

Observation 4a01011b-de54-4ebc-bc39-4e4c78c19010 · outbound

This paper cites Hivg: Hierarchical multimodal fine- grained modulation for visual grounding.ACMMM, 2024.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Hivg: Hierarchical multimodal fine- grained modulation for visual grounding.ACMMM, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.784340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.274556Z digest=sha256:320261abc7af10b0f844641e43a3ace5f7c8fb4a98db84bea4ba35f54864e921

Observation 44738b27-3913-4897-aada-5640ce390954 · outbound

This paper cites Oneref: Unified one-tower expression grounding and segmentation with mask referring modeling.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Oneref: Unified one-tower expression grounding and segmentation with mask referring modeling

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.768041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.279249Z digest=sha256:aabffcf40f7ce0734f5970104db1892f6689e808e6fa034eb87987d736d5f12e

Observation 836959df-9771-4384-a4bf-1ed401c0a9af · outbound

This paper cites Universal instance perception as object discovery and retrieval.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Universal instance perception as object discovery and retrieval

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T16:30:35.284451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:30:35.284451Z digest=sha256:a2fb1ee0a6c79540b8fc14bd0e97fdc1f60c27ff74cda2cc07b3f1662e690691

Observation 4e74217d-010e-451a-a28b-327d1e6c6241 · outbound

This paper cites Dynamic graph attention for referring expression comprehension.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Dynamic graph attention for referring expression comprehension

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.740637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.289174Z digest=sha256:17a0392d5323509ecfcfda6c0a05cdb78940820e34d1064cf141cf79296c4ee4

Observation 822e4ae3-7b4c-4ab9-af00-c3f06b6850fd · outbound

This paper cites A fast and accurate one- stage approach to visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination A fast and accurate one- stage approach to visual grounding

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.724371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.293880Z digest=sha256:e1ccdaf1bf8faa7bdef078dfcb0d01fcf86b546580de188a3e5f0464f5111a5e

Observation 48f7ddf1-77d5-4d11-a5f7-af02efef1280 · outbound

This paper cites Improving one-stage visual grounding by recursive sub- query construction.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Improving one-stage visual grounding by recursive sub- query construction

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.708922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.298759Z digest=sha256:d6c7f6c7d507c1d4cff2a046fff5b3e6a69858882f8b90ee657565cfaf9930af

Observation 1eae3787-7c6b-4640-a029-d8c6ddafd3e9 · outbound

This paper cites an unresolved cited work.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:30:35.692693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.303409Z digest=sha256:2603c7e1d576c1d52c21706cf5f2b628328aa95cc8070520ca0c4f91bb2c120d

Observation 1369d049-a50c-496a-9fd4-58214c734d32 · outbound

This paper cites Vi- sual grounding with multi-modal conditional adaptation.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Vi- sual grounding with multi-modal conditional adaptation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.676989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.308818Z digest=sha256:423233be21db458f5361a0aea94f6b5aa83bc0010b11b813f553b7101ac3b1ea

Observation 9efafb1b-51ab-4869-aaff-427580b3a811 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.659931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.313464Z digest=sha256:dcd4372d766a6d0e28563046407a1f1dcc97d344b7bb0d8085c7fa992050f142

Observation 444b17b7-9148-4eb2-ad9e-e584f7770cf2 · outbound

This paper cites Modeling context in referring expres- sions.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Modeling context in referring expres- sions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.643366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.318075Z digest=sha256:6aa22070a7d3d83992e911fa0eb0e98c7b3690e17ce67a7efbae7890e2974f0c

Observation 4e3d6e39-29d7-4a08-a3bb-298fb7fe551a · outbound

This paper cites Mattnet: Modular at- tention network for referring expression comprehension.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Mattnet: Modular at- tention network for referring expression comprehension

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.627068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.322938Z digest=sha256:a8a5b8b69b92bb2b19bbe40f83bff36dcb78ddd7887672d932f7066ccf7fd2dd

Observation 76690098-45a2-4bb0-b574-853c9768ab29 · outbound

This paper cites Rethinking diversified and discriminative proposal generation for visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Rethinking diversified and discriminative proposal generation for visual grounding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.610460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.327711Z digest=sha256:1ea2c1f390e42ce2c045ee2d0d6d44ea79d86ec254900c6dea30a5aa6618a8fd

Observation 91917c24-a7c3-4e1d-b42e-51240562bbe3 · outbound

This paper cites Grounding referring expressions in images by variational context.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Grounding referring expressions in images by variational context

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.592875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.332550Z digest=sha256:c700221ad5a588bf2d8998a1fbeeadf9a2b6d08fc7b8223dad610b98465f72cb

Observation 4b833189-8f1d-46d2-8710-e1503da078e0 · outbound

This paper cites A real-time global inference network for one-stage referring expression comprehension.TNNLS, 2021.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination A real-time global inference network for one-stage referring expression comprehension.TNNLS, 2021

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.574058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.337183Z digest=sha256:20561b622a95212b5375cecea04a913b89b5678ca59d7c72287ade285fd9fd9a

Observation fbb17a52-45fe-422f-8fe3-f112aac949b7 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Seqtr: A simple yet universal network for visual grounding

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.557219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.342036Z digest=sha256:5000ec040d5d06f28eb95462ca617f278451e534851e746705260075f4e312d5

Observation f4d939e1-67a4-4e38-91bb-269d62f33721 · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.arXiv, 2020.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Deformable detr: Deformable transformers for end-to-end object detection.arXiv, 2020

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.539668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.347098Z digest=sha256:a414c5d07bdcf7890ca3c3bd5583fe45e9b751489db94190bc078fa3ae5c693a

Observation 9e755d30-f0d0-43dc-ab66-0a18044358d4 · outbound

This paper cites Parallel attention: A unified framework for visual object discovery through dialogs and queries.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Parallel attention: A unified framework for visual object discovery through dialogs and queries

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.522471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.352020Z digest=sha256:1911c5de78e78ee3b188c1452b3dbf7cc488e51a251cbeebe9402ec8ac08e8d0

Observation c47b8905-6027-43b6-899e-3bc202ea8df5 · outbound

This paper cites Falip: Visual prompt as foveal attention boosts clip zero-shot performance.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination Falip: Visual prompt as foveal attention boosts clip zero-shot performance

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:30:35.505450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.356830Z digest=sha256:020344bf6af7331112ba42990f33e18536d66da5e71c4854bcbe2af4cc62c28b

Observation 61e0df0b-52e8-4f9d-9c0e-98efbd4a8ba7 · outbound

This paper cites St3: Accelerating multimodal large lan- guage model by spatial-temporal visual token trimming.

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination St3: Accelerating multimodal large lan- guage model by spatial-temporal visual token trimming

Reference 87

Resolution
verified exact
raw_fallback, observed 2026-08-15T16:30:35.469824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:30:35.361785Z digest=sha256:3f944ef3632b4be0e9e21538d0d2186ae553c0d5946cf2784984cef9a7a013a3

Pith citing papers

No inbound Pith citation observations are available.