Pith. sign in

Paper Citation Record · LEDGER

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues

As of 3 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2604.24036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24036 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:38:06.673737Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:38:06.673737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T21:41:15.213547Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact9
  • verified fuzzy30
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0160b1f-a493-4f02-a325-c6c6d50c4ffa · outbound

This paper cites Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:15.218723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:68e6ec342646bfb843df1c8c3fc941657ce977efc5de57bec7bd368422d00d47

Observation 51a0d5df-3597-41c9-9e7a-55d007143218 · outbound

This paper cites vehicle” “motorcycle.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues vehicle” “motorcycle

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.446195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:f9ecce3678b0fc0e5ff0c55ec8588efdb770ba7a969242c82ada282859b52087

Observation 313f9921-d4a7-4613-8f09-c53351899905 · outbound

This paper cites The pedestrians carrying umbrellas.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues The pedestrians carrying umbrellas

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.451353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:e8307ea55d494eef662d230e17edab88bb8bce74e123535f9e43ae277074b50a

Observation eb4f5b00-c628-4f00-976a-1f8831e0e61a · outbound

This paper cites These modules derive and integrate lingusitic semantic pri- ors, enhancing grounding robustness against occlusion and small objects in crowded scenes.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues These modules derive and integrate lingusitic semantic pri- ors, enhancing grounding robustness against occlusion and small objects in crowded scenes

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.459086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:0be3637684125a513cbbea72588a9d7da0c0e28d3df239deb9c5b1762cc55aac

Observation 69f883f4-cc5a-4f24-9031-9f3686e96ce6 · outbound

This paper cites Additionally, the supercomputing resource was partially supported by KSC (KSC-2025-CRE-0090).

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Additionally, the supercomputing resource was partially supported by KSC (KSC-2025-CRE-0090)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.463736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:5f070e8e98c361ea10833d8e215f7d0f9258388302b40bcc056b731bc191ab9a

Observation 5607226b-0c6e-4fb5-985c-eb6e444d81d7 · outbound

This paper cites an unresolved cited work.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:43:02.477340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:a4ed50804bf3262f855c2cb5df9a89df87ad7a583c163b6623fc9546e0b850be

Observation 9b7c3978-2d3a-44ea-8062-ed68e294e9b3 · outbound

This paper cites Qwen2.5-VL Technical Report.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Qwen2.5-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:15.245554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:c48279be599804942f17c60b186d1de47ccd4c5f88fdb1658f81e5ddf05061b4

Observation 2af1c8ad-4229-40d8-ba8e-fbd58212eb24 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:15.229293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:9825394b7006390031441907b64415e56054645be88e2cf099b2269e78392ef4

Observation 1594cc31-2fa4-4b68-ab46-ff4fe70cb395 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Visionllm: Large language model is also an open-ended decoder for vision-centric tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.536078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:d6937f373dc755777d31201bacbffb85528305b00e82fd4e400509fd4a8ec786

Observation 075897bc-3e6b-45bc-a436-4dc6bebf38a9 · outbound

This paper cites Kosmos-2: Grounding multimodal large lan- guage models to the world.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Kosmos-2: Grounding multimodal large lan- guage models to the world

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.489266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:842c355d320917bb3b5ce1384495287a9f68ed9989011e4720ad47d1c62e0c4d

Observation e05854a1-b585-478d-8d30-6cd280b73dc1 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Ferret: Refer and ground anything anywhere at any granularity

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.572753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:5f0e0cd5c18916e340b7ae5af871ecde585520e141689de48aec2743d8ded5bd

Observation b2911d23-037c-4e7a-ac5e-65966a6b8be9 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Groma: Localized visual tokenization for grounding multimodal large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.560532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:f1177d8173d136547a684caf504bb88aee16ff32daf0122a542c56147d6b4b90

Observation c0b42aca-156c-4d4e-a81d-bc1744630647 · outbound

This paper cites ChatRex: Taming Multimodal LLM for Joint Perception and Understanding.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.238967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:6d705115997a45285c29569e67a3de3d77d79b8cccd97a73f44553b7d54257db

Observation deb17694-42ca-4949-b855-cc9c97cd2f28 · outbound

This paper cites Modeling context in referring expressions.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Modeling context in referring expressions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.564282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:8fa6a710254c0275128ba099d876b270732c3ce735bd8a3b5f284aae7cd4b833

Observation 97cd706b-5a73-49eb-b496-2d4d3ea7d2c9 · outbound

This paper cites Generation and comprehension of unambigu- ous object descriptions.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Generation and comprehension of unambigu- ous object descriptions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.548023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:919ee31aa2d91dbca702dbe962c6e0d8a90d1cd8547b1b380d3825141f84179e

Observation 424a1474-5d98-4aab-9ebe-021720792dc2 · outbound

This paper cites CrowdHuman: A Benchmark for Detecting Human in a Crowd.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues CrowdHuman: A Benchmark for Detecting Human in a Crowd

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.263199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:891f744187eaf6074193483b2ee4175c5d464a81f4bf283fc72c92da3bced61d

Observation 6371b8be-e3dd-4598-819a-d2e945f019dc · outbound

This paper cites Visdrone-det2019: The vision meets drone object detection in image challenge results.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Visdrone-det2019: The vision meets drone object detection in image challenge results

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.555748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:609024002f285b0208441b7d8ebf84b72904df6dffe1e53c803f45f3bc55ed73

Observation 7e47a030-e1bc-4f9a-91ef-7531033e1482 · outbound

This paper cites The unmanned aerial vehicle benchmark: Object detection and tracking.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues The unmanned aerial vehicle benchmark: Object detection and tracking

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.552013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:54495f805d21088209bb1cb516ad6f82bc646acc1939594f14de8c271b6dc87d

Observation 16bb2646-2a81-4f26-b1fb-5ea59daa723f · outbound

This paper cites Refdrone: A challenging benchmark for referring expression comprehension in drone scenes.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Refdrone: A challenging benchmark for referring expression comprehension in drone scenes

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.255524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:a5c35a461b6d594d263ecf0bb57a39c553a3df33f788f368caa708da9834cf43

Observation 5f92e2cd-f9bf-4b65-92bb-076b74dbe8db · outbound

This paper cites Language can boost otherwise unseen ob- jects into visual awareness.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Language can boost otherwise unseen ob- jects into visual awareness

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.568845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:aeb3f079df9b40562d593573b98a584abdf57bb0345f08da0974a96b425cebd3

Observation 24d41887-ea99-42ba-a59d-dd05ed048dce · outbound

This paper cites Words jump-start vision: A label advan- tage in object recognition.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Words jump-start vision: A label advan- tage in object recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.544060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:49ee751e7688d7fb0cc879cef29f5709a9eadf686308a4417c963baf1200f624

Observation b74d33f4-b2de-44d5-97d3-2993e3db0e06 · outbound

This paper cites Weather-aware drone-view object detection via environmental context understanding.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Weather-aware drone-view object detection via environmental context understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.576533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:2db0661f189bd569e86d94d767a4b88fa0ba7c5a8b6ca7bf2951f427aa1432f0

Observation f0dcedd0-f893-41c4-8377-4992dee0d7f9 · outbound

This paper cites Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.271640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:57431e1cef45e7a4ac085c943f6448d7262bc8a46c29206e99b21ad0a21370c9

Observation b9fed8b5-a027-4db1-8d40-e578a2e802e5 · outbound

This paper cites Moai: Mixture of all intelligence for large language and vision models.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Moai: Mixture of all intelligence for large language and vision models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.524479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:e32b5ccc2ee99acac654e6f6c3cbcba09a1517964aced3148f2250de399cb9bb

Observation 5311179a-7c25-436c-aee7-3aca1bc96c42 · outbound

This paper cites Citypersons: A diverse dataset for pedestrian detection.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Citypersons: A diverse dataset for pedestrian detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.509617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:30e058649a74a3a85648fa61f3e7a5b3404162ecdc3110eb3233ea39420fab34

Observation f93928b3-f4c2-4bda-8db9-1aa041d2afa3 · outbound

This paper cites Widerperson: A diverse dataset for dense pedestrian detection in the wild.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Widerperson: A diverse dataset for dense pedestrian detection in the wild

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.441073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:4e2ae1599593b200a918a7f0fd0ead801f77155a5746680c06e63c05b485b469

Observation 24dec3c1-7d14-433d-a09a-39ed426da67e · outbound

This paper cites HazyDet: Open-Source Benchmark for Drone-View Object Detection with Depth-Cues in Hazy Scenes.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues HazyDet: Open-Source Benchmark for Drone-View Object Detection with Depth-Cues in Hazy Scenes

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.210278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:0034a7e36d49551425872ae249766198125da067647d30f8512bf410f1b2f4b9

Observation 0648ae53-a755-4d55-a2e0-bdcdd6f58c4b · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.481858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:55f26b1ae71dd291069ecea7a74a4b7ef82f2b4bac4cfdb0efee743437f05d7c

Observation 80c2269c-3dd6-4558-b319-dce2eee5d0fd · outbound

This paper cites A convnet for the 2020s.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues A convnet for the 2020s

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.500008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:f5566318bc81cac53beb89d3d8596db524d2438b8dfb2ec6e9c63a28a67ff565

Observation e40e722b-a530-41d3-ab55-797e51f27aea · outbound

This paper cites Vicuna: An open-source chatbot impress- ing gpt-4 with 90%* chatgpt quality.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Vicuna: An open-source chatbot impress- ing gpt-4 with 90%* chatgpt quality

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.505626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:c987810e83490207bb79e44a8025885bad5a49f78f579ec6c42ba472de6fddfb

Observation 53bd0878-64a5-4ce5-92ce-cd644cd29651 · outbound

This paper cites T-rex2: Towards generic object detection via text-visual prompt synergy.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues T-rex2: Towards generic object detection via text-visual prompt synergy

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.515368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:406a7a842e0ce7b4d538290d12f33faf8f769c427393d930568122c67b221cc8

Observation ad6756e3-f0fe-46da-848d-3f0252bd96d0 · outbound

This paper cites Detrs beat yolos on real-time object detection.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Detrs beat yolos on real-time object detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.519975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:c3168b2eace6bd04997953b24f36b9f2bd1cc98e74d358a2735091eb146d1327

Observation 6fd53837-d088-40b8-8399-cc30fc18e4f4 · outbound

This paper cites Mask r-cnn.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Mask r-cnn

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.531625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:56070103167f9fe44f65f9b2bbe3293413d414f0307fa1a08a1657c6d45bb20f

Observation a266671a-b802-4be4-82ed-250b452d2cde · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Gaussian Error Linear Units (GELUs)

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:15.279343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:9873f298eeb950972b350d2b261c0af598a6837a301fca264e3122bb0a9af829

Observation 27ad2295-0e18-44ee-a99b-70fb74035341 · outbound

This paper cites Clip-convnextlarge d 320.laion2b-s29b-b131k- ft-soup.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Clip-convnextlarge d 320.laion2b-s29b-b131k- ft-soup

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.485502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:495b0ec38556e2622e688abdca1c086d6607c72d2fd9f249886d8deb9c19da31

Observation c4c26015-5085-4464-95b8-8d89801e34dd · outbound

This paper cites Attention is all you need.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Attention is all you need

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.493347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:20738c99e1fdc71439d79b877f2be0a1c04d53d3481847313fbacdf79dff091e

Observation 6af14882-c6dc-47a1-95b6-13f418da74bb · outbound

This paper cites Microsoft coco: Common objects in context.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Microsoft coco: Common objects in context

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.539823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:b78fd11ab5362991e8aa3883568cd8f41303f5976cfadabf8853696952bb9eab

Observation 017d1ecf-4746-40f5-b112-dc4639dcbb57 · outbound

This paper cites Decoupled weight decay regularization.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Decoupled weight decay regularization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.467564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:af55930ef7dc2aa88c416dc7f8f0fab751340a706e10743cd1ba04f4d51b02b4

Observation 0119df1a-7a5d-42b1-8612-20e74770b16b · outbound

This paper cites Meteor: An automatic metric for mt evalu- ation with improved correlation with human judgments.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Meteor: An automatic metric for mt evalu- ation with improved correlation with human judgments

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.471494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:6f2e002faee68cb4442097f9e17dc86f1c60866fd1329df462b870dd02c088e1

Observation 84533ec3-5232-4497-a586-e40ab5bb527d · outbound

This paper cites Cider: Consensus-based image descrip- tion evaluation.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Cider: Consensus-based image descrip- tion evaluation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:43:02.436924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:d17829e8792ab9a004ba7e3cdf3e77b4ab3e232c9ab8cd27d24da047a229153c

Pith citing papers

Observation a0160b1f-a493-4f02-a325-c6c6d50c4ffa · inbound

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues cites this paper.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:15.218723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:68e6ec342646bfb843df1c8c3fc941657ce977efc5de57bec7bd368422d00d47