Pith. sign in

Paper Citation Record · LEDGER

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models

As of 23 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2504.14032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14032 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:03:26.439718Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:44:21.060148Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T16:44:21.321080Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0ffb65c-58df-4d85-bbc0-3b1021f6e9d5 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Coco- stuff: Thing and stuff classes in context

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:29.445301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.121419Z digest=sha256:e9f17e4922b2d479dc7437c0300a7e391be255f6b01126bf748c7d47943dd31b

Observation 01f35811-148b-41cb-b54e-8ffb14b462e7 · outbound

This paper cites Learning continuous image representation with local implicit image function.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Learning continuous image representation with local implicit image function

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.151439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.151439Z digest=sha256:2c7887d8971a6a0f223bdd281b705ee3e50a113246f9b4ee79e04fed836c8377

Observation 4e9c0e46-3301-4ed2-857b-ecc69915721d · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:29.333029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.157488Z digest=sha256:9cc9c1008771ddf06b1deb8754e6b448cde82cc40710e0572b6e41b338bb328a

Observation e56aec69-b1e2-4620-94ae-0e2809f1937a · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models The cityscapes dataset for semantic urban scene understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.163320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.163320Z digest=sha256:6c0765e0d522c53514397452974adca659f29e6572d359c291141d794c5ff0dc

Observation 022968e6-3d6d-46b1-82ef-2922d1d0438a · outbound

This paper cites Learning affinity- aware upsampling for deep image matting.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Learning affinity- aware upsampling for deep image matting

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:29.144251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.168883Z digest=sha256:1876203d1059ad165494041bae77c17baffa4a8a388ee09517dfab6966fe1b2c

Observation ca019ac8-593c-4049-9663-0a2d7ce38b8d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:29.124751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.174619Z digest=sha256:258ca26bd6bb65add26425f7cb7bea12819d425eeb77f0623aba706e1482e131

Observation 505ff348-5e1b-4845-8cb2-46dabc56123a · outbound

This paper cites Lanczos filtering in one and two dimen- sions.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Lanczos filtering in one and two dimen- sions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:29.036007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.185225Z digest=sha256:a162bd026dcd5a76d616939bc3b9465a44954c1075702f3a8d893d4638bbc125

Observation 5e150a97-d669-448f-b0c6-319e472ba37b · outbound

This paper cites A guide to convolution arithmetic for deep learning.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models A guide to convolution arithmetic for deep learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.193535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.193535Z digest=sha256:e6f3aeac6e2cffcf6e2b18efcf56c4453f23d859e20c52fb0c0ad032a9f3e41d

Observation fa18085d-94d2-407b-a4c5-1aae4ea1dd1b · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Depth map prediction from a single image using a multi-scale deep net- work

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.944119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.201122Z digest=sha256:4d02a1149ef027cab5568742b177fe133c9abca7fcbd66fa731c6091cbe02f71

Observation db8a2fdb-888a-4d9f-8712-ba01dc5645e7 · outbound

This paper cites Prob- ing the 3d awareness of visual foundation models.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Prob- ing the 3d awareness of visual foundation models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.920938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.210384Z digest=sha256:6e5b5db6493ecd943883d1d7077c3073e1e156ea83d98d9c20bf4cbea305dd78

Observation bd70ed5a-d8d0-4be5-9f14-d16827fad094 · outbound

This paper cites Single image 3d without a single 3d image.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Single image 3d without a single 3d image

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.772778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.291307Z digest=sha256:b2892a69bf4e84625c5cbac3adee124d7a6ec068eaee85af907c475bd2b583a0

Observation 4a049e1b-2bdb-42f3-b660-b3014ab6a82e · outbound

This paper cites Brandt, Axel Feld- mann, Zhoutong Zhang, and William T.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Brandt, Axel Feld- mann, Zhoutong Zhang, and William T

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.623835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.362854Z digest=sha256:844d61632849eb02d865a4348f4f63d9733a63eb03a4d533441cb3833f276f88

Observation c35f94f4-6859-442a-a8bc-92cd0f0b0cbc · outbound

This paper cites Unsupervised Semantic Segmentation by Distilling Feature Correspondences.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Unsupervised Semantic Segmentation by Distilling Feature Correspondences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.435956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.435956Z digest=sha256:c22b46beadfbc1e9e82033c46ef5b81350282bac7b696aab7aac2812973bdfd9

Observation f7ed419e-94a4-446c-96c1-f176c7880ba1 · outbound

This paper cites Semantic contours from inverse detectors.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Semantic contours from inverse detectors

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.602582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.469592Z digest=sha256:8a135df63839ee08a1e01e660813ee0066924b7462e8c5eeb07694a626ebf547

Observation c6a9d133-f79a-4a4b-a8b6-46e09aa0bf5a · outbound

This paper cites Guided image fil- tering.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Guided image fil- tering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.581830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.475420Z digest=sha256:ff9b7abe83eb611bc06ca077fd77139f807c7a11398050726824892a89f94baf

Observation 6d06c215-6ed2-42f5-b217-da2748f066e8 · outbound

This paper cites Renovating names in open-vocabulary segmenta- tion benchmarks.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Renovating names in open-vocabulary segmenta- tion benchmarks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.562333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.480880Z digest=sha256:ab81802420f50fc81612d740187f6d3c3d2e95843fd9e4f2ed888c7d22e22b7e

Observation 0de4ac1e-2248-4778-a4c4-f17a2c4e68e2 · outbound

This paper cites Space-time correspondence as a contrastive random walk.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Space-time correspondence as a contrastive random walk

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.540774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.486429Z digest=sha256:0b978c25315fe75617a5a0df6c7f3272cd7c977e00ec6ac68cdb8cc5ae773ac0

Observation 6c6f07e1-f108-4a21-b5d8-50cd5674bb90 · outbound

This paper cites NA VI: Category- agnostic image collections with high-quality 3d shape and pose annotations.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models NA VI: Category- agnostic image collections with high-quality 3d shape and pose annotations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.514550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.498426Z digest=sha256:bee29c526d3bbebed6a9f19797ec7efbccff23dfe3157e4746905fe9643f53e1

Observation 8af58bbf-caba-4d98-b1d2-1ae5393671d7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Adam: A Method for Stochastic Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.505689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.505689Z digest=sha256:2932c904823efa087b0db211e5d1ea17ac7e067a3f74e2cf12c8ca4df7156d52

Observation e2bc5a16-7675-47c3-a076-7c3e2114eb26 · outbound

This paper cites Pointrend: Image segmentation as rendering.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Pointrend: Image segmentation as rendering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.494028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.514720Z digest=sha256:9ff1fcee7a9c85ad77a8fd9c1ab9f8b2720a4ea3b160d71e4b53fb0cf0d99776

Observation 2cd2bed3-ead7-4009-aac3-6bbc9c52687e · outbound

This paper cites Segment any- thing.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Segment any- thing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.329290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.521279Z digest=sha256:002f8595f396f42da87b107b0b3aaed0e88a102ece2b0fdcf5650473ab4c885c

Observation 40d9a4f9-1c4b-4564-a45a-5086f753bfe8 · outbound

This paper cites Joint bilateral upsampling.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Joint bilateral upsampling

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.178239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.527528Z digest=sha256:c056b70a18026f72cc6e6b10fc88586b66e5490fb69f2b43fa5ba34cb3db2a79

Observation 232f6dbb-738e-454b-80a8-9364eab6cd5f · outbound

This paper cites Proxyclip: Proxy at- tention improves clip for open-vocabulary segmentation.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Proxyclip: Proxy at- tention improves clip for open-vocabulary segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.071109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.534166Z digest=sha256:62e8552ec47518a60a431fd086e377773b2ce079499bdfc3a953c1b24e627698

Observation 196ec7ff-9ef8-4a2e-add6-0660b46dd435 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Exploring plain vision transformer backbones for object de- tection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.540758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.540758Z digest=sha256:6a2ed54c1343a013fc23a390a3f63113afef13575edc1c9cba5b22d4dcdd728b

Observation 71fd4f56-99d8-411f-b3c8-8a27c20f7527 · outbound

This paper cites Vision transformer for nerf-based view synthesis from a single input image.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Vision transformer for nerf-based view synthesis from a single input image

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.031813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.546193Z digest=sha256:3e5ab5aa1f6d8bd8dc53e03c1a350cf319ac985c8d3e9478b3e1dbbb9098c0d2

Observation a01da6c6-ab06-47af-abcf-75b8e72ea4c5 · outbound

This paper cites Microsoft coco: Common objects in context.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Microsoft coco: Common objects in context

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:28.007710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.552333Z digest=sha256:e6b0847eb2164b47029caa8a425028c20a16be4ff6a7dc66b7e4969a485516da

Observation 3267e676-8e67-45c9-98d6-243ca5b93306 · outbound

This paper cites Simpleclick: Interactive image segmentation with sim- ple vision transformers.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Simpleclick: Interactive image segmentation with sim- ple vision transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.986748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.559631Z digest=sha256:887b3a27897b4f19d5718f7e160fa1353b9f98710bf48e689f86aeeb7dae51b5

Observation d3345656-a2ff-4fde-9e68-c1b78363431a · outbound

This paper cites Decoupled Weight Decay Regularization.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.567142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.567142Z digest=sha256:4185d7a175822f0c5c242c933b71027de41084c9a5365c02a891e66103c5c8a1

Observation 99c14532-b431-4d91-bd21-787b37665704 · outbound

This paper cites Index networks.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Index networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.897940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.574143Z digest=sha256:3720b43bbdbb47d0b3d12ea204192d597d3a7667302e2dd756422f9ff9991a33

Observation f8cb0c0e-4ef1-472e-8e5e-ed860884cc3b · outbound

This paper cites Fade: Fusing the assets of decoder and encoder for task-agnostic upsampling.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Fade: Fusing the assets of decoder and encoder for task-agnostic upsampling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.827640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.579280Z digest=sha256:d4079e699fbf6c1a068237bee9d78718e74601866a35fccaf60f1173c4ede006

Observation 263aa2cb-f4c1-46c8-b9b5-89d27f3e8e83 · outbound

This paper cites Sapa: Similarity-aware point affiliation for feature upsampling.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Sapa: Similarity-aware point affiliation for feature upsampling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.803507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.643814Z digest=sha256:1a0adc4f181a2aaeaa1aeedc1549071f57dd0554c5eb0afc53688c2cbc576bd4

Observation 354d878f-b330-4d04-8869-2a286ed67b26 · outbound

This paper cites A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.774355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.707674Z digest=sha256:a5f374cc9aaa2f8cdfb7085c6476bf776bbbfad2168c04118f6b5457b49d6e6a

Observation 30bd273d-5e16-416b-8ce6-3edf10bb59b9 · outbound

This paper cites Cubic spline interpola- tion.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Cubic spline interpola- tion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.742803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.827262Z digest=sha256:0b48027b7a71d2d4e2f610e92437817a3ffee11e13d34ceafcee1eafa17805c9

Observation c27c24ed-ea4d-4ca1-9a85-d6e405ece2e6 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- thesis.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Nerf: Representing scenes as neural radiance fields for view syn- thesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.719399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.884256Z digest=sha256:06463eda31500d2d202280668e1738806d1a0bbd12a05ddc3cea799d46afe5c3

Observation 55bd3c3b-4c8c-49ef-9e15-d6d694543b9a · outbound

This paper cites Learning deconvolution network for semantic segmentation.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Learning deconvolution network for semantic segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.666040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.891132Z digest=sha256:02af6a2c43bd725b860777a9b8819d73f96dd9ce52940cd6cbebbbe60d8ef2d6

Observation 2b03ef9e-9b34-47aa-b5ea-01314e784260 · outbound

This paper cites De- convolution and checkerboard artifacts.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models De- convolution and checkerboard artifacts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.543076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.898679Z digest=sha256:eb9b14a8313a11ef5fdebf9b98b15d3dc6142b6c391cebc73a3c348c39cfc267

Observation d727c67b-eb1c-4c69-af69-946e810b4e50 · outbound

This paper cites an unresolved cited work.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:03:27.510717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.905950Z digest=sha256:af9a2f896155e3ade245f0a1a52a5d4995e6d0a1ce19deef32bcee4b532e6b6d

Observation 13e9a62c-9180-43e9-9f91-3b27ff8cc12e · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models A benchmark dataset and evaluation methodology for video object segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.485097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.912617Z digest=sha256:94eb5ebcb3f3865a914a70b940756236970333cf3a2caa72f634dd0ceebeea9b

Observation 02277b65-c1f7-48de-b693-a3e2db9baac9 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models The 2017 DAVIS Challenge on Video Object Segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.918392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.918392Z digest=sha256:231cff69957b19d41858ab676c01a95dd416386f96d6848ecdb6d5ef18a6a3ff

Observation 9ada388d-e73a-436a-875d-ba7be166a87c · outbound

This paper cites Three pillars improving vision foundation model distillation for lidar.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Three pillars improving vision foundation model distillation for lidar

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.924585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.924585Z digest=sha256:510a6218d6d0b56e3ed4b66ce2960b71d0937049c9ee24c810aba01e33d0413f

Observation 9d262d93-7d02-412a-a2f7-2a67fe7921ae · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Learn- ing transferable visual models from natural language super- vision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.442923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.930490Z digest=sha256:fcb44825e30be72cee25f90ec5cf6425193504c3d2f2297824e2f50e49deec02

Observation 725bb191-03e2-42a4-adf0-3019693d319e · outbound

This paper cites Vi- sion transformers for dense prediction.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Vi- sion transformers for dense prediction

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.418808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.935289Z digest=sha256:e8a6e97d09085456e3b607a3b4968fa40cfbee931c14dbbd894faf179897f689

Observation ac334253-fc50-46a8-b987-182ee83d0360 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.395182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.941620Z digest=sha256:1adbb7d90364191a4541704812214de1b00e3ec161ba17b166680dac37853009

Observation 48d976be-5493-42cc-b6c4-1f3500cc6201 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Glamm: Pixel grounding large multimodal model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.368221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.948300Z digest=sha256:e3d94bfcb8ebdfa7d49a417fb35f03732511c66e38c18b6f038901c4712e0414

Observation 961460bf-1f9f-430f-baae-4abbf2619409 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models SAM 2: Segment Anything in Images and Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.953821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.953821Z digest=sha256:35b6b7944a5183fd1ad94a12c720df71c4c5bb463eae2e65392a8fc1d8e32e26

Observation 37be4766-e079-4087-b36f-db1d112c4781 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.959566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.959566Z digest=sha256:df932f8336a809c2d684f0f1e5dee0920006db89a32a58d6d12fea996494d48d

Observation 8dbe0da4-e5fc-47b1-9f8d-6040100805a8 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models U- net: Convolutional networks for biomedical image segmen- tation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.970639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.970639Z digest=sha256:3f5fc48868405ba1ad7041ce79071162f95f005412255815682c0c05d7998116

Observation 1da3f8cb-c7f6-41ef-b19a-68b61616c32c · outbound

This paper cites ” grabcut” interactive foreground extraction using iterated graph cuts.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models ” grabcut” interactive foreground extraction using iterated graph cuts

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.316736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.975569Z digest=sha256:0c61a7a33beaeb2a6df3429106ea84af108a6a503a92d9150e073cae35b82a8b

Observation 4b2eb3bf-ded7-4ce3-b308-8b776fcdd063 · outbound

This paper cites Is the deconvolution layer the same as a convolutional layer?.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Is the deconvolution layer the same as a convolutional layer?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.981802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.981802Z digest=sha256:c7727b719fecb20f95669780b342d37d35321f2b051c3abebb12d300ff1dbf1a

Observation 7e4b9436-0eaa-46e6-b6ed-8f48ef5a589e · outbound

This paper cites Adaptis: Adaptive instance selection network.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Adaptis: Adaptive instance selection network

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.296392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.987642Z digest=sha256:6ef6b418d2f8a081f74d5067fcce822b3e64cfdcf135d120a70c4699682a666d

Observation 86083929-f0fa-4ffa-8891-0d78ba74b985 · outbound

This paper cites Re- viving iterative training with mask guidance for interactive segmentation.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Re- viving iterative training with mask guidance for interactive segmentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.276563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:25.992611Z digest=sha256:dd451b4635576832936dcb252e2f6cc503d10be57be0c063fdfc606d05319908

Observation cdc25dd5-e865-4a1b-b8e6-d003968ebab9 · outbound

This paper cites Lift: A surprisingly simple lightweight feature transform for dense vit descriptors.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Lift: A surprisingly simple lightweight feature transform for dense vit descriptors

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:25.999077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:25.999077Z digest=sha256:aaaa55b770fb7adfe938f3c93aaa8eb3c9dffd236e1c1270d7939d969979366a

Observation 13d99416-c639-4449-8a2c-a7b10b47ae51 · outbound

This paper cites Splatter image: Ultra-fast single-view 3d recon- struction.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Splatter image: Ultra-fast single-view 3d recon- struction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:26.010065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:26.010065Z digest=sha256:081d1db97bfecc019893f4ea20e1dc9ff818983a3bbaf24ffa72563f205b737a

Observation 4c3cc74e-5ec7-4408-93ed-93ec6a2af364 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:26.077661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:26.077661Z digest=sha256:333858445a3a7db9db4f0b4b139a673d5adf6429fe96d9af340fa7e6a8730f63

Observation 295f7f4a-ef11-49ae-bc6c-3ea906ce0f26 · outbound

This paper cites Dino-tracker: Taming dino for self-supervised point tracking in a single video.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Dino-tracker: Taming dino for self-supervised point tracking in a single video

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:26.972172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.223795Z digest=sha256:1b2217828692f1242234c57717ffc20ce5a3ac504fc679e87d4d4af5ee85b393

Observation 298b4c20-b600-4c03-bf9d-81cabb0287c3 · outbound

This paper cites Carafe: Content-aware reassembly of fea- tures.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Carafe: Content-aware reassembly of fea- tures

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:26.736781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.348973Z digest=sha256:24490f3562f92a2d812b5c2c899643a81ec255eabcb5976206d62f5d6dee77bc

Observation 05ad44e8-54c2-43d7-a1c5-6ca5e87281bf · outbound

This paper cites Self-supervised trans- formers for unsupervised object discovery using normalized cut.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Self-supervised trans- formers for unsupervised object discovery using normalized cut

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:26.719209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.407722Z digest=sha256:55d5e68f12e55c38084af50e1f325be5b51124e3fca7555e7dc3640a699ae2f4

Observation 95033c19-ddf3-4674-b23a-ac402da74efd · outbound

This paper cites Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:26.414292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:26.414292Z digest=sha256:7ebaabd37652735a998a3fbe366b1b984afd7a32e2a9d469d1f57d18123c560d

Observation 7812cdc5-345c-4907-8ec8-5606f99aa4d7 · outbound

This paper cites Featuren- erf: Learning generalizable nerfs by distilling foundation models.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Featuren- erf: Learning generalizable nerfs by distilling foundation models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:26.689662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.420058Z digest=sha256:f5c8c7c13f4660cf7209edc01ab9d0ab5c991e44fac74c50a1f2a4beaee4acd9

Observation 158102a0-61ae-42f9-b3db-6794dd99c9e0 · outbound

This paper cites pixelnerf: Neural radiance fields from one or few images.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models pixelnerf: Neural radiance fields from one or few images

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:26.670225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.424921Z digest=sha256:d0d15ff9ba999ef235c646ccf6240cb2ae73c72dcaa37781738f30f17894e227

Observation 1261df93-e68b-472f-8fc0-b38887bbceff · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:26.429741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:26.429741Z digest=sha256:8a198e8fc96ce14ccbeb7f03bf3fa8ad29558b1d528adfc84f550309e55accad

Observation 8041424f-202d-4b2b-94e3-88a92b90848c · outbound

This paper cites Sigmoid loss for language image pre-training.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Sigmoid loss for language image pre-training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:26.641626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.434700Z digest=sha256:60e91fa4121e4b4f3889dcb47c9bc54100787827bbabf5573049c3d742e1f79f

Observation 0d108903-f6f4-4f17-baff-916114e47eaf · outbound

This paper cites LoftUp: earning a Coordinate-Based Feature Upsampler for Vi- sion Foundation Models.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models LoftUp: earning a Coordinate-Based Feature Upsampler for Vi- sion Foundation Models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:26.622513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.439718Z digest=sha256:3139cd5ece317f8f3ba4b0fd8b38bc5133d8e7817812a3e27c97c9b68fb9aee1

Observation e9bbc884-d312-4ede-b569-6e8492229de2 · outbound

This paper cites 1, 2, 3, 4, 6, 7, 12.

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models 1, 2, 3, 4, 6, 7, 12

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:03:27.163802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:03:26.004379Z digest=sha256:5c16904ff7e0a545b7452d20e21b50ef23116ba7348dc4db150a36d010a79b0b

Pith citing papers

Observation 2a3c4294-09b6-4365-b009-02afd379bed9 · inbound

Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation cites this paper.

Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:44:21.324779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:44:21.060148Z digest=sha256:19780a7216495fca795396c098334def7fa848905187fb7c262b4a0098c661ac

Observation 7fad0f2a-d470-45f4-aa9d-30d5c9227b70 · inbound

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders cites this paper.

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:41.230471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:41.230471Z digest=sha256:2258fa6cf8192b5a5506a3a996ccb92a717cb450972ae086940f8dc3d860a92e