Pith. sign in

Paper Citation Record · LEDGER

Video Depth without Video Models

As of 14 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 1 inbound Pith citation observation for arXiv:2411.19189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19189 v2

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:31:01.975283Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:15:20.254402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:24:02.091990Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact1
  • verified fuzzy45
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 507e7022-fdaf-4084-88b5-f9c5722ccc6e · outbound

This paper cites Bidirectional attention network for monocular depth estimation.

Video Depth without Video Models Bidirectional attention network for monocular depth estimation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.621125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.621125Z digest=sha256:ab26a537d7aefd09dcaadba1b8deb2f974fcf43af50f12de1fcac3eba5b26fdf

Observation 204de18b-90a8-4a39-8cea-91ab1ba0cd9b · outbound

This paper cites AdaBins: Depth estimation using adaptive bins.

Video Depth without Video Models AdaBins: Depth estimation using adaptive bins

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.625444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.625444Z digest=sha256:8e8fc526139957d772f2cc04f48cee5bda0938f85f538c5e8f3d8702c0c7d9a6

Observation 1ab953cd-a1aa-4f3c-8a2c-1a1caa5a53f7 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Video Depth without Video Models ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.629852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.629852Z digest=sha256:c6a527d55bd21313db7f55415416a24bc5a00b6708a20fcff2b0119f8fc5fba6

Observation 75bbbc40-938c-48f2-979b-ed448ffca992 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Video Depth without Video Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.634186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.634186Z digest=sha256:85606eb63cca24a848b8ae67ff8a98ab04056cba376c1870b999c8f14ee29d33

Observation b9a9e055-e7eb-4c60-b1b9-fcb1c021889c · outbound

This paper cites Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

Video Depth without Video Models Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.639244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.639244Z digest=sha256:1ba54bbd560f9980a92d432382cf74ac2bbb327164f7ae5bc990249963445400

Observation 1e554693-4b2d-4ed4-bd83-59bd4c5d1ebb · outbound

This paper cites Pix2Video: Video editing using image diffusion.

Video Depth without Video Models Pix2Video: Video editing using image diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.643511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.643511Z digest=sha256:7b19b862061bb4eaa38e82dc8f552d995bfc320b3cd5faac5fed73683629acd9

Observation 11d57d43-b082-48d9-8075-ab932d32cc00 · outbound

This paper cites Self-supervised learning with geometric constraints in monocular video: Connecting flow, depth, and camera.

Video Depth without Video Models Self-supervised learning with geometric constraints in monocular video: Connecting flow, depth, and camera

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.647548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.647548Z digest=sha256:12d4d86303e32cb95b711cf2c71f93cc1fb084aa3769560fa3f7f653a09fa94b

Observation 6829c104-18d1-4dac-87aa-366847d3ad27 · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Niessner.

Video Depth without Video Models Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Niessner

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.651830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.651830Z digest=sha256:d4df7a1d10be42e2eeb51bbc611aee0af72ad4750a03c00c852fee278db7e4af

Observation f7b13c93-1053-488e-986c-b8b61c30a104 · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

Video Depth without Video Models Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.656010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.656010Z digest=sha256:a709fb5fdc43bf1396d76014785f295781fe6b51c0afb86a64a761b884e0157d

Observation 0fdcab84-ca89-4c24-8f79-b2d7267b7803 · outbound

This paper cites Warped diffusion: Solving video inverse problems with image diffusion models.

Video Depth without Video Models Warped diffusion: Solving video inverse problems with image diffusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.659951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.659951Z digest=sha256:036fa335a9b974153b7d2df1e3e5902cd0fca6451ce3e8b0e15e7841debbf7a1

Observation b82abd2c-4468-4e0a-8398-10b027644a9b · outbound

This paper cites DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation.

Video Depth without Video Models DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.664345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.664345Z digest=sha256:eb394c208f7e1e568f6ae067bbcc536cc613873659b9d91549a68021ce980b9a

Observation cfe8e210-1d49-439a-a892-b2b25ce751c7 · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

Video Depth without Video Models Depth map prediction from a single image using a multi-scale deep net- work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.668578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.668578Z digest=sha256:d5aa9e364a48067748596e4804e0173509ed8be682ef4ae543274ff2aa2cd600

Observation 60ad2e89-56ff-40b4-b4d5-10f499920978 · outbound

This paper cites Deep ordinal regression net- work for monocular depth estimation.

Video Depth without Video Models Deep ordinal regression net- work for monocular depth estimation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:03.013999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.672397Z digest=sha256:5b923c5751adad633b43e5ae7994d9b1780673643bdfa2d29bd9664d8dd9093f

Observation c62ae05e-c8e3-41e5-9dec-ab40d8093018 · outbound

This paper cites GeoWiz- ard: Unleashing the diffusion priors for 3d geometry estima- tion from a single image.

Video Depth without Video Models GeoWiz- ard: Unleashing the diffusion priors for 3d geometry estima- tion from a single image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:03.002314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.676214Z digest=sha256:f2ea3a36900a75d30186ecaab6668adad077708c93e5b920b9b4fcb10d3b3ed2

Observation 952a6f5c-9836-4b1f-af69-3721799602e4 · outbound

This paper cites Multi-view stereo: A tutorial.

Video Depth without Video Models Multi-view stereo: A tutorial

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.679964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.679964Z digest=sha256:fcd70d5ea8a3ec622f8e48e5523c3a2af855289bdc0e8c486d77953e1f93d0b2

Observation 40640050-410c-421b-8f50-23ac2eee121b · outbound

This paper cites Fine-tuning image-conditional diffusion models is easier than you think.

Video Depth without Video Models Fine-tuning image-conditional diffusion models is easier than you think

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.683906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.683906Z digest=sha256:9bd4ac685d776328475b1fc5ad76099ce5fd3971edeec3f65379c0f8cca2b0ab

Observation d6ee4879-9902-4fcf-8e99-c9ae7a550b04 · outbound

This paper cites AliceVision Meshroom: An open- source 3d reconstruction pipeline.

Video Depth without Video Models AliceVision Meshroom: An open- source 3d reconstruction pipeline

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.983143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.687574Z digest=sha256:37b884abe9680ed53421767b7e6f1b117403542426a637c06c89697ee7966d9d

Observation 22feaac2-643b-48ef-9212-fff47be37735 · outbound

This paper cites DepthFM: Fast Monocular Depth Estimation with Flow Matching.

Video Depth without Video Models DepthFM: Fast Monocular Depth Estimation with Flow Matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.691346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.691346Z digest=sha256:9c46d53f5690d70f18aec4e4bf64bbdc3927cd13e522629428733a20dcdb20b7

Observation 435fd2d8-7dcd-49a0-908e-1f7b79d138f9 · outbound

This paper cites 3d packing for self-supervised monocular depth estimation.

Video Depth without Video Models 3d packing for self-supervised monocular depth estimation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.971137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.695822Z digest=sha256:b030c346036f379b6c9d4907be233d6b2faca1f0fd00f84f36967cd4986af855

Observation ff25ae2c-843c-42cd-a2b6-92a68d290066 · outbound

This paper cites Towards zero-shot scale-aware monoc- ular depth estimation.

Video Depth without Video Models Towards zero-shot scale-aware monoc- ular depth estimation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.959016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.699642Z digest=sha256:ffd1b9b5f38c4696e06adcc7c431bfdfa3ef7045a513255a08cf2e32cf50b01b

Observation 862a587c-9c33-450c-b771-d1576e9dc811 · outbound

This paper cites Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction.

Video Depth without Video Models Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.703813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.703813Z digest=sha256:19ce421f2b69b4187521ea4e7c8c6e0e570253c12cce16b8a6e0bcc551b0bdfd

Observation 4f34e0e8-3cdf-49a8-9e8a-5117ac170b38 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Video Depth without Video Models Denoising diffu- sion probabilistic models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.707761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.707761Z digest=sha256:59f0ef4ab775d9b5e0e55e5494666d7120a84470fdd01442085711897f6849da

Observation 4c04a0b5-1a9d-448a-ac4d-eca2b0ac8199 · outbound

This paper cites Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation.

Video Depth without Video Models Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.711328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.711328Z digest=sha256:b737037a7148115cdf7798f6bf1c61854f748019a0ee2a4f48e108d76f40d815

Observation 34e551eb-4817-44e1-92bf-6c88593ff22c · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

Video Depth without Video Models DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.715293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.715293Z digest=sha256:45549cd47ed014e47da3acfbff9c6146765a6a0b80572ab8041bbff4645dc214

Observation e8561de6-6fe2-41b2-a210-212a9823cac0 · outbound

This paper cites Repurpos- ing diffusion-based image generators for monocular depth estimation.

Video Depth without Video Models Repurpos- ing diffusion-based image generators for monocular depth estimation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.939756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.719107Z digest=sha256:c82e6d6e7ddeb46c8608415da0f492cf1d7f59bc0620fcc4a7b32467e40feb89

Observation fd429453-75fd-4c59-825c-c08c4c19135e · outbound

This paper cites 3d Gaussian splatting for real-time radiance field rendering.

Video Depth without Video Models 3d Gaussian splatting for real-time radiance field rendering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.927041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.722828Z digest=sha256:f531873e6dbb72cfa3d41ff205f6d429f05f0da02442a3219d60d09c2baf6a7a

Observation a67928a0-0eed-4e55-a7ae-720c6e0b10ad · outbound

This paper cites Text2Video-Zero: Text- to-image diffusion models are zero-shot video generators.

Video Depth without Video Models Text2Video-Zero: Text- to-image diffusion models are zero-shot video generators

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.915224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.726671Z digest=sha256:cc33146839563a9e70364657bb7867cfedca936e1c54bf32f1bc811cc32af2bc

Observation e8225162-637c-416d-b9d3-864408c7f9a1 · outbound

This paper cites EscherNet: A generative model for scalable view synthesis.

Video Depth without Video Models EscherNet: A generative model for scalable view synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.901859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.730724Z digest=sha256:c3032a204ed2cd8ecc1678eb4cd9e39e98c3f1b3aba4263381b95d7ce201ffbb

Observation 9bbbb42c-cefa-4371-8225-c70b9d9647fa · outbound

This paper cites Robust consistent video depth estimation.

Video Depth without Video Models Robust consistent video depth estimation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.886932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.735055Z digest=sha256:7ba7ecd286ca7a610d20eb16cee41523ff813cfe3a11511c9c18066fd6660153

Observation f87ee3bb-adc8-4bb9-aa32-f7a57d5d3b0e · outbound

This paper cites Solving Video Inverse Problems Using Image Diffusion Models.

Video Depth without Video Models Solving Video Inverse Problems Using Image Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.738592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.738592Z digest=sha256:666c8268ed52ec8e9354708c341c609f84ed635659802823b3458014aab0a34e

Observation 380c524c-704b-4acb-ba37-ba04eea4b98f · outbound

This paper cites From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation.

Video Depth without Video Models From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.742634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.742634Z digest=sha256:24d99737615654dd20d704b80fc2572396420ce513f1506024ca145da6ab2d4e

Observation d8f62468-4713-4128-b130-07e087c01a4e · outbound

This paper cites MegaDepth: Learning single- view depth prediction from internet photos.

Video Depth without Video Models MegaDepth: Learning single- view depth prediction from internet photos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.872664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.746844Z digest=sha256:aa1ce1bf669b1b8704160dbbb8590a06b266b4da7293bc757e13b45d449491d1

Observation eabffd05-2720-4d0c-bf5c-9f3cb901bc34 · outbound

This paper cites BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation.

Video Depth without Video Models BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.750683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.750683Z digest=sha256:d44198a7493cd243b8a8612858db40cbcbfd8b04f8a9b7bd1984b80c6dd727db

Observation b7fa91fa-e083-4195-ac68-b31b3b90d185 · outbound

This paper cites DepthFormer: Exploiting long-range correlation and local information for accurate monocular depth estimation.

Video Depth without Video Models DepthFormer: Exploiting long-range correlation and local information for accurate monocular depth estimation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.858021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.755108Z digest=sha256:fa51e43ec2eb1f8346dc993ecd624660bf7608d0b1066e3b8faffe07cf325d72

Observation 481c72b1-cec9-468c-86de-c1b221da4619 · outbound

This paper cites Temporally consistent online depth estimation in dynamic scenes.

Video Depth without Video Models Temporally consistent online depth estimation in dynamic scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.845020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.759587Z digest=sha256:66d9750bc08520fd68415a5f51580b422d981360bfa36dc98a35738dccd9fd96

Observation f9ebe0e3-dfb4-44f5-b95d-fc5cc9f4b944 · outbound

This paper cites Patch- Fusion: An end-to-end tile-based framework for high- resolution monocular metric depth estimation.

Video Depth without Video Models Patch- Fusion: An end-to-end tile-based framework for high- resolution monocular metric depth estimation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.832504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.763665Z digest=sha256:ff3d59d00acbc4a0e1c9e6e94df5d666f468d9c26d3a2c9ca9d5f3217d7f8800

Observation d64d3662-0bb2-4520-af55-3763bba33551 · outbound

This paper cites PatchRe- finer: Leveraging synthetic data for real-domain high- resolution monocular metric depth estimation.

Video Depth without Video Models PatchRe- finer: Leveraging synthetic data for real-domain high- resolution monocular metric depth estimation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.820438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.767980Z digest=sha256:4e2dccf5c038223bcf8372657c23383f1d152783025f5c590462cb62657becd3

Observation ad86ee3f-4ea5-408c-a561-38c47e2393cd · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

Video Depth without Video Models Common diffusion noise schedules and sample steps are flawed

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.808032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.772663Z digest=sha256:7bc5be0452b90a43e31c0b8076065a08dda961b8a846c2c89e5d52a2ebdd5e2b

Observation e08582c1-c91b-41fa-9d3e-b2392fe6da4c · outbound

This paper cites V A-DepthNet: A variational approach to sin- gle image depth prediction.

Video Depth without Video Models V A-DepthNet: A variational approach to sin- gle image depth prediction

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.794253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.776864Z digest=sha256:d8cfbe58458f9f3f601b667ed7038424b4ae6e3efbf46154d0e60b721deae710

Observation ea80b74e-2656-48ed-86f5-41870ada40b8 · outbound

This paper cites Video-P2P: Video editing with cross-attention control.

Video Depth without Video Models Video-P2P: Video editing with cross-attention control

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.778536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.781319Z digest=sha256:5d913b011520207f1c591f558de8e7e5d40d88b1e6dc909de0e951037411754e

Observation 2af15d2e-7aee-494f-9ff5-b44e70f2dc13 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

Video Depth without Video Models SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.785289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.785289Z digest=sha256:e6e58a960e486e054fa4ea7f657919aed2d79382c37987dfe6eef91ccda46c19

Observation 9be976d2-92a6-4e6e-9e8a-32cedc83b05b · outbound

This paper cites Decoupled weight decay regularization.

Video Depth without Video Models Decoupled weight decay regularization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.789391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.789391Z digest=sha256:81c63074520d7d07265c58419f300b198146131c6639a55d743bebe8b7dfc034

Observation 3e086bdf-840e-4fe8-b9fd-8ab3d8323262 · outbound

This paper cites Consistent video depth estimation.ACM Transactions on Graphics, 39(4), 2020.

Video Depth without Video Models Consistent video depth estimation.ACM Transactions on Graphics, 39(4), 2020

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.758481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.793484Z digest=sha256:41dc39897d9bde435bfb2a065ef7f95daee7128d91182047738e8c6ebc5f9729

Observation 2c6974da-0113-444c-bd31-ca9e15c17c33 · outbound

This paper cites NeRF: Representing scenes as neural radiance fields for view syn- thesis.

Video Depth without Video Models NeRF: Representing scenes as neural radiance fields for view syn- thesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.797151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.797151Z digest=sha256:5a96c1f264525c8deac6b055215845bd388ee98ded42827e4f61fcaf89976007

Observation 8fcb55ee-7cea-4a89-ad43-05c5ab084d2e · outbound

This paper cites All in tokens: Uni- fying output space of visual tasks via soft token.

Video Depth without Video Models All in tokens: Uni- fying output space of visual tasks via soft token

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.739508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.800748Z digest=sha256:baf898d2dad0904bc8cd62fcbc9959c53d188c5d6705b4b938ad738ca3ba9a50

Observation c2f00492-7546-43d7-9bdb-3708338816a4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Video Depth without Video Models DINOv2: Learning Robust Visual Features without Supervision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.804394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.804394Z digest=sha256:694ce84e682b03fb1bd68581667b044a8dc129657bb70afbcc797aee7e85b32a

Observation 38b7728a-73cd-40ae-af85-7a841339a714 · outbound

This paper cites ReFusion: 3d reconstruc- tion in dynamic environments for RGB-D cameras exploit- ing residuals.

Video Depth without Video Models ReFusion: 3d reconstruc- tion in dynamic environments for RGB-D cameras exploit- ing residuals

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.727163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.808185Z digest=sha256:b2a6656728a88cf3bae8a8976ac3733bd57af3bf017fbae61f38466e2bf949f6

Observation d49004a7-ec77-4ef3-af93-f532648ede2a · outbound

This paper cites P3Depth: Monocular depth estimation with a piecewise planarity prior.

Video Depth without Video Models P3Depth: Monocular depth estimation with a piecewise planarity prior

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.714725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.811873Z digest=sha256:bea8a06b8c600575214d4fa9a1b5fcd60a28bfacea5c93657c391f40461a7486

Observation e7e2d022-3458-4da9-a0d2-451a6a01bb24 · outbound

This paper cites UniDepth: Universal monocular metric depth estimation.

Video Depth without Video Models UniDepth: Universal monocular metric depth estimation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.815699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.815699Z digest=sha256:e05dbe8025f9fdb9c603137ab30f1c5db8ec89079900cdcda3a8638ac16cb60c

Observation c4d69034-2dee-4d07-8773-7a0640579482 · outbound

This paper cites UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.

Video Depth without Video Models UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.819538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.819538Z digest=sha256:a1aff699d38468a1d8480b7c3051762df8a08d47f3e26c7ab41871315ed60d19

Observation f12d81f2-dcf5-47c4-948f-4d4f11cf43e8 · outbound

This paper cites FateZero: Fus- ing attentions for zero-shot text-based video editing.

Video Depth without Video Models FateZero: Fus- ing attentions for zero-shot text-based video editing

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.694337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.823762Z digest=sha256:fdb63a8523b8194c7b5121645e37d0797859b1ce20293babd5b992957d07d81b

Observation 92630959-ae0c-4582-9385-b03606a3ffbf · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Video Depth without Video Models Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.681680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.827426Z digest=sha256:0f27fca1dc0ee8392fca58d599585f0014744d22186e01745702132f999906ed

Observation 997222fb-748c-42a3-a83b-44b755b595e9 · outbound

This paper cites Susskind.

Video Depth without Video Models Susskind

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.668767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.831122Z digest=sha256:33d6cc67b8298afbc4d4289e239d24cfeef44f42001b03641e330146e79ef8d1

Observation c4065ab6-b964-44a3-bf10-e06634455472 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Video Depth without Video Models High-resolution image syn- thesis with latent diffusion models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.834637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.834637Z digest=sha256:6b5bafc1c09f37513622e6b1b9cccc834e18518213179ea2d8a1bd49d331036b

Observation 8dee4387-6d9d-43ae-b03c-bc87f325f838 · outbound

This paper cites Monocular Depth Estimation using Diffusion Models.

Video Depth without Video Models Monocular Depth Estimation using Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.838842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.838842Z digest=sha256:4692cef19cc2b56e70b636649c69a8212324d41da6f68764cc343850bc79d65c

Observation 71f0d90f-0eee-4148-be45-cd37460a1245 · outbound

This paper cites Structure- from-motion revisited.

Video Depth without Video Models Structure- from-motion revisited

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.842806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.842806Z digest=sha256:088b0037dc0d894e1df843b32902fb9e9972628dbef948a0425ec0f18abd9280

Observation 8d395caf-8812-4c3a-bae8-6b7db2335741 · outbound

This paper cites LAION-5B: An open large-scale dataset for train- ing next generation image-text models.

Video Depth without Video Models LAION-5B: An open large-scale dataset for train- ing next generation image-text models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.639852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.846818Z digest=sha256:17f46e785848f19ed0689cf107dc97082f073538ddc65e44f0e337da8e3760d3

Observation 3df7feef-382a-4e39-80d7-41a017370b5a · outbound

This paper cites Learning Temporally Consistent Video Depth from Video Diffusion Priors.

Video Depth without Video Models Learning Temporally Consistent Video Depth from Video Diffusion Priors

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.850501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.850501Z digest=sha256:3b8e65f9583d73541814c67486d472bb3fcff0c064f1b64c72113ecdfef44778

Observation c30ea8b9-020f-4cda-a46e-9ff8417922a1 · outbound

This paper cites Denois- ing diffusion implicit models.

Video Depth without Video Models Denois- ing diffusion implicit models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.855178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.855178Z digest=sha256:afa98e1a086d7c7c0f755f2bae2b291fea0bd418d2be2ef1eb63515b33fd1090

Observation 62dfb33b-894b-4d7d-959a-352469cf5a12 · outbound

This paper cites Consistent Direct Time-of-Flight Video Depth Super-Resolution.

Video Depth without Video Models Consistent Direct Time-of-Flight Video Depth Super-Resolution

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:31:02.073424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.858949Z digest=sha256:c2e0ed0e6feebab5c0fb8fed7b31699ae5d54149c5e344ec1884ffd1ae69f125

Observation f6ebea8d-1e95-4ce6-bc31-eb07784e1f51 · outbound

This paper cites Deepv2d: Video to depth with differentiable structure from motion.

Video Depth without Video Models Deepv2d: Video to depth with differentiable structure from motion

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.620872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.863140Z digest=sha256:1b7e835ae401a6f81871f1af2fccc6002fdc3ab248a1031ecc79aa8ee5c10515

Observation 0361a183-0daf-42c1-bce3-1b93728a0e41 · outbound

This paper cites TartanAir: A dataset to push the limits of visual SLAM.

Video Depth without Video Models TartanAir: A dataset to push the limits of visual SLAM

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.607933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.866993Z digest=sha256:d624f0aa510ff523c8ac601b2cbec2df1901a3df67d6e12114817c671ef5a4fc

Observation 4c62d0ce-8de2-4cb9-a30a-b8821dd201db · outbound

This paper cites Less is more: Consistent video depth estimation with masked frames modeling.

Video Depth without Video Models Less is more: Consistent video depth estimation with masked frames modeling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.870791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.870791Z digest=sha256:dd80fe7825c5751c0d2254834cd3822e95a561faee11037fc3efe73a91746102

Observation 1d4eea3b-0414-4bdd-acd2-d3e4e11e8beb · outbound

This paper cites Neural video depth stabilizer.

Video Depth without Video Models Neural video depth stabilizer

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.588590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.874792Z digest=sha256:e0bdffe1d72c03aed5c6f079628a97ce1d645f25e4642af032e84aa46fe0e458

Observation fc7e1971-26d4-4f21-809b-bf92c0b5a383 · outbound

This paper cites NVDS+: Towards efficient and versatile neu- ral stabilizer for video depth estimation.

Video Depth without Video Models NVDS+: Towards efficient and versatile neu- ral stabilizer for video depth estimation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.576043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.878801Z digest=sha256:c02541cdd67655878fa2b08d88eef0a0644b6f9dbebe260cc4a2b0181df56a34

Observation adc8c878-75e9-47bb-afe0-7984c6043c7f · outbound

This paper cites Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation.

Video Depth without Video Models Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.882839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.882839Z digest=sha256:59f3aa63cc704153a42452a8e3e0a8e9666d7f2b4f372161a11853dda24240b1

Observation 914bfc38-a4fe-4180-baa8-6b680157748c · outbound

This paper cites What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?.

Video Depth without Video Models What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.887217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.887217Z digest=sha256:aaa0bc0d6bcc3f3baf370b7321b9888dc0d8c81213c36afa9f9f548aae1af549

Observation 5edbce83-a6ab-40b9-b495-a17193892467 · outbound

This paper cites GMFlow: Learning optical flow via global matching.

Video Depth without Video Models GMFlow: Learning optical flow via global matching

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.554131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.891212Z digest=sha256:93c2c1534c313c7ee9ec551cfe194362f65648705ceb44fe17e3bf1b661e0394

Observation d82e66a5-ba70-49cf-8cc8-d1764c459633 · outbound

This paper cites Transformer-based attention networks for con- tinuous pixel-wise prediction.

Video Depth without Video Models Transformer-based attention networks for con- tinuous pixel-wise prediction

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.542700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.899685Z digest=sha256:683e118f38db09c827911cbf59c7e27ded9cc29614ef7ea4e8361c620227bb13

Observation 37055c38-c71b-49cb-a01b-db8052a962ff · outbound

This paper cites Depth Any Video with Scalable Synthetic Data.

Video Depth without Video Models Depth Any Video with Scalable Synthetic Data

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.903771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.903771Z digest=sha256:239acefa4c066ee073a561e94de55a6847a9bc48371966b3444b8ef091db2fc4

Observation c001dabc-b9ce-4193-b837-2d24e45889a1 · outbound

This paper cites Depth Anything: Unleashing the power of large-scale unlabeled data.

Video Depth without Video Models Depth Anything: Unleashing the power of large-scale unlabeled data

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.531134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.908272Z digest=sha256:b0197270a6aaa89b1f81f341c731420186226189523ea26292e7d2890935502a

Observation f656238b-e34b-4dd8-a0c3-67b7e9065a5e · outbound

This paper cites Depth Anything V2.

Video Depth without Video Models Depth Anything V2

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.912048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.912048Z digest=sha256:9b235bf1fc837338922b39464654507700f4e975ab88f65c876218475f837fa5

Observation 4ada5868-a836-4b21-b93a-e6f1f4d06e91 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

Video Depth without Video Models Rerender a video: Zero-shot text-guided video-to-video translation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.518859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.916309Z digest=sha256:75b3f18aa697119cf5df0edb72f068dd6db4e1daa8e36fb0790087e58f6bea81

Observation df9d140d-52f4-4c24-b285-c596c080867f · outbound

This paper cites MVSNet: Depth inference for unstructured multi- view stereo.

Video Depth without Video Models MVSNet: Depth inference for unstructured multi- view stereo

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.506753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.920169Z digest=sha256:d88a746d286fe5375d9383a508e4edba279177d14c6f0780631ea49d9ea0c638

Observation b92f36a0-5e7d-4bb1-add5-4d1ba24db125 · outbound

This paper cites MAMo: Leveraging memory and attention for monocular video depth estimation.

Video Depth without Video Models MAMo: Leveraging memory and attention for monocular video depth estimation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.494777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.924254Z digest=sha256:299422c9896634ffddfd3580eb06821a56c1fd492ee7a1e5fa264fd00c083643

Observation 6516fa3c-291d-4ed3-9cf8-3eaa0994df44 · outbound

This paper cites FutureDepth: Learning to predict the future improves video depth estimation.

Video Depth without Video Models FutureDepth: Learning to predict the future improves video depth estimation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.482662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.928282Z digest=sha256:aab85422b5c25a91cda7d85158a0e219e51e6cfe86d5473a59fec28c8f70134f

Observation e33824b0-77e0-4717-83bc-21c33eb92127 · outbound

This paper cites DiverseDepth: Affine-invariant Depth Prediction Using Diverse Data.

Video Depth without Video Models DiverseDepth: Affine-invariant Depth Prediction Using Diverse Data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.932136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.932136Z digest=sha256:9d9e510c9605b7f6b1fed94b48e3713f99eb30c62651a9b40ac234b3029ffd8f

Observation 6b45c223-9bff-4909-b076-6bbe2b5d3fa6 · outbound

This paper cites Met- ric3D: Towards zero-shot metric 3d prediction from a single image.

Video Depth without Video Models Met- ric3D: Towards zero-shot metric 3d prediction from a single image

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.470501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.936092Z digest=sha256:f3c8bb1ccdd6dc107b8227a37291dbc7e873f71d3d16d7771e4a167426715be1

Observation 5ad4e55d-affb-4aaf-b72d-91abd1cd4616 · outbound

This paper cites NeWCRFs: Neural window fully-connected CRFs for monocular depth estimation.

Video Depth without Video Models NeWCRFs: Neural window fully-connected CRFs for monocular depth estimation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.456618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.939684Z digest=sha256:20a458a9c853305b46a5d1fd35ff90fd216d38849805b5ea7231f6b755d33e2a

Observation 00272e04-f534-4c59-9f72-b0c00e4637d0 · outbound

This paper cites Exploiting temporal consistency for real-time video depth estimation.

Video Depth without Video Models Exploiting temporal consistency for real-time video depth estimation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.444924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.943636Z digest=sha256:8be24e126c3830a943c7db5df8c5e6c47e562d1e6b58d97cd23d9881eb6015de

Observation f309b9ae-c06f-4312-9e26-58c0eb195f34 · outbound

This paper cites BetterDepth: Plug-and-play dif- fusion refiner for zero-shot monocular depth estimation.

Video Depth without Video Models BetterDepth: Plug-and-play dif- fusion refiner for zero-shot monocular depth estimation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.431214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.947360Z digest=sha256:10a623cb4235f78640cd984b1bbeed388faebf67b3c9caf019f2eedd7016e866

Observation 080af1e0-f9bf-4313-ae56-78081d158b03 · outbound

This paper cites Consistent depth of moving objects in video.

Video Depth without Video Models Consistent depth of moving objects in video

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.417580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.951224Z digest=sha256:961e27687b5ab8f19bff911a036bfde17110f8698bedb439a9020ec350ee7fb5

Observation 8fdf208e-210a-45c0-8063-9465219d8c2a · outbound

This paper cites Towards consistent video edit- ing with text-to-image diffusion models.

Video Depth without Video Models Towards consistent video edit- ing with text-to-image diffusion models

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.404731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.955157Z digest=sha256:ec93aedd35a72147f5e77f9fb53ba92ef91fb70a51d57e4c915ce01bd60ce9a6

Observation 0a3e5fbf-5e3d-485e-92b1-4bd17bd1b5d4 · outbound

This paper cites Unleashing Text-to-Image Diffusion Models for Visual Perception.

Video Depth without Video Models Unleashing Text-to-Image Diffusion Models for Visual Perception

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.959283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.959283Z digest=sha256:93941430987b4bb76ed6c84f2e4be584ce6ac1dc6ddfce1094d5e4b6adc03049

Observation 3833efa8-c59a-4798-9eae-ba395ee413d5 · outbound

This paper cites Discrete cosine transform network for guided depth map super-resolution.

Video Depth without Video Models Discrete cosine transform network for guided depth map super-resolution

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.963767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.963767Z digest=sha256:582a9866d72118d7718eb8102448ca6a54fa5b23c67eaf163e4527c8afa072c1

Observation 7796da87-de8d-4487-ad04-8d4727b1f475 · outbound

This paper cites DDFM: Denoising diffusion model for multi-modality image fusion.

Video Depth without Video Models DDFM: Denoising diffusion model for multi-modality image fusion

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.381613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.967642Z digest=sha256:25f26ad9bdf63fb88c7cf142e424e512503df48d7a314bfb97d445051610bd0d

Observation 60fe359e-4536-48dc-827b-f9cb8b9ff1d3 · outbound

This paper cites PointOdyssey: A large-scale synthetic dataset for long-term point tracking.

Video Depth without Video Models PointOdyssey: A large-scale synthetic dataset for long-term point tracking

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:31:02.368286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.971313Z digest=sha256:605c69ef7eb770046fe6ad39304ecb76ea3fcae3d75e44e1ae93cef10d1328d6

Observation 5d474f9b-87f7-44cf-bfc7-7ffb4ba00e1e · outbound

This paper cites num-frames.

Video Depth without Video Models num-frames

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T10:31:02.355165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T10:31:01.975283Z digest=sha256:5d7ab7f2eadee11f961ea9a1cdb8ebad7e08c60eed40490421bf6fa6ceeec46a

Pith citing papers

Observation 163a352c-fcf8-4785-b463-ba548262d900 · inbound

Stabilizing Streaming Video Geometry via Dynamic Feature Normalization cites this paper.

Stabilizing Streaming Video Geometry via Dynamic Feature Normalization Video Depth without Video Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:02.093483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T23:15:20.254402Z digest=sha256:7a00adbb82f26c0a32f464a19c328cfd52a0585dc84036663ef059e1a87e5b48