Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

Motif Channel Opened in a White-Box: Stereo Matching via Motif Correlation Graph

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read MoCha-V2 claims that a wavelet-domain Motif Correlation Graph uncovers recurring geometric structures across feature channels, and restoring them yields 1st place on Middlebury and 2nd on KITTI 2012 Reflective.

desk verdict A competent incremental extension of MoCha-Stereo with real but modest benchmark gains, yet the white-box motif claim does not survive contact with the method as written. read the letter →

arxiv 2411.12426 v2 pith:UHWECMK6 submitted 2024-11-19 cs.CV

classification cs.CV
keywords stereomatchingmotifcorrelationgraphwhite-boxwavelettransformdisparityestimationinterpretabilityedgedetailattentionmechanism
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MoCha-V2 claims that convolutional stereo networks lose edge detail because feature channels are activated unevenly, and that this lost geometry can be recovered by mining recurrent patterns across channels. The paper introduces a Motif Correlation Graph that counts, at each spatial position, which feature channel is the nearest neighbour of which other channel in the wavelet domain, and uses these counts to weight an average of 3x3 patches. Inverse-wavelet transforming this weighted average and multiplying it back into the features restores recurring geometric structures, which is what the authors call a white-box motif channel. If correct, this gives stereo matching a state-of-the-art edge-sensitive disparity estimate, ranked 1st on Middlebury Bad 1.0 all and 2nd on KITTI 2012 Reflective at submission, with an interpretable attention mechanism rather than a learned black-box one.

What carries the argument

The Motif Correlation Graph (MCG) is a directed graph built at each spatial location after a two-level Haar wavelet transform: for every 3×3 patch position, nodes are the feature channels in one of Ng groups, and an edge from node c to node c' carries weight equal to the Euclidean distance between their patch sequences. Node weights are incremented when another node's nearest neighbour is that node, with ties split evenly. The weighted channel average forms a motif, which after inverse wavelet transform and element-wise multiplication with the original features repairs lost edge information. The paper argues this is interpretable because the graph is constructed from distances and counts rather than learned parameters, and that it acts as both channel and spatial attention.

What would settle it

Train MoCha-V2 identically but replace the MCG node weights with fixed random weights (or a plain average) and compare Middlebury Bad 1.0; if the error does not rise materially, the motif-mining mechanism is not the cause of the reported gains. A complementary check is to measure whether channels with high MCG weight are actually the ones whose patches reappear at multiple spatial locations in the image, rather than merely being centrally placed in feature space.

Watch

Extended reading notes

Core claim

The central claim is that recurring geometric structures in stereo feature maps can be identified without learned attention weights. MoCha-V2 splits each feature channel into 3×3 patches, converts them into one-dimensional sequences, applies a two-level Haar wavelet transform, and builds a directed graph per spatial location whose edge weights are Euclidean distances between sequences across channels. Each node's weight is incremented when another channel's nearest neighbour points to it, so a channel that many others are closest to becomes a 'motif'. The weighted average of patches, followed by inverse wavelet transform and element-wise multiplication with the original features, yields the restored motif features that feed the correlation volume. On this mechanism the method reports state-of-the-art results: 1st on Middlebury (Bad 1.0 all), 2nd on KITTI 2012 Reflective, and improved zero-shot performance on Driving Stereo.

Load-bearing premise

The load-bearing premise is that a weighted average of 3×3 patches at the same spatial location across feature channels, weighted by wavelet-domain nearest-neighbour counts, captures recurring geometric structures whose inverse wavelet transform restores lost edge detail.

Editorial extensions

If this is right

  • Edge-sensitive disparity maps improve on thin structures such as streetlights, signage, ropes, and object contours, as shown on KITTI and ETH3D.
  • The parameter-free motif mining can be inspected: the node weights and graph edges visualize which channels recurrently encode a texture.
  • Including low-frequency motifs, via the wavelet approximation coefficients, contributes to accuracy beyond high-frequency-only mining.
  • The network remains accurate with fewer update iterations, cutting inference time while staying competitive.
  • Replacing the learned motif attention of the earlier MoCha-Stereo with MCG improves both accuracy and speed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the cross-channel nearest-neighbour counts really track recurring textures, the same graph construction could be plugged into other dense prediction heads (optical flow, monocular depth) wherever edges are lost; the paper does not test this.
  • The white-box claim is partial: the paper states that the rest of the deep pipeline remains a trained black box, so the safety benefit applies only to the attention stage.
  • A sharper test of the mechanism would replace MCG weights with random or uniform weights and measure the drop on Middlebury; the paper does not run this control.
  • The name 'motif' borrows from time-series analysis, but the operation is a spatial-location-wise cross-channel pooling rather than a spatial repetition search; the interpretability claim stands or falls on whether the two coincide in practice.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes MoCha-V2, a stereo matching network that extends the authors' prior MoCha-Stereo system. The key novelty is the Motif Correlation Graph Attention (MCGA), described as a white-box, parameter-free mechanism that identifies recurring geometric structures ('motifs') in feature channels, using Euclidean distances between wavelet-domain feature subsequences, and restores lost edge detail via inverse wavelet transform and element-wise multiplication. The paper reports state-of-the-art results on Scene Flow, Middlebury (rank 1 at submission time), KITTI 2012 Reflective (rank 2), KITTI 2015, ETH3D, and zero-shot Driving Stereo, with ablations and a speed/accuracy comparison against the conference version.

Significance. If the motif-mining interpretation were supported, the paper would offer a novel, interpretable attention mechanism for stereo matching with strong empirical results. The benchmark evaluation is thorough: five test sets, zero-shot generalization, an iteration ablation, and a comparison with the conference version are all reported, and the code is publicly available. However, the current description of the Motif Correlation Graph does not implement the stated motif-mining: it pools across channels at fixed spatial locations rather than detecting spatial recurrence across the image. The empirical gains in Tables 1-6 and the ablation in Table 7 are credible, and the module may act as a useful parameter-free channel-attention-like mechanism, but the paper's central 'white-box motif' contribution is not established by the described mechanism.

major comments (3)
  1. [3.3, Eq. (2)] The Motif Correlation Graph computes Euclidean distances d(s_{c,j}, s_{c',j}) between subsequences located at the same spatial position j in different channels c and c', and Equation 2 forms m_j as a weighted average of channel patches at that single position. No comparison is made between positions j and j' within a channel, so the graph cannot detect patterns that recur across the image. This contradicts the definition of a motif in Section 2.1 (Eq. 1), which requires the most similar pair of subsequences at different positions in a series. Consequently, the abstract's and introduction's claim that MCG 'captures recurring geometric structures' is unsupported by the described mechanism. The module is best characterized as a cross-channel attention or pooling operation; the authors should either modify the mechanism to compare across spatial positions or provide evidence (e.g., an analysis or visualization) that cross-channel similarity at a pixel correlates with spatial texture recurrence, and adjust the paper's claims accordingly.
  2. [3.3, Eq. (3)] The restoration mechanism is stated without support: the claim is that element-wise multiplication of the inverse-wavelet-transformed motif map with the original feature map 'restores' geometric structures. No derivation or analysis is given for why this operation recovers lost edge detail. The ablation in Table 7 shows an EPE improvement from 0.409 to 0.394 when adding MCGA, but it does not isolate the contributions of the wavelet transform, the graph weighting, and the gating multiplication. The paper should provide a component-wise ablation or a feature-level analysis to support the restoration claim, or it should be reframed as a heuristic gating mechanism.
  3. [3.3, 'stitch them sequentially'] The construction of the new feature map m_g from the k motif patches is underspecified. The text says the motifs are 'expanded into 3x3 features and stitched sequentially into a new feature map', but the spatial correspondence between the flattened patch index j and the locations in the inverse-wavelet-transformed map is not defined, nor are the sliding-window stride and boundary conditions. Without these details, the module cannot be reproduced exactly as described, which is a serious issue for a component that is claimed to be white-box and interpretable.
minor comments (6)
  1. [3.2] The feature notation 'fl,i(fr,i) ∈ RCi×H/i×H/i' is malformed; it should likely read RCi×(H/i)×(W/i) to be dimensionally consistent with the text.
  2. [3.6, Eq. (9)] The symbol 'HEF(o)' is a typo for 'HFE(o)', and the expression for d_k mixes LFE, HFE, and LMC in a way that is not clearly derived from the preceding definitions; please clarify the operations and correct the notation.
  3. [Table 4] In the caption, 'ercentage' should be 'percentage', and the description of the KITTI 2012 columns is confusing because the table contains both KITTI 2015 and KITTI 2012 metrics; please separate or more clearly label the captions for the two benchmarks.
  4. [References [41] and [44]] References [41] and [44] are the same work (Dau and Keogh, 'Matrix Profile V'); they should be merged to avoid duplicate citations.
  5. [3.5, Eq. (6)] The sentence 'this process happened at 1/25−t resolution' is unclear; it likely means 1/2^{5-t} resolution, but as written it is ambiguous and should be rewritten.
  6. [Figure 9] The visualization of the Motif Correlation Graph does not show how the graph nodes correspond to specific feature channels or spatial locations; without a clear mapping, the figure does not substantiate the interpretability claim.

Circularity Check

0 steps flagged · score 1.0 of 10

Low circularity: external benchmarks anchor the claims; the 'motif' labelling issue is an interpretability gap, not a circular derivation.

full rationale

The paper's central benchmark claims are evaluated on external test servers (Middlebury, KITTI, ETH3D, Scene Flow, Driving Stereo), so the reported SOTA numbers cannot reduce to the method's own inputs by construction. The MCG operation in Eq. 2-3 is self-referential in the sense that the 'motif' m_j is a weighted average of the same feature subsequences s_{c,j} that it gates, but this is an architectural feature transform, not a fitted constant disguised as a prediction; there is no external motif quantity being predicted from fitted data. The heavy reliance on the authors' prior conference version [15] is not load-bearing for the new claim: the sliding-window and REMP components are described in the text and independently ablated in Table 7, and the improvement of MCG over MCA is measured rather than assumed. The skeptic's objection that MCG compares patches across channels at a fixed position, rather than across positions, is a correctness/interpretability concern about whether the term 'motif' matches the time-series definition in Eq. 1; it does not amount to a derivation that reduces to its own inputs. The conclusion's limitation statement candidly concedes the method is not fully white-box, which is consistent with an empirical contribution rather than a circular one. No circular step is exhibited.

Assumptions & free parameters 4 free parameters · 5 assumptions · 2 invented entities

The method relies on several domain assumptions about what constitutes a motif and how it restores detail; these are not derived from first principles. The architectural choices (Ng, window size, wavelet level) are hand-set. The MCG and motif channels are new constructs without external falsifiable handles.

free parameters (4)
  • Number of channel groups Ng = 8
    Chosen manually following [15], [31], [32] to balance computation and performance; affects MCG construction and correlation volume.
  • Sliding window size = 3x3
    Patch size for motif extraction, chosen by hand; not derived from data.
  • Wavelet decomposition level = 2
    Two-level Haar wavelet chosen by hand to separate high- and low-frequency components.
  • Number of update iterations = 22
    Training setting; ablations show performance varies with iteration count from 1 to 32.
assumptions (5)
  • domain assumption Euclidean distance between wavelet-domain feature subsequences across channels identifies motifs.
    Invoked in Section 3.3 to build the Motif Correlation Graph; the paper provides no evidence that this distance corresponds to semantically recurring geometric structures.
  • ad hoc to paper Element-wise multiplication of the inverse-wavelet-transformed motif map with the original feature map restores lost geometric structures.
    Equation 3 in Section 3.3; this is the core recovery mechanism, asserted without derivation or validation.
  • domain assumption Sequential stitching of motif patches preserves the spatial layout of the original feature map.
    Section 3.3 describes flattening 3x3 patches and stitching them into a new feature map; the spatial correspondence is assumed but never stated.
  • standard math Haar wavelet transform is invertible and separates high- and low-frequency components.
    Standard signal processing; used in Section 3.3.
  • domain assumption The LSTM update operator and loss functions follow prior work [15], [31].
    The iterative update in Section 3.5 is adopted from existing methods; the paper does not re-derive it.
invented entities (2)
  • Motif Correlation Graph (MCG)
    purpose: A directed graph over feature channels at each spatial location to identify and weight recurring feature patterns, used to build motif channels for stereo matching.
    It is an algorithmic construct with no external falsifiable prediction; its usefulness is only indirectly supported by benchmark performance.
  • Motif channel
    purpose: A feature channel that encapsulates repeated geometric information, produced by applying MCG and inverse wavelet transform to the original features.
    Defined in Section 3.3; there is no independent measurement of 'motifness' outside the network's own representations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Motif Channel Opened in a White-Box: Stereo Matching via Motif Correlation Graph." pith.science (2026). https://pith.science/paper/UHWECMK6

@misc{pith2026241112426,
  author       = {Pith},
  title        = {Pith review of: Motif Channel Opened in a White-Box: Stereo Matching via Motif Correlation Graph},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UHWECMK6}},
  note         = {Machine review of arXiv:2411.12426}
}
read the original abstract

Real-world applications of stereo matching, such as autonomous driving, place stringent demands on both safety and accuracy. However, learning-based stereo matching methods inherently suffer from the loss of geometric structures in certain feature channels, creating a bottleneck in achieving precise detail matching. Additionally, these methods lack interpretability due to the black-box nature of deep learning. In this paper, we propose MoCha-V2, a novel learning-based paradigm for stereo matching. MoCha-V2 introduces the Motif Correlation Graph (MCG) to capture recurring textures, which are referred to as ``motifs" within feature channels. These motifs reconstruct geometric structures and are learned in a more interpretable way. Subsequently, we integrate features from multiple frequency domains through wavelet inverse transformation. The resulting motif features are utilized to restore geometric structures in the stereo matching process. Experimental results demonstrate the effectiveness of MoCha-V2. MoCha-V2 achieved 1st place on the Middlebury benchmark at the time of its release. Code is available at https://github.com/ZYangChen/MoCha-Stereo.

Figures

Figures reproduced from arXiv: 2411.12426 by the authors.

Figure 1
Figure 1. In the deep learning process, feature channels are extended to [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Pipeline of our previous work, MoCha-Stereo [15]. MoCha-Stereo first constructs the Motif Channel Correlation Volume (MCCV). MCCV is [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Architecture overview of MoCha-V2. We improved the method of obtaining Motif Channels in MoCha-Stereo [15] by introducing the Motif [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Visual comparisons with state-of-the-art stereo methods [38], [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Visual comparisons with state-of-the-art stereo methods [38], [64] on the KITTI 2012 test set. In the first row of images, the streetlight presents [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Visual evaluations using the KITTI 2015 test set in contrast to the SOTA techniques [38], [64]. In the first row, UCFNet and Selective-IGEV [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Visual comparisons with SOTA stereo methods [61], [69] on [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Visual comparisons with Selective-IGEV on the Driving Stereo dataset under various weather conditions reveal the robust performance of [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Visualization of Motif Correlation Graphs computed from a pair [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Visualization of feature channels. We utilize Principal Com [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hadamard Attention Recurrent Transformer: A Strong Baseline for Stereo Matching Transformer

    cs.CV 2025-01 conditional novelty 6.0 of 10

    HART, a linear-complexity stereo transformer using Hadamard product attention with the Dense Attention Kernel, reports SOTA EPE 0.42 on Scene Flow and 1st place on KITTI 2012 reflective areas at submission.

  2. BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A single network that iteratively aligns monocular features with stereo hypotheses reduces zero-shot stereo depth error by over 40% on Middlebury and ETH3D.

Reference graph

Works this paper leans on

81 extracted references · 75 canonical work pages · cited by 2 Pith papers

  1. [1]

    Object scene flow for autonomous vehicles,

    M. Menze and A. Geiger, “Object scene flow for autonomous vehicles,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 3061– 3070

  2. [2]

    Bounding boxes, segmentations and object coordinates: How important is recognition for 3d scene flow estimation in autonomous driving scenarios?

    A. Behl, O. Hosseini Jafari, S. Karthik Mustikovela, H. Abu Alhaija, C. Rother, and A. Geiger, “Bounding boxes, segmentations and object coordinates: How important is recognition for 3d scene flow estimation in autonomous driving scenarios?” inInt. Conf. Comput. Vis., 2017, pp. 2574–2583

  3. [3]

    Driv- ingstereo: A large-scale dataset for stereo matching in autonomous driving scenarios,

    G. Yang, X. Song, C. Huang, Z. Deng, J. Shi, and B. Zhou, “Driv- ingstereo: A large-scale dataset for stereo matching in autonomous driving scenarios,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 899–908

  4. [4]

    Diver- gent stereo for robot navigation: Learning from bees,

    J. Santos-Victor, G. Sandini, F. Curotto, and S. Garibaldi, “Diver- gent stereo for robot navigation: Learning from bees,” in IEEE Conf. Comput. Vis. Pattern Recog. IEEE, 1993, pp. 434–439

  5. [5]

    Stereo correspondence and reconstruction of endoscopic data challenge,

    M. Allan, J. Mcleod, C. Wang, J. C. Rosenthal, Z. Hu, N. Gard, P . Eisert, K. X. Fu, T. Zeffiro, W. Xia et al., “Stereo correspondence and reconstruction of endoscopic data challenge,” arXiv preprint arXiv:2101.01133, 2021

  6. [6]

    Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,

    Z. Li, X. Liu, N. Drenkow, A. Ding, F. X. Creighton, R. H. Taylor, and M. Unberath, “Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,” in Int. Conf. Comput. Vis., 2021, pp. 6197–6206

  7. [7]

    Colonoscopy 3d video dataset with paired depth from 2d-3d registration,

    T. L. Bobrow, M. Golhar, R. Vijayan, V . S. Akshintala, J. R. Garcia, and N. J. Durr, “Colonoscopy 3d video dataset with paired depth from 2d-3d registration,” Med. Image Anal., vol. 90, p. 102956, 2023

  8. [8]

    Weakly supervised learn- ing of deep metrics for stereo reconstruction,

    S. Tulyakov, A. Ivanov, and F. Fleuret, “Weakly supervised learn- ing of deep metrics for stereo reconstruction,” in Int. Conf. Comput. Vis., 2017, pp. 1339–1348

Show all 81 references
  1. [9]

    Mvsnet: Depth in- ference for unstructured multi-view stereo,

    Y. Yao, Z. Luo, S. Li, T. Fang, and L. Quan, “Mvsnet: Depth in- ference for unstructured multi-view stereo,” in Eur. Conf. Comput. Vis., 2018, pp. 767–783

  2. [10]

    Multi- view photometric stereo revisited,

    B. Kaya, S. Kumar, C. Oliveira, V . Ferrari, and L. Van Gool, “Multi- view photometric stereo revisited,” in IEEE Winter Conf. Appl. Comput. Vis., 2023, pp. 3126–3135

  3. [11]

    Semantic stereo for incidental satellite images,

    M. Bosch, K. Foster, G. Christie, S. Wang, G. D. Hager, and M. Brown, “Semantic stereo for incidental satellite images,” in IEEE Winter Conf. Appl. Comput. Vis. IEEE, 2019, pp. 1524–1532

  4. [12]

    Hmsm-net: Hierarchical multi- scale matching network for disparity estimation of high-resolution satellite stereo images,

    S. He, S. Li, S. Jiang, and W. Jiang, “Hmsm-net: Hierarchical multi- scale matching network for disparity estimation of high-resolution satellite stereo images,” ISPRS Ann. Photogrammetry, Remote Sens. Spatial Inf. Sciences , vol. 188, pp. 314–330, 2022

  5. [13]

    Surface depth estimation from multi-view stereo satellite images with distribution contrast network,

    Z. Chen, W. Li, Z. Cui, and Y. Zhang, “Surface depth estimation from multi-view stereo satellite images with distribution contrast network,” IEEE J. Sel. T opics Appl. Earth Observ. Remote Sens. , 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, NOVEMBER 2024 12

  6. [14]

    Stereo matching by training a convolu- tional neural network to compare image patches,

    J. ˇZbontar and Y. LeCun, “Stereo matching by training a convolu- tional neural network to compare image patches,” J. Mach. Learn. Res., vol. 17, no. 65, pp. 1–32, 2016

  7. [15]

    Mocha-stereo: Motif channel attention network for stereo matching,

    Z. Chen, W. Long, H. Yao, Y. Zhang, B. Wang, Y. Qin, and J. Wu, “Mocha-stereo: Motif channel attention network for stereo matching,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2024

  8. [16]

    Efficient belief prop- agation for early vision,

    P . F. Felzenszwalb and D. P . Huttenlocher, “Efficient belief prop- agation for early vision,” Int. J. Comput. Vis. , vol. 70, pp. 41–54, 2006

  9. [17]

    Processing by semiglobal matching and mutual information,

    H. H. Stereo, “Processing by semiglobal matching and mutual information,” IEEE T rans. Pattern Anal. Mach. Intell. , vol. 30, no. 2, pp. 328–341, 2007

  10. [18]

    Linear stereo matching,

    L. De-Maeztu, S. Mattoccia, A. Villanueva, and R. Cabeza, “Linear stereo matching,” in Int. Conf. Comput. Vis. IEEE, 2011, pp. 1708– 1715

  11. [19]

    Stochastic stereo matching over scale,

    S. T. Barnard, “Stochastic stereo matching over scale,” Int. J. Comput. Vis., vol. 3, no. 1, pp. 17–32, 1989

  12. [20]

    Learning and feature selection in stereo matching,

    M. S. Lew, T. S. Huang, and K. Wong, “Learning and feature selection in stereo matching,” IEEE T rans. Pattern Anal. Mach. Intell., vol. 16, no. 9, pp. 869–881, 1994

  13. [21]

    Stereo matching with transparency and matting,

    R. Szeliski and P . Golland, “Stereo matching with transparency and matting,” Int. J. Comput. Vis. , vol. 32, no. 1, pp. 45–61, 1999

  14. [22]

    Stereo matching using belief propagation,

    J. Sun, N.-N. Zheng, and H.-Y. Shum, “Stereo matching using belief propagation,” IEEE T rans. Pattern Anal. Mach. Intell., vol. 25, no. 7, pp. 787–800, 2003

  15. [23]

    Efficient deep learning for stereo matching,

    W. Luo, A. G. Schwing, and R. Urtasun, “Efficient deep learning for stereo matching,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 5695–5703

  16. [24]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556 , 2014

  17. [25]

    Gradient-based learning applied to document recognition,

    Y. LeCun, L. Bottou, Y. Bengio, and P . Haffner, “Gradient-based learning applied to document recognition,” Proc. IEEE , vol. 86, no. 11, pp. 2278–2324, 1998

  18. [26]

    On the synergies between machine learning and binocular stereo for depth estimation from images: a survey,

    M. Poggi, F. Tosi, K. Batsos, P . Mordohai, and S. Mattoccia, “On the synergies between machine learning and binocular stereo for depth estimation from images: a survey,” IEEE T rans. Pattern Anal. Mach. Intell., vol. 44, no. 9, pp. 5314–5334, 2021

  19. [27]

    End-to-end learning of geometry and context for deep stereo regression,

    A. Kendall, H. Martirosyan, S. Dasgupta, P . Henry, R. Kennedy, A. Bachrach, and A. Bry, “End-to-end learning of geometry and context for deep stereo regression,” in Int. Conf. Comput. Vis., 2017, pp. 66–75

  20. [28]

    Pyramid stereo matching network,

    J.-R. Chang and Y.-S. Chen, “Pyramid stereo matching network,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 5410–5418

  21. [29]

    Raft-stereo: Multilevel recurrent field transforms for stereo matching,

    L. Lipson, Z. Teed, and J. Deng, “Raft-stereo: Multilevel recurrent field transforms for stereo matching,” inInt. Conf. 3D Vision. IEEE, 2021, pp. 218–227

  22. [30]

    Raft: Recurrent all-pairs field transforms for optical flow,

    Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” in Eur. Conf. Comput. Vis. Springer, 2020, pp. 402–419

  23. [31]

    Iterative geometry encod- ing volume for stereo matching,

    G. Xu, X. Wang, X. Ding, and X. Yang, “Iterative geometry encod- ing volume for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 21 919–21 928

  24. [32]

    Group-wise correlation stereo network,

    X. Guo, K. Yang, W. Yang, X. Wang, and H. Li, “Group-wise correlation stereo network,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 3273–3282

  25. [33]

    Segstereo: Exploit- ing semantic information for disparity estimation,

    G. Yang, H. Zhao, J. Shi, Z. Deng, and J. Jia, “Segstereo: Exploit- ing semantic information for disparity estimation,” in Eur. Conf. Comput. Vis., 2018, pp. 636–651

  26. [34]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P . Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Med. Image Comput. Comput. Intervention . Springer, 2015, pp. 234–241

  27. [35]

    Multi-dimensional manifolds consistency regularization for semi-supervised remote sensing semantic segmentation,

    Y. Lu, Y. Zhang, Z. Cui, W. Long, and Z. Chen, “Multi-dimensional manifolds consistency regularization for semi-supervised remote sensing semantic segmentation,” Knowl. Based Syst. , p. 112032, 2024

  28. [36]

    Edgestereo: An effective multi-task learning network for stereo matching and edge detection,

    X. Song, X. Zhao, L. Fang, H. Hu, and Y. Yu, “Edgestereo: An effective multi-task learning network for stereo matching and edge detection,” Int. J. Comput. Vis. , vol. 128, no. 4, pp. 910–930, 2020

  29. [37]

    Cbam: Convolutional block attention module,

    S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” in Eur. Conf. Comput. Vis., 2018, pp. 3–19

  30. [38]

    Selective-stereo: Adaptive frequency information selection for stereo matching,

    X. Wang, G. Xu, H. Jia, and X. Yang, “Selective-stereo: Adaptive frequency information selection for stereo matching,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2024, pp. 19 701–19 710

  31. [39]

    Visualizing and understanding con- volutional networks,

    M. D. Zeiler and R. Fergus, “Visualizing and understanding con- volutional networks,” in Eur. Conf. Comput. Vis. Springer, 2014, pp. 818–833

  32. [40]

    Grad-cam: Visual explanations from deep networks via gradient-based localization,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Int. Conf. Comput. Vis. , 2017, pp. 618–626

  33. [41]

    Matrix profile v: A generic technique to incorporate domain knowledge into motif discovery,

    H. A. Dau and E. Keogh, “Matrix profile v: A generic technique to incorporate domain knowledge into motif discovery,” in ACM Special Interest Group Knowl. Discovery Data Mining , 2017, pp. 125– 134

  34. [42]

    Matrix profile i: all pairs similarity joins for time series: a unifying view that includes motifs, discords and shapelets,

    C.-C. M. Yeh, Y. Zhu, L. Ulanova, N. Begum, Y. Ding, H. A. Dau, D. F. Silva, A. Mueen, and E. Keogh, “Matrix profile i: all pairs similarity joins for time series: a unifying view that includes motifs, discords and shapelets,” in Int. Conf. Data Mining . Ieee, 2016, pp. 1317–1322

  35. [43]

    Matrix profile ii: Exploiting a novel algorithm and gpus to break the one hundred million barrier for time series motifs and joins,

    Y. Zhu, Z. Zimmerman, N. S. Senobari, C.-C. M. Yeh, G. Funning, A. Mueen, P . Brisk, and E. Keogh, “Matrix profile ii: Exploiting a novel algorithm and gpus to break the one hundred million barrier for time series motifs and joins,” in Int. Conf. Data Mining . IEEE, 2016, pp. 739–748

  36. [44]

    Matrix profile v: A generic technique to incorporate domain knowledge into motif discovery,

    H. A. Dau and E. Keogh, “Matrix profile v: A generic technique to incorporate domain knowledge into motif discovery,” 2017, pp. 125–134

  37. [45]

    Los: Local structure-guided stereo matching,

    K. Li, L. Wang, Y. Zhang, K. Xue, S. Zhou, and Y. Guo, “Los: Local structure-guided stereo matching,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 746–19 756

  38. [46]

    Spatial trans- former networks,

    M. Jaderberg, K. Simonyan, A. Zisserman et al. , “Spatial trans- former networks,” Adv. Neural Inform. Process. Syst. , vol. 28, 2015

  39. [47]

    Squeeze-and-excitation networks,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 7132–7141

  40. [48]

    Fcanet: Frequency channel attention networks,

    Z. Qin, P . Zhang, F. Wu, and X. Li, “Fcanet: Frequency channel attention networks,” in Int. Conf. Comput. Vis. , 2021, pp. 783–792

  41. [49]

    Distribution- decouple learning network: an innovative approach for single image dehazing with spatial and frequency decoupling,

    Y. Wu, W. Li, Z. Chen, H. Wen, Z. Cui, and Y. Zhang, “Distribution- decouple learning network: an innovative approach for single image dehazing with spatial and frequency decoupling,” Vis. Comput., pp. 1–16, 2024

  42. [50]

    High- frequency stereo matching network,

    H. Zhao, H. Zhou, Y. Zhang, J. Chen, Y. Yang, and Y. Zhao, “High- frequency stereo matching network,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 1327–1336

  43. [51]

    Im- agenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Im- agenet: A large-scale hierarchical image database,” in IEEE Conf. Comput. Vis. Pattern Recog. Ieee, 2009, pp. 248–255

  44. [52]

    Efficientnet: Rethinking model scaling for convolutional neural networks,

    M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in Int. Conf. Mach. Learn. PMLR, 2019, pp. 6105–6114

  45. [53]

    On the theory of orthogonal function systems,

    A. Haar, “On the theory of orthogonal function systems,” Mathe- matische Annalen, vol. 69, no. 3, pp. 331–371, 1910

  46. [54]

    Wavelet based transition region extraction for image segmentation,

    P . Parida and N. Bhoi, “Wavelet based transition region extraction for image segmentation,” Future Comput. Inform. J. , vol. 2, no. 2, pp. 65–78, 2017

  47. [55]

    Feature distribution normalization network for multi-view stereo,

    Z. Chen, Y. Zhao, J. He, Y. Lu, Z. Cui, W. Li, and Y. Zhang, “Feature distribution normalization network for multi-view stereo,” Vis. Comput., pp. 1–13, 2024

  48. [56]

    Learning for disparity estimation through feature constancy,

    Z. Liang, Y. Feng, Y. Guo, H. Liu, W. Chen, L. Qiao, L. Zhou, and J. Zhang, “Learning for disparity estimation through feature constancy,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 2811–2820

  49. [57]

    Fast r-cnn,

    R. Girshick, “Fast r-cnn,” in Int. Conf. Comput. Vis., 2015, pp. 1440– 1448

  50. [58]

    Any-stereo: Arbitrary scale disparity esti- mation for iterative stereo matching,

    Z. Liang and C. Li, “Any-stereo: Arbitrary scale disparity esti- mation for iterative stereo matching,” in Assoc. Advancement Artif. Intell., vol. 38, no. 4, 2024, pp. 3333–3341

  51. [59]

    Mc-stereo: Multi-peak lookup and cascade search range for stereo matching,

    M. Feng, J. Cheng, H. Jia, L. Liu, G. Xu, and X. Yang, “Mc-stereo: Multi-peak lookup and cascade search range for stereo matching,” arXiv preprint arXiv:2311.02340 , 2023

  52. [60]

    Hadamard attention recurrent transformer: A strong baseline for stereo matching transformer,

    Z. Chen, Y. Zhang, W. Li, B. Wang, Y. Wu, Y. Zhao, and C. Chen, “Hadamard attention recurrent transformer: A strong baseline for stereo matching transformer,” arXiv preprint arXiv:2501.01023 , 2025

  53. [61]

    Unifying flow, stereo and depth estimation,

    H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,” IEEE T rans. Pattern Anal. Mach. Intell. , 2023

  54. [62]

    Uncertainty guided adaptive warping for robust and efficient stereo matching,

    J. Jing, J. Li, P . Xiong, J. Liu, S. Liu, Y. Guo, X. Deng, M. Xu, L. Jiang, and L. Sigal, “Uncertainty guided adaptive warping for robust and efficient stereo matching,” in Int. Conf. Comput. Vis. , 2023, pp. 3318–3327

  55. [63]

    Croco v2: JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, NOVEMBER 2024 13 Improved cross-view completion pre-training for stereo matching and optical flow,

    P . Weinzaepfel, T. Lucas, V . Leroy, Y. Cabon, V . Arora, R. Br´egier, G. Csurka, L. Antsfeld, B. Chidlovskii, and J. Revaud, “Croco v2: JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, NOVEMBER 2024 13 Improved cross-view completion pre-training for stereo matching and optical ...

  56. [64]

    Digging into uncertainty-based pseudo-label for robust stereo matching,

    Z. Shen, X. Song, Y. Dai, D. Zhou, Z. Rao, and L. Zhang, “Digging into uncertainty-based pseudo-label for robust stereo matching,” IEEE T rans. Pattern Anal. Mach. Intell., 2023

  57. [65]

    Detail preserving coarse- to-fine matching for stereo matching and optical flow,

    Y. Deng, J. Xiao, S. Z. Zhou, and J. Feng, “Detail preserving coarse- to-fine matching for stereo matching and optical flow,”IEEE T rans. Image Process., vol. 30, pp. 5835–5847, 2021

  58. [66]

    Practical stereo matching via cascaded recurrent network with adaptive correlation,

    J. Li, P . Wang, P . Xiong, T. Cai, Z. Yan, L. Yang, J. Liu, H. Fan, and S. Liu, “Practical stereo matching via cascaded recurrent network with adaptive correlation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16 263–16 272

  59. [67]

    Rethinking training strategy in stereo matching,

    Z. Rao, Y. Dai, Z. Shen, and R. He, “Rethinking training strategy in stereo matching,” IEEE T rans. Neural Netw. Learn. Syst. , vol. 34, no. 10, pp. 7796–7809, 2022

  60. [68]

    Stereo risk: A continuous modeling approach to stereo match- ing,

    C. Liu, S. Kumar, S. Gu, R. Timofte, Y. Yao, and L. Van Gool, “Stereo risk: A continuous modeling approach to stereo match- ing,” arXiv preprint arXiv:2407.03152 , 2024

  61. [69]

    Adaptive multi-modal cross-entropy loss for stereo matching,

    P . Xu, Z. Xiang, C. Qiao, J. Fu, and T. Pu, “Adaptive multi-modal cross-entropy loss for stereo matching,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 5135–5144

  62. [70]

    Robust synthetic-to-real transfer for stereo matching,

    J. Zhang, J. Li, L. Huang, X. Yu, L. Gu, J. Zheng, and X. Bai, “Robust synthetic-to-real transfer for stereo matching,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2024, pp. 20 247–20 257

  63. [71]

    A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,

    N. Mayer, E. Ilg, P . Hausser, P . Fischer, D. Cremers, A. Dosovitskiy, and T. Brox, “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 4040–4048

  64. [72]

    Are we ready for autonomous driving? the kitti vision benchmark suite,

    A. Geiger, P . Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012

  65. [73]

    Joint 3d estimation of vehicles and scene flow,

    M. Menze, C. Heipke, and A. Geiger, “Joint 3d estimation of vehicles and scene flow,” ISPRS Ann. Photogrammetry, Remote Sens. Spatial Inf. Sciences , vol. 2, p. 427, 2015

  66. [74]

    High-resolution stereo datasets with subpixel-accurate ground truth,

    D. Scharstein, H. Hirschm ¨uller, Y. Kitajima, G. Krathwohl, N. Ne ˇsi´c, X. Wang, and P . Westling, “High-resolution stereo datasets with subpixel-accurate ground truth,” in Pattern Recog. German Conf. Springer, 2014, pp. 31–42

  67. [75]

    A multi-view stereo benchmark with high-resolution images and multi-camera videos,

    T. Schops, J. L. Schonberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 3260–3269

  68. [76]

    Decoupled weight decay regularization,

    I. Loshchilov, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017

  69. [77]

    Tartanair: A dataset to push the limits of visual slam,

    W. Wang, D. Zhu, X. Wang, Y. Hu, Y. Qiu, C. Wang, Y. Hu, A. Kapoor, and S. Scherer, “Tartanair: A dataset to push the limits of visual slam,” in Int. Conf. Intell. Robots Syst. IEEE, 2020, pp. 4909–4916

  70. [78]

    Falling things: A synthetic dataset for 3d object detection and pose estimation,

    J. Tremblay, T. To, and S. Birchfield, “Falling things: A synthetic dataset for 3d object detection and pose estimation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 2038–2041

  71. [79]

    In- stereo2k: a large real dataset for stereo matching in indoor scenes,

    W. Bao, W. Wang, Y. Xu, Y. Guo, S. Hong, and X. Zhang, “In- stereo2k: a large real dataset for stereo matching in indoor scenes,” Sci. China Inf. Sci. , vol. 63, pp. 1–11, 2020

  72. [80]

    Hierarchical deep stereo matching on high-resolution images,

    G. Yang, J. Manela, M. Happold, and D. Ramanan, “Hierarchical deep stereo matching on high-resolution images,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 5515–5524

  73. [81]

    A naturalistic open source movie for optical flow evaluation,

    D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black, “A naturalistic open source movie for optical flow evaluation,” in Eur. Conf. Comput. Vis. Springer, 2012, pp. 611–625. Ziyang Chen (Student Member, IEEE) is cur- rently working toward the master’s degree with the College...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.