Pith. sign in

REVIEW 4 major objections 5 minor 30 references

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read For ERCP dual scopes, role-asymmetric evidence routing outperforms symmetric fusion

desk verdict Novel role-asymmetric dual-scope routing with a solid benchmark, but the 'consistently outperforms' claim rests on single runs; referee it. read the letter →

arxiv 2608.03211 v1 pith:LO7QF7JY submitted 2026-08-04 cs.CV cs.RO

classification cs.CVcs.RO
keywords surgicalvideopredictionworldmodelrole-asymmetricmulti-viewERCPMother-Childendoscopygeometry-guidedroutinggenerationflowmatching
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper studies future-frame prediction for Mother–Child ERCP, where a wide-view duodenoscope and a near-field cholangioscope observe the same procedure without calibrated stereo. It argues that cross-scope evidence should be routed asymmetrically: geometric motion cues from the Mother view condition Child dynamics, while Child appearance is written into Mother prediction only where a learned geometric placement can establish valid spatial correspondence. The proposed CrossScope model keeps separate per-view experts and communicates through zero-initialized residual writes, and a new paired benchmark with phantom and real-world ERCP episodes shows consistent gains over independent prediction, input fusion, and symmetric cross-attention. The headline numbers are a 33.0% lower Child endpoint error and a 16.0% lower Mother FID versus Cosmos-H-Surgical.

What carries the argument

The load-bearing mechanism is a read/write routing contract between two view-specific DiT experts: cross-scope communication is only allowed as zero-initialized residual writes, and each route reads a pre-injection snapshot. The geometric interface is a predicted Mother-plane Child-tip trajectory from a Pose Readout, and a coverage gate—a 7x7 canvas with a learned affine footprint—determines per Mother slot whether Current Child appearance, causal History appearance, both, or neither may be written. M2C encodes the trajectory as 12-dimensional transition descriptors plus coarse endpoint conditions; C2M translates pose-aligned Child latents into Mother-domain residuals only at geometry-licens

What would settle it

Run the C2M route on held-out ERCP episodes with strong zoom, wide-angle distortion, or synthetic warps of the Child view that break the affine placement assumption, and compare Mother FID against the MoT-only variant. If Mother FID no longer improves or worsens, the geometry licensing contract is the source of the gain; if gains persist, the learned footprint is more flexible than the affine model suggests.

Watch

Extended reading notes

Core claim

The discovery is that future dynamics in a two-scope system are better predicted by role-asymmetric evidence routing than by symmetric fusion. CrossScope maintains distinct DiT experts for the Mother and Child views, and restricts cross-view communication to two directional residual paths. Mother-to-Child (M2C) conditions the Child stream on a Mother-plane trajectory of the Child tip, with fine transition descriptors and coarse group-level motion; Child-to-Mother (C2M) pose-aligns the observed Child anchor and causal history onto a 7x7 Mother-plane canvas, and writes translated appearance residuals only where geometric coverage licenses them. The learned geometry is a near-affine map with a

Load-bearing premise

The central premise is that the Child view's appearance can be mapped into the Mother plane by a learned near-affine transform with a fixed footprint and bounded correction, so a 7x7 coverage canvas determines where Child evidence is geometrically admissible; if non-rigid tissue, lens distortion, or dynamic zoom violates this contract, the geometry gate licenses misaligned writes and the C2M gains could vanish or reverse.

Editorial extensions

If this is right

  • If the routing principle holds, multi-observer world models for other cooperative endoscopes or laparoscopes should be designed with role asymmetry rather than symmetric fusion.
  • The benchmark protocol—paired synchronized episodes with view-specific structural, localization, and motion metrics—provides a template for evaluating dual-scope prediction beyond frame fidelity.
  • The causal-history mechanism means that long-horizon prediction can exploit strictly-earlier observations from the other view without leaking future evidence.
  • The geometry gate shows that explicit spatial licensing can be more effective than learned cross-attention for transferring appearance between uncalibrated views.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The routing pattern suggests a general design rule for multi-observer prediction: decide what evidence each target needs before deciding how to fuse inputs—applicable to robot teleoperation with a global camera and a wrist camera as much as to ERCP.
  • Because the paper releases phantom videos, splits, and evaluation code, other groups can test whether the 33% endpoint improvement persists under different backbone sizes and denoising schedules, isolating the routing contract as the causal factor.
  • Prior single-view surgical world models may have failed partly by treating all evidence as symmetric; revisiting them with role-aware routing could close part of the gap without new backbone capacity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces CrossScope, a dual-stream world model for joint Mother–Child ERCP video prediction. It formulates role-asymmetric communication: the Mother-to-Child (M2C) path routes Mother-plane motion cues to the Child stream, while the Child-to-Mother (C2M) path routes pose-aligned Child appearance to the Mother stream only where a geometric coverage gate licenses it. The authors propose a causal read/write contract, a paired dual-scope benchmark (300 phantom and 52 real-world episodes), and evaluate frame fidelity, structural consistency, target localization, and trajectory accuracy against single-view baselines, early fusion, and symmetric cross-attention. The central claim is that role-asymmetric, geometry-licensed evidence routing outperforms independent prediction, input fusion, and symmetric exchange.

Significance. The paper addresses a realistic and under-studied problem: multi-observer future prediction with non-interchangeable views. The architecture is well specified, the read/write invariant is clean, and the same-backbone ablations help attribute gains to the two directional routes. Strengths include the detailed appendices, pose-readout control diagnostics, episode-disjoint splits, and a commitment to release phantom data and evaluation code. If the empirical claims hold, the contribution is useful for surgical video generation and multi-view world models. However, the current evidence is under-powered: every quantitative comparison is a single training run (App. B.1), the real-world set has only seven episodes from one cohort, and the task evaluators are used without independent validation. The central claim of consistent superiority therefore rests on point estimates whose variance is not reported.

major comments (4)
  1. [Appendix B.1 and Tables 1, 2, 6, 9] All headline comparisons are based on one training run per model with seed 42 and one generation seed. The paper states this explicitly but draws strong conclusions such as 'consistently outperforms'. Video diffusion training is high-variance, and several decisive gaps are small in absolute value (Table 1 Child FID 17.402 vs 17.431; Mother SSIM 0.973 vs 0.969; Table 2 Child PSNR 34.868 vs 35.113). Without multiple seeds, confidence intervals, or a significance test, the central empirical claim is not statistically supported. This is load-bearing because the contribution is an empirical architecture claim, not a theorem.
  2. [Real-World Evaluation and Table 9] The real-world evidence is seven held-out episodes from a single cohort, also from a single training run. Claims of 'consistent gains across both roles' rely on small differences (e.g., Child PSNR 18.627 vs 18.314 in Table 9) with no per-episode breakdown or interval. The sample is too small and too homogeneous to support generalization beyond the single cohort. The conclusion properly acknowledges this limitation, but the main-text phrasing 'consistent gains across both roles' overstates the evidence.
  3. [Appendix D.2 and Table 8] The task evaluators (YOLO11x papilla detector, DeepLabV3-based mask segmenter) are frozen but their accuracy and stability are not reported. The endpoint and centerline errors are computed from detected box centers, so detector jitter or systematic bias on hard frames could dominate the measured differences. The paper does not provide validation of the evaluators on a labeled subset, nor confidence intervals from detection uncertainty. This matters because the route-attribution claims (e.g., M2C reducing endpoint error by 18.5%) depend on precise trajectory measurements.
  4. [Eq. (6) and Appendix A.3] The C2M placement contract assumes a near-affine map between Child-local coordinates and the Mother plane, with a single learned global footprint A and a bounded correction δ(u). This assumption is not directly validated. If the Mother–Child relationship deviates substantially due to non-rigid tissue, lens distortion, or dynamic zoom, the coverage gate may license misaligned appearance, and the C2M gains could vanish or reverse. The qualitative coverage visualizations (Figures 5 and 7) are suggestive but not a quantitative test. A concrete validation would be to compare learned placements against manual correspondences on held-out windows, or to compare a per-window affine estimate against the fixed global A.
minor comments (5)
  1. [Figure 3] The architecture figure is dense and some labels (EVAE, DVAE, ST Attention, AdaLN) are not fully defined in the caption or the main text. Adding a legend or a short description of each block would improve reproducibility and readability.
  2. [Abstract/Introduction] The phrase 'consistently outperforms' appears several times, but the body evidence is point estimates. Softening this language in the abstract until variance estimates are available would better match the results.
  3. [Appendix A.3, Eq. (9)] The definition of the trajectory-ROI weight ρ_{k,s} is vague ('weight slots by the predicted-trajectory region of interest'). The exact formula or normalization should be given so the selection rule is reproducible.
  4. [Section Dual-Scope Benchmark, Figure 2] The text says 596,190 synchronized time points for the phantom split and 34,800 for real-world, while Figure 2 says '1.262M denotes both views counted as individual frames.' Explain the relationship between synchronized pairs and individual frames to avoid reader confusion.
  5. [References] Some references use incomplete author lists or non-standard formatting (e.g., 'WanTeam 2025'), and the arXiv numbers for future-dated works (e.g., 2608.03211) are unusual. Please ensure the bibliography follows the journal style consistently.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: all headline metrics are held-out, and the routing choices are trained rather than derived from the claimed conclusions.

full rationale

CrossScope's central claim is empirical: role-asymmetric, geometry-licensed evidence routing improves dual-scope future prediction over independent prediction, input fusion, and symmetric exchange. The supporting comparisons are all end-to-end evaluations on episode-disjoint held-out data: 'Phantom tests use 860 causal windows from 15 episodes' and 'real-world evaluation uses a patient- and case-disjoint 45/7 split within one cohort'; every generated output is scored by frozen evaluators (e.g., fixed papilla detector, fixed segmentation segmenter). The pose and routing modules are trained on the training split via Eq. (8), but none of the reported test metrics is a training target in the same split: Mother/Child FID, PSNR, SSIM, LPIPS, mask IoU, box IoU/recall, and trajectory errors are measured on generated future frames against target clips using fixed checkpoints. The directional ablations in Table 2 are not tautological: adding M2C could in principle have affected the Mother view or failed to help Child motion, and adding C2M could have failed to reduce Mother FID; the observed specialization is an empirical outcome, not a definitional equivalence. The comparison against Symmetric Cross-Attention MoT and Early-Fusion Wan2.2 gives the role-asymmetry claim falsifiable content. The only plausible self-citations (e.g., Pan et al. 2026 as background for long-horizon surgical world models) are contextual and not load-bearing; no uniqueness theorem or architectural ansatz is imported from prior work by the same authors. The single-training-run and lack of variance estimates noted by the skeptic is a statistical-support limitation, not circularity, because the predictions still come from held-out inference rather than from fitted values. I find no step in which a reported 'prediction' reduces by construction to an input, fit, or self-citation.

Assumptions & free parameters 8 free parameters · 7 assumptions · 0 invented entities

The central claim rests on a set of domain assumptions about the geometric relationship between Mother and Child scopes, the reliability of mask-derived pose supervision, and the validity of fixed task evaluators. No new physical entities are introduced. Free parameters are hand-chosen hyperparameters that control the C2M placement and training balance; they do not encode the target outcome.

free parameters (8)
  • Coverage threshold tau = 0.05
    Hand-chosen in Appendix A.3 and Table 4; decides whether a Mother-plane slot is geometry-licensed for C2M writes.
  • Recency half-life h = 200 frames
    Hand-chosen in Eq. 9 and Table 4; controls age-weighting of causal history candidates.
  • Footprint scale = 0.1
    Hand-chosen in Table 4; sets the Child-to-Mother projection footprint in the 7x7 canvas.
  • Kernel width = 0.15
    Hand-chosen in Table 4; sets the coverage kernel width for C2M support.
  • Canvas size = 7x7
    Hand-chosen spatial resolution of C2M placement; coarse licensing granularity.
  • Outer loss weights (lambda_M, lambda_C, lambda_P, lambda_aux) = 1, 1, 1, 0.1
    Weights in Eq. 8; hand-selected, not tuned on the test set.
  • Pose loss sub-weights = 1, 0.5, 0.05, 0.2, 1
    Weights in L_pose (Appendix B); hand-selected.
  • Route warm-up steps = M2C 3000, C2M 500
    Delays residual learning to stabilize training; hand-chosen in Table 4.
assumptions (7)
  • standard math Flow matching on non-anchor latent groups yields a valid generative objective (Eq. 8, Appendix B).
    Uses prior theoretical guarantees of flow matching (Lipman et al. 2023).
  • domain assumption A Mother-plane Child-tip trajectory can be predicted from the Mother pre-injection state and anchor (Eq. 3).
    This is the basis for Pose Readout and M2C; the paper supervises it with mask-derived poses but assumes it is learnable and useful.
  • domain assumption Child appearance projects into the Mother plane as an affine footprint with bounded correction, pi(p,u) = (x,y) + Rot(theta)(Au + delta(u)) (Eq. 6).
    Load-bearing for C2M placement; if the real geometry deviates from near-affine (non-rigid tissue, zoom), the placement contract fails.
  • domain assumption The mask-derived pose extraction (largest component, centerline, tip/base) accurately reflects the Child-tip pose (Appendix C.2).
    Pose Readout targets and C2M ground truth come from this procedure; systematic errors propagate.
  • domain assumption Unsupported Mother-plane slots should receive zero residual, and the coverage gate in Eq. 7 correctly licenses evidence.
    Core of the role-asymmetric contract; the alternative that the Mother expert should sometimes fill missing geometry from appearance is untested.
  • domain assumption The benchmark's episode-disjoint synchronization and clinician-verified annotations are correct (Appendix C).
    All evaluation relies on paired, synchronized frames and annotations.
  • domain assumption The frozen task evaluators (DeepLabV3 segmenter, YOLO11 papilla detector) produce reliable targets and measurements (Appendix D.2).
    Mask IoU, box IoU, recall, and trajectory errors depend on these checkpoints.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction." pith.science (2026). https://pith.science/paper/LO7QF7JY

@misc{pith2026260803211,
  author       = {Pith},
  title        = {Pith review of: CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LO7QF7JY}},
  note         = {Machine review of arXiv:2608.03211}
}
read the original abstract

Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother--Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship. Unlike conventional multi-view fusion that assumes symmetric information exchange, we formulate \textbf{role-asymmetric dual-scope future prediction}, where cross-view evidence is selectively transferred according to the prediction target and its underlying spatial requirements. We propose \textbf{CrossScope}, a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions. CrossScope learns two complementary communication directions: geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. This design allows each scope to contribute task-relevant evidence without compromising its view-specific representation. To evaluate this problem, we establish a paired dual-scope benchmark comprising synchronized phantom and real-world ERCP episodes, with evaluations assessing visual fidelity, structural preservation, target localization, and motion consistency. Experiments demonstrate that CrossScope consistently outperforms strong surgical video generation baselines, validating the importance of role-aware evidence routing for multi-observer visual world modeling.

Figures

Figures reproduced from arXiv: 2608.03211 by the authors.

Figure 1
Figure 1. CrossScope for role-asymmetric dual-scope future [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Dual-scope benchmark composition. The bench [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. CrossScope architecture. Separate DiT experts expose read-only pre-injection states. Frame [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Qualitative future prediction across views. Independent examples are shown for a real-world Mother sequence (left) [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: C2M spatial-support diagnostic. A held-out phan [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Real-world quantitative comparison on seven held [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Geometry-licensed C2M evidence placement. Four held-out phantom windows from the 10,000-step C2M-only [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: View-specific annotations in the Dual-Scope Benchmark. Each column is a synchronized Mother–mask–Child triplet; [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Visual validation of Pose Readout. Four held-out phantom windows at step 10,000 are shown over their nine temporally [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Output-level Child papilla trajectories for nine selected held-out phantom episodes at step 10,000. The left and [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

30 extracted references · 21 canonical work pages

  1. [1]

    World Journal of Gastroenterology , volume =

    Peroral Cholangioscopy in the New Millennium , author =. World Journal of Gastroenterology , volume =. 2011 , doi =

  2. [2]

    2018 , eprint =

    World Models , author =. 2018 , eprint =. doi:10.48550/arXiv.1803.10122 , url =

  3. [3]

    Bruce, Jake and Dennis, Michael D. and Edwards, Ashley and Parker-Holder, Jack and Shi, Yuge and Hughes, Edward and Lai, Matthew and Mavalankar, Aditi and Steigerwald, Richie and Apps, Chris and Aytar, Yusuf and Bechtle, Sarah Maria Elisabeth and Behbahani, Feryal and Chan, Stephanie C. Y. and Heess, Nicolas and Gonzalez, Lucy and Osindero, Simon and Ozai...

  4. [4]

    and Li, Wuyang and Liu, Xinyu and Chen, Zhen and Shao, Jing and Yuan, Yixuan , booktitle =

    Li, Chenxin and Liu, Hengyu and Liu, Yifan and Feng, Brandon Y. and Li, Wuyang and Liu, Xinyu and Chen, Zhen and Shao, Jing and Yuan, Yixuan , booktitle =. 2024 , doi =

  5. [5]

    2025 , doi =

    Chen, Tong and Yang, Shuya and Wang, Junyi and Bai, Long and Ren, Hongliang and Zhou, Luping , booktitle =. 2025 , doi =

  6. [6]

    Data Engineering in Medical Imaging , series =

    Surgical Vision World Model , author =. Data Engineering in Medical Imaging , series =. 2025 , doi =

  7. [7]

    SurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue Dynamics

    Pan, Wentao and Li, Wuyang and Liu, Shengyuan and others , year =. doi:10.48550/arXiv.2606.19889 , url =. 2606.19889 , archiveprefix =

  8. [8]

    doi:10.48550/arXiv.2603.13024 , note =

    Rapuri, Sampath and Seenivasan, Lalithkumar and Schneider, Dominik and Soberanis-Mukul, Roger and He, Yufan and Ding, Hao and Xu, Jiru and Yu, Chenhao and Jing, Chenyan and Guo, Pengfei and Xu, Daguang and Unberath, Mathias , year =. doi:10.48550/arXiv.2603.13024 , note =. 2603.13024 , archiveprefix =

Show all 30 references
  1. [9]

    2026 , eprint =

    From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation , author =. 2026 , eprint =. doi:10.48550/arXiv.2605.08712 , url =

  2. [10]

    2022 , doi =

    Medical Image Computing and Computer Assisted Intervention -- MICCAI 2022 , series =. 2022 , doi =

  3. [11]

    and Liu, Yifan and Liu, Hengyu and Wang, Cheng and Yu, Weihao and Yuan, Yixuan , booktitle =

    Li, Chenxin and Feng, Brandon Y. and Liu, Yifan and Liu, Hengyu and Wang, Cheng and Yu, Weihao and Yuan, Yixuan , booktitle =. 2024 , doi =

  4. [12]

    Deep Generative Models , series =

    Interactive Generation of Laparoscopic Videos with Diffusion Models , author =. Deep Generative Models , series =. 2024 , doi =

  5. [13]

    2024 , doi =

    Huang, Yiming and Cui, Beilei and Bai, Long and Guo, Ziqi and Xu, Mengya and Islam, Mobarakol and Ren, Hongliang , booktitle =. 2024 , doi =

  6. [14]

    2025 , doi =

    Biagini, Diego and Navab, Nassir and Farshad, Azade , booktitle =. 2025 , doi =

  7. [15]

    doi:10.48550/arXiv.2410.17751 , url =

    Yeganeh, Youssef and Lazuardi, Reza and Shamseddin, Ahmad and others , year =. doi:10.48550/arXiv.2410.17751 , url =. 2410.17751 , archiveprefix =

  8. [16]

    2025 , eprint =

    How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment , author =. 2025 , eprint =. doi:10.48550/arXiv.2511.01775 , url =

  9. [17]

    2025 , doi =

    Sivakumar, Ssharvien Kumar and Frisch, Yannik and Ghazaei, Ghazal and Mukhopadhyay, Anirban , booktitle =. 2025 , doi =

  10. [18]

    2026 , doi =

    Stilz, Florian and Karaoglu, Mert and Tristram, Felix and Navab, Nassir and Busam, Benjamin and Ladikos, Alexander , journal =. 2026 , doi =

  11. [19]

    , booktitle =

    Cheng, Ho Kei and Schwing, Alexander G. , booktitle =. 2022 , doi =

  12. [20]

    2025 , eprint =

    Xiao, Zeqi and Lan, Yushi and Zhou, Yifan and Ouyang, Wenqi and Yang, Shuai and Zeng, Yanhong and Pan, Xingang , booktitle =. 2025 , eprint =

  13. [21]

    2026 , eprint =

    Scaling Video Pretraining for Surgical Foundation Models , author =. 2026 , eprint =. doi:10.48550/arXiv.2603.29966 , url =

  14. [22]

    doi:10.48550/arXiv.2602.05638 , url =

    Wu, Jinlin and Holm, Felix and Chen, Chuxi and others , year =. doi:10.48550/arXiv.2602.05638 , url =. 2602.05638 , archiveprefix =

  15. [23]

    npj Digital Medicine , year =

    Large-scale Self-supervised Video Foundation Model for Intelligent Surgery , author =. npj Digital Medicine , year =. doi:10.1038/s41746-026-02403-0 , url =

  16. [24]

    2025 , issn =

    Liang, Weixin and Yu, Lili and Luo, Liang and Iyer, Srinivasan and Dong, Ning and Zhou, Chunting and Ghosh, Gargi and Lewis, Mike and Yih, Wen-tau and Zettlemoyer, Luke and Lin, Xi Victoria , journal =. 2025 , issn =. 2411.04996 , archiveprefix =

  17. [25]

    and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =

    Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =. 2022 , eprint =. doi:10.48550/arXiv.2106.09685 , url =

  18. [26]

    2023 , eprint =

    Flow Matching for Generative Modeling , author =. 2023 , eprint =. doi:10.48550/arXiv.2210.02747 , url =

  19. [27]

    2017 , eprint =

    Heusel, Martin and Ramsauer, Hubert and Unterthiner, Thomas and Nessler, Bernhard and Hochreiter, Sepp , booktitle =. 2017 , eprint =. doi:10.48550/arXiv.1706.08500 , url =

  20. [28]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year =

    The Unreasonable Effectiveness of Deep Features as a Perceptual Metric , author =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year =. doi:10.48550/arXiv.1801.03924 , url =. 1801.03924 , archiveprefix =

  21. [29]

    2412.03603 , archiveprefix =

    Kong, Weijie and Tian, Qi and Zhang, Zijian and Min, Rox and Dai, Zuozhuo and others , year =. 2412.03603 , archiveprefix =

  22. [30]

    doi:10.48550/arXiv.2512.23162 , url =

    He, Yufan and Guo, Pengfei and Xu, Mengya and Li, Zhaoshuo and Myronenko, Andriy and Imans, Dillan and Liu, Bingjie and Yang, Dongren and Gu, Mingxue and Ji, Yongnan and Jin, Yueming and Zhao, Ren and Shen, Baiyong and Xu, Daguang , year =. doi:10.48550/arXiv.2512.23162 , url ...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.