Pith. sign in

REVIEW 4 major objections 5 minor 36 references

Neural-Network-Enhanced Metalens Camera for High-Definition, Dynamic Imaging in the Long-Wave Infrared Spectrum

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A neural-network-enhanced metalens camera restores long-wave infrared video to near-commercial quality.

desk verdict A plausible cycleGAN extension for metalens LWIR video, but the unregistered two-camera evaluation and a unit error in the resolution claim leave the headline numbers shaky. read the letter →

arxiv 2411.17139 v1 pith:Q4NCSC2X submitted 2024-11-26 eess.IV cs.CV

classification eess.IVcs.CV
keywords long-waveinfraredmetalenssingletimagingcycle-GANhigh-frequencyenhancementwavelettransformvideocomputational
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a single flat metalens, normally too aberration-prone for sharp infrared imaging, can produce high-definition video when paired with a specially designed neural network. The authors build a compact camera that couples a silicon metalens with a High-Frequency-Enhancing Cycle-GAN, which learns to restore the frequency content lost to chromatic aberration and noise. They report dynamic imaging at 125 frames per second, an End Point Error of 12.58, PSNR of 30.62, SSIM of 0.69, and FID of 0.42 on recorded video, with spatial resolution claimed to be on par with a commercial infrared lens camera. The significance is that such a system could replace bulky multi-element infrared lenses with thin, lightweight optics, simplifying fabrication and enabling compact thermal imaging.

What carries the argument

The central object is the High-Frequency-Enhancing (HFE) Cycle-GAN, a bidirectional cyclic generative adversarial network augmented with a high-frequency adversarial learning module. The module applies a two-dimensional Haar discrete wavelet transform to the generator's output, discards the low-frequency approximation band, reconstructs a high-frequency-only image, and feeds it to a separate high-frequency discriminator. This creates a feedback loop that pushes the generator to recover frequencies lost by the metalens, supplementing the full-frequency adversarial loss and cycle-consistency loss. The optical front end is a 7 mm-diameter silicon metalens with focal length 7 mm operating at 9.5 μm, designed to correct spherical aberration but still suffering chromatic aberration that the network compensates for.

What would settle it

Take the same paired recordings, register the metalens and commercial frames geometrically and radiometrically to subpixel accuracy, and recompute PSNR, SSIM, and FID; if the HFE Cycle-GAN's advantage largely disappears or its output matches the commercial camera only in global statistics, the claimed resolution recovery is not genuine frequency restoration. A complementary test is to measure the output MTF on a calibrated slit target, since true restoration should sharpen the edge profile rather than merely add texture.

Watch

Extended reading notes

Core claim

The central claim is that the High-Frequency-Enhancing Cycle-GAN, when integrated into a singlet metalens LWIR camera, restores high-frequency detail to a level comparable with commercial infrared lenses. The network uses two cyclic generators and three discriminators, including a dedicated high-frequency discriminator that operates on wavelet-decomposed image bands. By extracting the HH, HL, and LH bands via Haar wavelet transform, zeroing the LL band, and reconstructing a high-frequency-only image for adversarial feedback, the generator is forced to reproduce sharp edges and fine textures that the metalens suppresses. The authors demonstrate this with a resolution calibration board, ablation against plain Cycle-GAN, quantitative video metrics, optical-flow smoothness evaluation, and a subjective study in which over 90% of participants preferred the enhanced output.

Load-bearing premise

The paper treats each frame from a commercial infrared camera as the true scene for the metalens frame taken at the same moment, but it does not report aligning the two cameras pixel-by-pixel; if the views are misaligned, the network could be learning to imitate the commercial camera's look rather than physically restoring lost detail.

Editorial extensions

If this is right

  • If the claim holds, a singlet metalens paired with this network can deliver commercial-grade long-wave infrared video, shrinking camera size and cost.
  • The reported 8 ms per-frame processing time implies real-time performance, making the approach usable for dynamic thermal imaging and surveillance.
  • The method's reliance on learned frequency restoration means it could generalize to other imperfect optical systems, provided a diverse paired dataset captures their degradation.
  • The addition of a high-frequency discriminator yields measurable gains over plain Cycle-GAN, suggesting that frequency-domain adversarial training is a practical enhancement for image-to-image translation.
  • The optical-flow-based smoothness metric and subjective evaluation offer a way to assess temporal consistency in video enhancement, beyond per-frame metrics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A decisive test of whether the network genuinely restores optical detail rather than imitating the commercial camera's style would be to register the metalens and commercial frames to subpixel accuracy and recompute the metrics; if the advantage largely vanishes, the method may be learning a camera-to-camera style transfer.
  • The claimed resolution recovery could be validated by measuring the modulation transfer function of the output on a calibrated edge target; genuine restoration should sharpen the edge profile, while hallucinated texture would not improve the physical MTF.
  • The same high-frequency adversarial loop could be applied to other chromatic-aberration-limited flat optics, such as visible-light metalenses, but training would need a dataset matched to that system's specific frequency loss.
  • Because the network is trained on grayscale 256x192 frames, scaling to larger detectors or to multispectral infrared would require re-training and may hit memory or generalization limits not addressed in this work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript reports a long-wave infrared imaging system that combines a singlet metalens with a neural network, the High-Frequency-Enhancing (HFE) Cycle-GAN, to restore image detail lost due to chromatic aberration and other metalens defects. The network is trained on paired low-quality metalens video and high-quality commercial infrared camera video, using cycle-consistent adversarial training plus a high-frequency adversarial loss based on Haar wavelet decomposition. The authors claim dynamic imaging at 125 frames per second, with reported PSNR 30.62, SSIM 0.69, FID 0.42, and EPE 12.58, and state that the enhanced metalens camera achieves spatial resolution on par with a commercial infrared lens camera. The paper includes comparison with Cycle-GAN, Mocycle-GAN, and RISTN on five test videos, a subjective evaluation with 152 participants, and a video smoothness evaluation based on optical flow.

Significance. If the claimed performance holds, the integration of a lightweight metalens with a learning-based high-frequency restoration module would be a practical step toward compact, low-cost LWIR imaging systems for real-time video. The paper's novelty lies in the high-frequency adversarial feedback via wavelet decomposition, which is a plausible and clearly motivated architectural addition. The authors also propose an optical-flow-based video smoothness metric, which is a useful complement to frame-level metrics. However, the validation pipeline has load-bearing gaps: the paired-camera evaluation lacks demonstrated pixel-level registration, the resolution claim contains an apparent unit error, and the quantitative metrics are reported without error bars or a clear validation split. The central contribution is therefore not yet convincingly established.

major comments (4)
  1. [Image Qualification / Supporting Information, Camera Setup] The paired-camera evaluation does not establish pixel-level geometric, radiometric, or temporal registration between the metalens camera and the commercial infrared camera, and the manuscript explicitly states that 'the lack of a one-to-one pixel match between the ground truth infrared image and the enhanced image makes it difficult to accurately assess image quality using PSNR and SSIM.' Without registration, PSNR and SSIM comparisons to the commercial camera cannot be interpreted as measuring restoration of frequency loss; they may instead reflect spatial remapping or camera-to-camera style transfer. The authors should either provide a registration procedure and report residual alignment error, or use metrics that do not require pixel correspondence (e.g., distribution-based or perceptual metrics) as the primary evidence for the restoration claim.
  2. [Resolution Calibration] The claim of 'an angular resolution of 0.002°' for resolving a 2 mm slit at a distance of 1 m is not arithmetically consistent: 2 mm at 1 m subtends approximately 0.115° (or 0.002 radians), which is 57 times larger than 0.002°. This unit error directly affects the resolution claim and the 'on par with that of the commercial infrared lens camera' assertion. The authors should correct the angular resolution value and, ideally, back the 'on par' statement with a quantitative comparison such as a measured modulation transfer function or a contrast-based resolution criterion.
  3. [Image Qualification, Figures 4–6] All PSNR, SSIM, and FID results are reported as averages without error bars, confidence intervals, or significance tests across the 15 or 100 sampled frames, so it is not possible to judge whether the reported improvements (e.g., 8.76% PSNR and 15.79% SSIM over Cycle-GAN in Figure 4) are statistically reliable. The authors should report per-frame distributions or standard deviations and, where possible, run a paired significance test across the test videos.
  4. [Supporting Information, Hyperparameter Optimization] The high-frequency loss weight ω1 was selected 'based on the network's performance' on PSNR, SSIM, and FID (Supporting Information Figure S3), which are the same metrics used for the final evaluation, and no separate validation split is described. This creates a risk of overfitting to the test metrics, making the reported improvements partially a product of hyperparameter selection. The authors should document a validation set that is disjoint from the test videos and report the chosen hyperparameters without referencing test-set performance.
minor comments (5)
  1. [Network Architecture, Equation (2)] The layer indexing in Equation (2) is ambiguous: Equation (2.3) uses Conv2D2 after DeConv, while Table 1 lists DeConv separately and does not list Conv2D2 for upsampling; the authors should align the layer descriptions with Table 1.
  2. [Image Qualification, Figure 5 caption] The text 'RINST' appears to be a typo for 'RISTN'; please correct it.
  3. [Video Smoothness, Subjective evaluation] Table 3 reports the percentage of participants who chose each method, but no information is given on whether the 152 participants were screened for infrared image experience or whether the differences across videos are statistically significant; a brief explanation of the participant pool and a significance test would improve the presentation.
  4. [Dataset] The manuscript states that the dataset contains 19,715 video frame pairs from thirty clips, split 9:1 into train and test, but it is not stated whether the five test videos are disjoint from the training clips; please clarify the split at the video level to avoid frame-level leakage.
  5. [Data availability] The data availability statement indicates that the dataset is not publicly available; given that the paper's central claims depend on the paired-camera dataset, at minimum the trained model and a representative sample of paired frames should be released to support reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the main results are held-out test evaluations against external baselines, and the central claims do not reduce to the training inputs by construction.

full rationale

The paper's derivation chain is a supervised image-restoration pipeline: the HFE Cycle-GAN is trained on paired VLQ/VHQ frames and evaluated on a 9:1 held-out split against Cycle-GAN, Mocycle-GAN, and RISTN. The reported PSNR/SSIM/FID and EPE are therefore test-set measurements of a trained model, not quantities that reduce by definition to the training inputs. The hyperparameter ω1 is tuned in the Supporting Information using PSNR/SSIM/FID, which is a selection effect, but the central comparison against the Cycle-GAN ablation and other external baselines is independent of that tuning, and no equation in the paper defines the claimed performance in terms of the loss weights. The self-citation to the authors' earlier CycleGAN work is used only to set baseline loss weights and to motivate the gap; it is not load-bearing. The admitted lack of one-to-one pixel match between GT and enhanced images (Image Qualification) is an evaluation-validity concern, not a circularity of reasoning.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The central claim rests on a learned mapping between two distinct cameras, several standard machine-learning assumptions, and hand-tuned loss weights. No new physical entities are introduced.

free parameters (5)
  • ω1 (high-frequency adversarial loss weight) = 5
    Set by manual hyperparameter search (Supporting Information Figure S3) based on PSNR, SSIM, and FID; not derived from theory.
  • ω2 (full-frequency adversarial loss weight) = 10
    Taken as a standard CycleGAN value from the authors' prior work; used without re-derivation.
  • ω3 (cycle-consistency loss weight) = 5
    Taken as a standard CycleGAN value from the authors' prior work.
  • Learned weights of generators and discriminators = not released
    All reported enhancement results depend on weights fit to 19,715 frame pairs; final weights are not provided.
  • Metalens unit cell design (period 4 μm, height 5.8 μm, diameter 0.5 to 3.5 μm) = fixed by design
    Chosen to give 2π phase coverage and 69.3% average transmittance; not fitted to image data but central to the physical system.
assumptions (6)
  • domain assumption The commercial infrared camera images (VHQ) are treated as ground truth for metalens images, assuming the two cameras capture the same scene with aligned fields of view and negligible temporal offset.
    End-to-end training and PSNR/SSIM/FID comparisons rely on this correspondence; no registration or synchronization procedure is reported. Location: Supporting Information, Camera Setup.
  • domain assumption The frequency degradation of the metalens is primarily deterministic and can be learned from a finite paired dataset.
    If the degradation varies with scene content, temperature, or noise in ways not represented in 30 clips, the learned mapping will not generalize to unseen scenes. Location: Introduction and Dataset paragraphs.
  • ad hoc to paper A Haar wavelet split into LL/HL/LH/HH captures the perceptually relevant high-frequency detail for metalens restoration.
    The choice of wavelet and the decision to zero out the LL band is a design choice introduced for this paper; no comparison to other frequency decompositions is provided. Location: Equations (4) and (5).
  • domain assumption Cycle-consistency loss preserves content and prevents mode collapse in this image-to-image translation setting.
    Standard CycleGAN assumption; the paper does not provide evidence that cycle consistency holds on metalens images with strong chromatic aberration. Location: Equations (7) and (8).
  • domain assumption RAFT optical flow, trained on visible-light video, yields valid motion fields for grayscale IR metalens videos and is an appropriate smoothness metric.
    The EPE evaluation and smoothness conclusions depend on this assumption; no validation of RAFT on infrared imagery is provided. Location: Video Smoothness section.
  • domain assumption Inference on an unspecified hardware platform completes in 8 ms per frame.
    The '125 fps' claim is not supported by hardware details or measurement methodology. Location: Conclusion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Neural-Network-Enhanced Metalens Camera for High-Definition, Dynamic Imaging in the Long-Wave Infrared Spectrum." pith.science (2026). https://pith.science/paper/Q4NCSC2X

@misc{pith2026241117139,
  author       = {Pith},
  title        = {Pith review of: Neural-Network-Enhanced Metalens Camera for High-Definition, Dynamic Imaging in the Long-Wave Infrared Spectrum},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q4NCSC2X}},
  note         = {Machine review of arXiv:2411.17139}
}
read the original abstract

To provide a lightweight and cost-effective solution for the long-wave infrared imaging using a singlet, we develop a camera by integrating a High-Frequency-Enhancing Cycle-GAN neural network into a metalens imaging system. The High-Frequency-Enhancing Cycle-GAN improves the quality of the original metalens images by addressing inherent frequency loss introduced by the metalens. In addition to the bidirectional cyclic generative adversarial network, it incorporates a high-frequency adversarial learning module. This module utilizes wavelet transform to extract high-frequency components, and then establishes a high-frequency feedback loop. It enables the generator to enhance the camera outputs by integrating adversarial feedback from the high-frequency discriminator. This ensures that the generator adheres to the constraints imposed by the high-frequency adversarial loss, thereby effectively recovering the camera's frequency loss. This recovery guarantees high-fidelity image output from the camera, facilitating smooth video production. Our camera is capable of achieving dynamic imaging at 125 frames per second with an End Point Error value of 12.58. We also achieve 0.42 for Fr\'echet Inception Distance, 30.62 for Peak Signal to Noise Ratio, and 0.69 for Structural Similarity in the recorded videos.

Figures

Figures reproduced from arXiv: 2411.17139 by the authors.

Figure 1
Figure 1. The Neural-Network-Enhanced (NNE) metalens camera. (a) The NNE metalens camera configuration includes an infrared metalens, the HFE Cycle-GAN for high-frequency enhancement, and an infrared detector. (b) illustrates portrait images captured by commercial infrared cameras and naked metalenses, along with a comparison of their frequency comparison. (c) The infrared metalens attached to a 3D printed component. (d) and … view at source ↗
Figure 3
Figure 3. Resolution quantification of the HFE Cycle-GAN. (a) depictsthe calibration board image captured by the visible light camera, naked metalenses, and commercial infrared camera, along with the image reconstructed by the HFE Cycle-GAN. (b) depicts the reconstructed the metalenses images of human faces, upper limbs, and outdoor scenes. (c) depicts the average frequency intensity of 200 video frames. The darker green curv… view at source ↗
Figure 4
Figure 4. The evaluation of Video 1. (a) illustrates five consecutive frames from Video 1 [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: The evaluation of Video 2. (a) presents five frames from Video 2, which showcases the reconstruction results of an outdoor scene, including building edges and foliage. (b) and (c) display PSNR and SSIM across 15 frames, and (d) shows the average metrics over 100 frames…
Figure 6
Figure 6. Figure 6: The evaluation of Video 3. (a) presents five frames from Video 3, which showcases the reconstruction results of roads and trees. (b) and (c) display PSNR and SSIM across 15 frames, and (d) shows the average metrics over 100 frames. Video Smoothness High image quality i…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 33 canonical work pages

  1. [1]

    Largest aperture metalens of high numerical aperture and polarization independence for long-wavelength infrared imaging[J]

    Li J, Wang Y, Liu S, et al. Largest aperture metalens of high numerical aperture and polarization independence for long-wavelength infrared imaging[J]. Optics Express, 2022, 30(16): 28882-28891

  2. [2]

    Thermography of Asteroid and Future Applications in Space Missions

    Okada, T. Thermography of Asteroid and Future Applications in Space Missions. Appl. Sci. 2020, 10, 2158. https://doi.org/10.3390/app10062158

  3. [3]

    Novel Methodology for Condition Monitoring of Gear Wear Using Supervised Learning and Infrared Thermography

    Resendiz-Ochoa, E.; Saucedo-Dorantes, J.J.; Benitez-Rangel, J.P.; Osornio-Rios, R.A.; Morales-Hernandez, L.A. Novel Methodology for Condition Monitoring of Gear Wear Using Supervised Learning and Infrared Thermography. Appl. Sci. 2020, 10, 506. https://doi.org/10.3390/app10020506

  4. [4]

    Thermal Infrared Small Ship Detection in Sea Clutter Based on Morphological Reconstruction and Multi-Feature Analysis

    Li, Y.; Li, Z.; Zhu, Y.; Li, B.; Xiong, W.; Huang, Y. Thermal Infrared Small Ship Detection in Sea Clutter Based on Morphological Reconstruction and Multi-Feature Analysis. Appl. Sci. 2019, 9, 3786. https://doi.org/10.3390/app9183786

  5. [5]

    & Tang, J

    He, Y., Song, B. & Tang, J. Optical metalenses: fundamentals, dispersion manipulation, and applications. Front. Optoelectron. 15, 24 (2022). https://doi.org/10.1007/s12200-022-00017-4

  6. [6]

    Doublet Metalens with Simultaneous Chromatic and Monochromatic Correction in the Mid-Infrared

    Zhou, Y.; Gan, F.; Wang, R.; Lan, D.; Shang, X.; Li, W. Doublet Metalens with Simultaneous Chromatic and Monochromatic Correction in the Mid-Infrared. Sensors 2022, 22, 6175. https://doi.org/10.3390/s22166175

  7. [7]

    Aberrations of flat lenses and aplanatic metasurfaces

    Aieta F, Genevet P, Kats M, Capasso F. Aberrations of flat lenses and aplanatic metasurfaces. Opt Express. 2013;21:31530-9

  8. [8]

    Metasurface Enabled Multi‐Target and Multi‐Wavelength Diffraction Neural Networks[J]

    Chi H, Zang X, Zhang T, et al. Metasurface Enabled Multi‐Target and Multi‐Wavelength Diffraction Neural Networks[J]. Laser & Photonics Reviews, 2024: 2401178

Show all 36 references
  1. [9]

    Metasurfaces designed by a bidirectional deep neural network and iterative algorithm for generating quantitative field distributions[J]

    Zhu Y, Zang X, Chi H, et al. Metasurfaces designed by a bidirectional deep neural network and iterative algorithm for generating quantitative field distributions[J]. Light: Advanced Manufacturing, 2023, 4(2): 104- 114

  2. [10]

    Terahertz multi-foci metalens enabling high-accuracy intensity distributions and polarization-dependent images based on inverse design[J]

    Lu B, Zang X, Zhang T, et al. Terahertz multi-foci metalens enabling high-accuracy intensity distributions and polarization-dependent images based on inverse design[J]. Applied Physics Letters, 2024, 124(12)

  3. [11]

    Multiwavelength metasurfaces through spatial multiplexing

    Arbabi E, Arbabi A, Kamali SM, Horie Y, Faraon A. Multiwavelength metasurfaces through spatial multiplexing. Sci Rep. 2016;6:32803

  4. [12]

    Composite functional metasurfaces for multispectral achromatic optics

    Avayu O, Almeida E, Prior Y, Ellenbogen T. Composite functional metasurfaces for multispectral achromatic optics. Nat Commun. 2017;8:14992

  5. [13]

    Broadband achromatic optical metasurface devices

    Wang S, Wu PC, Su V-C, Lai Y-C, Chu CH, Chen J-W, Lu S-H, Chen J, Xu B, Kuan C-H, Li T, Zhu S, Tsai DP. Broadband achromatic optical metasurface devices. Nat Commun. 2017;8:187

  6. [14]

    Integrated resonant unit of Metasurfaces for broadband efficiency and phase manipulation

    Hsiao H-H, Chen VH, Lin RJ, Wu PC, Wang S, Chen BH, Tsai DP. Integrated resonant unit of Metasurfaces for broadband efficiency and phase manipulation. Adv Opt Mater. 2018;6:1800031

  7. [15]

    Broadband lightweight flat lenses for long-wave infrared imaging

    Meem M, Banerji S, Majumder A, Vasquez FG, Sensale-Rodriguez B, Menon R. Broadband lightweight flat lenses for long-wave infrared imaging. Proc Natl Acad Sci U S A. 2019;116:21375-8

  8. [16]

    Restoration of infrared metalens images with deep learning[J]

    Li R, Wei J, Wang L, et al. Restoration of infrared metalens images with deep learning[J]. Optics Communications, 2024, 552: 130069

  9. [17]

    Video super-resolution based on deep learning: a comprehensive survey[J]

    Liu H, Ruan Z, Zhao P, et al. Video super-resolution based on deep learning: a comprehensive survey[J]. Artificial Intelligence Review, 2022, 55(8): 5981-6035

  10. [18]

    3DSRnet: Video Super-resolution using 3D Convolutional Neural Networks[J]

    Kim S Y, Lim J, Na T, et al. 3DSRnet: Video Super-resolution using 3D Convolutional Neural Networks[J]. arXiv e-prints, 2018: arXiv: 1812.09079

  11. [19]

    Action recognition method based on a novel keyframe extraction method and enhanced 3D convolutional neural network[J]

    Tian Q, Li S, Zhang Y, et al. Action recognition method based on a novel keyframe extraction method and enhanced 3D convolutional neural network[J]. International Journal of Machine Learning and Cybernetics, 2024: 1-17. 18

  12. [20]

    Learning temporal coherence via self-supervision for GAN -based video generation[J]

    Chu M, Xie Y, Mayer J, et al. Learning temporal coherence via self-supervision for GAN -based video generation[J]. ACM Transactions on Graphics (TOG), 2020, 39(4): 75: 1-75: 13

  13. [21]

    Jointly harnessing prior structures and temporal consistency for sign language video generation[J]

    Suo Y, Zheng Z, Wang X, et al. Jointly harnessing prior structures and temporal consistency for sign language video generation[J]. ACM Transactions on Multimedia Computing, Communications and Applications, 2024, 20(6): 1-18

  14. [22]

    Video super-resolution via bidirectional recurrent convolutional networks[J]

    Huang Y, Wang W, Wang L. Video super-resolution via bidirectional recurrent convolutional networks[J]. IEEE transactions on pattern analysis and machine intelligence, 2017, 40(4): 1015-1028

  15. [23]

    Residual invertible spatio-temporal network for video super- resolution[C]//Proceedings of the AAAI conference on artificial intelligence

    Zhu X, Li Z, Zhang X Y, et al. Residual invertible spatio-temporal network for video super- resolution[C]//Proceedings of the AAAI conference on artificial intelligence. 2019, 33(01): 5981-5988

  16. [24]

    Recycle-gan: Unsupervised video retargeting[C]//Proceedings of the European conference on computer vision (ECCV)

    Bansal A, Ma S, Ramanan D, et al. Recycle-gan: Unsupervised video retargeting[C]//Proceedings of the European conference on computer vision (ECCV). 2018: 119-135

  17. [25]

    Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling[J]

    Shi X, Huang Z, Wang F Y, et al. Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling[J]. arXiv preprint arXiv:2401.15977, 2024

  18. [26]

    Flownet: Learning optical flow with convolutional networks[C]//Proceedings of the IEEE international conference on computer vision

    Dosovitskiy A, Fischer P, Ilg E, et al. Flownet: Learning optical flow with convolutional networks[C]//Proceedings of the IEEE international conference on computer vision. 2015: 2758-2766

  19. [27]

    Flownet 2.0: Evolution of optical flow estimation with deep networks[C]//Proceedings of the IEEE conference on computer vision and pattern recognition

    Ilg E, Mayer N, Saikia T, et al. Flownet 2.0: Evolution of optical flow estimation with deep networks[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 2462- 2470

  20. [28]

    Raft: Recurrent all-pairs field transforms for optical flow[C]//Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16

    Teed Z, Deng J. Raft: Recurrent all-pairs field transforms for optical flow[C]//Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. Springer International Publishing, 2020: 402-419

  21. [29]

    Mocycle-gan: Unpaired video-to-video translation[C]//Proceedings of the 27th ACM international conference on multimedia

    Chen Y, Pan Y, Yao T, et al. Mocycle-gan: Unpaired video-to-video translation[C]//Proceedings of the 27th ACM international conference on multimedia. 2019: 647-655

  22. [30]

    Frame-recurrent video super-resolution[C]//Proceedings of the IEEE conference on computer vision and pattern recognition

    Sajjadi M S M, Vemulapalli R, Brown M. Frame-recurrent video super-resolution[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 6626-6634

  23. [31]

    Temporally coherent gans for video super-resolution (tecogan)[J]

    Chu M, Xie Y, Leal-Taixé L, et al. Temporally coherent gans for video super-resolution (tecogan)[J]. arXiv preprint arXiv:1811.09393, 2018, 1(2): 3

  24. [32]

    Real-time video super-resolution with spatio-temporal networks and motion compensation[C]//Proceedings of the IEEE conference on computer vision and pattern recognition

    Caballero J, Ledig C, Aitken A, et al. Real-time video super-resolution with spatio-temporal networks and motion compensation[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 4778-4787

  25. [33]

    Y., Park, T., Isola, P., & Efros, A

    Zhu, J. Y., Park, T., Isola, P., & Efros, A. A. (2017). Unpaired image-to-image translation using cycle- consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision (pp. 2223-2232)

  26. [34]

    H., Chen, K., & Wang, S

    Li, S., Han, B., Yu, Z., Liu, C. H., Chen, K., & Wang, S. (2021). I2V-GAN: Unpaired Infrared-to-Visible Video Translation. In Proceedings of the 29th ACM International Conference on Multimedia (pp. 1249-1258)

  27. [35]

    High-fidelity GAN inversion by frequency domain guidance[J]

    Liu F, Shao M, Wang F, et al. High-fidelity GAN inversion by frequency domain guidance[J]. Computers & Graphics, 2023

  28. [36]

    Wavelet knowledge distillation: Towards efficient image-to-image translation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Zhang L, Chen X, Tu X, et al. Wavelet knowledge distillation: Towards efficient image-to-image translation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 12464-12474. 19 For Table of Contents Use Only Neural -Network -Enhanced Meta...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.