REVIEW 3 major objections 3 minor 59 references
Translationese is a rational response to the cognitive difficulty of the translation task and can be partly predicted from measures of that difficulty.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 22:28 UTC pith:Y37OOXIM
load-bearing objection We only have the abstract for the translationese claim; the supplied full text is a different paper (h-transform visual generation), so the operationalizations cannot be checked. the 3 major comments →
Translationese as a Rational Response to Translation Task Difficulty
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Observable translationese, operationalised as a segment-level translatedness score from an automatic classifier, can be partly predicted from quantifiable translation-task difficulty that comprises source-text complexity and cross-lingual transfer difficulty. The link is clearer for English-to-German than the reverse; transfer difficulty contributes more than source complexity in most settings; and source-text syntactic complexity plus translation-solution entropy are the most consistent predictors across language pairs and written/spoken modes.
What carries the argument
Translation task difficulty, defined as source-text complexity plus cross-lingual transfer difficulty and measured chiefly by LLM-surprisal (information-theoretic) metrics together with syntactic and semantic alternatives; this difficulty is used to predict the automatic classifier’s segment-level translatedness score.
Load-bearing premise
The automatic classifier’s translatedness score truly indexes cognitive-load-driven translationese, and LLM surprisal plus translation-solution entropy truly measure the cognitive difficulty of translating rather than shared surface cues or model artefacts.
What would settle it
If, after matching or controlling for genre, register and surface features that both the classifier and the surprisal metrics can exploit, the difficulty predictors no longer account for variance in human ratings of cognitive effort or of “how translated” a segment feels, the central claim fails.
If this is right
- Segments whose sources are syntactically complex or whose possible translations have high entropy should show stronger translationese.
- Cross-lingual transfer difficulty, more than raw source complexity, should drive divergence from original target-language norms.
- Information-theoretic surprisal features can stand in for traditional syntactic features in written translation but not necessarily in spoken mode.
- Directionality matters: English-to-German yields clearer difficulty–translationese links than German-to-English.
- Quality estimation or post-editing tools can use these difficulty predictors to anticipate where translationese (and related effort) will appear.
Where Pith is reading between the lines
- If cognitive load is the common cause, interventions that lower source complexity or ease transfer (glossaries, constrained decoding, better MT suggestions) should reduce classifier-detected translationese.
- The same difficulty metrics should also predict process measures such as pauses, fixations or post-editing time if the load account is general.
- Spoken mode’s lack of benefit from information-theoretic features suggests real-time production constraints can override surprisal-based difficulty.
- Without independent human validation of both the classifier and the surprisal measures, the predictive link risks being partly circular through shared surface cues.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript (as represented by its abstract for arXiv:2603.12050) claims that translationese is a rational response to cognitive load in the translation task. Translationese is operationalised as segment-level translatedness scores from an automatic classifier; task difficulty is split into source-text complexity and cross-lingual transfer difficulty, measured mainly by LLM surprisal and related information-theoretic quantities, with syntactic and semantic alternatives. On a bidirectional English–German corpus (written and spoken), the authors report that difficulty partly predicts translatedness, more strongly for English→German; transfer difficulty usually outweighs source complexity; information-theoretic features match or beat traditional ones in written mode only; and source syntactic complexity plus translation-solution entropy are the strongest predictors across pairs and modes.
Significance. If the operationalisations are valid and the partial predictive relationships hold with adequate effect sizes and controls, the paper would offer a unified, testable cognitive-load account of translationese that links translation studies to information-theoretic difficulty measures. That would matter for translation-process research, MT evaluation, and bilingual production models. The abstract alone, however, does not establish those conditions, and the supplied full-text body is an unrelated computer-vision paper (Weighted h-Transform Sampling, arXiv:2603.12057), so the claimed contribution cannot be verified from the materials provided for review.
major comments (3)
- The full manuscript text supplied for review is not the translationese paper: it is “Coarse-Guided Visual Generation via Weighted h-Transform Sampling” (arXiv:2603.12057). None of the classifier training, surprisal definitions, regression specifications, corpus construction, or EN–DE written/spoken results exist in the provided body. Load-bearing claims (partial explanation, relative contribution of transfer vs source difficulty, superiority of syntactic complexity and translation-solution entropy) therefore cannot be checked. Review of the central claim is blocked until the correct full text is provided.
- From the abstract alone: the claim that translatedness is a rational response to cognitive load rests on two operationalisations whose independence is not shown—(i) an automatic translatedness classifier and (ii) LLM-surprisal / translation-solution entropy difficulty metrics. If both are driven by shared surface or model-based cues (syntax, register, LLM-like distributions), the reported prediction can be circular. The abstract does not report independence tests, residualisation, or out-of-model controls; these are required for the causal/cognitive interpretation.
- Abstract results are only directional (“partly explained,” “especially English-to-German,” “for most experiments”). Without effect sizes, R² / variance explained, confidence intervals, or model comparison against genre/register baselines, it is impossible to judge whether the partial explanation is scientifically meaningful or merely statistically detectable. This is load-bearing for the unified-account claim.
minor comments (3)
- Abstract does not name the classifier architecture, training data, or how translatedness scores are calibrated against human judgments of translationese.
- “Translation-solution entropy” is listed as a strongest predictor but is not defined in the abstract; a one-sentence operational definition would help readers assess construct validity.
- Bidirectional EN–DE design is a strength; the abstract should state whether spoken and written subcorpora are balanced for domain and whether mode is modelled as a fixed effect or analysed separately only.
Circularity Check
No verifiable circularity: abstract treats translatedness and difficulty as separate operationalizations; full text is a mismatched CV paper so no derivation chain can be checked.
full rationale
The translationese abstract claims an empirical prediction: segment-level translatedness (automatic classifier) is partly explained by independently operationalized task difficulty (source-text complexity plus cross-lingual transfer, mainly LLM surprisal, plus syntactic/semantic alternatives). Difficulty is not defined as translatedness, nor is a parameter fitted to translatedness then re-reported as a prediction of the same quantity. No equations, classifier training details, surprisal definitions, or regression specs appear in the supplied materials. The CACHEABLE full manuscript is an unrelated paper (Weighted h-Transform Sampling, arXiv 2603.12057) and contains none of the load-bearing operationalizations. Under the rule that circularity may be claimed only when a specific reduction can be quoted, no circular step can be exhibited. Shared-cue confounds between classifier and surprisal would be a validity concern, not circularity by construction. Score 0; steps empty.
Axiom & Free-Parameter Ledger
free parameters (2)
- Classifier and regression hyperparameters (unspecified)
- Weighting / relative contribution of source vs transfer difficulty
axioms (4)
- domain assumption Segment-level automatic translatedness scores validly measure translationese as studied in the literature.
- domain assumption LLM surprisal and related information-theoretic metrics index cognitive load of source processing and cross-lingual transfer.
- domain assumption English–German bidirectional written and spoken subcorpora are representative enough to support claims about language-pair and mode effects.
- standard math Standard statistical / predictive modeling assumptions for relating difficulty features to translatedness hold.
invented entities (1)
-
Translation task difficulty as dual source-text + cross-lingual transfer construct
no independent evidence
read the original abstract
Translations systematically diverge from texts originally produced in the target language, a phenomenon widely referred to as translationese. Translationese has been attributed to production tendencies (e.g. interference, simplification), socio-cultural variables, and language-pair effects, yet a unified explanatory account is still lacking. We propose that translationese reflects cognitive load inherent in the translation task itself. We test whether observable translationese can be predicted from quantifiable measures of translation task difficulty. Translationese is operationalised as a segment-level translatedness score produced by an automatic classifier. Translation task difficulty is conceptualised as comprising source-text and cross-lingual transfer components, operationalised mainly through information-theoretic metrics based on LLM surprisal, complemented by established syntactic and semantic alternatives. We use a bidirectional English-German corpus comprising written and spoken subcorpora. Results indicate that translationese can be partly explained by translation task difficulty, especially in English-to-German. For most experiments, cross-lingual transfer difficulty contributes more than source-text complexity. Information-theoretic indicators match or outperform traditional features in written mode, but offer no advantage in spoken mode. Source-text syntactic complexity and translation-solution entropy emerged as the strongest predictors of translationese across language pairs and modes.
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2111.13606 (2021)
Batzolis, G., Stanczuk, J., Schönlieb, C.B., Etmann, C.: Conditional image gener- ation with score-based diffusion models. arXiv preprint arXiv:2111.13606 (2021)
Pith/arXiv arXiv 2021
-
[2]
arXiv preprint arXiv:2410.02073 (2024)
Bochkovskii, A., Delaunoy, A., Germain, H., Santos, M., Zhou, Y., Richter, S.R., Koltun, V.: Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073 (2024)
Pith/arXiv arXiv 2024
-
[3]
In: Proceedings of the Computer Vision and Pattern Recognition Conference
Burgert, R., Xu, Y., Xian, W., Pilarski, O., Clausen, P., He, M., Ma, L., Deng, Y., Li, L., Mousavi, M., et al.: Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 13–23 (2025)
2025
-
[4]
IEEE Transactions on Computational Imaging3(1), 84–98 (2016)
Chan, S.H., Wang, X., Elgendy, O.A.: Plug-and-play admm for image restoration: Fixed-point convergence and applications. IEEE Transactions on Computational Imaging3(1), 84–98 (2016)
2016
-
[5]
arXiv preprint arXiv:2108.02938 (2021)
Choi, J., Kim, S., Jeong, Y., Gwon, Y., Yoon, S.: Ilvr: Conditioning method for denoising diffusion probabilistic models. arXiv preprint arXiv:2108.02938 (2021)
Pith/arXiv arXiv 2021
-
[6]
arXiv preprint arXiv:2209.14687 (2022)
Chung, H., Kim, J., Mccann, M.T., Klasky, M.L., Ye, J.C.: Diffusion posterior sam- pling for general noisy inverse problems. arXiv preprint arXiv:2209.14687 (2022)
Pith/arXiv arXiv 2022
-
[7]
Advances in Neural Information Processing Systems35, 25683–25696 (2022)
Chung, H., Sim, B., Ryu, D., Ye, J.C.: Improving diffusion models for inverse problems using manifold constraints. Advances in Neural Information Processing Systems35, 25683–25696 (2022)
2022
-
[8]
Advances in neural informa- tion processing systems34, 17695–17709 (2021)
De Bortoli, V., Thornton, J., Heng, J., Doucet, A.: Diffusion schrödinger bridge with applications to score-based generative modeling. Advances in neural informa- tion processing systems34, 17695–17709 (2021)
2021
-
[9]
Advances in neural information processing systems34, 8780–8794 (2021)
Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis. Advances in neural information processing systems34, 8780–8794 (2021)
2021
-
[10]
Doob, J.L., Doob, J.: Classical potential theory and its probabilistic counterpart, vol. 262. Springer (1984)
1984
-
[11]
In: Forty-first international conference on machine learning (2024)
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al.: Scaling rectified flow transformers for high-resolution image synthesis. In: Forty-first international conference on machine learning (2024)
2024
-
[12]
In: Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision
Gao, Z., Song, J., Zhang, Z., Deng, J., Patras, I.: Frequency-guided diffusion for training-free text-driven image translation. In: Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision. pp. 19195–19205 (2025)
2025
-
[13]
Advances in neural information processing systems30(2017)
Heusel,M.,Ramsauer,H.,Unterthiner,T.,Nessler,B.,Hochreiter,S.:Ganstrained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems30(2017)
2017
-
[14]
Advances in neural information processing systems33, 6840–6851 (2020)
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Advances in neural information processing systems33, 6840–6851 (2020)
2020
-
[15]
arXiv preprint arXiv:2207.12598 (2022)
Ho, J., Salimans, T.: Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)
Pith/arXiv arXiv 2022
-
[16]
arXiv preprint arXiv:2410.21276 (2024)
Hurst, A., Lerer, A., Goucher, A.P., Perelman, A., Ramesh, A., Clark, A., Os- trow, A., Welihinda, A., Hayes, A., Radford, A., et al.: Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)
Pith/arXiv arXiv 2024
-
[17]
In: The Twelfth International Conference on Learning Representations (2023)
Ju, X., Zeng, A., Bian, Y., Liu, S., Xu, Q.: Pnp inversion: Boosting diffusion-based editing with 3 lines of code. In: The Twelfth International Conference on Learning Representations (2023)
2023
-
[18]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Karras, T., Laine, S., Aila, T.: A style-based generator architecture for generative adversarial networks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 4401–4410 (2019) 16 Y. Wang, Z. Jiang, Z. Wang, L. Chen
2019
-
[19]
Advances in neural information processing systems35, 23593–23606 (2022)
Kawar, B., Elad, M., Ermon, S., Song, J.: Denoising diffusion restoration models. Advances in neural information processing systems35, 23593–23606 (2022)
2022
-
[20]
arXiv preprint arXiv:2505.23145 (2025)
Kim, J., Hong, Y., Ye, J.C.: Flowalign: Trajectory-regularized, inversion-free flow- based image editing. arXiv preprint arXiv:2505.23145 (2025)
Pith/arXiv arXiv 2025
-
[21]
arXiv preprint arXiv:2503.08136 (2025)
Kim, J., Kim, B.S., Ye, J.C.: Flowdps: Flow-driven posterior sampling for inverse problems. arXiv preprint arXiv:2503.08136 (2025)
Pith/arXiv arXiv 2025
-
[22]
arXiv preprint arXiv:2412.08629 (2024)
Kulikov, V., Kleiner, M., Huberman-Spiegelglas, I., Michaeli, T.: Flowedit: Inversion-free text-based editing using pre-trained flow models. arXiv preprint arXiv:2412.08629 (2024)
Pith/arXiv arXiv 2024
-
[23]
Labs, B.F.: Flux.https://github.com/black-forest-labs/flux(2024)
2024
-
[24]
1 kontext: Flow match- ing for in-context image generation and editing in latent space
Labs, B.F., Batifol, S., Blattmann, A., Boesel, F., Consul, S., Diagne, C., Dock- horn, T., English, J., English, Z., Esser, P., et al.: Flux. 1 kontext: Flow match- ing for in-context image generation and editing in latent space. arXiv preprint arXiv:2506.15742 (2025)
Pith/arXiv arXiv 2025
-
[25]
In: Proceedings of the IEEE/CVF international conference on computer vision
Liang, J., Cao, J., Sun, G., Zhang, K., Van Gool, L., Timofte, R.: Swinir: Image restoration using swin transformer. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 1833–1844 (2021)
2021
-
[26]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Ling, L., Sheng, Y., Tu, Z., Zhao, W., Xin, C., Wan, K., Yu, L., Guo, Q., Yu, Z., Lu, Y., et al.: Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 22160–22169 (2024)
2024
-
[27]
arXiv preprint arXiv:2210.02747 (2022)
Lipman, Y., Chen, R.T., Ben-Hamu, H., Nickel, M., Le, M.: Flow matching for generative modeling. arXiv preprint arXiv:2210.02747 (2022)
Pith/arXiv arXiv 2022
-
[28]
arXiv preprint arXiv:2209.03003 (2022)
Liu, X., Gong, C., Liu, Q.: Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003 (2022)
Pith/arXiv arXiv 2022
-
[29]
Interna- tional journal of computer vision60(2), 91–110 (2004)
Lowe, D.G.: Distinctive image features from scale-invariant keypoints. Interna- tional journal of computer vision60(2), 91–110 (2004)
2004
-
[30]
Entropy22(8), 802 (2020)
Maoutsa, D., Reich, S., Opper, M.: Interacting particle solutions of fokker–planck equations through gradient–log–density estimation. Entropy22(8), 802 (2020)
2020
-
[31]
arXiv preprint arXiv:2108.01073 (2021)
Meng, C., Song, Y., Song, J., Wu, J., Zhu, J.Y., Ermon, S.: Sdedit: Image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073 (2021)
Pith/arXiv arXiv 2021
-
[32]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Murez, Z., Kolouri, S., Kriegman, D., Ramamoorthi, R., Kim, K.: Image to image translation for domain adaptation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4500–4509 (2018)
2018
-
[33]
arXiv preprint arXiv:2304.07193 (2023)
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al.: Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023)
Pith/arXiv arXiv 2023
-
[34]
arXiv preprint arXiv:2403.12036 (2024)
Parmar, G., Park, T., Narasimhan, S., Zhu, J.Y.: One-step image translation with text-to-image models. arXiv preprint arXiv:2403.12036 (2024)
Pith/arXiv arXiv 2024
-
[35]
arXiv preprint arXiv:2307.01952 (2023)
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952 (2023)
Pith/arXiv arXiv 2023
-
[36]
In: International conference on machine learning
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PmLR (2021)
2021
-
[37]
In: Proceedings of the IEEE/CVF interna- tional conference on computer vision
Ren, M., Delbracio, M., Talebi, H., Gerig, G., Milanfar, P.: Multiscale structure guided diffusion for image deblurring. In: Proceedings of the IEEE/CVF interna- tional conference on computer vision. pp. 10721–10733 (2023) Coarse-guided-Gen 17
2023
-
[38]
Cambridge university press (2000)
Rogers,L.C.G.,Williams,D.:Diffusions,Markovprocesses,andmartingales,vol.2. Cambridge university press (2000)
2000
-
[39]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10684–10695 (2022)
2022
-
[40]
Advances in neural information processing systems37, 6246– 6266 (2024)
Shen, F., Tang, J.: Imagpose: A unified conditional framework for pose-guided person generation. Advances in neural information processing systems37, 6246– 6266 (2024)
2024
-
[41]
arXiv preprint arXiv:2511.08633 (2025)
Singer, A., Rotstein, N., Mann, A., Kimmel, R., Litany, O.: Time-to-move: Training-free motion controlled video generation via dual-clock denoising. arXiv preprint arXiv:2511.08633 (2025)
arXiv 2025
-
[42]
arXiv preprint arXiv:2010.02502 (2020)
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502 (2020)
Pith/arXiv arXiv 2010
-
[43]
Advances in neural information processing systems32(2019)
Song, Y., Ermon, S.: Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems32(2019)
2019
-
[44]
arXiv preprint arXiv:2011.13456 (2020)
Song, Y., Sohl-Dickstein, J., Kingma, D.P., Kumar, A., Ermon, S., Poole, B.: Score- based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456 (2020)
Pith/arXiv arXiv 2011
-
[45]
arXiv preprint arXiv:2203.08382 (2022)
Su, X., Song, J., Meng, C., Ermon, S.: Dual diffusion implicit bridges for image- to-image translation. arXiv preprint arXiv:2203.08382 (2022)
Pith/arXiv arXiv 2022
-
[46]
SIAM (2005)
Tarantola, A.: Inverse problem theory and methods for model parameter estima- tion. SIAM (2005)
2005
-
[47]
In: European conference on computer vision
Teed, Z., Deng, J.: Raft: Recurrent all-pairs field transforms for optical flow. In: European conference on computer vision. pp. 402–419. Springer (2020)
2020
-
[48]
arXiv preprint arXiv:1812.01717 (2018)
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., Gelly, S.: Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717 (2018)
Pith/arXiv arXiv 2018
-
[49]
arXiv preprint arXiv:2503.20314 (2025)
Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.W., Chen, D., Yu, F., Zhao, H., Yang, J., et al.: Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 (2025)
Pith/arXiv arXiv 2025
-
[50]
arXiv preprint arXiv:2205.12952 (2022)
Wang, T., Zhang, T., Zhang, B., Ouyang, H., Chen, D., Chen, Q., Wen, F.: Pretraining is all you need for image-to-image translation. arXiv preprint arXiv:2205.12952 (2022)
Pith/arXiv arXiv 2022
-
[51]
IEEE transactions on pattern analysis and machine intelligence43(10), 3365–3387 (2020)
Wang, Z., Chen, J., Hoi, S.C.: Deep learning for image super-resolution: A survey. IEEE transactions on pattern analysis and machine intelligence43(10), 3365–3387 (2020)
2020
-
[52]
arXiv preprint arXiv:2406.03293 (2024)
Yang, X., Chen, C., Yang, X., Liu, F., Lin, G.: Text-to-image rectified flow as plug-and-play priors. arXiv preprint arXiv:2406.03293 (2024)
Pith/arXiv arXiv 2024
-
[53]
arXiv preprint arXiv:2408.06072 (2024)
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al.: Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072 (2024)
Pith/arXiv arXiv 2024
-
[54]
Yu, J., Wang, Y., Zhao, C., Ghanem, B., Zhang, J.: Freedom: Training-free energy- guidedconditionaldiffusionmodel.In:ProceedingsoftheIEEE/CVFInternational Conference on Computer Vision. pp. 23174–23184 (2023)
2023
-
[55]
arXiv preprint arXiv:2409.19365 (2024)
Zhan, Z., Chen, D., Mei, J.P., Zhao, Z., Chen, J., Chen, C., Lyu, S., Wang, C.: Conditional image synthesis with diffusion models: A survey. arXiv preprint arXiv:2409.19365 (2024)
Pith/arXiv arXiv 2024
-
[56]
In: Proceedings of the IEEE/CVF international conference on computer vision
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 3836–3847 (2023) 18 Y. Wang, Z. Jiang, Z. Wang, L. Chen
2023
-
[57]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 586–595 (2018)
2018
-
[58]
arXiv preprint arXiv:2405.15885 (2024)
Zheng, K., He, G., Chen, J., Bao, F., Zhu, J.: Diffusion bridge implicit models. arXiv preprint arXiv:2405.15885 (2024)
Pith/arXiv arXiv 2024
-
[59]
Zhou, L., Lou, A., Khanna, S., Ermon, S.: Denoising diffusion bridge models. arXiv preprint arXiv:2309.16948 (2023) Coarse-guided-Gen 19 Appendix This appendix is organized as follows: –Section A.1 gives the proof of equivalent marginal distributions between the reversed-SDE and its corresponding PF-ODE. –Section A.2 demonstrates the detailed derivation o...
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.