{"id":"ed597fca-9329-4e60-a87a-dd449f1c421f","arxiv_id":"2501.01629","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An Indonesian-to-English Manhwa translation pipeline using fine-tuned YOLOv5xu, Tesseract OCR, and MarianMT reports component-level F1, CER/WER, and BLEU/METEOR scores.","lead":"This paper builds a pipeline that detects speech bubbles, reads Indonesian text with OCR, translates it to English, and overlays the translation back onto comic panels. It matters because it attempts to automate a slow manual task for a low-resource language pair, Indonesian to English.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported BLEU/METEOR scores are not shown to be computed on OCR output from Manhwa panels, so the end-to-end translation claim rests on visual inspection rather than a measured pipeline evaluation.","rationale":"The reader's weakest assumption concerns whether the OCR evaluation is representative of real Manhwa panels. I agree that this is a serious gap: Section III-B states images were 'already of satisfactory quality' and does not describe ground-truth creation, sample size, or page selection, so CER 3.1% and WER 8.6% may reflect only easy, clean bubbles. However, the single most load-bearing condition for the paper's strongest claim is the connection between the machine-translation evaluation and the actual OCR output. The OCR and translation stages are not independent in a cascading pipeline: OCR errors become translation input, and Manhwa dialogue style differs from the OpenSubtitles/Identic corpora used for fine-tuning. If BLEU 0.27 and METEOR 0.61 were computed on a held-out portion of those clean parallel corpora, they say nothing about how MarianMT performs on OCR-transcribed comic dialogue. The end-to-end result is then supported only by visual inspection in Section IV-D and Figure 3, which is not a measured quantity. This does not contradict the reader's conditional verdict; it reinforces it. The requested concrete test is minimal and should be feasible with modest annotation effort, and it would settle whether the pipeline-level claim is currently supported. I therefore keep the verdict at CONDITIONAL and see no reason to move it to accept or reject based on the present evidence.","tokens_in":4975,"tokens_out":2606,"duration_ms":29043,"concrete_test":"Build a held-out end-to-end test set of 50-100 Manhwa panels with human reference translations of the speech-bubble text, run the full pipeline (bubble detection, Tesseract OCR, MarianMT translation), and compute BLEU/METEOR and CER/WER on this real pipeline output. If the translation scores drop substantially from Table III, or if OCR errors make the comparison unstable, then the component metrics cannot be used to claim end-to-end feasibility.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the three components work well enough to support an end-to-end Indonesian-to-English Manhwa translation pipeline. Section IV-C reports BLEU 0.27 and METEOR 0.61 for MarianMT, but the paper never states what test set produced these numbers. Section II says a dataset was built from OpenSubtitles and Identic and split 80/10/10, with training over 30000 lines; no statement connects the MT test set to actual speech-bubble OCR text. Since OCR introduces errors and Manhwa dialogue is conversational and stylized, BLEU/METEOR measured on clean general-domain parallel text do not establish translation quality on the pipeline's real input. Section VI itself acknowledges 'challenges with informal language' and 'the lack of a dedicated Manhwa dataset.' Moreover, Section IV-D's only end-to-end evidence is 'visual inspection of translated samples,' which is not a quantitative check that OCR errors, bubble-boundary errors, and translation errors compound acceptably. The reader flagged OCR ground-truth representativeness as the weakest assumption; that is real, but the more load-bearing gap is that even if OCR is accurate on clean panels, the translation scores are not demonstrated to be measured on OCR-extracted Manhwa text. Without that link, the strongest claim of a working pipeline is not actually supported by the reported numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an automated Indonesian-to-English Manhwa translation pipeline combining fine-tuned YOLOv5xu for speech bubble detection, Tesseract OCR with the Indonesian language model, fine-tuned MarianMT for machine translation, and an overlay stage for reintegrating translated text into panels. The authors report component-level scores on a small Webcomics dataset and on OpenSubtitles/Identic parallel data: detection F1 90.7%, mAP@0.5:0.95 88.9%, OCR CER 3.1% and WER 8.6%, and MT BLEU 0.27 / METEOR 0.61. They conclude that the pipeline successfully automates Manhwa translation, with end-to-end success supported by visual inspection of sample outputs. The paper also compares its components against prior work and claims consistently superior performance.","tokens_in":5228,"tokens_out":4539,"duration_ms":43538,"significance":"If the reported component numbers were firmly established and linked to a measured end-to-end evaluation, the paper would provide a useful practical baseline for low-resource comic translation. The authors choose a relevant language pair and use standard, publicly available datasets, with held-out splits that avoid circular evaluation. However, the current evidence does not support the central claim of a working end-to-end pipeline: the MT scores are not tied to OCR-extracted Manhwa text, the OCR evaluation protocol is under-specified, the comparative baselines are not matched on datasets, and the only end-to-end evidence is visual inspection. The contribution is therefore a plausible component-level proof of concept rather than a validated system, and the paper would need substantial additional evaluation to justify its main conclusion.","major_comments":[{"comment":"The BLEU 0.27 and METEOR 0.61 scores are not shown to be computed on OCR-extracted Manhwa speech-bubble text. The dataset section describes only an 80/10/10 split of combined OpenSubtitles and Identic data for training/validation/testing, and no statement connects that test set to the pipeline's actual input. Since OCR errors and the informal, stylized dialogue of Manhwa differ from clean general-domain parallel text, these scores do not establish translation quality on the pipeline's real input. Please report MT metrics on the OCR output of a held-out set of Manhwa panels, or clearly justify why the clean-text test set is representative, and state the test-set size for every reported metric.","section":"§IV-C and §II"},{"comment":"The OCR evaluation lacks a specified protocol: the paper does not state how many bubbles or panels were tested, how the ground truth was created, which chapters or sources were sampled, or whether the evaluated text was actually stylized Manhwa text. Given the pipeline target, the reported CER 3.1% and WER 8.6% are strikingly low compared to published comic OCR results (e.g., the segmentation-free method in [11] reports CER 22.78% and WER 39.30%), and the claim that images were 'of satisfactory quality' is not quantified. Please provide the ground-truth collection procedure, the number of test samples, and error bars or a per-sample distribution, so that the reader can assess representativeness.","section":"§III-B and §IV-B"},{"comment":"The comparative claims in Section V are not supported because the baselines are evaluated on different datasets and tasks: CO-DETR on COCO, YOLOX-L on Manga109-s, Rigaud et al. on comic books (not necessarily Manhwa), and Dwiastuti on IWSLT spoken-language data. A comparison across different test sets cannot establish 'consistently superior performance' as claimed. Please either run the baseline methods on the same evaluation data used for the proposed pipeline or explicitly reframe the numbers as contextual references rather than as a head-to-head comparison.","section":"§V"},{"comment":"The dataset description is internally inconsistent: the paper states that the Roboflow Webcomics dataset contains 538 images, but then reports a split of 465 training and 118 validation images, which sums to 583. Additionally, no test split is described for the detection model, so the F1 score, mean precision, recall, and mAP values in Table I lack a clearly defined evaluation set. Please reconcile the dataset size and specify the exact training/validation/test split used when generating Table I.","section":"§II and §IV-A"},{"comment":"The end-to-end claim rests solely on 'visual inspection of translated samples' with no quantitative or structured human evaluation. Since the central contribution is the complete pipeline, the accumulation of detection, OCR, and translation errors should be measured on a held-out set of panels, for example through task-specific metrics (e.g., final-panel text accuracy) or a structured human rating protocol. Without such evidence, the reported component scores cannot substantiate the statement that the pipeline 'successfully integrates all steps.'","section":"§IV-D"}],"minor_comments":[{"comment":"There is a typo in Section I: 'detecting and extracting speech bubbles bubbles' should read 'detecting and extracting speech bubbles.'","section":"§I"},{"comment":"The Tesseract citation is inconsistent: the abstract and introduction cite Tesseract as [2], but Section III-B cites reference [7], which is actually the automatic manga text detection paper by Zhang et al. Please unify the citation for Tesseract.","section":"§III-B and references"},{"comment":"The notation 'Mean mAP@0.5' is confusing; the text later refers to mAP@0.5:0.95, and the relationship between mAP@0.5 = 96.3%, mean mAP@0.5 = 88.9%, and F1 = 90.7% should be clarified. Also, no confidence intervals or error bars are reported for any metric in Tables I–III.","section":"Table I"},{"comment":"The text references 'Fig. 7' as an example of the final translated image, but the manuscript contains only Figures 2 and 3. Please add the figure or correct the cross-reference.","section":"§IV-D"},{"comment":"The BLEU score is reported as 0.27 in Table III but as 27% in Section V-C; please use a single consistent scale when comparing with the 23% baseline.","section":"§V-C and Table III"},{"comment":"The paper does not describe the fine-tuning procedure for MarianMT (e.g., hyperparameters, number of epochs, base model checkpoint) or state whether code and trained models will be released. Providing these details is important for reproducibility.","section":"§II and §III-C"}],"recommendation":"major_revision","confidential_remarks":"This manuscript has the character of an undergraduate project report rather than a fully developed research paper. The topic is within the journal's scope and the component-level work is not without interest, but the central claim of a working end-to-end pipeline is not yet supported by the evidence. The main revision burden is to connect the MT evaluation to the actual OCR output, specify the OCR evaluation protocol, and either perform matched comparisons or temper the comparative claims. With those additions, a revised version could be publishable as a short application-oriented paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: this is a course-project pipeline paper, and it is exactly as modest and as limited as it looks. The one genuinely new thing is the specific combination — YOLOv5xu fine-tuned on webcomic bubble detection, Tesseract with Indonesian, MarianMT fine-tuned on OpenSubtitles+Identic — applied to Indonesian-to-English Manhwa. I don't know of a prior pipeline for that language pair, and the authors don't cite one. The system description is clear, the component numbers are plausible, and the paper is honest about the absence of a dedicated Manhwa dataset and about informal-language difficulties.\n\nThe soft spots are concentrated in the evaluation. The stress-test note is right: Section IV-C's BLEU 0.27 / METEOR 0.61 are never tied to OCR output from Manhwa panels. They look like they were computed on the clean parallel test split from OpenSubtitles/Identic. That means the end-to-end claim — that the pipeline translates comics acceptably — rests on visual inspection of a few panels, not on a measured pipeline evaluation. For a course project that is forgivable; as a published claim it is a real gap. The OCR numbers have the same problem: no description of ground-truth creation, test panel selection, or sample size, and CER 3.1% is surprisingly clean for stylized fonts. Section V's comparisons are not apples-to-apples: CO-DETR on COCO, YOLOX-L on Manga109-s, and the OCR baseline on different comic data are different tasks and domains; saying 'superior performance' overstates what the numbers show. The authors do mention dataset differences, but the tone of the section is more confident than the evidence.\n\nWhat is good: the work is clearly explained, the pipeline is a sensible integration, and the limitations are acknowledged up front. It could serve as a starting point for someone wanting a baseline for low-resource comic translation. It is not reproducible as submitted — no code, no data, no error bars — but it is not incoherent or deceptive.\n\nMy take: this is a fine student report that should not be treated as a rigorous research result. If it comes to a venue, I would not desk-reject it out of hand for a workshop or student-track venue; the integration is useful and the evaluation gaps are fixable with a modest revision that reports OCR ground truth, links MT scores to actual OCR output, and drops or re-frames the non-comparable comparisons. For a main conference, I'd send it to reviewers only if the authors release artifacts and address those gaps. As is, it is not referee-worthy at a strong venue.","headline":"A clear course-project pipeline for Indonesian Manhwa translation with plausible component scores, but the end-to-end translation claim is not actually measured and the comparison section overreaches.","tokens_in":5722,"tokens_out":2735,"would_cite":false,"duration_ms":25031,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A three-stage pipeline that chains fine-tuned YOLOv5xu bubble detection, Tesseract OCR, and MarianMT translation can translate Indonesian Manhwa panels into English automatically.","keywords":["Indonesian Manhwa translation","speech bubble detection","YOLOv5","Tesseract OCR","MarianMT","low-resource language","text overlay","machine translation"],"falsifier":"Take a held-out set of Indonesian Manhwa panels spanning multiple series, artists, and image qualities (including stylized fonts, overlapping text, and low resolution), run the pipeline on them, and compute OCR CER/WER against human transcription and translation BLEU/METEOR against professional references. If CER and WER rise substantially compared to the reported 3.1% and 8.6% on ordinary panels, or if the end-to-end translation loses speech bubbles, the claim of strong component performance is not representative.","tokens_in":4788,"feed_emoji":"🖼️","tokens_out":6737,"duration_ms":52852,"temperature":0.7,"pith_summary":"The paper claims that a fully automatic pipeline can translate Indonesian Manhwa (Korean comics translated into Indonesian) into English by chaining three fine-tuned components: YOLOv5xu to find speech bubbles, Tesseract OCR with the Indonesian language model to read the text, and MarianMT to translate it, with the translated text laid back over the panels. The authors report component-level results: F1 90.7% for bubble detection, 3.1% character error rate and 8.6% word error rate for OCR, and BLEU 0.27 with METEOR 0.61 for translation. If these numbers hold, the work shows that a practical, low-cost baseline for low-resource comic translation can be assembled from existing models and a relatively small dataset (538 annotated images for detection, plus existing parallel corpora for translation). The significance is in the workflow: tasks that typically take days by hand are reduced to hours, and the approach is presented as transferable to other underrepresented language pairs.","feed_headline":"Automated pipeline translates Indonesian Manhwa to English","feed_subtitle":"Fine-tuned bubble detection, OCR, and translation cut per-chapter work from days to hours","key_machinery":"The pipeline itself is the central mechanism: (1) fine-tuned YOLOv5xu detects speech bubbles and gives bounding boxes; (2) Tesseract OCR with the Indonesian language model (OEM 3, PSM 6) reads the extracted bubbles, with grayscale preprocessing; (3) fine-tuned MarianMT, trained on Identic and OpenSubtitles, translates the text; (4) OpenCV and Pillow overlay the translated text back into the bubble shapes. The load-bearing identity is the coupling of these components; each step's output becomes the next step's input, so the reported component scores are only as strong as the assumption that the test panels are representative.","core_discovery":"The central discovery is that a domain-specific, three-stage pipeline automatically translates Indonesian Manhwa panels to English while preserving context and artistic layout, despite the low-resource setting. The paper shows that fine-tuning a pre-trained YOLOv5xu on a 538-image webcomics dataset yields reliable speech-bubble detection (F1 90.7%), that applying Tesseract's Indonesian model directly to extracted bubbles gives low character error (3.1% CER, 8.6% WER) when panels are clean, and that fine-tuning MarianMT on Identic plus OpenSubtitles produces translations that retain meaning (METEOR 0.61) better than n-gram fidelity (BLEU 0.27) would suggest. The complete pipeline, including text overlay with OpenCV and Pillow, produced translated panels that the authors judge to keep context and meaning, demonstrating feasibility of automating Manhwa translation for a low-resource language pair.","pith_inferences":["The paper does not measure end-to-end quality on an entire chapter; a fair evaluation would track how errors propagate from missed bubbles to OCR mistakes to translation drift, since a missed bubble removes dialogue entirely.","The OCR ground-truth procedure is undocumented, so the 3.1% CER may not generalize to the stylized fonts and noisy backgrounds common in Manhwa; a controlled font/background study would define the boundary of the claim.","The component comparisons in Section V are indirect, pitting each component against different baselines on different benchmarks; a direct head-to-head on the same test set would be needed to assert superiority.","The pipeline's structure (detection → OCR → MT → overlay) is language-agnostic; the same recipe could be tried for other Southeast Asian languages, but the MarianMT step would need a similarly matched conversational parallel corpus."],"forward_implications":["A practical baseline exists for automating low-resource comic translation: each component works well enough to support the end-to-end workflow.","Translators and scanlation teams can cut per-chapter time from days to hours by using such a pipeline, though human editing would still be needed.","A relatively small detection dataset (538 images) is sufficient to fine-tune a detector for speech bubbles in Manhwa, suggesting similar results are achievable for other comic styles.","Combining a formal parallel corpus (Identic) with a conversational one (OpenSubtitles) is a viable recipe for fine-tuning a translation model for dialogue-heavy content.","The same architecture could be adapted to other low-resource language pairs by swapping the OCR language pack and the parallel corpus."],"supporting_citations":[{"why":"Supplies the YOLOv5xu object-detection model that is fine-tuned for speech-bubble detection.","marker":"[1]"},{"why":"Supplies the Tesseract OCR engine used with the Indonesian language model for text extraction.","marker":"[2]"},{"why":"Supplies the MarianMT neural translation model fine-tuned on Indonesian-English parallel data.","marker":"[3]"},{"why":"Provides the 538-image Webcomics text-selection dataset used to fine-tune YOLOv5xu for bubble detection.","marker":"[4]"},{"why":"Provides the formal Indonesian-English parallel corpus used alongside OpenSubtitles to fine-tune MarianMT.","marker":"[5]"},{"why":"Provides the conversational subtitle corpus that gives the translation model coverage of informal, dialogue-style text.","marker":"[6]"},{"why":"Serves as the general object-detection baseline (CO-DETR on COCO) that the paper compares its bubble-detection mAP against.","marker":"[9]"},{"why":"Serves as a comic-domain detection baseline (YOLOX-L on Manga109-s) used for comparison with the fine-tuned YOLOv5xu.","marker":"[10]"},{"why":"Serves as the comic-book OCR baseline whose CER and WER are compared with the Tesseract results.","marker":"[11]"},{"why":"Serves as the Indonesian-English NMT baseline whose BLEU and METEOR are compared with the fine-tuned MarianMT.","marker":"[12]"}],"fun_headline_variants":["Automated pipeline turns Indonesian Manhwa into English","Indonesian Manhwa translation: three-stage AI pipeline","AI reads and translates Indonesian Manhwa bubbles","Low-resource Manhwa translation automated end-to-end","Manhwa translation from Indonesian to English via AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported OCR error rates assume the test panels are clean and the stylized font difficulty is low, because the paper does not describe how OCR ground truth was selected, which panels were tested, or how many samples produced the 3.1% CER and 8.6% WER.","fun_headline_variants_meta":{"raw":{"variants":["Automated pipeline turns Indonesian Manhwa into English","Indonesian Manhwa translation: three-stage AI pipeline","AI reads and translates Indonesian Manhwa bubbles","Low-resource Manhwa translation automated end-to-end","Manhwa translation from Indonesian to English via AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000263,"raw_usage":{"total_tokens":1559,"prompt_tokens":866,"completion_tokens":693,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":617}},"tokens_in":482,"tokens_out":693,"duration_ms":6204,"temperature":1.0,"reasoning_tokens":617,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:23:10.165296+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out set of Indonesian Manhwa panels spanning multiple series, artists, and image qualities (including stylized fonts, overlapping text, and low resolution), run the pipeline on them, and compute OCR CER/WER against human transcription and translation BLEU/METEOR against professional references. If CER and WER rise substantially compared to the reported 3.1% and 8.6% on ordinary panels, or if the end-to-end translation loses speech bubbles, the claim of strong component performance is not representative.","supporting_citations":[{"cited_title":"Tesseract OCR Engine,","cited_arxiv_id":null,"evidence_quote":"Supplies the Tesseract OCR engine used with the Indonesian language model for text extraction."},{"cited_title":"Marian: Fast neural machine translation in c++,","cited_arxiv_id":null,"evidence_quote":"Supplies the MarianMT neural translation model fine-tuned on Indonesian-English parallel data."},{"cited_title":"Webcomics text selection dataset,","cited_arxiv_id":null,"evidence_quote":"Provides the 538-image Webcomics text-selection dataset used to fine-tune YOLOv5xu for bubble detection."},{"cited_title":"Identic: The indonesian-english parallel corpus,","cited_arxiv_id":null,"evidence_quote":"Provides the formal Indonesian-English parallel corpus used alongside OpenSubtitles to fine-tune MarianMT."},{"cited_title":"Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles,","cited_arxiv_id":null,"evidence_quote":"Provides the conversational subtitle corpus that gives the translation model coverage of informal, dialogue-style text."},{"cited_title":"Detrs with collaborative hybrid assign- ments training,","cited_arxiv_id":null,"evidence_quote":"Serves as the general object-detection baseline (CO-DETR on COCO) that the paper compares its bubble-detection mAP against."},{"cited_title":"Usb: Universal-scale object detection benchmark,","cited_arxiv_id":null,"evidence_quote":"Serves as a comic-domain detection baseline (YOLOX-L on Manga109-s) used for comparison with the fine-tuned YOLOv5xu."},{"cited_title":"Segmentation-free speech text recognition for comic books,","cited_arxiv_id":null,"evidence_quote":"Serves as the comic-book OCR baseline whose CER and WER are compared with the Tesseract results."},{"cited_title":"English-Indonesian neural machine translation for spoken language domains,","cited_arxiv_id":null,"evidence_quote":"Serves as the Indonesian-English NMT baseline whose BLEU and METEOR are compared with the fine-tuned MarianMT."}],"review_version":1}