{"id":"4bd1e1ee-1923-4376-8c2b-6c88c5849abf","arxiv_id":"2605.27978","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"ABot-OCR is a new end-to-end VLM for direct image-to-Markdown transcription using a custom data engine and structure-constrained RL optimization, reporting SOTA scores of 92.81/93.30 on OmniDocBench v1.5/v1.6.","lead":"ABot-OCR is an end-to-end vision-language model that converts a page image directly into clean Markdown in one forward pass. This could simplify document processing by removing the need for multi-step pipelines that often fail when individual modules break.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"SOTA claim on OmniDocBench rests on unreported experimental details and comparisons","rationale":"The reader's UNVERDICTED verdict stems directly from the absence of the full manuscript, which is the precise barrier to validating the strongest claim. This matches the load-bearing issue identified here; no additional internal inconsistency is detectable from the abstract alone.","tokens_in":1662,"tokens_out":256,"duration_ms":25021,"concrete_test":"Supply the full manuscript sections on experiments, OmniDocBench protocol, and result tables; verify that the reported scores match the public benchmark definition and that pipeline baselines were evaluated under identical conditions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim states specific scores (92.81 / 93.30) as SOTA among end-to-end systems and narrowing the gap to pipelines, but the provided text supplies neither the results table, the exact evaluation metric, the list of compared systems, nor any numbers for the pipeline baselines. The data engine and Decoupled Heterogeneous Document Optimization are invoked as enabling factors yet receive no technical description, equations, or ablation data. Without these elements the numerical claim cannot be checked for correctness, fairness of protocol, or absence of leakage.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces ABot-OCR, an end-to-end vision-language model that transcribes document page images directly into clean Markdown in a single forward pass. It describes a dedicated data engine for large-scale structurally consistent supervision and proposes Decoupled Heterogeneous Document Optimization, a structure-constrained reinforcement learning method intended to improve textual accuracy and enforce markup well-formedness beyond supervised fine-tuning. The paper claims state-of-the-art scores of 92.81 and 93.30 on OmniDocBench v1.5 and v1.6 among end-to-end systems, substantially narrowing the gap to strong pipeline baselines, along with robust multilingual performance across ten languages.","tokens_in":1764,"tokens_out":427,"duration_ms":38222,"significance":"If the performance claims hold under detailed scrutiny, the work would represent a meaningful step toward simplifying document parsing by replacing error-prone modular pipelines with a unified model. The combination of large-scale data engineering and RL-based structure enforcement is a plausible direction for improving fidelity in layout-sensitive tasks. The multilingual results, if quantified, would further support generalizability claims.","major_comments":[{"comment":"Abstract: The central empirical claims—that ABot-OCR achieves SOTA scores of 92.81 (v1.5) and 93.30 (v1.6) among end-to-end systems and narrows the gap to pipeline baselines—are stated without any results table, list of compared systems, definition of the evaluation metric, baseline scores, or experimental protocol. These numbers are load-bearing for the paper's primary contribution yet cannot be verified or reproduced from the provided text.","section":"Abstract"},{"comment":"Abstract: Assertions regarding the superiority of Decoupled Heterogeneous Document Optimization over supervised fine-tuning and the enabling role of the dedicated data engine are presented without technical descriptions, equations, ablation studies, or implementation details. These elements are required to substantiate how the proposed components produce the reported gains.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and the specific comments on the abstract. We agree that the abstract should better support its central claims for readers who may not immediately consult the full text. We will revise the abstract accordingly while preserving its brevity.","responses":[{"response":"We agree the abstract as written does not define the metric or list baselines. The full manuscript contains a results section with the comparison table, the OmniDocBench metric definition, and the experimental protocol. To address the concern directly in the abstract, we will add a short clause referencing the evaluation benchmark and the fact that scores are reported against both end-to-end and pipeline systems. This will make the claims verifiable from the abstract alone.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The central empirical claims—that ABot-OCR achieves SOTA scores of 92.81 (v1.5) and 93.30 (v1.6) among end-to-end systems and narrows the gap to pipeline baselines—are stated without any results table, list of compared systems, definition of the evaluation metric, baseline scores, or experimental protocol. These numbers are load-bearing for the paper's primary contribution yet cannot be verified or reproduced from the provided text."},{"response":"The abstract is a high-level summary; the technical description, equations for the structure-constrained RL objective, ablation studies, and implementation details of both the data engine and Decoupled Heterogeneous Document Optimization appear in Sections 3 and 4 of the manuscript. We will revise the abstract to include one additional sentence that briefly characterizes the data engine and the RL method, directing readers to the relevant sections for the supporting evidence.","revision_made":"yes","referee_comment":"[Abstract] Abstract: Assertions regarding the superiority of Decoupled Heterogeneous Document Optimization over supervised fine-tuning and the enabling role of the dedicated data engine are presented without technical descriptions, equations, ablation studies, or implementation details. These elements are required to substantiate how the proposed components produce the reported gains."}],"tokens_in":1376,"tokens_out":443,"duration_ms":15485,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core pitch is an end-to-end vision-language model that maps a page image to clean Markdown in one pass, backed by a custom data engine and a reinforcement learning procedure called Decoupled Heterogeneous Document Optimization. The report says this beats other end-to-end systems on OmniDocBench v1.5 and v1.6 (92.81 and 93.30) and narrows the gap to pipeline methods, plus it handles ten languages.\n\nWhat stands out as new is the specific pairing of the data engine with that named RL step for enforcing markup structure. The goal of cutting out modular error accumulation is clear enough on paper.\n\nThe problem is that none of the claims can be checked. There are no results tables, no listed baselines, no ablation numbers, no description of the data engine, and no equations or pseudocode for the RL method. The abstract just states the scores and calls the evaluations extensive. Without those pieces the SOTA assertion and the superiority over supervised fine-tuning stay unverified.\n\nThis is a short technical report, not a full methods paper. It would interest people already working on document parsing who want to test an end-to-end alternative if code or weights appear later. Right now there is not enough substance to discuss in a reading group or to cite. It does not look ready for peer review either; the central numerical claims need the missing experimental record before a referee could do anything useful with them.","headline":"ABot-OCR claims SOTA scores on OmniDocBench with an end-to-end VLM plus RL, but the abstract supplies zero tables, baselines, or method details to support any of it.","tokens_in":2254,"tokens_out":377,"would_cite":false,"duration_ms":31643,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"An end-to-end vision-language model transcribes document images to Markdown in one pass and achieves state-of-the-art end-to-end scores on OmniDocBench.","keywords":["end-to-end OCR","vision-language model","document to Markdown","reinforcement learning","OmniDocBench","multilingual text recognition","data engine"],"falsifier":"Running ABot-OCR on the OmniDocBench benchmarks and finding scores that do not reach 92.81 or 93.30, or do not narrow the gap to pipelines, would challenge the reported results.","tokens_in":2571,"feed_emoji":"📄","tokens_out":689,"duration_ms":36836,"temperature":0.7,"pith_summary":"The paper presents ABot-OCR, which processes a full page image through a single forward pass of a vision-language model to produce clean Markdown output. This design removes the need for separate modules that can introduce errors at each step. Training relies on a dedicated data engine for consistent large-scale examples and a new reinforcement learning technique called Decoupled Heterogeneous Document Optimization to ensure both accurate text and proper formatting. Results on OmniDocBench v1.5 and v1.6 show top scores among end-to-end methods at 92.81 and 93.30, reducing the difference from strong pipeline approaches. The model also performs well on text recognition in ten languages.","feed_headline":"End-to-end model scores 93.30 on document OCR benchmark","feed_subtitle":"It outputs clean Markdown from page images in one pass and narrows the gap to pipeline baselines.","key_machinery":"Decoupled Heterogeneous Document Optimization, a structure-constrained reinforcement learning method that improves textual accuracy and enforces markup well-formedness beyond supervised fine-tuning.","core_discovery":"ABot-OCR is an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By developing a dedicated data engine for large-scale, structurally consistent supervision and proposing Decoupled Heterogeneous Document Optimization as a structure-constrained reinforcement learning method, the approach sharpens textual accuracy and enforces markup well-formedness. On the OmniDocBench v1.5 and v1.6 benchmarks, ABot-OCR achieves state-of-the-art scores of 92.81 and 93.30 among all end-to-end systems, substantially narrowing the performance gap relative to strong pipeline baselines, with additional confirmation from multilingual evaluations.","pith_inferences":["The single-pass design could reduce computational overhead in high-volume document processing.","Similar techniques might apply to other tasks requiring structured output from images, such as chart understanding.","Future work could test the model on more varied document types beyond the current benchmarks."],"forward_implications":["Removes the error accumulation typical of modular pipelines by using a single forward pass.","Achieves higher scores than previous end-to-end systems on standard document benchmarks.","Demonstrates strong performance across ten diverse languages in text recognition.","Enables direct production of well-formed Markdown without additional post-processing steps."],"fun_headline_variants":["ABot-OCR scores 93.30 on OmniDocBench end-to-end","ABot-OCR transcribes images to Markdown in single pass","End-to-end model scores 92.81 and 93.30 on benchmarks","Single-pass ABot-OCR outputs clean Markdown from pages"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The dedicated data engine provides large-scale, structurally consistent supervision sufficient for the single model to handle complex document layouts effectively.","fun_headline_variants_meta":{"raw":{"variants":["ABot-OCR scores 93.30 on OmniDocBench end-to-end","ABot-OCR transcribes images to Markdown in single pass","End-to-end model scores 92.81 and 93.30 on benchmarks","Single-pass ABot-OCR outputs clean Markdown from pages"]},"model":"grok-4.3","cost_usd":0.006949,"raw_usage":{"total_tokens":3135,"prompt_tokens":657,"num_sources_used":0,"completion_tokens":76,"cost_in_usd_ticks":69490500,"prompt_tokens_details":{"text_tokens":657,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2402,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":657,"tokens_out":76,"duration_ms":28476,"temperature":1.0,"reasoning_tokens":2402,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T13:26:09.462195+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running ABot-OCR on the OmniDocBench benchmarks and finding scores that do not reach 92.81 or 93.30, or do not narrow the gap to pipelines, would challenge the reported results.","supporting_citations":[],"review_version":1}