{"id":"9aa86cbf-a8e4-4fdb-af1b-8cc6c0a1160c","arxiv_id":"2412.15260","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A pilot study finds GPT-4o extracts 73% of fields from photos of a lease form, with accuracy dropping from 98% on typed copies to 60% on low-quality handwritten photos.","lead":"This paper tests whether OpenAI's GPT-4o can read photographs of a filled-in Ontario lease form and pull out structured details like names and addresses. It reports promising accuracy, but the dataset is tiny, and errors grow when handwriting is messy or photos are low quality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Per-condition accuracy ordering is built on one image per cell; reported format differences are within noise and could flip with a single additional document.","rationale":"The paper's contribution is an empirical demonstration that a multi-modal LLM can extract structured fields from photos of legal forms. The most valuable and novel part is not the overall 73%, which is a pilot average, but the differential analysis across scenarios and formats in Tables 1-2; it is these differences that inform claims about image quality, handwriting, and the digital divide in Section 4.4. If those differentials are statistical noise, the paper's practical conclusions reduce to 'GPT-4o did reasonably on 15 images,' a much weaker and less actionable claim. The reader's CONDITIONAL verdict correctly targeted this. The proposed repeated-sampling test is cheap because the data-generating process, filling out a standard Ontario lease form and photographing it, is under the authors' control, and the repository is public. An alternative concern, absence of an OCR baseline, is real but less load-bearing because the claim is that multimodal LLMs can do this rather than that they beat existing methods. Similarly, the model's tendency to correct uncommon names is interesting but not the linchpin. The sampling fragility is the linchpin because all quantitative effect claims flow through it; a single additional image per condition could change the reported ordering, and the test would reveal whether the ordering is stable.","tokens_in":8493,"tokens_out":4544,"duration_ms":40080,"concrete_test":"Use the public repository to run a repeated-sampling experiment: create 10-20 independently filled lease forms per condition (S1-S3 x typed/neat/sloppy x modern/old phone), draw new names and addresses that are not reused across samples, and evaluate GPT-4o with the same prompt. Report per-condition proportions with confidence intervals clustered on form image. If the 95% intervals for neat-HD vs neat-SD or sloppy-HD vs sloppy-SD overlap, or if the format ordering is not reproduced when scenarios are re-randomized, the RQ2 conclusions should be stated as anecdotal rather than as measured effects. A single counterfactual check: randomly swap one of the three images between neat-HD and neat-SD and recompute Table 2; if the 0.74/0.69 gap changes sign, the reported ordering is fragile.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims for RQ2 (effects of image quality and handwriting) rest on a 3-scenario x 5-format design with one photographed form per cell. Each format-level accuracy in Table 2 is therefore a proportion over 3 images x 14 fields = 42 observations, but these observations are not independent: an unfavorable camera angle, one smudged field, or a model prior that 'corrects' an uncommon name can shift several fields in the same image at once. The headline ordering neat-HD 0.74 vs neat-SD 0.69 and sloppy-HD 0.64 vs sloppy-SD 0.60 corresponds to a difference of about two correct fields out of 42. Under effective sample sizes of a few images, such gaps are within sampling noise, so the qualitative narrative that image quality and neatness monotonically drive performance is not supported by the reported numbers. The authors explicitly call this an initial investigation (Section 5), but RQ2 is answered with apparent effect sizes rather than framed as hypothesis generation. The lack of confidence intervals, repeated forms, or cluster-aware error bars makes the most load-bearing part of the paper's evidence unverifiable from the data as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an initial study of using GPT-4o to extract 14 structured fields from photographs and screenshots of a standard Ontario residential tenancy agreement. Three scenarios of increasing expected difficulty (common names, less common names, and confusable/missing values) are crossed with five formats (typed PDF, neat and sloppy handwriting photographed with a modern phone, and neat and sloppy handwriting photographed with an old phone), yielding 15 images total. The overall field-level accuracy is 73%, with format accuracies of 98%, 74%, 69%, 64%, and 60% (typed HD, neat HD, neat SD, sloppy HD, sloppy SD) and scenario accuracies of 89%, 71%, and 59%. The authors conclude that the results are promising for access-to-justice applications while noting limitations such as image quality and the model's tendency to correct uncommon names to more common ones.","tokens_in":8677,"tokens_out":4925,"duration_ms":44420,"significance":"If the quantitative results were robust, this would be a useful contribution to an important problem: helping laypeople and self-represented litigants extract information from paper documents using multimodal LLMs. The paper's strengths are its clearly described experimental pipeline, the decision to use a real legal form, the release of code and data, and a scoring protocol (exact match modulo capitalization) that is transparent and easy to reproduce. The temperature-0 setting and the use of 14 well-defined fields make the measurements straightforward to audit. The main significance is as a proof-of-concept that multimodal LLMs can perform this extraction task at all, which the paper does demonstrate; however, the quantitative comparisons between formats and scenarios are not statistically supported.","major_comments":[{"comment":"The experimental design uses exactly one image per scenario-format combination (15 images total). Each format-level average in Table 2 is computed from 42 field-level observations, but these observations are clustered within only three images, and the differences between adjacent formats are very small: neat-HD (0.74) versus neat-SD (0.69) corresponds to about two correct fields out of 42, and sloppy-HD (0.64) versus sloppy-SD (0.60) is even smaller. With an effective sample size of a few images per condition, these gaps are within sampling noise. The paper nevertheless draws substantive conclusions in §4.3 (\"the quality of the results decreased when working with the printed versions\") and §4.4 (the digital-divide discussion) from this ordering. The authors should either add repeated images per condition (e.g., multiple filled copies and multiple photographs) with cluster-aware confidence intervals, or explicitly reframe the format-comparison results as exploratory and avoid making quantitative comparative claims about image quality.","section":"§3.2/§3.4, Tables 1–2"},{"comment":"The scenario variable is confounded. S2 differs from S1 in at least three ways: less common names, an additional tenant, and a missing field; S3 additionally includes names that resemble common names and multiple missing fields. Consequently, the scenario-level accuracy differences in Table 1 (89%, 71%, 59%) cannot be attributed to any single factor such as \"complex data\" or \"missing fields\" as listed in RQ2. Since RQ2 asks specifically about the effects of these factors, the current design cannot answer it. A factorial design varying one factor at a time (e.g., name frequency with all fields filled, or missing fields with common names) would be needed to support the claims in §4.2 about the causes of the performance drop.","section":"§3.2, Table 1, RQ2"},{"comment":"The field-level interpretation in §4.2 also ignores the small denominators. For example, rental_unit_street_number has an average accuracy of 0.27, which is computed over 15 images total, and the perfect scores for rental_unit_city_town and rental_unit_province are based on the same 15 images; a single misclassification would change these estimates by roughly 7 percentage points. The claim that the model benefits from pre-training to \"guess\" city and province names \"even if the image quality is lacking\" is a plausible hypothesis but not one that can be supported by these counts. The field-level results should be presented as observations subject to high uncertainty, not as established findings about which fields are intrinsically easy or hard.","section":"§4.2, Table 1"}],"minor_comments":[{"comment":"The text says \"poor lightning conditions\" in items 4 and 5; \"lighting\" is the correct word (also appears in §4.4).","section":"§3.2"},{"comment":"The caption contains typos: \"Referst to\" should be \"refers to,\" and the definition of T/P would be clearer as \"T: target value, P: predicted value.\"","section":"Figure 4 caption"},{"comment":"The caption contains typos: \"informationd\" and \"targer\" should be \"information\" and \"target.\"","section":"Figure 5 caption"},{"comment":"The claim that the model \"had no trouble locating the information on the page\" is supported only by the informal observation that erroneous outputs share some letters or numbers. A quantitative localization metric (e.g., whether the model's errors are attributable to reading a correctly identified field versus mislocating the field) would make this claim testable; as stated, it is an anecdotal inference.","section":"§4.1"},{"comment":"The scoring procedure says capitalization is ignored, but the paper does not state whether other formatting variations (e.g., extra spaces, punctuation, or different postal-code separators) were normalized. Clarifying this would improve reproducibility.","section":"§3.4"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-style paper (AI4AJ at JURIX), and as a proof-of-concept it is reasonable. For a journal version, the statistical limitations are substantial: with one image per condition, the paper's quantitative comparative claims about image quality and handwriting neatness cannot be supported. The confounded scenario design compounds the problem. The authors should either collect more data or reframe the paper's contributions as a demonstration of feasibility and a source of hypotheses. I would not reject outright because the core capability result (that GPT-4o can extract many fields from photographed forms) is credible and the dataset/code release is useful, but the paper needs a major revision to align its claims with its evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a transparent pilot that shows GPT-4o can extract a decent portion of structured data from photos of a lease form, but the per-format accuracy ordering it reports (typed > neat > sloppy, HD > SD) is built on one image per cell and is within noise. Read it as a proof-of-concept, not as an estimate of effect sizes.\n\nThe paper does a few things right. It picks a real, representative form (Ontario standard lease) and designs three scenarios of increasing difficulty—common names, uncommon names, missing fields. It varies the input format from a typed PDF to neat and sloppy handwriting photographed on a modern and an old phone. It reports per-field and per-format accuracy, shows qualitative examples of successes and failures, and releases code and data. The finding that the model 'corrects' uncommon names to more common ones (Jame to Jane, Wane to Wayne) is a genuine observation that suggests a failure mode distinct from traditional OCR. The authors are explicitly cautious in their language, calling it an initial investigation.\n\nThe soft spot is statistical. Each cell in Table 2 covers three images, and each image has 14 fields, so each percentage is a proportion over 42 observations that are not independent. A single unfavorable photo or a single field where the model pattern-matches a common name can move several fields at once. The difference between neat HD (74%) and neat SD (69%) is two correct fields out of 42; between sloppy HD and sloppy SD it is one. Without repeated images per condition, cluster-aware error bars, or baselines like OCR-plus-text-LLM, the ordering in Table 2 should be treated as hypothesis-generating, not as measured fact. The paper does not do that framing—it answers RQ2 with apparent effect sizes, even though it acknowledges the data are preliminary.\n\nFor whom: people working on access-to-justice tools or document understanding for legal forms. It is a useful existence proof that multimodal LLMs can locate and read handwritten fields, and it includes enough detail to reproduce the setup. It is also a good teaching example of why one-image-per-condition designs cannot support comparative claims. I would bring it to a reading group, but I would not cite the specific accuracy numbers without more data.\n\nMy recommendation: yes, send it to peer review. It is not fatally flawed, it is just small. A serious referee could prompt the authors to add more samples, confidence intervals, and a baseline, and the paper would be stronger for it. As it stands, it is a straightforward pilot that is honest about its scope.","headline":"A transparent pilot showing multimodal LLMs can read handwritten form fields, but its per-format accuracy ordering is noise-level with one image per cell.","tokens_in":9190,"tokens_out":3066,"would_cite":false,"duration_ms":28921,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a small test, GPT-4o extracts 73% of fields correctly from images of a paper lease form, and the model almost always finds the right field even when it misreads a value.","keywords":["multimodal large language models","legal documents","government forms","access to justice","information extraction","handwriting recognition","GPT-4o","form parsing"],"falsifier":"Take the three scenarios and five formats, generate roughly 20 independently photographed versions of each cell, run the same prompt and matching rule, and check whether the 73% average and the format ordering (98, 74, 69, 64, 60) reproduce within a few points; if the per-condition averages scatter widely or the ordering flips, the central quantitative claim is an artifact of the single-image design.","tokens_in":8288,"feed_emoji":"📸","tokens_out":9216,"duration_ms":79733,"temperature":0.7,"pith_summary":"The paper asks whether a multimodal large language model—a model that can take images as well as text as input—can replace manual transcription as the bridge between paper legal documents and AI-based legal help. Using 15 images of a standard Ontario residential lease form, covering three scenarios of increasing difficulty and five capture formats, it reports that GPT-4o extracts the correct value for 73% of 14 tracked fields on average: nearly perfect from typed PDFs (98%), and progressively worse with neat modern-phone photos (74%), neat old-phone photos (69%), sloppy modern-phone photos (64%), and sloppy old-phone photos (60%). The authors read this as evidence that the image modality is usable for access-to-justice tools, with the clear caveat that image quality and uncommon names are the main failure points.","feed_headline":"AI reads photographed lease forms at 73% field accuracy","feed_subtitle":"From 98% on typed PDFs down to 60% on sloppy old-phone photos, a photo could replace typing in legal intake.","key_machinery":"The mechanism is a zero-shot extraction pipeline: a base64-encoded image of the form is sent to GPT-4o (gpt-4o-2024-08-06) with a system prompt specifying a JSON schema for 14 fields and instructing the model to extract values exactly as they appear, using '-' for missing values, with temperature set to 0. Success is scored as exact match to the gold labels, ignoring only capitalization. The design deliberately varies scenario difficulty (common versus uncommon names, missing fields) and capture format (typed PDF, neat or sloppy handwriting, modern or old phone) so that the accuracy differences across the 15 images can be traced to image quality and the model's prior expectations about names rather than to OCR-style preprocessing.","core_discovery":"The central claim is that a single general-purpose multimodal LLM can locate and read the relevant fields on a photographed paper form without document-specific tuning. Across the three scenarios, 89%, 71%, and 59% of fields matched the gold standard, and the errors are not random: the model nearly always finds the correct field on the page, but misreads values, especially short numeric strings such as street numbers (27% correct) and uncommon names, which it tends to correct toward more common spellings (Jame to Jane, Wane to Wayne). The authors conclude that the capability is strong enough to motivate integration into access-to-justice systems, particularly with a human in the loop to verify extracted values.","pith_inferences":["If this result extends beyond leases, the same prompting pattern could be applied to court claim forms, benefit applications, and certificates; the paper's design suggests the main barrier is not the document type but the legibility of handwriting and the length of numeric strings.","The name-correction failure points toward a trade-off: a model with a strong language prior will repair messy text but also overwrite legitimate minority names, so an access-to-justice deployment should probably surface a confidence score for name fields rather than silently submitting them.","A direct comparison against traditional OCR plus a spelling-correcting language model would clarify whether end-to-end vision is the right technology; the paper's data predicts OCR would win on street numbers and the LLM would win on contextual fields like city and province.","A natural next step is a larger benchmark with repeated images per condition and confidence intervals; without it, the reported per-condition ordering (98%, 74%, 69%, 64%, 60%) is best read as suggestive rather than settled."],"forward_implications":["A photo-based intake system could let a user photograph a paper lease and have the landlord's name, tenant names, rental address, and condominium flag extracted automatically, removing the need to type those facts into a computer.","For electronic documents (typed PDFs), the 98% field accuracy suggests near-term practical deployment is plausible.","Because the model almost always locates the right field even when it misreads it, an interface that asks the user to confirm highlighted values could catch most errors cheaply.","Field-level differences imply that handwritten or low-resolution numeric fields such as street number, postal code, and condominium flag should be treated as high-risk and routed to human verification.","The gap between modern-phone and old-phone capture (74% versus 60% on sloppy forms) implies that deployment must account for the hardware of the target population, not just model capability."],"supporting_citations":[{"why":"Supplies the underlying GPT-4 model family that the paper's multimodal extraction experiments use.","marker":"[18]"},{"why":"Shows LLMs can map laypeople's factual descriptions to legal issues, the text-based pipeline that this paper extends with image input.","marker":"[15]"},{"why":"Demonstrates LLM-driven automated drafting of interactive legal applications, the form-filling use case the image extraction is meant to feed.","marker":"[9]"},{"why":"Presents the JusticeBot methodology for laypeople's legal rights, an access-to-justice system that could consume the extracted form data.","marker":"[7]"},{"why":"Introduces RateMyPDF for court-form usability analysis, an example of a legal-form tool that could be paired with image-based extraction.","marker":"[16]"}],"fun_headline_variants":["AI reads handwritten legal forms at 73% accuracy","Multimodal LLM extracts fields from photographed legal docs","Legal AI from photos: 73% field match, numbers weakest","LLM reads paper leases and court forms, aims to help laypeople"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Each per-condition accuracy figure is computed from a single photographed form (15 images total, 14 fields each), so the paper's percentages and difficulty ordering are assumed to generalize from one example per cell.","fun_headline_variants_meta":{"raw":{"variants":["AI reads handwritten legal forms at 73% accuracy","Multimodal LLM extracts fields from photographed legal docs","Legal AI from photos: 73% field match, numbers weakest","LLM reads paper leases and court forms, aims to help laypeople"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1370,"prompt_tokens":893,"completion_tokens":477,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":406}},"tokens_in":509,"tokens_out":477,"duration_ms":5472,"temperature":1.0,"reasoning_tokens":406,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:30:34.985481+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the three scenarios and five formats, generate roughly 20 independently photographed versions of each cell, run the same prompt and matching rule, and check whether the 73% average and the format ordering (98, 74, 69, 64, 60) reproduce within a few points; if the per-condition averages scatter widely or the ordering flips, the central quantitative claim is an artifact of the single-image design.","supporting_citations":[{"cited_title":"Westermann, S","cited_arxiv_id":null,"evidence_quote":"Shows LLMs can map laypeople's factual descriptions to legal issues, the text-based pipeline that this paper extends with image input."},{"cited_title":"Weaving Pathways for Justice with GPT: LLM-driven automated drafting of interactive legal applications","cited_arxiv_id":"2312.09198","evidence_quote":"Demonstrates LLM-driven automated drafting of interactive legal applications, the form-filling use case the image extraction is meant to feed."},{"cited_title":"Steenhuis, B","cited_arxiv_id":null,"evidence_quote":"Introduces RateMyPDF for court-form usability analysis, an example of a legal-form tool that could be paired with image-based extraction."}],"review_version":1}