{"id":"1cb08d7e-cfd2-4e53-9b71-e96d3b0119a2","arxiv_id":"2605.29897","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"ExCAM is a new explainable metric for cultural awareness in LLMs, trained on the ExCAM40k dataset derived from nine existing benchmarks with added synthetic errors, achieving up to 80% error detection accuracy.","lead":"The paper introduces ExCAM, a metric that detects, rates, and explains cultural errors in AI-generated text from instruction-output pairs. This could help make AI systems fairer across cultures by enabling detailed, actionable feedback instead of overall scores.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Synthetic errors in ExCAM40k may not match real cultural errors LLMs produce in free text","rationale":"The load-bearing assumption identified by the reader is exactly the one required for the reported accuracy to support the broader claim about free-text evaluation. Because the paper's evaluation remains internal to the synthetic construction, the concern stands and keeps the verdict at UNVERDICTED.","tokens_in":1648,"tokens_out":288,"duration_ms":14566,"concrete_test":"Sample 200 real LLM generations on the same cultural topics as the source benchmarks, have two independent annotators label cultural errors, then run the trained ExCAM model on these pairs and report precision/recall; if detection rate drops below 60% the synthetic-to-real generalization claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (up to 80% error detection on a balanced test set) is measured exclusively on instruction-output pairs created by taking existing benchmarks and injecting synthetic errors. The central claim that ExCAM provides a useful metric for free-text cultural evaluation therefore requires that the synthetic error distribution is representative of the errors LLMs actually make when generating unconstrained text. No independent human validation of that match is described in the abstract, and the reformatting step from QA-style benchmarks to free-text pairs introduces an additional distributional shift whose effect on error realism is untested.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces ExCAM, an explainable metric for identifying, rating, and explaining cultural errors in LLM instruction-output pairs. It constructs the ExCAM40k dataset from nine existing benchmarks by reformatting them and augmenting with synthetic errors, then reports that ExCAM achieves the highest error detection rate with up to 80% accuracy on a balanced test set, outperforming baselines including GPT-5.","tokens_in":1771,"tokens_out":365,"duration_ms":19386,"significance":"If the synthetic errors prove representative of real LLM cultural errors in free-text generation, ExCAM would offer a scalable, fine-grained alternative to human-annotated benchmarks for cultural awareness evaluation, addressing the noted scarcity of free-text metrics.","major_comments":[{"comment":"Abstract: the central claim of up to 80% accuracy supplies no information on training procedure, model architecture, error types, baseline implementations, or statistical significance, rendering the performance result unevaluable.","section":"Abstract"},{"comment":"Dataset construction: training and testing both use the same synthetic-error-augmented ExCAM40k; without human validation that the injected errors match the distribution of cultural errors LLMs produce in unconstrained free-text generation, the reported accuracy cannot be interpreted as evidence of genuine generalization rather than fitting to the augmentation process. The reformatting step from QA-style benchmarks to free-text pairs introduces an additional untested distributional shift.","section":"Dataset Construction (ExCAM40k)"}],"minor_comments":[{"comment":"The abstract states that nine existing benchmarks are used but neither names them nor cites the original sources.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment below and indicate where revisions will be made to improve clarity and acknowledge limitations.","responses":[{"response":"We agree that the abstract is too concise and omits key methodological details needed to evaluate the 80% accuracy claim. In the revised manuscript, we will expand the abstract to briefly describe the ExCAM model architecture and training procedure on ExCAM40k, the categories of synthetic cultural errors, the specific baselines including GPT-5, and any statistical significance results.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim of up to 80% accuracy supplies no information on training procedure, model architecture, error types, baseline implementations, or statistical significance, rendering the performance result unevaluable."},{"response":"We acknowledge this limitation. ExCAM40k is built by reformatting nine existing benchmarks and injecting synthetic errors to enable scalable training and evaluation, as large-scale human annotation of real free-text cultural errors is resource-intensive. Training and testing occur on this augmented dataset to measure detection of the defined error types. However, we agree that this setup does not constitute direct evidence of generalization to real LLM free-text outputs or fully account for reformatting shifts. We will revise the paper to add an explicit limitations section discussing these points, clarify that results are specific to the synthetic test set, and outline future work involving human validation on real generations.","revision_made":"partial","referee_comment":"[Dataset Construction (ExCAM40k)] Dataset construction: training and testing both use the same synthetic-error-augmented ExCAM40k; without human validation that the injected errors match the distribution of cultural errors LLMs produce in unconstrained free-text generation, the reported accuracy cannot be interpreted as evidence of genuine generalization rather than fitting to the augmentation process. The reformatting step from QA-style benchmarks to free-text pairs introduces an additional untested distributional shift."}],"tokens_in":1287,"tokens_out":431,"duration_ms":20795,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"ExCAM claims to be the first dedicated explainable metric for cultural errors in free-text LLM outputs. The authors build ExCAM40k by reformatting nine existing benchmarks into instruction-output pairs and injecting synthetic errors, then report up to 80% error detection accuracy on a balanced test set, beating several baselines including GPT-5.\n\nThe useful piece is the focus on free text plus explainability. Most cultural benchmarks stay in QA format, and the paper correctly notes the cost of human annotation, so an automated metric that also explains errors could be practical if it works.\n\nThe bigger problems are the evaluation setup and missing details. All results sit on the same synthetic-error data, with no human validation that the injected errors match the cultural mistakes models actually produce in unconstrained generation. The abstract supplies no information on model architecture, training procedure, error taxonomy, baseline implementations, or statistical tests, so the performance number cannot be assessed. This leaves open the possibility that the model is mainly learning the synthetic construction process rather than cultural patterns.\n\nThe work is aimed at people building or auditing culturally aware LLMs who need scalable evaluation tools. A reader in that area could borrow the dataset construction idea, but would treat the accuracy claim as preliminary until the synthetic-to-real gap is checked.\n\nI would send it for peer review so referees can look at the methods section and any added validation experiments rather than desk-rejecting it on the abstract alone.","headline":"ExCAM's 80% accuracy is measured only on synthetic errors whose realism for actual LLM free-text mistakes is untested.","tokens_in":2243,"tokens_out":362,"would_cite":false,"duration_ms":21342,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ExCAM is the first metric to identify, rate, and explain cultural errors in LLM instruction-output pairs.","keywords":["explainable cultural awareness","LLM metrics","cultural errors","error detection","synthetic dataset","free text evaluation","cultural fairness"],"falsifier":"A study applying ExCAM to a collection of real-world LLM outputs containing human-verified cultural errors and measuring detection accuracy would falsify the result if it is substantially lower than 80%.","tokens_in":2573,"feed_emoji":"🌐","tokens_out":601,"duration_ms":30585,"temperature":0.7,"pith_summary":"This paper presents ExCAM, a metric designed to evaluate the cultural awareness of large language models in free text generation. Creating cultural benchmarks has been expensive due to human annotations, and explainable metrics for free text are rare. The authors build ExCAM40k by enhancing nine existing benchmarks with synthetic errors and train ExCAM on it. ExCAM outperforms baselines including GPT-5, reaching 80% accuracy in detecting errors on a balanced test set. This development supports more accessible and transparent evaluation of cultural fairness in AI outputs.","feed_headline":"ExCAM achieves 80% accuracy detecting cultural errors in AI text","feed_subtitle":"Outperforms GPT-5 with explanations for mistakes in language model generations","key_machinery":"ExCAM, the explainable cultural awareness metric for detecting and explaining cultural errors in LLM outputs.","core_discovery":"ExCAM is, to our knowledge, the first dedicated evaluation metric that identifies, rates and explains cultural errors in instruction-output pairs. To train and evaluate ExCAM, we introduce ExCAM40k, a dataset comprised of nine existing benchmarks that we reformat and enhance with synthetic errors. Compared to several baselines, including GPT-5, ExCAM achieves the highest error detection rate with up to 80% accuracy on a balanced test set. Therefore, ExCAM opens the pathway towards fine-grained and explainable cultural evaluation of free text.","pith_inferences":["The method of augmenting benchmarks with synthetic errors could be used for other types of model evaluation.","ExCAM might be adapted to evaluate cultural awareness in languages other than those in the dataset.","Future work could test if using ExCAM during training improves model performance on cultural tasks."],"forward_implications":["Enables fine-grained evaluation of cultural awareness beyond question answering.","Provides explanations for detected errors to aid understanding.","Lowers the barrier to creating cultural evaluation benchmarks.","Supports generalizability of LLM applications across cultures."],"fun_headline_variants":["ExCAM detects AI cultural errors with 80 percent accuracy","ExCAM rates and explains cultural mistakes in AI outputs","ExCAM40k enables error detection at up to 80 percent accuracy","ExCAM outperforms GPT-5 in spotting cultural errors in text"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The synthetic errors added to create the ExCAM40k dataset from existing benchmarks accurately represent real cultural errors that LLMs make in free text generation tasks.","fun_headline_variants_meta":{"raw":{"variants":["ExCAM detects AI cultural errors with 80 percent accuracy","ExCAM rates and explains cultural mistakes in AI outputs","ExCAM40k enables error detection at up to 80 percent accuracy","ExCAM outperforms GPT-5 in spotting cultural errors in text"]},"model":"grok-4.3","cost_usd":0.005597,"raw_usage":{"total_tokens":2673,"prompt_tokens":653,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":55974500,"prompt_tokens_details":{"text_tokens":653,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1952,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":653,"tokens_out":68,"duration_ms":16626,"temperature":1.0,"reasoning_tokens":1952,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T07:40:18.372179+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A study applying ExCAM to a collection of real-world LLM outputs containing human-verified cultural errors and measuring detection accuracy would falsify the result if it is substantially lower than 80%.","supporting_citations":[],"review_version":1}