{"id":"679b49f9-cb2c-40bb-9ed8-ec6b50f7ce15","arxiv_id":"2508.10001","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A new Hinglish political fact-checking benchmark (HiFACT) and a graph-aware model (HiFACTMix) are described, with the abstract claiming improved accuracy over multilingual baselines.","lead":"This paper introduces HiFACT, a benchmark of 1,500 real-world Hinglish political claims from 28 Indian state chief ministers, each annotated with evidence and a veracity label. It also proposes HiFACTMix, a graph-aware retrieval-augmented model that combines multilingual encoding, claim-evidence alignment, graph reasoning, and explanation generation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The empirical claims are uncheckable: the supplied full text is corrupted, and the abstract lacks baseline names, effect sizes, label-quality metrics, and faithfulness metrics, so the HiFACTMix accuracy and justification claims cannot be assessed.","rationale":"Good-faith reading: the paper wants to contribute a benchmark and a stronger model. The reader's verdict was UNVERDICTED because the full text is corrupted. My pass agrees and sharpens the point: the load-bearing condition is not any single architectural equation but the inspectability of the evaluation. The abstract's phrase 'outperformed accuracy' carries no effect size, baseline list, variance, or statistical test. The benchmark's usefulness likewise depends on annotation quality, and the claimed faithful justifications require a faithfulness measure; none appears in the abstract. The corrupted full text cannot supply these details, so the central claim is unsupported in the only material available. This is not an internal inconsistency; I found no legible contradiction between the model description and results. It is an evidentiary gap. I would not move the verdict to REJECT, because a clean source might contain the missing tables and protocols; I would keep it unverdictable pending inspection.","tokens_in":19297,"tokens_out":4172,"duration_ms":45278,"concrete_test":"Download the LaTeX/source for arXiv:2508.10001 and inspect the experiments section. The decisive check is whether the main results table names current multilingual or code-mixed fact-checking baselines, reports confidence intervals or significance tests for accuracy, and provides an inter-annotator agreement or faithfulness metric; if the clean text contains these, the concern is resolved, and if it does not, the headline accuracy and faithfulness claims are unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is empirical: HiFACTMix is more accurate than state-of-the-art multilingual baselines on a new 1,500-claim benchmark and provides faithful justifications. For that claim to hold, three conditions must be true: (i) the HiFACT labels and evidence annotations are accurate and unbiased; (ii) the evaluation uses named, properly configured baselines with clearly defined metrics; and (iii) 'faithful justification' is measured by an actual faithfulness protocol rather than by generated text alone. None of these conditions is checkable in the material supplied. The abstract contains no numbers, no baseline names, no inter-annotator agreement, and no faithfulness metric; the rest of the text is largely mojibake, and the visible tables do not contain a legible comparison. This is a verification blocker, not an accusation of error: I cannot distinguish a correct paper with a corrupted PDF from an unsupported one. The only responsible status, therefore, is UNVERDICTED rather than acceptance or rejection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HiFACT, a new benchmark of 1,500 Hinglish political claims made by 28 Indian state chief ministers, with veracity labels and textual evidence, and HiFACTMix, a graph-aware retrieval-augmented fact-checking model combining multilingual encoding, claim-evidence semantic alignment, evidence graph construction, graph neural reasoning, and natural language explanation generation. The abstract claims that HiFACTMix outperforms state-of-the-art multilingual baselines in accuracy and provides faithful justifications for its verdicts. However, the supplied full text is largely corrupted and unreadable, so the experimental evidence, model details, and evaluation protocol cannot be independently verified.","tokens_in":19501,"tokens_out":2318,"duration_ms":24315,"significance":"If the claims are correct, the paper would deliver the first code-mixed Hinglish political fact-checking benchmark and a model that improves over multilingual baselines while generating explanations, which is a useful step for low-resource and code-mixed fact-checking. The benchmark alone, with 1,500 real-world political claims from a linguistically diverse setting, could be a valuable resource for the community. However, because the supporting evidence is not readable in the submitted manuscript, the significance is entirely conditional on a complete and corrected resubmission.","major_comments":[{"comment":"The central empirical claim in the abstract—'HiFACTMix outperformed accuracy in comparison to state of art multilingual baselines models'—is not accompanied by any quantitative result, baseline name, evaluation metric, confidence interval, or dataset statistic. The full text supplied for review is corrupted mojibake, so no table, figure, or equation can be read to verify the comparison. This is load-bearing because the entire contribution rests on this comparative claim, and it is currently uncheckable.","section":"Abstract and full text"},{"comment":"The manuscript does not provide, in any legible form, the annotation procedure, veracity label definition, evidence collection protocol, inter-annotator agreement, or label distribution for the 1,500-claim HiFACT benchmark. Because the model is evaluated against these labels, the accuracy claim is only as strong as the label quality; without this information the benchmark's validity cannot be assessed.","section":"HiFACT dataset description (unreadable in the supplied text)"},{"comment":"The abstract states that HiFACTMix 'provides faithful justifications for its verdicts,' but no definition of faithfulness, no human evaluation, and no faithfulness metric (e.g., entailment-based or token-overlap based) is reported in any readable portion of the manuscript. Generated natural language explanations alone are not evidence of faithfulness, so this claim is currently unsupported.","section":"Faithfulness claim in the abstract"}],"minor_comments":[{"comment":"The sentence 'HiFACTMix outperformed accuracy in comparison to state of art multilingual baselines models' is grammatically awkward and should be rewritten, e.g., 'HiFACTMix achieves higher accuracy than state-of-the-art multilingual baselines.'","section":"Abstract"},{"comment":"The paper should explicitly name the multilingual baselines used (e.g., mBERT, XLM-R, MuRIL) and report their hyperparameters, fine-tuning procedures, and any per-class results, so that the comparison is reproducible.","section":"Experimental setup (unreadable)"},{"comment":"The corrupted text contains the string 'arXiv:2508.10000v1' while the submitted paper is numbered 2508.10001; the authors should verify the arXiv identifier and ensure the manuscript header is correct.","section":"Header of the full text"},{"comment":"All figure and table captions are unreadable in the supplied version; the authors should embed high-resolution figures and ensure captions and table entries are rendered properly.","section":"Figures and tables"}],"recommendation":"uncertain","confidential_remarks":"The supplied full text is corrupted to the point that no substantive technical evaluation is possible. I recommend asking the authors to resubmit a readable version with the complete experimental details, including baseline configurations, dataset annotation statistics, and a faithfulness evaluation protocol. Additionally, there is a circularity risk: the model is evaluated on a benchmark constructed by the same authors, so the paper should address why the benchmark labels and evidence are trustworthy, ideally with inter-annotator agreement and an external validation step."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: HiFACTMix targets a real gap—code-mixed Hinglish political claim verification—but the version in front of me can't be assessed. The full text is largely mojibake, and the abstract gives no effect sizes, no baseline names, no dataset statistics, and no faithfulness metric. That's a verification blocker, not proof of error, but it means every empirical claim in the abstract is currently uncheckable.\n\nWhat is genuinely new is the benchmark concept: 1,500 factual claims from 28 Indian chief ministers, in Hinglish, with evidence annotations and veracity labels. If the dataset is solid, that is a useful contribution to code-mixed NLP, which is underserved in fact-checking. The model side—multilingual encoding, claim-evidence alignment, evidence graph construction, graph neural reasoning, and explanation generation—is a plausible pipeline, though nothing in the abstract distinguishes it from a fairly standard graph-aware RAG architecture.\n\nThe soft spots are proportional to what we can actually see. The central comparative claim, 'outperformed accuracy in comparison to state of art multilingual baselines,' is supported by zero numbers. The same-team benchmark plus model carries a tuning risk, but their claim to compare against external multilingual baselines is the right check; the abstract just doesn't tell us which baselines or how they were configured. The 'faithful justifications' claim also needs a real protocol—token-level attribution or entailment-based evaluation—rather than generated text alone. None of this is visible.\n\nI also want to be careful: I'm not accusing the authors of anything. A corrupt PDF submission happens; the underlying work might be fine. But as a reviewer, I can't distinguish a strong paper from a weak one on this material.\n\nMy recommendation: ask the authors to resubmit a clean PDF with a full experimental section—baselines named, accuracy with confidence intervals, inter-annotator agreement, and a faithfulness evaluation. If that version arrives, it deserves a serious referee. In its current form, I would not send it out; there is nothing for referees to evaluate.","headline":"A Hinglish fact-checking benchmark with a graph-aware model sounds genuinely useful, but this version is unreadable: the full text is corrupted and the abstract reports no numbers, so the central claims are unverifiable.","tokens_in":19988,"tokens_out":2435,"would_cite":false,"duration_ms":27933,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces HiFACT, a 1,500-claim Hinglish benchmark, and HiFACTMix, a graph-aware model that it says beats multilingual baselines while explaining its verdicts.","keywords":["Hinglish","code-mixed NLP","political fact-checking","claim verification","benchmark dataset","graph neural networks","retrieval-augmented generation","low-resource languages"],"falsifier":"Re-annotate a random sample of roughly 100 to 200 HiFACT claims with independent annotators and measure label agreement, then re-run the baselines with the same training data and hyperparameter budget as HiFACTMix; if inter-annotator agreement is low, or the accuracy gap shrinks to statistical noise, the central claim would fail.","tokens_in":19131,"feed_emoji":"🗳️","tokens_out":5421,"duration_ms":52871,"temperature":0.7,"pith_summary":"This paper tries to establish that political fact-checking can work directly in Hinglish, the code-mixed Hindi-English used widely in Indian public speech. It introduces HiFACT, a benchmark of 1,500 real-world political claims from 28 Indian state chief ministers, each annotated with veracity labels and textual evidence. It then proposes HiFACTMix, a graph-aware retrieval-augmented model that encodes claims and evidence multilingually, aligns them semantically, builds an evidence graph, reasons over it with graph neural networks, and generates natural-language justifications. The paper's central empirical claim is that HiFACTMix beats state-of-the-art multilingual baselines in accuracy and offers faithful explanations for its verdicts. If true, this would give researchers a benchmark and a model for fact-checking low-resource code-mixed political discourse rather than only high-resource monolingual text.","feed_headline":"Graph model beats baselines at Hinglish fact-checking","feed_subtitle":"A 1,500-claim benchmark from Indian chief ministers plus a model that explains its verdicts targets code-mixed Hindi-English.","key_machinery":"The load-bearing object is the evidence graph inside HiFACTMix. After a multilingual encoder represents the claim and candidate evidence, a semantic-alignment step scores how each evidence piece relates to the claim, and those pieces become nodes in a graph connected by relevance or consistency edges. A graph neural network then reasons over this graph to produce the verdict, and the same graph feeds the natural-language explanation generator. The graph is what lets the model aggregate support and contradiction across multiple evidence pieces instead of relying on one retrieved sentence, and it also gives the explanation generator a concrete structure to describe.","core_discovery":"On its own terms, the paper's discovery is a working combination of a new benchmark and a new model. HiFACT supplies 1,500 verifiable claims made by chief ministers in Hinglish, with evidence and labels, while HiFACTMix is the proposed system for verifying those claims. The paper claims that this system achieves higher accuracy than existing state-of-the-art multilingual baselines on this benchmark, and that the natural-language justifications it generates are faithful, meaning they reflect the evidence used to reach the verdict. The central contribution is thus a demonstrated proof-of-concept that code-mixed, politically grounded fact verification can be benchmarked and automated.","pith_inferences":["Editorial inference: the graph-aware retrieval architecture is a natural candidate for other code-mixed varieties such as Bangla-English or Taglish, because the method does not depend on a clean monolingual corpus.","Editorial inference: the 1,500-claim, 28-speaker design opens the door to measuring whether verification accuracy shifts from one chief minister or region to another, an analysis the abstract does not promise.","Editorial inference: a direct test of faithful justification would be to present only the generated explanations to readers, without the verdict labels, and check whether the readers' implied verdicts match the model's labels; the abstract reports accuracy but not this agreement.","Editorial inference: if the benchmark is released with its evidence annotations, it could support studies of how code-mixed wording changes the difficulty of verification compared with an English-translated version of the same claims."],"forward_implications":["Fact-checking can be built directly on code-mixed text instead of being forced through an English translation, preserving the wording in which claims actually spread.","A model that reasons over an evidence graph can combine supporting and contradicting pieces across several documents, so a verdict reflects a body of evidence rather than one matched sentence.","The HiFACT benchmark gives future systems a fixed resource for comparing Hinglish and other code-mixed fact-verification methods on real political claims.","If explanations are generated from the same evidence graph that produces the verdict, a user or auditor can check a justification directly against the underlying evidence.","The same graph-aware architecture is a candidate for other low-resource and code-mixed political settings, not just India."],"supporting_citations":[],"fun_headline_variants":["New benchmark and graph model verify Hinglish political claims","Graph-aware model beats baselines on Hinglish claim verification","Code-mixed claims from Indian CMs get a new fact-checking benchmark","HiFACTMix: Graph reasoning for explainable code-mixed fact-checking","1,500 chief minister claims in Hinglish with evidence and verdicts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The performance claim rests on the assumption that the veracity labels and evidence annotations in HiFACT are accurate and unbiased, and that the multilingual baselines were implemented and evaluated fairly.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark and graph model verify Hinglish political claims","Graph-aware model beats baselines on Hinglish claim verification","Code-mixed claims from Indian CMs get a new fact-checking benchmark","HiFACTMix: Graph reasoning for explainable code-mixed fact-checking","1,500 chief minister claims in Hinglish with evidence and verdicts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1420,"prompt_tokens":923,"completion_tokens":497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":402}},"tokens_in":539,"tokens_out":497,"duration_ms":4768,"temperature":1.0,"reasoning_tokens":402,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:36:35.825326+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of roughly 100 to 200 HiFACT claims with independent annotators and measure label agreement, then re-run the baselines with the same training data and hyperparameter budget as HiFACTMix; if inter-annotator agreement is low, or the accuracy gap shrinks to statistical noise, the central claim would fail.","supporting_citations":[],"review_version":2}