{"id":"09feed1b-eda3-4cd3-a981-d09e9e832e11","arxiv_id":"2608.08512","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Across nine LLMs, version resolution on amended customs documents tops out at 68.5% accuracy, with detection of a non-governing version at 26.7% and context misalignment at 27.62% even with gold documents.","lead":"This paper presents TIDE, an expert-verified benchmark of 3,050 question-answer pairs over 644 Bangladesh customs documents, and tests nine large language models on resolving which version of an amended rule applies on a queried date. It finds that the best model reaches only 68.5% accuracy and that models often trust a confident wrong version over the supplied authoritative text.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Event Sorting is scored with a graded metric under ICL despite Table 3 claiming exact match, so the headline 68.5% macro is not auditable.","rationale":"The reader correctly identified the parametric knowledge-cutoff problem as a weakness, and the paper's Limitations section already concedes that low parametric scores may indicate absent documents rather than version-resolution failure; that concern is real but it does not attack the gold-context central claim. I find the Event Sorting scoring inconsistency more load-bearing because it sits directly in the headline macro average and in the claim of a unified evaluation protocol. The contradiction is internal: Section 3.2 and Table 3 say exact match, while Appendix A.10.2, A.10.1, and A.11 say graded-under-ICL or describe Event Sorting as under re-scoring. If the ICL Event Sorting number is graded, the headline 68.5% is overstated or at least not comparable with the other settings. The benchmark construction, expert verification, and SME agreement with the LLM council provide genuine independent support, so this is a fixable evaluation error rather than a reason to reject. The existing CONDITIONAL verdict stands, with the condition now more specific: reconcile the Event Sorting metric and recompute the affected macro averages before the headline numbers are used.","tokens_in":43780,"tokens_out":6476,"duration_ms":72406,"concrete_test":"Re-score the 207 Event Sorting items in the ICL condition with the exact-match criterion defined in Section 3.2 and Appendix A.6, using the released model outputs, and recompute the ICL macro average in Table 3 exactly as was done for the parametric and RAG columns. If the corrected Event Sorting ICL score reproduces 66.23, then the appendix statements are the error and must be fixed; if it lands near the ~57% matched-set figure implied by Section A.10.1, then Table 3, the 68.49% macro average, and the ICL-vs-RAG discussion all need revision. The authors should also publish the exact-match per-model Event Sorting scores for all three settings so the unified protocol is fully auditable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is an internal contradiction about how Event Sorting is scored. Section 3.2 and the Table 3 footnote state that event ordering is reported as exact-match accuracy in every setting. Appendix A.10.2 contradicts this: its Table 16 footnote says 'Event sorting uses a graded ordering metric under ICL but exact match elsewhere, so this row is not on a common scale.' Section A.10.1 also describes the ICL-RAG event-sorting gap using graded pairwise-order accuracy, with matched-set numbers that imply the exact-match ICL score is well below the 66.23 shown in Table 3. Section A.11 further says 'Event sorting is under re-scoring and excluded' from the RAG ablation. If the ICL Event Sorting score in Table 3 is graded rather than exact, then the headline macro average of 68.49% mixes different metrics across settings, and the central claims of a 'unified protocol' and of reading-the-correct-text being necessary-but-not-sufficient are not currently auditable from the published tables. This is not a challenge to the benchmark's construction or to the qualitative finding that version resolution remains hard; it is a concrete scoring inconsistency that affects the paper's headline number and the ICL-vs-RAG comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TIDE, a benchmark of 3,050 expert-verified QA pairs over 644 official Bangladesh customs instruments (1969–2025), across eight task types that are designed to test version resolution in evolving documents. The authors evaluate nine recent LLMs under three knowledge-access settings (parametric, in-context learning with gold documents, and retrieval-augmented generation), using a three-judge LLM council with a hard date gate for free-form answers. The central empirical claims are that the best macro-averaged accuracy is only 68.49% (GPT-5, ICL), that models find correct versions more readily than they reject incorrect ones, and that providing the correct text improves but does not solve version resolution. The paper also reports task-specific failures such as low Context Misalignment accuracy and a large ICL-vs-RAG gap on Event Sorting, which is attributed to retrieval incompleteness.","tokens_in":44032,"tokens_out":7328,"duration_ms":69987,"significance":"The dataset construction is a genuine contribution: the corpus is drawn from official legal documents, all QA pairs were verified by two subject-matter experts with high inter-expert agreement, the evaluation protocol is described in enough detail to be reproduced, and the deterministic scoring equation plus the date gate is a principled way to separate factual correctness from temporal correctness. If the reported results are sound, TIDE fills a real gap by testing formal amendment semantics rather than treating time as a mere annotation. The paper also ships the prompt templates, licensing details, and a reproducibility-oriented appendix, which are valuable for the community. However, the headline numbers are currently not auditable because of an internal contradiction about how Event Sorting is scored across settings, and the parametric setting relies on a cutoff assumption that is contradicted by the paper's own model table.","major_comments":[{"comment":"Section 3.2 and the Table 3 footnote state that Event Sorting is reported as exact-match accuracy in every setting, but Appendix A.10.2 (Table 16, footnote) states that 'Event sorting uses a graded ordering metric under ICL but exact match elsewhere, so this row is not on a common scale,' and Appendix A.11 excludes Event Sorting from the RAG ablation because it is 'under re-scoring.' Appendix A.10.1 further reports the same ICL–RAG event-sorting comparison as 0.57 versus 0.28 while also giving graded pairwise-order values of 0.81 versus 0.67, which is consistent only if the ICL number in Table 3 is not on the exact-match scale. Since Event Sorting is one of the eight tasks entering the headline macro average of 68.49% (Table 3), the 'unified protocol' claim and the ICL-vs-RAG comparison are not auditable from the published tables. Please re-score Event Sorting with a single metric across all three settings and recompute all affected macro averages and conclusions.","section":"Section 3.2 / Table 3 / Appendix A.10.2, Table 16"},{"comment":"The parametric setting is justified by the statement in Section 3.1 that models were selected 'whose cut-off dates cover the date of our most recent document.' This is the load-bearing assumption for interpreting the Parametric column of Table 3 as a measure of version-resolution failure rather than missing documents. Table 8, however, lists Gemini-3.5-Flash and Claude-Sonnet-4.5 with reliable knowledge cutoffs of January 2025, while Appendix A.4 contains benchmark items referencing instruments dated 2025-02-17 and 2025-06-03 (Temporal MCQ Example 1 and Event Sorting Example 1). The open-weight models are listed with undisclosed cutoffs, so their coverage cannot be confirmed. The Limitations section explicitly concedes that 'A low parametric score might indicate an absent document rather than a version resolution failure,' but this concession is not reflected in the main interpretation of the parametric scores. Please either restrict parametric analyses to items dated before every model's disclosed cutoff, report a cutoff-aware analysis, or reframe the parametric column as a model-specific lower bound that is not directly comparable across models.","section":"Section 3.1 / Section 3.3 / Table 8 / Limitations"},{"comment":"The reliability of the LLM council is central to the free-form task scores, but the reported human agreement is not item-level agreement on the booleans that Equation (2) actually consumes. Section 3.2 says that two SMEs rescored the council's decisions and that 'their average agreement with the council is 94.21%,' while Appendix A.6 shows this value is the mean of two rubric-point totals (92.93 and 95.49 out of 100) from a rubric that awards points for verdict accuracy, reasoning, and clarity. A rubric-point total can be high even when the experts disagree with the council on a subset of the final verdicts, and those verdicts are the exact inputs to the deterministic scorer. Please report item-level agreement (for example, percent agreement or Cohen's kappa on the final verdict booleans) so that the validation statistic matches the decision-relevant quantity.","section":"Section 3.2 / Appendix A.6, Tables 10 and 11"}],"minor_comments":[{"comment":"The abstract states that resolving a version from an implicit date reaches 59.7%, but Table 3 shows Claude-Sonnet-4.5 reaching 68.03% strict accuracy on Relative-Time QA under ICL; if 59.7% refers to a different model or aggregation, the abstract should say so.","section":"Abstract / Table 3, Relative-Time QA row"},{"comment":"The model is called GPT-OSS-20B in Table 4 but GPT-OSS-120B in Section 3.1, Table 1, and Table 3; the name should be consistent.","section":"Table 4 vs. Section 3.1 and Table 3"},{"comment":"The table rows are visually compressed: for example, the Temporal MCQ and Event Sorting rows run together in the stimulus column ('4 options 136 128 41Event Sorting'), and the Event Sorting stimulus value '3.5 events' is unclear; use explicit column separators and state whether this is a mean, median, or range.","section":"Table 2"},{"comment":"The appendix says pairs rated 1 are discarded and new pairs are written on the same thread theme, and that after repairs every retained pair meets the score-3 criterion; please state explicitly that the replacement pairs were also reviewed by both SMEs and whether the 94.89% figure is before or after these replacements.","section":"Appendix A.3"},{"comment":"The paper would be easier to assess if it explicitly discussed the family-level overlap in which Gemini models are used for parsing, entity and concept extraction, QA generation, retrieval embeddings, as a council judge, and as a subject model; the SME verification mitigates the gold-data risk, but a short paragraph on the residual risk would be useful.","section":"Section 2.2 / Section 3.2 / Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"I think the benchmark itself is a solid contribution and the central qualitative message—that version resolution over evolving documents remains hard even with gold context—is likely to survive correction. The Event Sorting scoring inconsistency and the parametric cutoff contradiction must be fixed before the headline numbers can be trusted; both are fixable within the manuscript's scope. I would not reject on the basis of the Gemini involvement given the expert verification, but I would require item-level SME agreement statistics and a unified Event Sorting metric in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me get to the point. The dataset is the real contribution: 3,050 QA pairs over 644 official customs instruments, two SME verifiers with strong agreement, per-item provenance, and careful licensing. That part is solid and worth publishing. But the paper currently contradicts itself on how Event Sorting is scored, and the contradiction lands on the headline number.\n\nTable 3 says Event Sorting is exact-match. Table 16's footnote says ICL uses a graded ordering metric while other settings use exact match. Section A.11 says Event Sorting is 'under re-scoring' and excluded from the RAG ablation. All three cannot be true. If ICL Event Sorting is graded, the 68.49% macro average mixes different metrics across settings, and the 'unified protocol' claim is not auditable. If it is exact-match, then the footnote is wrong. Either way, the authors must fix this before the numbers can be trusted. The stress-test note is right.\n\nThe parametric setting has a smaller, related problem. Section 3.1 says models were selected whose cutoffs cover the most recent document, but Table 8 lists Gemini-3.5-Flash and Claude-Sonnet-4.5 with reliable cutoffs of January 2025, before the corpus's later instruments. The Limitations section honestly acknowledges that low parametric scores might mean missing documents, so the paper is not hiding it. But the contradiction should be reconciled, and the parametric column of Table 3 read with that caveat.\n\nWhat is genuinely new: separating version resolution from encyclopedic temporal QA. TIDE is the first benchmark, as far as I can see, built on official amendment texts that state what they replace and when. The eight task types are well chosen, and the main qualitative finding — models find correct versions more often than they reject wrong ones, with Context Misalignment staying under 27% even with gold documents — is credible and matters for legal and compliance deployment. The hard date gate in the council scoring is a good design.\n\nThe Gemini overlap deserves more prominent disclosure: it generates, extracts, embeds, judges, and is evaluated as a subject. The independent SME verification mitigates this, but the family-level dependency is real and should be stated up front. The repository needs to be checked during review — the current paper doesn't give enough evidence that the code and data run.\n\nThis paper deserves a serious referee. The benchmark is valuable, the qualitative finding is likely robust to the scoring fixes, and the inconsistencies are fixable without new experiments. I would not quote the 68.5% headline as it stands, but I would cite the benchmark and use it. The authors need one careful revision pass, not a rejection.","headline":"TIDE's dataset is a solid, expert-verified contribution to testing version resolution in evolving documents, but the Event Sorting scoring inconsistency makes the headline 68.5% and the ICL-vs-RAG comparison unauditable as printed.","tokens_in":44590,"tokens_out":5880,"would_cite":true,"duration_ms":54871,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper presents TIDE, a benchmark of 3,050 expert-verified question-answer pairs over 644 official Bangladesh customs instruments issued between 1969 and 2025, and claims that even the strongest of nine large language models resolves…","keywords":["version resolution","evolving documents","temporal question answering","LLM benchmark","customs law","retrieval-augmented generation","temporal reasoning","knowledge cutoff"],"falsifier":"Run the parametric probe with a model whose training-corpus membership can be independently verified, for example a model trained on a checked-in snapshot of the TIDE corpus, and see whether the version-sensitive scores stay in the reported near-zero range; if they rise, the low parametric scores reflect absent documents rather than failed version resolution. A second check: on the 29.2% of items that stayed wrong even with the gold document, add an explicit instruction to state which clause governs the queried date and see whether accuracy moves, since improvement would show the reported ceiling is a prompting artifact rather than a reasoning limit.","tokens_in":43597,"feed_emoji":"📜","tokens_out":7693,"duration_ms":70385,"temperature":0.7,"pith_summary":"The paper argues that answering questions about laws, tax codes, and other documents that are amended over time requires a distinct skill, version resolution, which existing temporal-question benchmarks never test. TIDE is positioned as the first benchmark built for that skill: 3,050 question-answer pairs over authentic customs instruments where the correct answer changes with the queried date. Under one evaluation protocol across nine large language models, the best macro-averaged accuracy reaches only 68.5%, resolving the rule in force from an implied date works 59.7% of the time, and detecting that a supplied version does not govern the query reaches only 26.7%. The stakes are concrete because any deployment of language models in legal, tax, or compliance settings depends on knowing which version of a rule applies, and the paper shows that today's models cannot be trusted to do this even when the right documents are in front of them.","feed_headline":"Best LLM resolves dated rules only 68.5% of the time","feed_subtitle":"New benchmark over 644 evolving customs laws shows models can't tell which rule version applies on a given date.","key_machinery":"The load-bearing object is the evolving concept thread: a cluster of dated provisions across multiple instruments that all speak about the same regulatory topic, stored as a time-ordered list of instruments and the values they set. Version resolution means mapping a query date to the value in force on that date, and every generated question is written against a full thread history rather than a snapshot of one document. Three supporting mechanisms carry the results: the entity-clustering and alias-judgement pipeline that assembles threads from 644 code-mixed bilingual documents; the three knowledge-access settings (parametric memory, gold context, and top-10 retrieval) that isolate where failure occurs; and the scoring protocol, a three-judge LLM council whose verdicts pass through a hard date gate that scores an answer zero whenever it commits to a wrong or contradictory date even if the meaning is correct. The date gate lets the paper separate correct meaning from correct time, which reveals that once grounding is supplied, the residual error flips from wrong substance with a right date to right substance with a wrong date.","core_discovery":"The paper's central claim is that version resolution over evolving documents is a distinct and presently unsolved capability. In an evolving document, an amendment is itself an official text that names what it replaces and when it takes effect, so several versions of the same rule can be simultaneously correct, each for its own validity period; existing temporal QA datasets treat time only as an annotation and therefore never test this. TIDE operationalises the claim with 3,050 QA pairs over 644 official customs instruments, and the evaluation shows two things: reading the correct text is necessary but not sufficient, since even with the gold documents in the prompt the best model scores 68.49%, and models are systematically asymmetric, finding correct versions far more readily than rejecting incorrect ones, while tending to follow a confident parametric answer over the supplied authoritative text. The hardest failures sit exactly where deployment risk is highest, with a fluent wrong premise overriding correct evidence that appears in the same prompt.","pith_inferences":["If version resolution is a general skill rather than a corpus-specific one, the same eight-task design should transfer to income tax, VAT, medical guidelines, and software documentation; porting the pipeline to another jurisdiction or domain is a direct test of that generality that the paper does not run.","The monotonic-increase bias the paper observes, where models assume regulatory values only ever rise and therefore misorder or misdate downward amendments, suggests a broader prior in parametric knowledge; a small targeted probe of downward revisions across domains could quantify how widespread it is.","The date-gated scoring protocol could serve as a reusable evaluation standard for any time-sensitive QA, since it cleanly separates meaning errors from temporal errors; a lighter deployed variant might rely on the deterministic date gate alone without the three-judge council.","The reported 5.5% regression, where the gold document overturns a correct closed-book answer, implies that giving a model the authoritative source can actively harm accuracy when dual-calendar dates are misread, so document-grounded pipelines need an explicit disagreement mechanism between memory and evidence."],"forward_implications":["Any LLM deployment in legal, customs, tax, or compliance workflows that answers date-sensitive questions without a version check will produce superseded answers; the context-misalignment task shows a fluent wrong premise overrides correct evidence even when the authoritative text sits above it in the same prompt.","Retrieval completeness is the current bottleneck for amendment-wide reasoning: event-sorting accuracy for the best model collapses from 66.23% under gold context to 31.42% under a standard semantic retriever, so any retriever that cannot recall the full amendment history caps model performance.","The 29.2% of items that stay wrong even when the gold document is supplied define a reasoning ceiling that retrieval cannot address, so closing the gap requires better version-resolution behaviour rather than more context.","Detecting that a provision was altered (about 79% detection on average) is far easier than identifying what was altered (about 19% identification on average) without grounding, so verification tasks should be scored on identification rather than on a binary verdict alone."],"supporting_citations":[{"why":"Closest prior benchmark of evolving knowledge; TIDE must be distinguished from it because it probes closed-book memory and treats superseded values as outdated rather than as concurrently valid versions.","marker":"Nakshatri et al., 2025"},{"why":"Supplies the misaligned-context and implicit-date task precedents over Wikidata facts that TIDE's context-misalignment and relative-time tasks extend.","marker":"Zhu et al., 2025"},{"why":"Early closed-book, open-book, and reasoning three-setting protocol that the unified parametric, gold-context, and retrieval evaluation adapts.","marker":"Tan et al., 2023"},{"why":"Grounds the parametric-setting rationale that knowledge freezes at the training cutoff, which is why the paper selects models by cutoff date.","marker":"Cheng et al., 2024"},{"why":"Sourced claim that semantic retrievers match meaning rather than time, cited as the reason the RAG scores are a lower bound without temporal re-ranking.","marker":"Zhang et al., 2025"},{"why":"Version-aware retrieval work that ignores amendment semantics, cited as motivation for testing retrieval over genuinely amending documents.","marker":"Huwiler et al., 2025"},{"why":"Provides the Likert-scale rating methodology used for the two-expert verification of the generated QA pairs.","marker":"Amidei et al., 2019"},{"why":"The LLM-as-judge evaluation setup that the three-judge council extends with the date-gate scoring.","marker":"Kim et al., 2026b"}],"fun_headline_variants":["LLMs only 68.5% accurate on evolving legal documents","TIDE benchmark: Best LLM scores 68.5% on dated rule questions","Models fail to reject wrong law versions in new benchmark","Amended laws trip up LLMs: top accuracy only 68.5%","New test: LLMs can't consistently select in-force rule versions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every evaluated model's training data actually contains the full 1969 to 2025 corpus, so low closed-book scores measure version-resolution failure rather than simply missing documents; the paper admits this cannot be verified because providers do not disclose training data, and two model cards list knowledge cutoffs earlier than the most recent benchmark documents.","fun_headline_variants_meta":{"raw":{"variants":["LLMs only 68.5% accurate on evolving legal documents","TIDE benchmark: Best LLM scores 68.5% on dated rule questions","Models fail to reject wrong law versions in new benchmark","Amended laws trip up LLMs: top accuracy only 68.5%","New test: LLMs can't consistently select in-force rule versions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000613,"raw_usage":{"total_tokens":2892,"prompt_tokens":1032,"completion_tokens":1860,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":1765}},"tokens_in":648,"tokens_out":1860,"duration_ms":12893,"temperature":1.0,"reasoning_tokens":1765,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:33:26.297547+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the parametric probe with a model whose training-corpus membership can be independently verified, for example a model trained on a checked-in snapshot of the TIDE corpus, and see whether the version-sensitive scores stay in the reported near-zero range; if they rise, the low parametric scores reflect absent documents rather than failed version resolution. A second check: on the 29.2% of items that stayed wrong even with the gold document, add an explicit instruction to state which clause governs the queried date and see whether accuracy moves, since improvement would show the reported ceiling is a prompting artifact rather than a reasoning limit.","supporting_citations":[],"review_version":1}