{"id":"4d6a0147-02bd-41c6-9090-7656645be239","arxiv_id":"2605.22079","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Presents Ishigaki-IDS-Bench, the first benchmark for LLM-based IDS generation from BIM requirements, with baseline results showing max 65.6% Facet F1 and 33.1% content pass rate across 10 models.","lead":"This paper introduces Ishigaki-IDS-Bench, the first public benchmark with 166 expert-authored examples for generating formal IDS XML files from BIM information requirements in Japanese and English. A smart generalist might read it to see how current AI models perform on specialized construction data standards that affect real building projects.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Gold IDS correctness rests solely on the six experts with no reported validation or agreement metrics.","rationale":"The reader's weakest_assumption exactly isolates the evaluation dependency that is least secured by the provided description. No other internal inconsistency (e.g., metric definitions or dataset size) appears more load-bearing once the gold-standard quality is granted.","tokens_in":1802,"tokens_out":299,"duration_ms":21388,"concrete_test":"Select 20 scenarios at random; have one additional independent BIM/IDS expert (unaffiliated with the six authors) re-author the gold IDS from the same input requirements and compare against the released gold files for structural and content discrepancies; if disagreement exceeds 15% on either metric, recompute the LLM leaderboard using the new gold set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline performance numbers (65.6% Facet F1, 33.1% Content pass rate) are computed against 83 expert-authored gold IDS files. The paper states these were produced by six BIM/IDS experts but supplies no inter-annotator agreement figures, external review process, or independent cross-check against buildingSMART/IFC conventions. Because both the IDSAuditTool Content score and the facet-level F1 treat the gold as ground truth, any systematic authoring inconsistencies would directly distort the reported LLM gaps. The abstract and metadata description give no indication that such checks were performed or reported.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces Ishigaki-IDS-Bench, the first publicly released benchmark for generating Information Delivery Specification (IDS) XML files from natural-language BIM information requirements. It contains 166 examples spanning 83 scenarios authored in Japanese and English by six BIM/IDS experts, each paired with a gold IDS file and metadata on input format, turn setting, target IFC versions, and construction domain. Evaluation uses two stages: formal validity via the buildingSMART IDSAuditTool (Processability, Structure, Content) and content fidelity via facet-level macro-F1 against the gold files. Across 10 LLMs in zero-shot, the best results are 65.6% Facet F1 (GPT-5.5) and 33.1% Content pass rate (Claude Opus 4.5). The benchmark is released on Hugging Face (DOI 10.57967/hf/8873) and evaluation code on Zenodo (DOI 10.5281/zenodo.20550510).","tokens_in":1938,"tokens_out":566,"duration_ms":13844,"significance":"If the gold-standard IDS files are reliable and representative, the benchmark would provide a useful domain-specific resource that incorporates IFC vocabulary conformance and external-validator agreement requirements absent from general structured-generation benchmarks. The dual-metric evaluation and public release of data plus code would support reproducible research on LLM use in BIM compliance tasks and establish initial performance baselines for this specialized generation problem.","major_comments":[{"comment":"Abstract: The 83 scenarios and corresponding gold IDS files are described as authored by six BIM/IDS experts, but the manuscript supplies no inter-annotator agreement figures, external review process, or independent cross-check against buildingSMART/IFC conventions. Because both the IDSAuditTool Content score and the facet-level F1 treat the gold files as ground truth, the absence of such validation directly affects the reliability of the headline performance numbers (65.6% Facet F1, 33.1% Content pass rate).","section":"Abstract"},{"comment":"Benchmark construction section (referenced in abstract): No information is given on the criteria used to select the 83 scenarios, their distribution across construction domains and IFC versions, or how edge cases were handled. This information is required to evaluate whether the benchmark is representative of practical use cases.","section":"Benchmark construction"}],"minor_comments":[{"comment":"Abstract: The relationship between the 83 scenarios and 166 examples should be clarified (e.g., whether each scenario contributes parallel Japanese and English versions).","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the reliability of the gold-standard annotations and the representativeness of the benchmark. We address each major comment below and outline the revisions we will make.","responses":[{"response":"We agree that additional documentation of the gold-standard creation process is necessary to support the benchmark's credibility. The current manuscript does not report inter-annotator agreement statistics or a formal external validation step. In the revised manuscript we will add a new subsection under benchmark construction that describes the authoring protocol used by the six experts, the consensus process employed, and any cross-checks performed against buildingSMART guidelines and IFC conventions.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The 83 scenarios and corresponding gold IDS files are described as authored by six BIM/IDS experts, but the manuscript supplies no inter-annotator agreement figures, external review process, or independent cross-check against buildingSMART/IFC conventions. Because both the IDSAuditTool Content score and the facet-level F1 treat the gold files as ground truth, the absence of such validation directly affects the reliability of the headline performance numbers (65.6% Facet F1, 33.1% Content pass rate)."},{"response":"We acknowledge that the manuscript lacks explicit details on scenario selection and coverage. The 83 scenarios were chosen by the expert authors to span common practical BIM use cases. In the revision we will expand the benchmark construction section with (i) the explicit selection criteria, (ii) a table or breakdown showing distribution across construction domains and target IFC versions, and (iii) a description of how edge cases were identified and incorporated.","revision_made":"yes","referee_comment":"[Benchmark construction] Benchmark construction section (referenced in abstract): No information is given on the criteria used to select the 83 scenarios, their distribution across construction domains and IFC versions, or how edge cases were handled. This information is required to evaluate whether the benchmark is representative of practical use cases."}],"tokens_in":1600,"tokens_out":442,"duration_ms":24753,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper releases Ishigaki-IDS-Bench, the first public dataset for testing LLMs on turning BIM information requirements into valid IDS XML files, along with evaluation code.\n\nIt does some things solidly. The benchmark covers 83 scenarios in both Japanese and English, ties outputs to real IFC versions and the buildingSMART IDSAuditTool for scoring, and reports clear baselines across ten models in zero-shot settings. Highest facet F1 reaches 65.6% and content pass rate only 33.1%. Open release on Hugging Face and Zenodo makes it usable for follow-up work.\n\nThe soft spot sits in the gold files. Six experts created the 83 reference IDS files, yet the paper supplies no inter-annotator agreement figures, no external cross-check against IFC conventions, and no description of how scenarios were chosen or edge cases handled. Because both the validator scores and the facet F1 treat these files as ground truth, any systematic authoring differences would directly affect the claimed LLM gaps.\n\nThis is for researchers working on domain-specific structured generation in architecture, engineering, and construction, or anyone building LLM tools around BIM standards. It will not reshape the wider field but fills a narrow, practical hole.\n\nSend it for peer review. The data release and concrete task definition are worth referee time even if the methods section needs added detail on annotation quality.","headline":"First public benchmark for IDS generation from BIM requirements with data and code released, but gold labels from six experts have no reported agreement or validation checks.","tokens_in":2444,"tokens_out":359,"would_cite":false,"duration_ms":25396,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Ishigaki-IDS-Bench is the first public benchmark for turning BIM information requirements into valid IDS XML files, where even top LLMs reach only 65.6 percent facet F1 and 33.1 percent content pass rate.","keywords":["Information Delivery Specification","BIM","IDS generation","LLM benchmark","buildingSMART","IFC","structured generation","facet F1"],"falsifier":"Re-running the evaluation after independent expert re-validation of all 83 gold IDS files and finding that any model's content pass rate rises above 50 percent.","tokens_in":2721,"feed_emoji":"📊","tokens_out":776,"duration_ms":19149,"temperature":0.7,"pith_summary":"The paper introduces Ishigaki-IDS-Bench to test large language models on generating Information Delivery Specifications from natural-language BIM requirements. The benchmark supplies 166 examples across 83 scenarios written in Japanese and English by six domain experts, each paired with a gold-standard IDS file and metadata on input format, IFC version, and construction domain. Evaluation splits into two checks: formal validity scored by the buildingSMART IDSAuditTool on processability, structure, and content, plus content fidelity measured by macro-F1 over individual facets against the gold files. Results across ten models in zero-shot settings show the best facet F1 at 65.6 percent and best content pass rate at 33.1 percent. The dataset and evaluation code are released publicly so others can measure progress on this structured-generation task that demands both IFC vocabulary compliance and external-validator agreement.","feed_headline":"Benchmark shows LLMs top at 65.6% facet accuracy on IDS generation","feed_subtitle":"First public test of turning BIM requirements into validator-approved XML files finds content pass rates stuck at 33 percent.","key_machinery":"Ishigaki-IDS-Bench, a dataset of 83 expert-authored scenarios each linked to a gold IDS file and evaluated by a two-stage protocol of IDSAuditTool validity plus facet macro-F1.","core_discovery":"Ishigaki-IDS-Bench is the first publicly released benchmark for IDS generation from BIM information requirements. The benchmark contains 166 examples spanning 83 practical scenarios authored in Japanese and English by six BIM/IDS experts, each paired with a gold IDS file. Evaluation proceeds in two stages: formal validity scored by the buildingSMART IDSAuditTool along Processability, Structure, and Content, and content fidelity scored by facet-level macro-F1 against the gold IDS. Across 10 LLMs in zero-shot, the highest Facet F1 is 65.6 percent, achieved by GPT-5.5, while the highest Content pass rate is only 33.1 percent, achieved by Claude Opus 4.5.","pith_inferences":["The benchmark's bilingual design may reveal whether Japanese or English inputs produce systematically different error patterns in IFC vocabulary handling.","Similar dual-stage benchmarks could be created for other machine-checkable construction documents such as COBie or BCF to test the generality of the evaluation approach.","If models improve on this set, the same scenarios could serve as seed data for supervised fine-tuning aimed at IFC property-set conventions."],"forward_implications":["Current LLMs still fail to produce IDS files that both pass the official validator and match expert facet content at high rates.","Benchmarks for structured output must now incorporate domain-specific vocabulary checks and external-validator agreement rather than pure syntactic correctness.","Progress on IDS generation can be tracked quantitatively as new models are tested on the released 166-example set.","The dual evaluation protocol separates format compliance from semantic fidelity, allowing targeted diagnosis of model errors.","The public release under open licenses enables direct comparison of future methods against the reported zero-shot baselines."],"fun_headline_variants":["Ishigaki-IDS-Bench tests LLMs on generating IDS from BIM requirements","First IDS benchmark shows 65.6% facet F1 by GPT-5.5","LLMs max out at 65.6% F1 and 33.1% content pass on IDS task","New benchmark exposes LLM content pass rates at 33.1% for IDS","Experts release 166-example IDS dataset for BIM requirements"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 83 scenarios and corresponding gold IDS files authored by the six experts are accurate, representative of practical use cases, and correctly validated against IFC and buildingSMART conventions.","fun_headline_variants_meta":{"raw":{"variants":["Ishigaki-IDS-Bench tests LLMs on generating IDS from BIM requirements","First IDS benchmark shows 65.6% facet F1 by GPT-5.5","LLMs max out at 65.6% F1 and 33.1% content pass on IDS task","New benchmark exposes LLM content pass rates at 33.1% for IDS","Experts release 166-example IDS dataset for BIM requirements"]},"model":"grok-4.3","cost_usd":0.006964,"raw_usage":{"total_tokens":3304,"prompt_tokens":822,"num_sources_used":0,"completion_tokens":106,"cost_in_usd_ticks":69637000,"prompt_tokens_details":{"text_tokens":822,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2376,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":822,"tokens_out":106,"duration_ms":19534,"temperature":1.0,"reasoning_tokens":2376,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T17:32:36.094185+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-running the evaluation after independent expert re-validation of all 83 gold IDS files and finding that any model's content pass rate rises above 50 percent.","supporting_citations":[],"review_version":2}