{"id":"6ce18c22-4d12-4035-84ef-b6cb3e242d23","arxiv_id":"2605.21773","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HIDBench unifies DARPA-E3, DARPA-E5, and NodLink datasets with a data pipeline to benchmark LLMs for host-based intrusion detection, showing high precision on simple logs but sharp drops in MCC and rises in false positives on complex noisy data.","lead":"This paper introduces HIDBench, a benchmark that unifies three public system log datasets and converts raw host telemetry into inputs for testing large language models on intrusion detection. Smart generalists should read it to see how LLMs perform on noisy, imbalanced real-world security data and where they fall short.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Data construction pipeline may introduce artifacts altering detection difficulty across datasets","rationale":"The reader's weakest assumption directly identifies the pipeline fidelity issue as load-bearing for the degradation claim. Full-text review does not remove this risk, as the abstract's emphasis on the pipeline remains the critical unverified link; no other internal inconsistency or stronger evidence (e.g., machine-checked elements) overrides it.","tokens_in":1773,"tokens_out":332,"duration_ms":28430,"concrete_test":"Reconstruct one dataset (e.g., DARPA-E5) using the paper's pipeline steps versus a minimal variant (raw log lines concatenated to token limit, no extra filtering); re-evaluate the same LLM and compare MCC/FPR delta. A difference >0.15 in MCC indicates the pipeline confounds the complexity claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim—that LLMs achieve high precision on simpler logs but degrade sharply (MCC <0.5, rising FPR) on noisier/complex ones—rests on the pipeline faithfully transforming raw telemetry from DARPA-E3, DARPA-E5, and NodLink into LLM inputs while preserving benign-malicious interactions. If the pipeline uses truncation, event selection, windowing, or formatting to fit context limits, it could artificially modulate task difficulty, making cross-dataset differences attributable to construction choices rather than inherent log properties. This assumption is least secure because the abstract positions the pipeline as central to realistic evaluation, yet no explicit check (e.g., against raw-input baselines or traditional detectors) is described to rule out artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces HIDBench, a new benchmark for assessing large language models (LLMs) on host-based intrusion detection (HIDS) from system logs. It unifies three public datasets (DARPA-E3, DARPA-E5, and NodLink), presents a data construction pipeline to convert raw host telemetry into LLM-compatible inputs, and evaluates frontier LLMs. The central empirical finding is that many models achieve high precision (often >0.8) on simpler datasets but degrade sharply on noisier and more complex logs, with MCC frequently dropping below 0.5 and false positive rates rising; the work also identifies distinct behavioral regimes such as conservative low-FPR detectors versus over-sensitive models.","tokens_in":1908,"tokens_out":443,"duration_ms":35806,"significance":"If the results hold after addressing pipeline validation, the benchmark fills an important gap in LLM evaluation for cybersecurity by focusing on fine-grained reasoning over noisy, imbalanced logs. The unification of multiple datasets and the identification of performance sensitivity to complexity provide actionable insights for deploying LLMs in HIDS, highlighting the need for robust system design rather than direct model use.","major_comments":[{"comment":"Abstract and data construction pipeline description: The central claim that performance degrades due to increasing log complexity (MCC <0.5, rising FPR) depends on the pipeline faithfully preserving benign-malicious interactions without introducing artifacts via truncation, event selection, windowing, or formatting to fit context limits. No explicit validation (e.g., raw-input baselines or comparisons to traditional detectors) is described to rule out construction choices as the source of cross-dataset differences; this assumption is load-bearing for attributing results to inherent data properties rather than pipeline design.","section":"Abstract; data construction pipeline"}],"minor_comments":[{"comment":"The abstract mentions 'distinct regimes' of model behavior but does not specify the exact metrics or thresholds used to classify conservative vs. over-sensitive detectors; adding a brief definition or reference to the relevant results table would improve clarity.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"Thank you for the constructive feedback on our manuscript. We have carefully considered the referee's major comment regarding the data construction pipeline and provide our response below, along with planned revisions to address the concerns.","responses":[{"response":"We agree that validating the pipeline is crucial to ensure that the observed performance differences across datasets can be attributed to variations in log complexity rather than artifacts introduced during data construction. Our pipeline applies consistent processing steps to all datasets to enable fair comparison, and the datasets themselves are established in the literature with known differences in noise and complexity levels. To address this point directly, we will revise the manuscript to include additional validation experiments. Specifically, we will report results from traditional HIDS approaches, such as signature-based or anomaly detection methods, applied to the same processed inputs. This will help demonstrate that the degradation in LLM performance on more complex datasets aligns with the inherent challenges of those datasets. We will also expand the description of the pipeline to detail how truncation and windowing were chosen to minimize information loss while respecting context constraints.","revision_made":"yes","referee_comment":"[Abstract; data construction pipeline] Abstract and data construction pipeline description: The central claim that performance degrades due to increasing log complexity (MCC <0.5, rising FPR) depends on the pipeline faithfully preserving benign-malicious interactions without introducing artifacts via truncation, event selection, windowing, or formatting to fit context limits. No explicit validation (e.g., raw-input baselines or comparisons to traditional detectors) is described to rule out construction choices as the source of cross-dataset differences; this assumption is load-bearing for attributing results to inherent data properties rather than pipeline design."}],"tokens_in":1433,"tokens_out":358,"duration_ms":39438,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"Colleague, the main takeaway is that this paper fills an actual gap by building HIDBench, which combines DARPA-E3, DARPA-E5, and NodLink into one evaluation setup for LLMs doing host-based intrusion detection from system logs. Prior LLM security benchmarks skipped this task, so the unification plus the pipeline that turns raw telemetry into model inputs counts as the concrete new piece rather than just another dataset wrapper. They evaluate several frontier models and report the expected pattern: high precision above 0.8 on cleaner logs, then MCC often below 0.5 and rising false positives once the logs get noisier and more complex. The breakdown into conservative versus over-sensitive detector regimes is a useful extra observation. The soft spot sits in the data construction pipeline. If truncation, windowing, or formatting steps change how benign and malicious events interact or how imbalanced the inputs become, then the cross-dataset degradation could partly trace to those choices instead of inherent log difficulty. The paper would be stronger with direct comparisons to traditional detectors on the same transformed data or explicit checks that the pipeline preserves the original interaction structure. This is aimed at people working on LLM applications in security or building practical detection systems; anyone testing models on log data would get value from the trends and the open benchmark. It deserves a serious referee because the empirical core is grounded in public datasets and the findings are falsifiable, even if the pipeline section needs more validation work.","headline":"HIDBench creates a new unified benchmark for LLMs on host log intrusion detection and documents clear performance drops with data complexity, but the construction pipeline needs tighter checks to confirm the trends reflect real log properties rather than processing choices.","tokens_in":2398,"tokens_out":377,"would_cite":false,"duration_ms":35864,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"HIDBench pipeline for LLM log segmentation and attack-graph reconstruction has no overlap with RS forcing chain","alignment":"orthogonal","rationale":"The paper's core machinery (attack-centric windowing of DARPA/NodLink logs, MEI/AGE/ACR stages, provenance-graph serialization) is a practical cybersecurity benchmark construction. RS theorems (reality_from_one_distinction, Jcost uniqueness via Aczél, AlexanderDuality D=3 forcing, 8-tick periodicity, phi-ladder constants) derive spacetime and constants from bare distinguishability; none of these structures appear in the log-processing pipeline or evaluation metrics. Domain mismatch is total.","tokens_in":58525,"confidence":"high","tokens_out":152,"duration_ms":12192,"cache_read_input_tokens":16512,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs achieve high precision on simple host logs for intrusion detection but degrade sharply as logs grow noisier and more complex.","keywords":["large language models","host-based intrusion detection","system logs","benchmark","cybersecurity","performance evaluation","intrusion detection"],"falsifier":"If the same models were tested on a deliberately more complex and noisier variant of the same log collections and still maintained MCC above 0.5 with stable false-positive rates, the claimed sensitivity to data complexity would be directly challenged.","tokens_in":2654,"feed_emoji":"🛡️","tokens_out":719,"duration_ms":28316,"temperature":0.7,"pith_summary":"The paper introduces HIDBench to test large language models on detecting intrusions from system logs, a task that demands fine-grained reasoning over large, noisy, and imbalanced data. It unifies three public datasets and supplies a pipeline to prepare raw telemetry for LLM input. Evaluation of frontier models shows they often reach precision above 0.8 on easier collections yet see Matthews correlation coefficient fall below 0.5 and false-positive rates climb on harder ones. A sympathetic reader would care because reliable host-based intrusion detection matters for everyday cybersecurity and knowing where current LLMs break helps decide when they can be trusted in practice.","feed_headline":"LLMs lose accuracy on complex noisy logs for intrusion detection","feed_subtitle":"Benchmark shows precision above 0.8 on simple datasets but MCC often below 0.5 as complexity and noise rise","key_machinery":"The HIDBench benchmark and its data construction pipeline, which unifies DARPA-E3, DARPA-E5, and NodLink datasets and converts raw host telemetry into LLM-compatible inputs for systematic evaluation under realistic settings.","core_discovery":"The paper claims that frontier LLMs exhibit substantial performance gaps across the unified datasets. While many models achieve high precision often above 0.8 on simpler datasets, performance degrades significantly as system logs become noisier and more complex, with MCC frequently dropping below 0.5 and false positive rates increasing sharply. Models fall into distinct regimes such as conservative detectors with low false positive rates and over-sensitive models that generate excessive alerts. The results indicate that LLMs hold strong potential for host-based intrusion detection yet their effectiveness remains highly sensitive to data complexity, making robust system design essential for可靠","pith_inferences":["Hybrid detectors that route simple logs to LLMs and complex logs to traditional rule-based methods could mitigate the observed drops.","The benchmark setup could be reused to compare fine-tuned versus zero-shot LLMs on the same log collections.","Similar evaluation pipelines might reveal whether the same complexity sensitivity appears in other log-driven security tasks such as anomaly detection."],"forward_implications":["LLMs can support host-based intrusion detection but need robust surrounding systems to handle varying data complexity.","Model behavior splits into conservative low-alert regimes and over-sensitive high-alert regimes depending on the input logs.","High precision on clean datasets does not predict success when logs contain more overlapping benign and malicious activity.","Deployment decisions for LLMs in intrusion detection must explicitly account for expected noise and imbalance levels."],"fun_headline_variants":["LLMs struggle on complex noisy logs for intrusion detection","HIDS benchmark exposes LLM limits on noisy system logs","LLM performance degrades as logs get noisier and complex","LLMs show high precision but MCC falls on complex logs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The data construction pipeline that transforms raw host telemetry into LLM-compatible inputs preserves the complex interactions between benign and malicious activities and does not introduce artifacts that artificially change detection difficulty.","fun_headline_variants_meta":{"raw":{"variants":["LLMs struggle on complex noisy logs for intrusion detection","HIDS benchmark exposes LLM limits on noisy system logs","LLM performance degrades as logs get noisier and complex","LLMs show high precision but MCC falls on complex logs"]},"model":"grok-4.3","cost_usd":0.009938,"raw_usage":{"total_tokens":4373,"prompt_tokens":742,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":99378000,"prompt_tokens_details":{"text_tokens":742,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3568,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":742,"tokens_out":63,"duration_ms":34363,"temperature":1.0,"reasoning_tokens":3568,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-22T08:48:55.701990+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If the same models were tested on a deliberately more complex and noisier variant of the same log collections and still maintained MCC above 0.5 with stable false-positive rates, the claimed sensitivity to data complexity would be directly challenged.","supporting_citations":[],"review_version":1}