{"id":"ed92b36c-a3db-48c5-bcae-617b6cea0038","arxiv_id":"2607.05641","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Vision models trained on chart images of time series often latch onto rendering style rather than temporal class structure, an effect the authors call visual encoding hijacking.","lead":"Chart-style images of time series make CNNs learn encoding-specific visual cues (line edges, bar heights, scatter density) more than pure temporal structure. The paper shows this bias with similarity, transfer, and attribution tests and finds attention guidance helps only when the bias is already clear.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The joint diagnostic suite still cannot separate hijacking from legitimate encoding-specific evidence, so the central framing overclaims what the metrics establish.","rationale":"The reader correctly isolates the definitional gap that the paper itself flags but does not close. The multi-metric design (CKA matrices, probe transfer, ESI, UMAP, HINT Table 1) is clean and the empirical patterns are consistent, so the work remains publishable after the authors either (a) supply a controlled synthetic test that can falsify the hijacking interpretation or (b) rephrase the central claim to “encoding dependence is pervasive and must be measured” without asserting that the dependence is illegitimate. No mathematical error or data issue is present; the load-bearing weakness is purely interpretive. Hence the verdict stays CONDITIONAL and confidence remains high.","tokens_in":9999,"tokens_out":466,"duration_ms":4299,"concrete_test":"Construct a synthetic UCR-style suite in which class labels are defined solely by a known temporal motif (e.g., a fixed-frequency sinusoid segment) while all other statistics are matched; render the identical series under the four encodings, recompute ESI_CKA, ESI_Probe and Grad-CAM overlap with the motif mask. If the joint diagnostics still flag high “hijacking” even when the only discriminative information is temporal, the separation claim fails; if they remain low, the original interpretation is supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s strongest claim is that low off-diagonal CKA, weak cross-encoding linear-probe transfer, and Grad-CAM/perturbation sensitivity jointly diagnose “visual encoding hijacking” (encoding-dependent shortcuts rather than temporal structure). Yet Sections 3.2, 4.2 and the Discussion repeatedly concede that the same pattern is also produced by legitimate encoding-specific evidence that faithfully reflects signal properties (line continuity, bar heights, scatter density). Because no ground-truth temporal-only baseline or controlled synthetic series is provided that would make the two cases diverge, the joint suite remains definitional rather than discriminative. The claim that chart design “must be treated as a representation/measurement choice” therefore rests on an untested equivalence between observed encoding dependence and illegitimate bias.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies whether CNNs trained on chart renderings of time series learn encoding-invariant temporal structure or encoding-dependent visual cues. It introduces “visual encoding hijacking,” diagnosed jointly via low cross-encoding CKA, weak linear-probe transfer, Grad-CAM attribution shifts, and perturbation sensitivity, and evaluates four chart encodings (line, area, bar, scatter) on 31 UCR datasets under a fixed CNN/training protocol. An Encoding Sensitivity Index (ESI) summarizes dataset-level dependence; HINT-style attention guidance is tested as a mitigation. The main claim is that chart design acts as an inductive bias that shapes learned representations, so chart-based TSC should be treated as a representation/measurement problem rather than a neutral modeling choice.","tokens_in":10208,"tokens_out":1431,"duration_ms":22253,"significance":"If the encoding-dependence results hold, the paper makes a useful contribution to image-based TSC and to machine graphical perception: it shows that visualization design is not a free parameter and supplies a reusable multi-metric diagnostic suite (CKA, cross-probe, ESI, Grad-CAM, controlled perturbations) under a clean, fixed-protocol design across 31 datasets. Strengths include the complementary metrics, the transparent caveats that perturbations/attributions alone do not prove task-irrelevance, the nuanced (not universal) HINT results, and the explicit link to human graphical perception. The work is significant as a diagnostic and framing paper even if the stronger “hijacking = illegitimate bias” reading is softened.","major_comments":[{"comment":"Central framing vs. evidence (Abstract, §1 definition of hijacking, §3.2 Chart Perturbations, §4.2, Discussion): The paper defines hijacking as encoding-dependent behavior diagnosed by the joint suite (low off-diagonal CKA, weak transfer, attribution/perturbation sensitivity) and titles/frames the work as encoding-induced bias. Yet the same sections repeatedly state that this pattern is also produced by legitimate encoding-specific evidence that faithfully reflects signal structure (line continuity, bar heights, scatter density), and that perturbations do not establish task-irrelevance. Without a control that separates the two cases (e.g., synthetic series with known temporal-only labels, or a temporal-only baseline that forces the metrics to diverge), the joint suite remains definitional rather than discriminative. Either add such a control or align title/abstract/claims with the weaker","section":"Abstract, §1, §3.2, §4.2, Discussion"},{"comment":"§3.2 “Generalization Across Encoding Families” promises a robustness check on five mathematical transforms (GASF, MTF, RP, CWT, STFT) under the same protocol, but §4 reports no corresponding CKA, probe, ESI, or accuracy results. This is a load-bearing gap for the claim that the effect is not an artifact of chart-style rendering alone. Report the results or remove the claim.","section":"§3.2, §4"},{"comment":"Fig. 6a reports max–min accuracy Δ = 0.023 across four encodings over 31 datasets, while the narrative emphasizes substantial encoding effects on what models learn. Representational divergence (CKA/UMAP/ESI) can matter even when accuracy is similar, but the paper should reconcile the small accuracy spread with the practical claim that encoding choice is consequential for modeling decisions, and clarify when representation dependence is harmful versus merely different.","section":"§4.2, Fig. 6a"},{"comment":"§3.3 / Table 1: HINT uses model-derived Grad-CAM saliency as the alignment target (self-supervised adaptation). Large degradations (Computers −36%, SonyAIBORobotSurface1 −14.14%) are reported but under-analyzed. If HINT is offered as evidence that attention guidance can reduce encoding-sensitive behavior when diagnostics agree, the paper needs a clearer account of when the self-derived target is reliable versus when it reinforces the wrong cues, and whether those failures undermine the diagnostic suite that selected the “encoding-divergent” cases.","section":"§3.3, Table 1, §4.3"}],"minor_comments":[{"comment":"Footnote 1 expands VEIL as “Visual Encoding Illusion in Learning,” while the title uses “Hijacking.” Align the expansion with the title term or explain the distinction.","section":"Footnote 1 / title"},{"comment":"Fig. 4 caption says “FordB and ECG5000” but the subcaptions are “UMAP- Yoga” and “UMAP- ECG5000”; fix the mismatch.","section":"Fig. 4"},{"comment":"Eq. (1) writes CKA(X,Y) with H_a, H_b introduced just above; the formula uses X,Y without stating X=H_a, Y=H_b. Minor notational cleanup.","section":"§3.2, Eq. (1)"},{"comment":"ESI_Probe = 1 − mean off-diagonal probe accuracy mixes a similarity-style index with an accuracy scale; state the accuracy range and whether chance-level baselines are subtracted so ESI values are comparable across multi-class datasets.","section":"§3.2 Sensitivity Indices"},{"comment":"Several citations appear as CoRR/arXiv or “VTBench” concurrent work; ensure final versions and venue details are updated where available.","section":"References"},{"comment":"Abstract says attention-guided training “appears to mitigate this effect when encoding sensitivity is consistently identified”; Table 1 also shows null and large negative deltas. Soften the abstract sentence to match the “diagnostic intervention, not reliable mitigation” stance in §4.3.","section":"Abstract, §4.3"}],"recommendation":"major_revision","confidential_remarks":"The empirical design is careful and the multi-metric suite is a real contribution; the main risk is rhetorical overclaim via the “hijacking/bias” framing relative to what the metrics can establish. If the authors reframe to encoding-as-measurement and either add a synthetic control or drop the illegitimacy claim, this is a solid journal paper. Scope fits CV/visualization/ML venues that care about representation bias and graphical perception; less of a pure TSC methods paper."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: on 31 UCR sets, the same CNN trained under identical protocol learns clearly encoding-dependent representations of the identical series—low off-diagonal CKA, weak linear-probe transfer, UMAP clustering by chart type, Grad-CAM on strokes/edges/density—and HINT only helps when those diagnostics already agree. That pattern is real and useful.\n\nWhat is actually new is the systematic multi-diagnostic framing (CKA + cross-encoding probes + ESI + attribution + controlled perturbations) applied to four ordinary chart families rather than a single exotic transform, plus the conditional HINT result. The experimental design is clean: fixed backbone, Adam defaults, early stopping, no per-dataset tuning. They are honest in the text that perturbations and saliency do not prove the cues are task-irrelevant. Citations cover the image-TSC and shortcut literature without obvious gaps; the anonymous repo is a plus.\n\nThe soft spot is definitional, not experimental. The paper repeatedly concedes (methods, results, discussion) that the same joint signature can arise from legitimate encoding-specific evidence that still faithfully reflects the signal. Without a temporal-only baseline or synthetic series that would make the two cases diverge, “visual encoding hijacking” remains a useful label for observed dependence rather than a demonstrated shortcut. Secondary issues: λ for HINT is unreported, no significance tests on ESI differences, and a few HINT degradations are large. Those are fixable and do not erase the empirical patterns.\n\nThis is for people building chart-based or vision-backbone TSC pipelines and for anyone interested in how machines read visualizations. It deserves a serious referee. I would engage with the diagnostics and the HINT conditionality; I would treat the stronger causal language as provisional until the distinction is tightened. Send it to review.","headline":"Clean multi-metric evidence that chart encodings reshape CNN features on UCR data; the “hijacking” label still outruns what the diagnostics can separate.","tokens_in":10818,"tokens_out":459,"would_cite":true,"duration_ms":10503,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Chart design can hijack what CNNs learn from time-series images, so models often latch onto rendering style instead of the signal itself.","keywords":["time-series classification","chart-based representations","CNNs","visual encoding hijacking","representation similarity","CKA","Grad-CAM","encoding sensitivity"],"falsifier":"A controlled experiment in which two encodings that differ only in task-irrelevant visual cues (for example stroke style or bar edge thickness) still produce high CKA, high probe transfer, and Grad-CAM maps focused on the same temporal regions would falsify the claim that those cues systematically hijack the representation.","tokens_in":10881,"feed_emoji":"📈","tokens_out":605,"duration_ms":5062,"temperature":0.7,"pith_summary":"When time series are drawn as line, area, bar, or scatter charts and fed to CNNs, the networks frequently learn features that depend on how the chart was drawn rather than on the underlying temporal class structure. The authors call this visual encoding hijacking and diagnose it with a joint suite of tests: representation similarity (CKA), linear-probe transfer across encodings, and Grad-CAM attribution plus controlled rendering perturbations. Across 31 UCR datasets the same series produce divergent feature spaces and weak cross-encoding transfer for many complex multi-class problems, while simple binary problems remain more invariant. Attention guidance (HINT) can reduce the sensitivity when the diagnostics agree that hijacking is present, but it helps little or hurts when they do not. The practical claim is that chart-based time-series classification is a representation and measurement choice, not a neutral modeling detail.","feed_headline":"Chart style hijacks what CNNs learn from time series","feed_subtitle":"Same signal, different plots: models often learn the drawing, not the data","key_machinery":"VEIL: a joint diagnostic framework that combines Centered Kernel Alignment (CKA) for representation similarity, cross-encoding linear probing for transferability, Grad-CAM attribution and controlled chart perturbations for evidence localization, plus scalar Encoding Sensitivity Indices (ESI) that summarize off-diagonal divergence.","core_discovery":"Visual encoding hijacking occurs: CNN representations of the same time series become encoding-dependent (low cross-encoding CKA, weak linear-probe transfer, attribution shifts toward rendering cues) rather than encoding-invariant temporal structure, so chart design must be treated as a representation and measurement decision rather than a simple modeling choice.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["CNNs learn chart style over time-series patterns","Visual encodings bias what vision models extract from plots","Same data different charts: CNNs hitch on rendering not signal","Chart design hijacks CNN representations of time series","Encoding choices shape what CNNs see in time-series charts"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The joint pattern of low cross-encoding similarity, weak transfer, and attribution or perturbation sensitivity is enough to label the behavior as illegitimate hijacking rather than legitimate encoding-specific evidence that still reflects true signal structure.","fun_headline_variants_meta":{"raw":{"variants":["CNNs learn chart style over time-series patterns","Visual encodings bias what vision models extract from plots","Same data different charts: CNNs hitch on rendering not signal","Chart design hijacks CNN representations of time series","Encoding choices shape what CNNs see in time-series charts"]},"model":"grok-4.5","effort":"low","cost_usd":0.003814,"raw_usage":{"total_tokens":1113,"prompt_tokens":668,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":38140000,"prompt_tokens_details":{"text_tokens":668,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":365,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":668,"tokens_out":80,"duration_ms":3060,"temperature":1.0,"reasoning_tokens":365,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T04:28:16.297974+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled experiment in which two encodings that differ only in task-irrelevant visual cues (for example stroke style or bar edge thickness) still produce high CKA, high probe transfer, and Grad-CAM maps focused on the same temporal regions would falsify the claim that those cues systematically hijack the representation.","supporting_citations":[],"review_version":1}