{"id":"695e8415-0316-4884-822a-0dc6f0c05d1b","arxiv_id":"2507.03315","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A concept bottleneck model built from polarimetric target decomposition plus a Kolmogorov-Arnold Network gives PolSAR classification with human-auditable concept predictions and symbolic decision formulas.","lead":"This paper proposes an interpretable PolSAR image classification model that uses physical scattering concepts as an intermediate layer and replaces the final neural network with a Kolmogorov-Arnold Network. It reports competitive accuracy on three satellite datasets while letting users inspect and edit the concepts driving each prediction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Concept labels in Table I are category-level prototypes hand-derived from six GF3 patches; if training supervision uses these fixed vectors rather than per-pixel PTD, concept prediction and intervention reduce to category classification and the interpretability claim is unsupported.","rationale":"The reader identified the concept labels as the weakest assumption, and my reading sharpens that concern: the manuscript appears to use category-level prototypes, not per-pixel PTD measurements, as supervision. That would make the concept bottleneck a re-encoded categorical classifier rather than a physically grounded intermediate representation, which directly undermines the paper's novelty claim of transforming high-dimensional features into human-comprehensible, physically verifiable concepts. The accuracy results and ablation study (Tables II-V) are still plausible engineering evidence, and the intervention example in Fig. 10 is suggestive, but neither can certify interpretability if the concept labels are class priors. I therefore keep the reader's CONDITIONAL verdict: the condition should explicitly require the authors to release or precisely specify the per-pixel concept-labeling algorithm and to validate label transfer across sensors and frequency bands. Credit is due for the three-dataset comparison, the KAN-vs-MLP ablation, and the explicit attempt at symbolic and intervention-based interpretability; the concern is about the missing link between those demonstrations and the actual supervision signal, not about the integrity of the experiments.","tokens_in":23468,"tokens_out":9426,"duration_ms":116309,"concrete_test":"Re-implement concept-label generation from the 9-dim T input: compute Cloude-Pottier H/alpha/A, Freeman-Durden proportions, and Huynen parameters per pixel; apply Eq. (9)-(10) and the equal-interval splits described in Section III-B; and compare the resulting per-pixel concept vectors with the Table I prototypes on GF3-SF, RS2-SF, and Oberpfaffenhofen. Report per-class agreement and, for Oberpfaffenhofen (L-band), whether the GF3-calibrated labels reproduce. If agreement is high (above 90% per class), the per-pixel grounding and transferability assumptions hold; if not, the concept supervision is a category prior and the interpretability claims need to be re-benchmarked.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the construction of the concept labels in Section III-B. The manuscript states that the authors manually selected 200x200 patches of six GF-3 classes, summarized concept labels based on domain knowledge and statistical analysis, and reports one fixed concept vector per class in Table I. It never specifies a per-pixel algorithm that converts the 9-dimensional T input into the c_n used in the expanded dataset D_c. If each training sample is assigned the Table I prototype of its class, then the concept loss in Eq. (13) is a second categorical classification loss, Fig. 8 demonstrates category recognition rather than scattering-mechanism grounding, and the Table VI formulas map class prototypes to class labels rather than measured PTD concepts to labels. The transfer to RS2-SF and Oberpfaffenhofen is asserted via band-robust scattering characteristics, but Table I has no row for Oberpfaffenhofen's class names (built-up, wood land, open area), and the Eq. (9)-(10) thresholds and H/alpha/A interval splits were calibrated on six GF-3 patches without L-band validation. This is not an accusation of fraud; per-pixel labeling may exist. But as written, the central interpretability claim depends on an unstated and untested labeling procedure, so the evidence does not yet establish that decisions are traceable through physically verified scattering mechanisms.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes PaCBM-KAN, a PolSAR image classification framework that couples a parallel concept bottleneck network with Kolmogorov-Arnold networks. Concepts are defined from polarimetric target decomposition (Cloude-Pottier, Freeman-Durden, Huynen) and are claimed to provide physically meaningful intermediate representations. The model is evaluated on GF3-SF, RS2-SF, and Oberpfaffenhofen, reporting competitive accuracy (AA 92.16%, 97.28%, and 95.51%, respectively). Interpretability is demonstrated through concept prediction visualization, KAN-based symbolic formulas for concept-to-label mapping, and concept intervention for correcting misclassifications. The central claim is that high-dimensional PolSAR features can be transformed into human-comprehensible, physically grounded concepts while preserving classification accuracy and enabling traceable decision-making.","tokens_in":23771,"tokens_out":5929,"duration_ms":73020,"significance":"If the interpretability claim is fully supported, the paper would be a meaningful contribution to interpretable PolSAR classification: it is, to my knowledge, the first to combine PTD-derived concept bottlenecks with KAN, and it reports accuracy competitive with or above several strong baselines while offering concept prediction and intervention. The paper also includes a useful ablation of CBM versus PaCBM and KAN versus MLP, as well as a hyperparameter study for lambda and KAN structure. However, three evidentiary gaps currently limit the significance: the per-pixel concept-label generation procedure is not specified, the symbolic formulas in Table VI are presented without any fidelity check, and the cross-dataset transfer of concept labels is asserted rather than validated. These gaps are fixable, so the work is promising but not yet established.","major_comments":[{"comment":"The per-pixel assignment of concept labels to training samples is not specified. The text describes manually selecting 200x200 patches of six GF-3 classes and summarizes concept labels in Table I, but no algorithm is given that converts the 9-dimensional T-matrix input of each training sample into the concept vector c_n used in D_c. If each sample is simply assigned the Table I prototype of its class, then the concept loss in Eq. (13) is a second category-level classification loss, and the concept predictions in Fig. 8 demonstrate category recognition rather than scattering-mechanism grounding. Please provide the exact per-pixel labeling procedure, or clarify how the Table I prototypes are expanded to every pixel, and validate the resulting labels on held-out patches.","section":"Section III-B, Table I, Eq. (13)"},{"comment":"The symbolic formulas in Table VI are presented as evidence of traceable decision-making, but the manuscript reports no fidelity metric showing how well these formulas reproduce the trained KAN outputs. A closed-form expression that deviates from the learned function by a large error cannot support the interpretability claim. Please report the approximation error (e.g., R-squared or mean absolute error) between the formula and the KAN logits, over the concept-value ranges encountered in the test set, and describe how the formulas were extracted.","section":"Table VI, Section IV-C-2"},{"comment":"The concept labels are derived solely from six 200x200 patches of GF-3 (C-band) data, yet the same labels are used for RS2-SF and Oberpfaffenhofen. The Oberpfaffenhofen classes (built-up, wood land, open area) do not appear in Table I, and the thresholds in Eqs. (9)-(10) and the H/alpha/A interval splits were calibrated on GF-3 without L-band validation. The claim that the scattering mechanisms are band-robust needs direct evidence: please provide per-dataset concept-label statistics or a concept-prediction accuracy evaluation on the validation/test splits of RS2-SF and Oberpfaffenhofen.","section":"Section IV-A, Table I"},{"comment":"All classification results are reported from a single run, with no error bars or repeated-seed statistics. The differences among the best-performing methods are often small (e.g., Table III OA 98.04% for HyCVNet versus 98.06% for PaCBM-j, and Kappa 97.20% versus 97.23%). The statements that PaCBM 'outperforms' the baselines are therefore not statistically supported. Please report means and standard deviations over at least three independent runs with different random seeds.","section":"Tables II-V"}],"minor_comments":[{"comment":"The sentence 'the features could be conceptualization' is ungrammatical; it should be 'the features could be conceptualized.'","section":"Abstract"},{"comment":"The text refers to 'HyCBNet' in the GF3-SF and Oberpfaffenhofen analyses, but the tables and references use 'HyCVNet'; please make the naming consistent.","section":"Section IV-B-1"},{"comment":"The Oberpfaffenhofen dataset name is spelled 'Oberpfaffenhofe' in the dataset description; the standard spelling is 'Oberpfaffenhofen.'","section":"Section IV-A"},{"comment":"Equation (6) is typeset in a way that is hard to read: the definitions of log3 and alpha_n are not clearly separated from the formula, and the notation for P_n and alpha_n deserves explicit definitions before the equation.","section":"Section III-B"},{"comment":"References [30] and [60] are duplicates (both are Alkhatib's HybridCVNet paper); renumber to cite the work only once.","section":"References"},{"comment":"The caption of Table VI does not define the concept indices c1...c33 or state that these correspond to the ordering in Table I; please add a mapping so the formulas can be interpreted.","section":"Section IV-C-2"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about concept-label circularity is, in my reading, valid and load-bearing: the manuscript never specifies how the manually summarized class-level prototypes become per-pixel training labels. This is not an accusation of misconduct, but it is a central omission that must be resolved before the interpretability claims are publishable. The paper also does not release code or data, which would materially help verify the labeling procedure. The contribution is within the journal's scope and the accuracy results are promising, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a decent engineering contribution—the first CBM for PolSAR that I know of—and the classification numbers are competitive. But the central interpretability story has a load-bearing gap: the paper never specifies how per-sample concept labels are generated. If the training supervision uses the Table I category prototypes, then the concept loss is just a second categorical loss and the interpretability evidence dissolves into category recognition.\n\nThe genuinely new piece is bringing the CBM/KAN machinery into PolSAR and grounding the concept vocabulary in PTD features. The parallel-branch design to mitigate CBM information loss is also reasonable, and the ablation on RS2-SF (baseline vs CBM vs PaCBM-MLP vs PaCBM-KAN) is clean and shows the expected ordering. Three datasets, sensible hyperparameter search, honest acknowledgment of KAN's parameter cost. That is real work and worth publishing as an application.\n\nThe soft spots, in order of severity.\n\nFirst, Section III-B and Table I. The concept labels are category-level prototypes built from six manually picked GF-3 patches. There is no per-pixel rule described for filling in c_n in D_c, no statement about how thresholds generalize across the C-band datasets, and no row at all for Oberpfaffenhofen's classes in Table I, even though the labels are used there. If c_n is just the class prototype, then Eq. (13) is a classification loss in disguise, Fig. 8 demonstrates category recognition, and the Table VI formulas map prototypes to labels. The stress-test note is right; I don't think this is fraud, but the paper as written does not establish physical traceability.\n\nSecond, the interpretability evidence is under-quantified. Table VI shows symbolic formulas with no fidelity metric; a formula can look clean and not reproduce the KAN outputs. There are no error bars or repeated runs anywhere. The concept prediction examples are qualitative, and the intervention demonstration is a single sample. These are fixable, and not fatal to the application claim, but they are real limits.\n\nBottom line: this is a serious paper for the PolSAR interpretability community and for anyone applying CBMs in a physical domain. It deserves peer review, but the referee should be asked to provide the exact concept-labeling pipeline, per-pixel labels or a defensible rule for assigning prototypes, and a fidelity check for the symbolic formulas. Without those, the interpretability claim is category-level classification in a CBM wrapper.","headline":"A useful first CBM/KAN application to PolSAR whose main interpretability claim is undercut by an unspecified concept-labeling procedure and missing fidelity checks.","tokens_in":24260,"tokens_out":3352,"would_cite":true,"duration_ms":41629,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims PolSAR deep classifiers can be opened up by routing features through physical scattering concepts, then mapping concepts to labels with a Kolmogorov-Arnold network that yields closed-form formulas, without losing…","keywords":["PolSAR image classification","concept bottleneck model","polarimetric target decomposition","scattering mechanism","Kolmogorov-Arnold network","interpretable deep learning","symbolic formula","remote sensing"],"falsifier":"Take any test pixel, compute its PTD parameters directly from its coherency matrix, and compare them with the concept scores the network predicts for that pixel; if pixels with high predicted 'double-bounce scattering dominant' do not actually have dominant Freeman–Durden double-bounce power, or if systematically mislabeled concept training patches are demonstrably the source of concept errors, then the bottleneck is not actually encoding the claimed physics.","tokens_in":23314,"feed_emoji":"🛰️","tokens_out":6220,"duration_ms":67585,"temperature":0.7,"pith_summary":"This paper tries to open the black box of deep-learning PolSAR image classification. It replaces the usual final feature-to-label mapping with two transparent stages: a concept bottleneck layer that predicts human-comprehensible labels such as 'double-bounce scattering dominant' or 'high polarization entropy,' and a Kolmogorov-Arnold network that combines those concept predictions into a class decision. Because the concept labels come from polarimetric target decomposition, the intermediate layer is tied to physically verifiable scattering mechanisms rather than arbitrary semantics. The authors show on three datasets that this design reaches accuracy comparable to or better than conventional deep models, and they report that the concept-to-label mapping can be written as closed-form symbolic formulas. If this holds, the model's reasoning is traceable end to end: raw polarimetric data to physical concepts to an explicit decision rule.","feed_headline":"PolSAR networks justify decisions with scattering formulas","feed_subtitle":"Deep features become readable scattering-based rules, with accuracy on par with black-box classifiers.","key_machinery":"The load-bearing object is the polarimetric concept label set, built from Cloude–Pottier ($H$, $\\alpha$, $A$), Freeman–Durden (surface, double-bounce, volume components), and Huynen (nine target parameters) decompositions. These decompositions turn the $3\\times 3$ coherency matrix into physically meaningful attributes, and the paper discretizes those attributes into binary concepts such as 'surface scattering dominant' and 'high polarization entropy.' The PaCBM structure runs the concept branch in parallel with the main classifier over a shared CNN and vision transformer encoder with cross-attention, so the concept bottleneck supervises feature learning without forcing all information through it. The KAN then maps predicted concepts to class logits via learned univariate spline functions, which is what enables the symbolic readout.","core_discovery":"The central claim is that interpretability does not have to be bought with accuracy: a Parallel Concept Bottleneck Network built on PTD-derived concept labels plus a Kolmogorov-Arnold network can classify PolSAR images while exposing the decision process. The parallel branch keeps a standard feature-to-label path for accuracy while a concept branch learns to predict PTD-based concepts, with a shared encoder and a weighted concept loss. Replacing the MLP with KAN lets the concept-to-label map be expressed as sums of B-spline univariate functions; the paper converts selected KAN nodes into human-readable formulas, such as the logit for 'Mountain' expressed as combinations of sine, exponential, polynomial, and linear terms in concept variables. The authors report average accuracies of 92.16% on GF3-SF, 97.28% on RS2-SF, and 95.51% on Oberpfaffenhofen, and demonstrate intervention: flipping a wrongly activated concept from 0.84 to 0 changes a misclassified sample's label to the correct 'Developed' class.","pith_inferences":["The same recipe could be applied to other physical priors: using interferometric phase or coherence descriptors as concepts would extend this style of auditability to InSAR and PolInSAR models.","A direct extension is to report per-concept prediction accuracy; if certain concepts are hard to predict, the symbolic formulas may be assigning weight to concept values the network does not actually estimate reliably.","Cross-sensor transfer audits are a natural next test: applying the GF-3-derived concept vocabulary to L-band data would reveal whether concept disagreements encode genuine frequency-dependent scattering or just label noise.","One could also test the faithfulness of the extracted formulas by removing individual concepts and comparing the formula's sensitivity with the network's actual response, which the paper does not report."],"forward_implications":["A user can audit any single classification by reading the predicted concept scores and checking them against physical expectations for that terrain type.","Because the concept-to-label stage is written as symbolic formulas, decision rules can be inspected, simplified, and compared across sensors without re-running the network.","Concept intervention becomes a practical debugging tool: manually correcting an erroneous concept activation should change the output to the intended class, as shown for a 'Developed' sample.","The parallel branch means concept supervision can regularize feature learning instead of forcing all information through the bottleneck, so interpretability does not necessarily cost accuracy.","If the concept vocabulary is band-robust as claimed, the same PTD-derived concepts can supervise classification at both C-band and L-band without rebuilding labels."],"supporting_citations":[{"why":"Defines the concept bottleneck model that the paper adapts to PolSAR by routing features through interpretable concepts.","marker":"[22]"},{"why":"Introduces Kolmogorov-Arnold networks with spline-based univariate functions, the mechanism used to replace the MLP and obtain symbolic formulas.","marker":"[25]"},{"why":"Freeman–Durden three-component decomposition provides surface, double-bounce, and volume scattering components used to build concept labels.","marker":"[19]"},{"why":"Cloude–Pottier entropy, anisotropy, and alpha angle provide the decomposition parameters that become concept labels.","marker":"[55]"},{"why":"Huynen decomposition supplies the nine target parameters used as additional physical concept labels.","marker":"[56]"},{"why":"The HyCVNet dual-branch CNN and vision transformer with cross-attention is the encoder architecture the PaCBM design adopts.","marker":"[60]"},{"why":"CBAM attention module is embedded in the CNN branch of the encoder.","marker":"[61]"},{"why":"Provides the GF3-SF and RS2-SF San Francisco datasets and their terrain categories used in experiments.","marker":"[58]"},{"why":"Supports the relative-proportion rule used to decide dominant, secondary, and weak scattering concept labels.","marker":"[59]"},{"why":"ProtoPNet is the interpretable baseline that the proposed method is compared against on all three datasets.","marker":"[66]"}],"fun_headline_variants":["PolSAR classification becomes transparent with physics-based concepts","Scattering mechanics explain PolSAR decisions without accuracy loss","Concept bottleneck plus KAN opens the black box in PolSAR","Human-readable formulas for PolSAR classification from scattering concepts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's interpretability evidence rests on the assumption that the concept labels it hand-built from six 200x200 patches of one GF-3 image are correct, complete, and valid for all three datasets; if those labels do not reflect real scattering behavior or omit concepts the network actually uses, the concept predictions and symbolic formulas explain the labels, not the model's true reasoning.","fun_headline_variants_meta":{"raw":{"variants":["PolSAR classification becomes transparent with physics-based concepts","Scattering mechanics explain PolSAR decisions without accuracy loss","Concept bottleneck plus KAN opens the black box in PolSAR","Human-readable formulas for PolSAR classification from scattering concepts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000473,"raw_usage":{"total_tokens":2390,"prompt_tokens":1028,"completion_tokens":1362,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":1298}},"tokens_in":644,"tokens_out":1362,"duration_ms":10613,"temperature":1.0,"reasoning_tokens":1298,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:13:30.270140+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any test pixel, compute its PTD parameters directly from its coherency matrix, and compare them with the concept scores the network predicts for that pixel; if pixels with high predicted 'double-bounce scattering dominant' do not actually have dominant Freeman–Durden double-bounce power, or if systematically mislabeled concept training patches are demonstrably the source of concept errors, then the bottleneck is not actually encoding the claimed physics.","supporting_citations":[{"cited_title":"Concept bottleneck models,","cited_arxiv_id":null,"evidence_quote":"Defines the concept bottleneck model that the paper adapts to PolSAR by routing features through interpretable concepts."},{"cited_title":"A three-component scattering model for polarimetric sar data,","cited_arxiv_id":null,"evidence_quote":"Freeman–Durden three-component decomposition provides surface, double-bounce, and volume scattering components used to build concept labels."},{"cited_title":"An entropy based classification scheme for land applications of polarimetric sar,","cited_arxiv_id":null,"evidence_quote":"Cloude–Pottier entropy, anisotropy, and alpha angle provide the decomposition parameters that become concept labels."},{"cited_title":"Phenomenological theory of radar targets,","cited_arxiv_id":null,"evidence_quote":"Huynen decomposition supplies the nine target parameters used as additional physical concept labels."},{"cited_title":"Polsar image classification using a hybrid complex- valued network (hybridcvnet),","cited_arxiv_id":null,"evidence_quote":"The HyCVNet dual-branch CNN and vision transformer with cross-attention is the encoder architecture the PaCBM design adopts."},{"cited_title":"Polsf: Polsar image datasets on san francisco,","cited_arxiv_id":null,"evidence_quote":"Provides the GF3-SF and RS2-SF San Francisco datasets and their terrain categories used in experiments."},{"cited_title":"Assessment of model-based polsar decompositions,","cited_arxiv_id":null,"evidence_quote":"Supports the relative-proportion rule used to decide dominant, secondary, and weak scattering concept labels."}],"review_version":1}