{"id":"435e8cc2-80b6-4474-ada8-052be41de41e","arxiv_id":"1908.07124","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"LAMA extends the self-organizing map with user-specified landmarks, pairs of data points and nodes, trained by alternating ordinary SOM updates with landmark-pulling updates to create user-intended nonlinear projections.","lead":"This paper proposes a landmark map (LAMA), a variant of the self-organizing map that lets a user pin selected data points to chosen grid nodes so the projection follows the user's intent. It matters because it offers a simple way to make unsupervised visualizations and human-computer interfaces more controllable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Final-iteration QEL rises in two of the paper's own LAMA experiments, so the claimed landmark retention is not demonstrated at the converged state.","rationale":"The reader's weakest assumption focuses on the existence and discoverability of LAMA's parameter settings. My concern is closely related but more specific: even in the paper's own reported experiments, the landmark-retention objective is not stably achieved at the final training step (LAMA2 and LAMA5 QEL increase at the end), and the evaluation of success is partly circular because QEL directly measures the landmark-driven training objective. This does not overturn the reader's CONDITIONAL verdict; it reinforces it. The paper convincingly describes a reasonable algorithmic extension of SOM and provides plausible qualitative demonstrations, but the central claim about reliable user-intended projection needs stronger endpoint evidence, independent success metrics, and statistical validation. Therefore the appropriate verdict remains CONDITIONAL, and no change to the reader's assessment is needed.","tokens_in":17256,"tokens_out":3158,"duration_ms":33468,"concrete_test":"Reproduce the LAMA5 formant experiment with Table 2 parameters for 100 random initializations. At t=0, t=tcenter=30000, and t=tmax=59999, record: (i) QEL, (ii) STE, and (iii) the fraction of runs in which every one of the five landmark data points has the designated landmark node as its nearest codebook vector. Also record the iteration at which QEL is minimal. If the visualized Figure 16 map is not the tmax state, or if fewer than 95% of runs satisfy all landmark assignments at tmax, then the claimed final-state landmark retention is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that after training is complete, the LAMA map simultaneously retains landmark relationships and smooth topology. The paper's own numerical results contradict this requirement for two of its five LAMA configurations: Section 3.3 states that the QEL of LAMA2 'increased toward the final iteration,' and Section 4.3 states that the QEL of LAMA5 'increased toward the end of learning,' meaning landmark nodes moved away from their assigned landmark data at the end of training. Yet the paper does not state which learning iteration produced the visualizations in Figures 8-11 and 16, so the displayed 'user-intended' projections may reflect a transient mid-training state rather than the final converged map. This matters because the central claim is about what LAMA learns, not about a momentary configuration during training. Furthermore, QEL is precisely the term minimized by the landmark-driven phase, so reporting low QEL is not independent evidence of success; the evaluation is circular with respect to the training objective. The only quantitative support for the central claim is therefore hand-selected visual inspection with no statistical characterization, and Section 5 explicitly concedes that parameter setting is manual and that missetting can cause twisting or wrinkling. The load-bearing premise—that a parameter setting exists, and is discoverable, which yields a final map satisfying both landmark constraints and topology—is not established by the provided evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LAMA, an extension of the self-organizing map in which a user designates landmark pairs, each pairing a data point with a specific output-space node. Training alternates a data-driven SOM phase with a landmark-driven phase that moves the landmark node's codebook vector toward its assigned landmark data. The paper evaluates LAMA on the UCI Zoo dataset with four landmark configurations and on an artificial formant dataset intended to associate tongue/jaw movement with cursor movement. The central claim is that, by setting a small number of landmarks, a user can obtain a nonlinear projection that preserves SOM-like topology while imposing chosen semantic constraints, leading to landmark-centered visualization and HCI applications.","tokens_in":17557,"tokens_out":4671,"duration_ms":50909,"significance":"If validated, LAMA would fill a real gap: existing supervised or semi-supervised SOM variants use class labels, whereas LAMA directly expresses the user's intended correspondence between data points and locations in the output space. The algorithm is clearly specified through explicit equations (Eqs. 1-18), the learning rules are transparent, and the Zoo examples do illustrate a qualitatively different view from a standard SOM. The main weakness is that the quantitative evidence is not independent of the training objective: QEL is the term minimized by the landmark-driven update, and the central success claim rests on hand-selected landmark placements and manually tuned parameters. With additional quantitative, final-state evaluation and a more principled parameter protocol, the contribution would be of interest to the SOM and visualization communities.","major_comments":[{"comment":"QEL is not an independent success measure. Eq. (14) measures the distance between each landmark data point and its assigned landmark codebook vector, which is precisely the quantity that Eq. (11) reduces during the landmark-driven phase. Therefore the low QEL values reported in Sections 3.3 and 4.3 largely confirm that the optimization objective was met, not that the resulting projection is semantically useful. I recommend adding an independent evaluation, such as testing on held-out landmark pairs, measuring how well the projected layout matches the intended semantic directions, or comparing against a baseline that optimizes the same objective without the SOM topology term.","section":"Section 2.4.2 (Eq. 14) and Section 5"},{"comment":"The paper does not establish that the reported visualizations come from the converged final state. Section 3.3 states that the QEL of LAMA2 'increased toward the final iteration,' and Section 4.3 states that the QEL of LAMA5 'increased toward the end of learning,' while Fig. 19 shows an increasing STE for LAMA5. Yet the training iteration that produced Figures 8-11 and 16 is not stated. Since the central claim concerns what LAMA learns, not a transient mid-training configuration, the visualization and all error indices should be reported for the final iteration t = tmax, and the final values with variability across the 100 runs should be given explicitly.","section":"Sections 3.3 and 4.3, Figures 8-11 and 16"},{"comment":"The parameter-setting premise is load-bearing but unsupported. The paper concedes that there is 'no strict rule for setting the parameters' and that missetting can cause the mesh grid to be twisted or wrinkled. Every reported LAMA variant uses a different manual combination of bmin, pth, ρb, ρmin, and other parameters, and the choice of landmarks is itself hand-picked. The manuscript therefore does not demonstrate that a typical user can discover a parameter setting that yields a smooth grid satisfying arbitrary landmark constraints. I ask for a sensitivity analysis over the parameter ranges, a proposed automatic or semi-automatic tuning rule, or at least multiple landmark configurations with fixed shared parameters to show that the result is not specific to the chosen hand-tuned runs.","section":"Section 5"},{"comment":"The formant experiment's success claim is supported only by visual inspection. The text says that the mesh grid has 'a line between /e/ and /o/' and 'a line between /u/ and /a/', but no numerical measure quantifies how well the tongue and jaw directions align with the horizontal and vertical axes of the output space. Because this is the paper's main demonstration of a user-intended projection for HCI, I recommend a quantitative alignment measure, such as the angle between the landmark-pair direction in codebook space and the corresponding output-space axis, evaluated over multiple random initializations.","section":"Section 4.2"}],"minor_comments":[{"comment":"The caption says 'The number of codebook vectors (N) is the same as that of location vectors,' but the number of codebook/location vectors is K, not N; N is the number of data points. Please correct this notation error.","section":"Fig. 2 caption"},{"comment":"The text after Eq. (8) refers to τb as the parameter controlling the distribution over steps, but Eq. (8) uses ρb. Please make the text and equations consistent.","section":"Section 2.2.2, Eq. (8)"},{"comment":"The formant experiment is called 'LAMA' in Table 2 and Section 4.2, but Section 4.3 introduces the name 'LAMA5' without warning. Use a single label consistently, or explicitly state that LAMA5 is the same configuration.","section":"Sections 4.2 and 4.3"}],"recommendation":"major_revision","confidential_remarks":"The core idea is simple and worth reporting, but the empirical section currently reads as a qualitative demonstration with hand-tuned parameters on two datasets. The QEL-based evidence is circular with respect to the training objective, and the final-iteration behavior of LAMA2 and LAMA5 is not reconciled with the visual claims. I would be willing to accept after the authors provide final-state quantitative metrics, an independent evaluation of the intended projection, and a more systematic treatment of parameter sensitivity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"LAMA is a real, clearly specified extension of SOM—user picks data-point/node pairs and training alternates data-driven and landmark-driven updates. That is new relative to the cited supervised/semi-supervised SOM work. What the paper does not do is demonstrate the central claim that the final converged map retains the landmark relationships. QEL is exactly the quantity the landmark phase minimizes, so low QEL is not independent evidence. More importantly, two of the paper's own experiments (LAMA2 on Zoo, LAMA5 on formants) show QEL rising toward the end of learning, and the paper never states which iteration produced the figures. The appealing landmark-aligned visualizations could be mid-training snapshots. That is a genuine gap, not a quibble.\n\nThe paper does some things well. The algorithm is precisely specified, the problem—user-intended projection rather than label-based classification—is genuinely distinct from earlier supervised SOMs, and the discussion is honest about manual parameter setting and about twisted or wrinkled grids when parameters are off. The Zoo examples are a plausible visual proof of concept, and the formant-to-cursor demo is a reasonable synthetic illustration. I also give the paper credit for not pretending the method is automatic.\n\nSoft spots, in order of importance. Convergence behavior is unanalyzed; the QEL rise suggests landmark retention can degrade precisely where the user cares about it. Evaluation is qualitative on two datasets, with learning parameters tuned per variant and only mean curves reported. No code is provided, and the formant data appears only as promised supplementary material. There is no comparison to a simpler approach, such as starting from a SOM and adding a soft penalty on selected nodes. Section 5 explicitly concedes that there is no strict rule for setting parameters, so the load-bearing premise—that a workable parameter setting exists and is discoverable for a user's landmarks—is not established.\n\nWho is this for: researchers working on SOM-based visualization or HCI mapping. I would take it seriously and send it to a referee rather than desk reject; the idea is concrete and new. But I would expect revision before trust: at minimum, final-iteration behavior, a reproducibility package, and at least one quantitative comparison.","headline":"Genuinely new SOM extension, but the landmark-retention claim is undercut by the paper's own convergence plots and the evaluation is mostly qualitative.","tokens_in":18084,"tokens_out":3832,"would_cite":false,"duration_ms":39485,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The landmark map extends the self-organizing map so that a few user-chosen data-to-node pairs steer the projection toward a designated semantic layout.","keywords":["landmark map","self-organizing map","semi-supervised learning","nonlinear projection","data visualization","topology preservation","human-computer interaction","formant analysis"],"falsifier":"Run LAMA many times on the artificial formant dataset with the paper's landmarks, but with a grid search over $p_{\\rm th}$, $b_{\\min}$, $b_{\\max}$, $t_{\\rm center}$, $\\rho_b$, $\\rho_{\\min}$, and $\\rho_{\\max}$; if every configuration that drives the landmark quantization error near zero also produces a wrinkled or self-intersecting mesh, with square topographic error well above the plain SOM's, then user-intended projections are not reliably attainable at the claimed level of generality.","tokens_in":17004,"feed_emoji":"🗺️","tokens_out":8320,"duration_ms":82737,"temperature":0.7,"pith_summary":"The paper introduces the landmark map (LAMA), an extension of the self-organizing map in which a few designated pairs of data points and output nodes act as landmarks. Training alternates a standard data-driven phase with a landmark-driven phase that pulls the landmark node's reference vector toward its paired data point and propagates that pull through the neighborhood. The claim is that this produces a user-intended nonlinear projection while retaining the SOM's topology-preserving fit to the whole dataset. Demonstrations on the Zoo dataset produce landmark-centered views, and an artificial formant experiment maps tongue and jaw movement to horizontal and vertical cursor movement. A sympathetic reader would care because it offers a way to impose semantic constraints on an unsupervised projection without using class labels.","feed_headline":"Few landmark pairs bend a self-organizing map into a chosen view","feed_subtitle":"A handful of data-to-node anchors lets users impose semantic directions, like mapping speech formants to cursor movement.","key_machinery":"The load-bearing mechanism is the alternating update with two independent neighborhood functions. A landmark node is an output-grid node with a location vector $v_k$ and a reference vector $w_k$ that is paired to a landmark data point $x'_m$. In the landmark-driven phase, winner selection is overridden ($k_l(m)=l_m$), so the update is $w_k \\leftarrow w_k + \\beta_k(t)(x'_m-w_k)$ with $\\beta_k(t)=b(t)\\exp(-\\|v_{k_l}-v_k\\|^2/2\\rho^2(t))$; here $b(t)$ is a Gaussian over training time with a peak at $t_{\\rm center}$. This pulls the designated codebook vector toward its landmark data and, more weakly, pulls output-space neighbors toward that data, bending the input-space mesh without tearing its adjacency.","core_discovery":"The paper claims that a self-organizing map can be made to honor user-chosen constraints by adding a small number of landmark nodes, each a fixed pair of a data point in input space and a grid node in output space. Training alternates, with probability $p_{\\rm th}$, the ordinary SOM update with a landmark-driven update: the paired node is forced to be the winner for its landmark data, and all codebook vectors are pulled toward that landmark data through a Gaussian neighborhood centered at the landmark node. Because the landmark-driven neighborhood has its own schedule, peaking in the middle of training and decaying differently from the data-driven neighborhood, the grid bends toward the landmarks while mostly retaining its ordering. On the Zoo dataset this produces landmark-centered views with one, two, three, and four landmarks; on the artificial formant dataset it maps tongue movement from /e/ to /o/ to horizontal cursor movement and jaw movement from /u/ to /a/ to vertical cursor movement. The numerical indices show that the data fit stays close to the SOM's while the landmark fit remains low in most configurations.","pith_inferences":["Inference: LAMA turns projection design into an interactive loop: after inspecting a current view, a user could add one more pairwise anchor and re-train; the paper does not describe such an interface, but the landmark formulation supports it.","Inference: The formant experiment is static and offline; a natural extension is an online version where landmark targets are fixed mean formant values and incoming formant estimates update the map continuously, which would test whether the intended cursor mapping persists under real speech variability.","Inference: If landmark fit and grid smoothness conflict in general, then a testable design principle is to place landmarks only where a preliminary map already puts them and to automate the schedule search; this would quantify how far from the unconstrained solution a user-intended projection can be pushed before wrinkles appear.","Inference: Because landmarks are point-to-node pairs rather than class labels, the same machinery could enforce invariance constraints in other unsupervised projections, for example pinning known stable or prototype points to fixed screen positions, a use the paper does not discuss."],"forward_implications":["A user can produce several deliberately different views of the same dataset by changing landmark placements, as shown by the four Zoo configurations: centered, between two anchors, inside a triangle of anchors, and inside a square of edge anchors.","The projection can align chosen semantic directions with output axes: in the formant experiment, moving from /e/ to /o/ in the input space corresponds to horizontal motion in the output space, and moving from /u/ to /a/ corresponds to vertical motion.","Landmark data do not have to be training points; the formant landmarks are externally defined mean values, so LAMA can pull a projection toward reference prototypes outside the observed set.","The cost of landmark constraints appears where the two objectives conflict: stronger landmark forcing can distort the codebook mesh and raise square topographic error, so parameter tuning by inspecting visualizations is part of the method as presented.","The quantization error on the data stays comparable to the plain SOM at the end of training, so landmark constraints can be added without abandoning the map's fit to the whole dataset."],"supporting_citations":[{"why":"Supplies the self-organizing map architecture, neighborhood learning, and online-update terminology that LAMA directly extends.","marker":"[1]"},{"why":"States the essential SOM projection properties that LAMA aims to retain while adding landmarks.","marker":"[2]"},{"why":"Defines the U-matrix visualization used to compare SOM and LAMA projections in the Zoo and formant experiments.","marker":"[6]"},{"why":"Provides the quantization and topographic error measures that LAMA adapts into QED, QEL, and STE.","marker":"[28]"},{"why":"Supplies the mean first and second formant values for the five Japanese vowels used as landmark data in the formant experiment.","marker":"[29]"},{"why":"Demonstrates a formant-controlled pointing device, the application that the LAMA projection is meant to extend.","marker":"[30]"}],"fun_headline_variants":["User-chosen anchors bend a SOM into a custom view","Landmark pairs steer self-organizing maps to chosen directions","SOM with a few landmarks yields user-intended nonlinear projections","Anchor points let users impose semantic directions on SOMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim assumes that for any user-chosen set of landmark pairs, some setting of the many learning parameters exists and can be found by manual tuning that keeps the map's grid smooth while placing the landmark nodes' reference points at their assigned data points.","fun_headline_variants_meta":{"raw":{"variants":["User-chosen anchors bend a SOM into a custom view","Landmark pairs steer self-organizing maps to chosen directions","SOM with a few landmarks yields user-intended nonlinear projections","Anchor points let users impose semantic directions on SOMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000313,"raw_usage":{"total_tokens":1787,"prompt_tokens":965,"completion_tokens":822,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":755}},"tokens_in":581,"tokens_out":822,"duration_ms":9190,"temperature":1.0,"reasoning_tokens":755,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:25:23.110155+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run LAMA many times on the artificial formant dataset with the paper's landmarks, but with a grid search over $p_{\\rm th}$, $b_{\\min}$, $b_{\\max}$, $t_{\\rm center}$, $\\rho_b$, $\\rho_{\\min}$, and $\\rho_{\\max}$; if every configuration that drives the landmark quantization error near zero also produces a wrinkled or self-intersecting mesh, with square topographic error well above the plain SOM's, then user-intended projections are not reliably attainable at the claimed level of generality.","supporting_citations":[{"cited_title":"Kohonen, Self-organizing maps, 3rd Edition, Springer series in informa- tion sciences, 30, Springer, 2001","cited_arxiv_id":null,"evidence_quote":"Supplies the self-organizing map architecture, neighborhood learning, and online-update terminology that LAMA directly extends."},{"cited_title":"Kohonen, Essentials of the self-organizing map, Neural Networks 37 (2013) 52–65","cited_arxiv_id":null,"evidence_quote":"States the essential SOM projection properties that LAMA aims to retain while adding landmarks."},{"cited_title":"Ultsch, U * -matrix : A tool to visualize clusters in high dimensional data, Tech","cited_arxiv_id":null,"evidence_quote":"Defines the U-matrix visualization used to compare SOM and LAMA projections in the Zoo and formant experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the quantization and topographic error measures that LAMA adapts into QED, QEL, and STE."},{"cited_title":"Nakagawa, H","cited_arxiv_id":null,"evidence_quote":"Supplies the mean first and second formant values for the five Japanese vowels used as landmark data in the formant experiment."},{"cited_title":"Uemi, A study of a human interface device controlled by formant fre- quencies for the disabled, in: International Conference on Cross-Cultural Design, Springer, 2013, pp","cited_arxiv_id":null,"evidence_quote":"Demonstrates a formant-controlled pointing device, the application that the LAMA projection is meant to extend."}],"review_version":1}