{"id":"4797bc53-d128-43fd-97f0-ac2fd58c7b38","arxiv_id":"2412.20937","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Semantic feature multiple access with GAI frame interpolation improves multi-user wireless video rates, but the claimed gains rely on a simulation-fitted interference factor.","lead":"A wireless base station combines semantic features of two video frames into one superimposed signal, and users regenerate the missing intermediate frame with a generative AI interpolation model. The paper claims this semantic feature multiple access scheme raises transmission rates by up to 66% over three baselines when pairing and power are optimized.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The rate gains hinge on an unspecified semantic interference factor rho, whose closed form and differentiability are never given; without it Eq. (9), the KKT system, and the headline numbers are not evaluable or reproducible.","rationale":"The single load-bearing assumption is that rho is a well-defined, deterministic, differentiable function of transmit power. The entire quantitative machinery—Eq. (9), the objective (10)-(11), constraints (17a), Lagrangian (18), KKT condition (27a) with rho'(pk), and extreme points (19)-(21)—depends on this function. The paper provides no definition. Fig. 5 is a simulated surface, not a model; the text says only that rho 'can be represented as a function of pk,1 and pk,2.' Without a closed form, no reader can reproduce the rate curves or check whether the reported 24.8%/45.8%/66.1% improvements come from physical semantic interference suppression or from the free parameter rho being chosen to make the metric favorable. This is not a typical parameter-estimation gap: rho is the only place where 'semantic level' interference enters the rate calculation, so its functional form is the substantive content of the contribution. The paper's own internal description is also inconsistent: Algorithm 3 states 'Fix P, and then implement Algorithm 1 to form the pairing set S,' but Algorithm 1 is the inter-group power allocation algorithm, not the Gale-Shapley pairing described in Section IV.A. This strengthens the case that the system as described is not reproducible as-is. I therefore agree with the reader's weakest assumption and verdict.","tokens_in":17290,"tokens_out":6835,"duration_ms":66245,"concrete_test":"Ask the authors to release the exact functional form of rho_k21(pk,1,pk,2) used for Fig. 5, including the interpolation rule and the derivative rho'(pk) used in (27a). Independently re-run Algorithms 1-2 with this released function and verify (i) the KKT conditions are satisfied at the output, and (ii) the sum-rate comparison in Figs. 6-7 applies the same rate formula to all baselines. If no function is released, or if the gains vanish under a common metric, the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equation (9) defines the SINR as gamma = p|h|^2/(rho_21(pk,1,pk,2) p_j|h|^2 + sigma^2), where rho is the 'semantic interference factor.' The paper states only that rho is a function of transmit power and shows a simulated surface in Fig. 5; it gives no closed form, no estimation algorithm, and no measurement protocol. This quantity is not incidental: it is the engine of every rate claim. The objective (10)-(11) is built on this modified SINR; the inter-group Lagrangian (18) writes rho_kji(pk); and the KKT stationarity condition (27a) contains the derivative rho'(pk). Algorithm 1 computes extreme points (19)-(21) using rho as an explicit function of pk. Without a specification of rho, the optimization problem (17) is not well posed, the KKT system cannot be evaluated, and the reported 24.8%/45.8%/66.1% gains cannot be reproduced. The problem is not merely that rho is hard to measure: because rho is the only mechanism that reduces the effective interference in Eq. (9), its functional form (e.g., whether rho<1 over the optimized power range) directly manufactures the rate advantage over standard Shannon-rate baselines. The paper's argument that standard SINR does not capture semantic performance motivates the modification, but a calibrated weight fitted to simulation and then optimized is circular unless the same metric is independently validated as an achievable rate for the actual semantic codec.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SFMA, a downlink semantic multiple access scheme in which a base station pairs users into groups of two, superimposes their semantic features, and transmits a combined signal; receivers reconstruct their frames and use a GAI-based video frame interpolation model to generate intermediate frames. To capture semantic-level interference, the authors modify the standard SINR by introducing a semantic interference factor rho_21(pk,1,pk,2), then formulate a sum-rate maximization problem with a temporal-gap penalty, which they decompose into user pairing (Gale-Shapley), inter-group power allocation (KKT-based), and intra-group power allocation (gradient descent). Simulations claim rate gains of up to 24.8%, 45.8%, and 66.1% over F-NOMA, O-JSCC, and OFDMA, together with favorable MS-SSIM and LPIPS interpolation quality.","tokens_in":17608,"tokens_out":4773,"duration_ms":47058,"significance":"If the modified SINR in Eq. (9) were a validated achievable-rate expression, the SFMA concept would be of interest to the semantic communication and multiple access communities: treating semantic interference as a power-dependent scaling factor and pairing users via stable matching is a plausible design direction, and the use of GAI interpolation to exploit temporal correlation is timely. The paper provides a clear system architecture, a structured three-step solution, and a concrete comparison setup. However, the central metric is not specified, and the current evidence does not establish that the reported gains are real, reproducible, or attributable to the proposed system rather than to the fitted interference factor.","major_comments":[{"comment":"The semantic interference factor rho_21(pk,1,pk,2) is never defined. The paper gives no closed form, no estimation algorithm, and no measurement protocol; Fig. 5 shows a simulated surface but without axis labels, numeric values, or the underlying formula. Since Eq. (9) defines the SINR used in the rate expression (10), the objective (12), the constraints (17a), and the KKT conditions (27), every rate claim in Section V-A depends on an unspecified quantity. The optimization problem (17) is therefore not well-posed as stated, and the reported gains cannot be reproduced by a reader.","section":"Section III-A, Eq. (9)"},{"comment":"The KKT stationarity condition (27a) contains the derivative rho'(pk) (written as rho_21'(pk) and rho_12'(pk)), and Lemma 1's extreme points (19)-(21) require evaluating rho at specific power values. Because no functional form or differentiability assumptions for rho are provided, these conditions cannot be checked or computed. Moreover, the paper does not establish the convexity or regularity conditions needed for the KKT conditions to characterize a global optimum; the derivation simply asserts the KKT system as the proof of Lemma 1.","section":"Section IV-B, Eq. (27a) and Lemma 1"},{"comment":"The claim that the objective in (23) is 'a sum of two concave logarithmic functions with respect to pk,1' is not established and is, in general, false. For fixed pk, rk,1(pk,1) has the form log(1 + A p1/(rho(pk - p1) + sigma^2)), which is not generally concave; for example, in the interference-dominated regime where the +1 is negligible, it behaves like log(c/(b - p1)), which is convex. Consequently, the gradient-descent step in Algorithm 2 is not guaranteed to converge to a maximum, and the decomposition of (22) into independent per-group problems requires a proof that is not supplied.","section":"Section IV-C, paragraph after Eq. (23)"},{"comment":"The headline gains (24.8%, 45.8%, 66.1%) are computed from the rate expression (10) that incorporates the fitted rho. Since rho is calibrated from the very system it describes (Fig. 5), comparing these rates against standard Shannon-rate baselines is circular unless the modified SINR is independently validated as an achievable rate for the actual semantic codec, for example, by relating rho to measured end-to-end distortion. The paper provides no such validation, and the comparisons in Figs. 6 and 7 report no error bars or multiple trials, so the claimed margins are not shown to be significant.","section":"Section V-A, Figs. 6-7"}],"minor_comments":[{"comment":"The final paragraph of the conclusion states that the proposed method yields 'significant improvements in terms of positioning accuracy,' which is unrelated to the transmission-rate results reported in the paper; this appears to be a leftover from another manuscript and should be corrected.","section":"Section VI, conclusion paragraph"},{"comment":"The sentence 'there will be random 2 users whose performances are the same' is unclear; it should say 'two of the users' or 'a randomly chosen pair of users,' and the comparison should be quantified.","section":"Section V-D, paragraph on three-user case"},{"comment":"Algorithm 1 sets pk,1 = pk,2 in each iteration, which is inconsistent with the earlier notation in which pk,1 and pk,2 follow a fixed power allocation factor eta; the algorithm should state how eta is chosen and how the extreme points (19)-(20) are evaluated under that constraint.","section":"Section IV-B, Algorithm 1, line 4"},{"comment":"The axes of Fig. 5 are not fully labeled and no colorbar or numerical scale is provided, so the reader cannot infer the range or behavior of rho from the figure.","section":"Fig. 5"},{"comment":"The subscripts of rho are used inconsistently: Eq. (18) writes rho_21(pk) and rho_12(pk), while the earlier definition in Eq. (9) uses rho_21(pk,1, pk,2); the relationship between the two-argument and one-argument forms should be stated explicitly.","section":"Eq. (18) and throughout Section IV-B"}],"recommendation":"reject","confidential_remarks":"The paper fits the journal's scope, and the system concept has merit, but as written the central quantity rho is unspecified, the optimization depends on its derivative, and the reported gains are not verifiable. I would not rule out a future revision that specifies rho through an estimation protocol and validates the modified SINR as an achievable rate, but that would require substantial new material rather than local corrections."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The SFMA concept is genuinely new and worth knowing, but the headline rate gains are not supported because the semantic interference factor rho that powers every rate calculation is never specified.\n\nWhat the paper does well: combining GAI-based video frame interpolation with semantic feature superposition in a NOMA-style downlink is a real and timely idea. The temporal-gap-aware user pairing is a sensible addition. The attention-based JSCC network is described concretely, and the PSNR, MS-SSIM, and LPIPS evaluations give at least some evidence that the codec and interpolation actually work. The Gale-Shapley pairing and KKT-based power allocation are standard machinery, but they are competently assembled. The three-user extension and the discussion of scalability show thought about practical limits.\n\nThe soft spot is load-bearing. Equation (9) defines the SINR through rho_21(pk,1,pk,2), a \"semantic interference factor\" that is shown only as a simulated surface in Fig. 5. There is no closed form, no estimation algorithm, no measurement protocol, and no discussion of its smoothness. This is not a minor omission: rho appears in the Lagrangian, in the KKT stationarity condition (27a) as rho'(pk), and in Lemma 1's extreme points. Without a specification for rho, the optimization problem (17) is not well posed and the reported 24.8%, 45.8%, and 66.1% gains cannot be reproduced. There is also a circularity concern. The rates in Eq. (10) are computed from an SINR calibrated on the very system it describes, so the gains are partly a consequence of the fitted rho rather than an independent prediction.\n\nSmaller issues: no error bars on any of the simulation curves, no code or data release, and a few internal inconsistencies (Algorithm 3 references \"Algorithm 1\" for pairing even though Algorithm 1 is the inter-group power allocation; the conclusion mentions \"positioning accuracy\" where it should say transmission quality). These are fixable in revision, but they add to the sense that the evaluation is a first pass rather than a rigorous one.\n\nWho this is for: researchers working on semantic multiple access or GAI-assisted video transmission. They should read it for the system concept and the simulation setup, but they should not take the rate numbers at face value. The paper deserves a serious referee, because the core idea is plausible and the flaw is repairable: if the authors supply a concrete characterization of rho (e.g., measured from a trained semantic codec at different power splits) and rerun the optimization with that definition, the contribution could become solid. I would send it to peer review rather than desk-reject, but I would insist on major revision.","headline":"The SFMA concept is genuinely new and worth knowing, but the headline rate gains are not supported because the semantic interference factor rho that powers every rate calculation is never specified.","tokens_in":18169,"tokens_out":2143,"would_cite":false,"duration_ms":22706,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A power-dependent semantic interference factor in the SINR equation lets one superimposed signal serve two video users, with simulated rate gains up to 66.1% over OFDMA.","keywords":["semantic communication","generative AI","multiple access","non-orthogonal multiple access","video frame interpolation","power allocation","semantic interference","user pairing"],"falsifier":"Run the trained semantic encoder-decoder over a grid of power pairs at fixed channels, measure the actual MSE-based SINR, and check whether a single choice of the semantic interference factor inserted into the modified SINR equation predicts the rates used in the optimization; if no such function fits the data, the reported gains and allocation policies do not follow.","tokens_in":17064,"feed_emoji":"📡","tokens_out":7668,"duration_ms":68733,"temperature":0.7,"pith_summary":"This paper claims that a base station can serve two paired video users at once by superimposing their semantic features into a single signal, provided interference between users is measured in semantic space rather than by raw power overlap. To make that measurement tractable, the authors insert a semantic interference factor into the standard SINR formula, so that each user's achievable rate depends on both users' transmit powers through that factor. They then decompose the joint user-pairing and power-allocation problem into a stable matching among users, an inter-group power allocation solved through KKT conditions, and an intra-group concavity-based power split. Simulations on CIFAR-10 with a generative video frame interpolator show the scheme outperforming fixed-power NOMA by 24.8%, orthogonal joint source-channel coding by 45.8%, and OFDMA by 66.1% in sum rate while preserving interpolation quality.","feed_headline":"Shared semantic signal lifts video rates up to 66%","feed_subtitle":"Pairing users with GAI frame interpolation keeps video quality high while one signal serves both.","key_machinery":"The load-bearing mechanism is the modified SINR equation, in which a semantic interference factor converts cross-user semantic confusion into a power-dependent weight in the denominator. Because this factor is assumed to depend on the two users' transmit powers, the rate expression becomes intrinsically coupled across users, which is what forces the three-stage solution: a stable matching algorithm for user pairing, a KKT-based computation of extreme power points for inter-group allocation, and a concave one-dimensional search for the intra-group split. The same semantic interference factor appears in the KKT derivative condition, so the entire optimization is carried by this single semantic quantity.","core_discovery":"The paper's central discovery is that the physical-layer SINR formula, which treats a superimposed user's signal purely as power interference, is the wrong performance model for semantic multiple access. Using the MSE between the original and reconstructed frames as the true SINR, the authors show a gap between measured semantic performance and the standard formula, and close that gap by introducing a semantic interference factor that scales the interfering user's power in the denominator of the SINR. With this modified SINR, the sum rate of a pair becomes a function of the group power budget through that factor, which justifies a two-level power allocation: first across groups, then within each group. The optimized SFMA system achieves the reported gains and, when paired users' frames are temporally close, the GAI interpolation model produces intermediate frames with MS-SSIM around 0.82 and LPIPS around 0.05, indicating that the spectral-efficiency gain does not come at the cost of video quality.","pith_inferences":["A natural next step, not pursued in the paper, is to estimate the semantic interference factor from the trained encoder-decoder by measuring MSE under many power pairs and fitting a function; the same optimization machinery would then apply at deployment time.","Because the semantic interference factor is architecture- and content-dependent, the reported rate gains are likely to shift for different video datasets or interpolation models; the method's general claim is the SINR structure, not the specific percentages.","The SINR modification suggests a general recipe for other semantic multiple-access schemes: replace physical interference weights with data-derived semantic weights, then reuse standard resource-allocation tools that assume a power-dependent SINR."],"forward_implications":["If SFMA works as claimed, a base station can double the number of video users served per resource block without splitting bandwidth, because the semantic decoder and the GAI interpolator jointly suppress the superimposed user's interference.","The power-dependent semantic interference factor turns user pairing and power control into one coupled design problem, so systems that fix power allocation (F-NOMA) or use orthogonal bandwidth (OFDMA, O-JSCC) are leaving throughput on the table.","The measured MS-SSIM and LPIPS results imply that video quality stays high when the temporal gap between paired users' frames is small, making the temporal gap a first-class resource to schedule rather than just a transmission artifact.","The three-user extension indicates that adding more users per group increases semantic interference and reduces fairness, so the two-user pairing design is a deliberate compromise between spectral efficiency and decoding complexity."],"supporting_citations":[{"why":"Supplies the cross-attention transformer video interpolation model used to generate intermediate frames from the reconstructed frames.","marker":"[24]"},{"why":"Supplies the stable-matching algorithm used to pair users according to preference values that trade off rate and temporal gap.","marker":"[26]"},{"why":"Supplies the SNU-FILM video dataset whose Easy/Medium/Hard/Extreme settings define the temporal gap scenarios used in evaluation.","marker":"[28]"},{"why":"Defines the baseline user-pairing strategy (pairing users with the most distinct channel conditions) used by all comparison schemes.","marker":"[29]"},{"why":"Provides the deep joint source-channel coding architecture on which the O-JSCC baseline and the SFMA encoder-decoder design are based.","marker":"[7]"},{"why":"Gives the standard capacity/SINR formula that the paper argues is inadequate for semantic communication and then modifies.","marker":"[25]"},{"why":"Supplies the multi-scale structural similarity metric used to assess interpolated frame quality.","marker":"[30]"},{"why":"Supplies the learned perceptual similarity metric used to assess semantic and visual quality of interpolated frames.","marker":"[31]"}],"fun_headline_variants":["Semantic multi-access lifts video rates up to 66%","GAI frame interpolation enables 66% rate boost","Rethinking SINR: semantic access gains 66%","Two users share one signal, video rates jump 66%","New semantic access: 66% higher rates with GAI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The semantic interference factor is treated as a known, deterministic, differentiable function of transmit power in the SINR equation and in the KKT derivative, but the paper provides no closed-form expression, no estimation algorithm, and no measurement procedure for it, so the rate gains and power allocations cannot be reproduced without that missing function.","fun_headline_variants_meta":{"raw":{"variants":["Semantic multi-access lifts video rates up to 66%","GAI frame interpolation enables 66% rate boost","Rethinking SINR: semantic access gains 66%","Two users share one signal, video rates jump 66%","New semantic access: 66% higher rates with GAI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2967,"prompt_tokens":1031,"completion_tokens":1936,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":1853}},"tokens_in":647,"tokens_out":1936,"duration_ms":12918,"temperature":1.0,"reasoning_tokens":1853,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:06:07.931158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained semantic encoder-decoder over a grid of power pairs at fixed channels, measure the actual MSE-based SINR, and check whether a single choice of the semantic interference factor inserted into the modified SINR equation predicts the rates used in the optimization; if no such function fits the data, the reported gains and allocation policies do not follow.","supporting_citations":[{"cited_title":"Cross-attention transformer for video interpolation,","cited_arxiv_id":null,"evidence_quote":"Supplies the cross-attention transformer video interpolation model used to generate intermediate frames from the reconstructed frames."},{"cited_title":"Machiavelli and the gale-shapley al- gorithm,","cited_arxiv_id":null,"evidence_quote":"Supplies the stable-matching algorithm used to pair users according to preference values that trade off rate and temporal gap."},{"cited_title":"Channel attention is all you need for video frame interpolation,","cited_arxiv_id":null,"evidence_quote":"Supplies the SNU-FILM video dataset whose Easy/Medium/Hard/Extreme settings define the temporal gap scenarios used in evaluation."},{"cited_title":"Impact of user pairing on 5g nonorthog- onal multiple-access downlink transmissions,","cited_arxiv_id":null,"evidence_quote":"Defines the baseline user-pairing strategy (pairing users with the most distinct channel conditions) used by all comparison schemes."},{"cited_title":"Weaver, The mathematical theory of communication","cited_arxiv_id":null,"evidence_quote":"Gives the standard capacity/SINR formula that the paper argues is inadequate for semantic communication and then modifies."},{"cited_title":"Multiscale structural similarity for image quality assessment,","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-scale structural similarity metric used to assess interpolated frame quality."}],"review_version":1}