{"id":"98620f3e-036e-4815-aed2-8dfd7efededc","arxiv_id":"2505.07118","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using LLM summarization and automatic axis extraction, the authors scale Japanese Diet members on political issues and track party shifts from 2000 to 2024, matching expert party orderings on two of three compared topics.","lead":"This paper tests an LLM-driven pipeline that reads Japanese Diet speeches, summarizes them, automatically finds the main political fault lines, and scores every representative and party on those fault lines over time. The authors report that their rankings of parties match expert rankings well on some topics, but less well on others.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation is restricted to a manually filtered subset of extracted axes (Sec. 4.3), and the reported correlations may reflect researcher selection rather than the full framework.","rationale":"The reader's verdict is CONDITIONAL, and I agree; my concern sharpens one of the reader's stated reasons rather than changing the verdict. The most load-bearing weakness is the gap between the framework's advertised automatic axis extraction and the evidence, which covers only a manually selected subset of axes. The paper itself flags this in Section 4.3, so this is not an external objection but an internal limitation. A single negative correlation on one of the three included axes already demonstrates instability; including the excluded axes would likely lower the aggregate performance. I would not move to REJECT because the method is transparent and the two positive correlations are substantive for the axes that were used; a conditional acceptance requiring full-axis evaluation is the appropriate stance. Secondary concerns include the possible contamination of the LLM-generated reference summaries by GPT's parametric knowledge of Japanese politics, and the absence of uncertainty quantification or significance tests for the rank correlations; these strengthen the case for the concrete test but are not needed to establish the primary objection.","tokens_in":30646,"tokens_out":6147,"duration_ms":61793,"concrete_test":"Re-run the scaling and Mielka comparison on every axis automatically extracted by GPT4o-mini for the five topics, dropping the Section 4.3 suitability filter. For each axis compute Spearman, Kendall, and LCS, and report the distribution (mean, median, min, max, number of negative correlations). If the mean Spearman over all axes falls materially below the reported 0.93/0.96 values, the headline accuracy is an artifact of the manual axis selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the framework 'robustly' and 'accurately' scales stances rests on the quantitative comparison in Table 4, which reports Spearman/Kendall/LCS against Mielka expert orderings for three topics. However, Section 4.3 explicitly states that although all automatically extracted axes were scaled, 'some of the extracted axes were not suitable for scaling as they were very vague or did not directly relate to the topic of interest,' and 'in section 5, we will only show the results where we decided that the extracted axes were appropriate.' No counts are given for how many axes were extracted versus used. The two high correlations (0.9286 for JSDF, 0.9642 for nuclear power) are therefore computed on a post-hoc selected subset, which can reintroduce the exact manual bias the paper claims to remove. The third included axis, consumption tax, already yields Spearman -0.7, and the explanation (one party inverted) shows how fragile the ordering is. Without evaluating the full set of extracted axes, the headline correlations are not attributable to the framework; they are attributable to researcher filtering. Additionally, the comparison is at the party-average level, not at the level of individual representatives, so the claim of scaling 'parliamentary representatives' is validated only indirectly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes KOKKAI DOC, an LLM-driven framework for scaling the issue stances of Japanese parliamentary representatives from Diet speech data. The framework adds three components to the earlier L(u)PIN method: (1) de-noising speech segments by summarizing them with GPT4o-mini before embedding with SBERT; (2) automatically extracting axes of political controversy by prompting GPT4o-mini with the stance summaries; and (3) a diachronic analysis that projects year-by-year party mean embeddings onto the extracted axes. The authors validate the cross-sectional scaling by comparing party-level orderings against expert orderings from the Japanese NPO Mielka using Spearman, Kendall, and LCS ratios, report qualitative PMI-based cluster analysis, and provide event-based interpretations of the diachronic trajectories. The paper also describes a public web application (kokkaidoc.com) that presents the results.","tokens_in":30860,"tokens_out":3195,"duration_ms":30276,"significance":"If the validation were sound, this would be a useful and relatively low-cost contribution to text-based political scaling: it automates axis selection, shows a plausible de-noising effect of LLM summarization, and extends the analysis to temporal dynamics. The authors are unusually transparent about costs and make their data pipeline and web application publicly available. The diachronic party trajectories for LDP, JCP, and Komeito over 2000-2024, interpreted through historical events, are a compelling qualitative demonstration. However, the current quantitative validation is too thin to support the paper's central claim of a 'robust and accurate' framework: it relies on a single external expert source, a hand-picked subset of axes, point estimates without uncertainty quantification, and party-average rather than representative-level comparisons. The paper's significance would be substantially higher if the validation were made comprehensive and if the claims were calibrated to the evidence.","major_comments":[{"comment":"The quantitative validation is performed on a hand-picked subset of extracted axes, and the paper gives no count of how many axes were extracted versus used. Section 4.3 states that 'some of the extracted axes were not suitable for scaling as they were very vague or did not directly relate to the topic of interest' and that 'in section 5, we will only show the results where we decided that the extracted axes were appropriate.' Because the two headline correlations in Table 4 (0.9286 for JSDF acknowledgement, 0.9642 for restarting nuclear plants) are computed on axes selected after seeing the results, the reported agreement with Mielka could reflect researcher filtering rather than the framework's automatic behavior. Please report the full set of extracted axes, the number used versus rejected, and the correlation values for all axes, including those deemed unsuitable.","section":"§4.3 and §5"},{"comment":"The quantitative evidence consists of three point estimates of rank correlation for three topics, with no confidence intervals, significance tests, or sensitivity analyses. The consumption tax row shows Spearman rho = -0.7; the text explains this by noting that 'the Reiwa party is placed on the opposite side of the ordering,' but with roughly seven to eight parties in the comparison, a single inverted party flips the sign. This negative result cannot be dismissed as a minor anomaly when the paper claims the methodology 'robustly' scales stances. Please report bootstrap or permutation-based confidence intervals for all correlations, state the number of parties n in each row, and discuss the consumption tax case as a substantive failure mode rather than an explanation of a single discrepant party.","section":"§6.1.1 and Table 4"},{"comment":"The axis endpoints are generated by GPT4o-mini from stance summaries that GPT4o-mini itself produced, and the projection of a legislator onto the line between those endpoints assumes that the reference summaries faithfully represent the true opposing endpoints of the controversy. Because the same model produces both the summaries and the axis text, a systematic bias in the model's summarization behavior (for example, a tendency toward generic or stylized pro/con statements) would contaminate both the embeddings and the axis direction, and the reported correlations with expert orderings cannot rule this out. Please include an ablation that uses a different model for summarization versus axis extraction, or a hand-coded axis validation, to demonstrate that the results are not simply an artifact of model self-consistency.","section":"§4.3"},{"comment":"The quantitative validation is performed at the party-average level: Table 4 compares the ordering of party means against Mielka's party-level expert orderings. However, the paper claims to scale 'parliamentary representatives' and the violin plots in Figures 10-19 show substantial within-party dispersion. The representative-level claim is therefore validated only indirectly. Please provide an individual-level evaluation where possible (for example, comparison with any available individual-level reference data or at least within-party rank consistency checks), or explicitly scope the claim as a party-level scaling method.","section":"§5.1, §6.1, and Table 4"}],"minor_comments":[{"comment":"The abstract and Section 7 contain the typo 'teh web application'; Figure 14's caption says 'Figtures'; Section 6.2 has 'assocaited' and Section 6.1 has 'dissimation.'","section":"Abstract and §7"},{"comment":"The caption of Figure 17 is garbled: 'Representatives who have a pro stance lower right on the upper side of the UMAP visualization' should be reworded to describe the pro and anti positions consistently with the other figure captions.","section":"Figure 17 caption"},{"comment":"Table 4 does not report the number of parties compared for each topic; this number is needed to interpret the rank correlations and is particularly important for the consumption tax row with its negative values.","section":"Table 4"},{"comment":"The paper should state how many times GPT4o-mini was sampled for each summary and for each axis extraction, since prompt randomness could affect both the embeddings and the generated axis endpoints; no variance information is reported.","section":"§4.2"},{"comment":"The abbreviation 'L(u)PIN' is used in the title and Section 4.1 but is not expanded or defined in the text; define it at first use. The Mielka reference is cited as a URL only and should include an access date and a description of the data collection methodology.","section":"§4.1 and References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads like an expanded arXiv preprint rather than a finished journal article. The central validation weakness—evaluation on a post-hoc selected subset of axes with point estimates only—is fixable, but the authors must either substantially expand the evaluation (full axis set, uncertainty quantification, representative-level checks) or materially weaken the 'robust and accurate' claim. The paper's fit is better with a methods-oriented political science or computational social science venue than with a general CS venue, though the public web application and reproducibility-oriented reporting are strengths. I recommend major revision rather than rejection because the core approach is plausible and the reported qualitative results are suggestive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nI read Kato and Cochrane's KOKKAI DOC. The thing to know up front: the framework is a genuine, transparent extension of their L(u)PIN method, and the headline validation is weaker than the abstract claims. The paper deserves a serious referee, but the \"robustly and accurately\" wording should be walked back until the full axis set is evaluated.\n\nWhat is new: LLM summarization of opinion-classified speech segments as a de-noising step, with a prompt-style comparison; GPT4o-mini extraction of controversy axes from the summaries, replacing manual axis selection; and diachronic party positioning over 2000-2024. They also report compute cost transparently and make data and code available. The diachronic plots align with known events (Fukushima, defense-budget shifts), which is a sensible sanity check. The qualitative PMI analysis is illustrative and mostly consistent with what we know about Japanese parties.\n\nThe main soft spot is exactly what the stress-test note says. Section 4.3 says all extracted axes were scaled, but only axes the authors deemed \"appropriate\" are shown in Section 5, with no count of how many axes were extracted versus used. The two high Spearman correlations (0.93 and 0.96) come from that hand-filtered subset. The third axis, consumption tax, gives -0.7, explained by one party's placement. That one negative result is not disqualifying, because expert rankings are also contestable, but it does show how fragile the ordering comparison is on a small number of parties. There are no confidence intervals or significance tests, and the comparison is at party-average level, so the claim about scaling individual representatives is only indirectly validated. The axis-reference summaries come from the same LLM that summarizes speeches, which could bias the projection; the external Mielka data are used only for validation, so this is not circular in the worst way, but it deserves scrutiny.\n\nThe literature review is competent. I do not see fabricated entities or hidden fitting to the validation data. The central scaling mechanism is not tested on the full output, so the paper's own quantitative evidence does not support \"robust\" yet. But the method is specified clearly and the failure mode is fixable: report all extracted axes, give denominators, add uncertainty quantification, and validate at the individual or within-party level where possible.\n\nWho this is for: political scientists using text-as-data, especially for Japanese politics, and NLP people working on LLM-assisted measurement. I would bring it to a reading group and would cite it as related work. Send it to peer review with the expectation of major revisions. The conditional verdict is right.\n\nRecommendation: engage.","headline":"Useful, transparent extension of L(u)PIN; the validation is too selective to support the 'robustly' claim, but the paper deserves serious review.","tokens_in":31388,"tokens_out":1863,"would_cite":true,"duration_ms":19642,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an LLM pipeline—summarizing opinion segments, embedding them, automatically extracting controversy axes, and projecting legislators onto each axis—scales Japanese Diet members' policy stances with high rank…","keywords":["LLM","political stance scaling","parliamentary speeches","Japanese politics","document embeddings","political text scaling","diachronic analysis","issue axes"],"falsifier":"Take a topic with a known expert ordering, replace the LLM-generated 'for' and 'against' reference summaries with paraphrases that say the opposite of the intended pole (e.g., a 'pro-nuclear' summary that lists only safety concerns), and check whether the party ordering from the projections flips accordingly; if the ordering does not respond, or if paraphrasing the reference summaries changes the Spearman correlation by more than the gap between the method and the baseline, the result depends on the exact LLM output rather than on the speeches.","tokens_in":30420,"feed_emoji":"🗳️","tokens_out":8707,"duration_ms":77947,"temperature":0.7,"pith_summary":"This paper claims that the policy stances of Japanese Diet members can be read off their parliamentary speeches alone, summarized and embedded by an LLM, without a researcher choosing anchor politicians by hand. The framework summarizes each member's opinion-bearing sentences into a consistent stance summary, embeds those summaries with a sentence-transformer, asks the same LLM to list axes of controversy within a broad topic along with pro and con positions, and then scores each legislator by projecting their embedding onto the line between the embeddings of the generated pro and con reference summaries. Compared with expert party orderings from a Japanese voter-information nonprofit, the method reports high Spearman correlations on most topics and exceeds the earlier L(u)PIN method on the two axes where both were measured (acknowledging the JSDF in the constitution: 0.9286 vs 0.8857; restarting nuclear power plants: 0.9642 vs 0.4857). The paper also projects per-year party averages onto the same axes to show party positions shifting over 2000–2024 in step with events such as the Fukushima disaster and the annexation of Crimea.","feed_headline":"LLM pipeline matches expert party rankings for Japan","feed_subtitle":"Projecting summarized Diet speeches onto LLM-extracted controversy axes beats the prior method against expert orderings.","key_machinery":"The mechanism that carries the argument is the projection of a legislator's average sentence embedding onto the line connecting the embeddings of two LLM-generated reference summaries, one for each pole of an automatically extracted controversy axis. The LLM first produces a labelled axis from the speech summaries (e.g., 'should Japan restart nuclear power plants'), then a prompt asks for a short speech a politician with a 'for' stance would give and one a politician with an 'against' stance would give; those two texts are embedded and define the scale. Each member's projected position on that line is the stance score, so the entire method reduces to consistency between the LLM's summarization, its axis extraction, and the geometric assumption that semantic opposition in embedding space aligns with real political opposition.","core_discovery":"The paper's central claim is that a largely automated LLM pipeline can replace manual anchor selection and produce reliable, issue-specific ideological scaling for an entire parliament. Opinion-based speech segments, isolated by a fine-tuned classifier, are summarized by an LLM into a uniform format; the summaries are embedded with a sentence-transformer model. The same LLM reads those summaries and proposes axes of controversy—such as whether the Self-Defense Forces should be written into the constitution or whether nuclear plants should restart—each with one 'for' and one 'against' description. Embeddings of these polar descriptions act as the two ends of a scale, and each legislator's score is the scalar projection of their embedding onto the line between the two ends. The authors argue that the resulting ordering of parties tracks expert judgements, that the summarization step removes the excessive dispersion seen in earlier embeddings, and that projecting each year's party-average embedding onto the same axes reveals credible diachronic movement.","pith_inferences":["If the central claim holds, the same recipe should transfer to any parliament with digitised transcripts and an LLM API; a straightforward test is applying it to another country's speeches and comparing with existing expert or manifesto-based estimates.","The negative Spearman result on consumption-tax reduction (−0.7, driven by one party placed on the wrong side) suggests the method can fail on specific axes while the overall LCS ratio looks acceptable; reporting per-party deviations would clarify whether such flips are noise or a systematic failure mode.","The diachronic series mixes genuine party movement with year-to-year changes in which members spoke and how much; the appendix's table of speech segments per party-year shows large fluctuations, so a stronger design would weight or control for speaker composition.","Since the references come from the LLM itself, the 'automatic' axes may encode the model's prior expectations about Japanese politics rather than structure discovered purely from the speeches; comparing axes extracted from the same summaries by two different LLMs would test this."],"forward_implications":["On the two axes where a direct baseline comparison is reported, the method's Spearman correlations with expert party orderings exceed L(u)PIN's: 0.9286 vs 0.8857 for acknowledging the JSDF in the constitution, and 0.9642 vs 0.4857 for restarting nuclear power plants.","Because axes are extracted from speech summaries rather than hand-picked, the pipeline can scale representatives on controversies the researcher may not have anticipated, which the paper presents as removing manual bias from axis selection.","Projecting per-year party averages onto the same axes yields a diachronic series for 2000–2024 that the paper aligns with events such as the 2011 Fukushima disaster, the 2014 Crimea invasion, and shifts in coalition government, demonstrating dynamic tracking rather than a static snapshot.","The framework is intended to operate end-to-end at low cost—the paper reports about 103 USD in API usage on a consumer laptop—and the results are published on a public web application aimed at Japanese voters."],"supporting_citations":[{"why":"Base method (L(u)PIN) that this framework extends; supplies the projection approach and the baseline correlations the new results must beat.","marker":"Kato et al., 2024"},{"why":"Expert party-order estimates used as the ground truth for the quantitative evaluation.","marker":"Mielka, 2024"},{"why":"Sentence-BERT embeddings that map summaries and reference speeches into the shared semantic space used for projection.","marker":"Reimers and Gurevych, 2019"},{"why":"Word-embedding approach to ideological placement in parliamentary corpora; source of the diachronic analysis idea.","marker":"Rheault and Cochrane, 2020"},{"why":"Prior LLM-based scaling (pairwise comparisons) that the paper contrasts with its representation-based method; a comparison point for the field.","marker":"Wu et al., 2023"},{"why":"Direct-prompting LLM scaling baseline showing the alternative the paper avoids (asking the LLM to output scores directly).","marker":"Mens and Gallego, 2023"}],"fun_headline_variants":["LLM auto-ranks Japanese parties from Diet speeches","AI finds political axes, ranks Japan's parliament","Speech summaries + LLM axes scale Diet ideology","LLM replaces manual anchors in political scaling","Diachronic Diet stance tracking via LLM axes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the LLM's automatically generated pro and con reference summaries for each extracted axis faithfully represent the true opposing positions in that controversy, so that the direction between their embeddings is the direction of real political disagreement; if those summaries are biased or off-topic, the projection scores are meaningless even when the embedding arithmetic is correct.","fun_headline_variants_meta":{"raw":{"variants":["LLM auto-ranks Japanese parties from Diet speeches","AI finds political axes, ranks Japan's parliament","Speech summaries + LLM axes scale Diet ideology","LLM replaces manual anchors in political scaling","Diachronic Diet stance tracking via LLM axes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1547,"prompt_tokens":987,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":487}},"tokens_in":603,"tokens_out":560,"duration_ms":5847,"temperature":1.0,"reasoning_tokens":487,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:23:55.443640+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a topic with a known expert ordering, replace the LLM-generated 'for' and 'against' reference summaries with paraphrases that say the opposite of the intended pole (e.g., a 'pro-nuclear' summary that lists only safety concerns), and check whether the party ordering from the projections flips accordingly; if the ordering does not respond, or if paraphrasing the reference summaries changes the Spearman correlation by more than the gap between the method and the baseline, the result depends on the exact LLM output rather than on the speeches.","supporting_citations":[],"review_version":1}