REVIEW 3 major objections 5 minor 43 references
Refined and Segmented Price Sentiment Indices from Survey Comments
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper constructs price sentiment indices from Japanese survey comments using LLM classification and segmentation, and claims higher correlations with CPI, CGPI, and SPPI than the prior word-count index.
desk verdict Solid LLM-based price sentiment index extension; the claimed win over the BoJ baseline is not yet established because the baseline row may come from a different sample and seasonal-adjustment scheme. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the four-step pipeline: (1) FinBERT filters comments about prices; (2) several LLMs classify each comment as rising, stable, falling, or not price-related, each returning a confidence and a reason; (3) a separate LLM synthesizes these outputs into one label; (4) comments are aggregated by month and segment into PSI = (Rise - Fall) / (Rise + Fall + Stable). Segmentation uses the survey's comment domain (household vs corporate trends) and a manual mapping of 169 respondent industries into manufacturing or non-manufacturing, yielding five specific indices plus a general one. The formula is the same one used by earlier work, so the claimed performance gain comes from the classification and segmentation rather than from a new aggregation formula.
What would settle it
Recompute the [7] word-count price sentiment index on the identical January 2001 to June 2024 sample and run the same lagged-correlation search; if the recomputed baseline matches the column shown in Table VI, then re-test the comparison. Alternatively, hold out 2025 onward data and check whether the LLM-based index still out-correlates the baseline out-of-sample.
Extended reading notes
Core claim
The central claim is that a price sentiment index (PSI) built from LLM-classified survey comments outperforms the prior word-based PSI for every official price index studied: the general PSI attains higher maximum lagged correlations than baseline [7] for core-core CPI, CPI goods, CPI services, CGPI, and SPPI (e.g., 0.635 vs 0.583 for core-core CPI at a 14-month lead), and the segmented consumer PSIs improve further on the three consumer indices. The paper also reports Granger causality from the general PSI to core-core CPI, CGPI, and SPPI at the 1% level with 12-month lags, which supports the interpretation that the index leads rather than merely tracks official prices. The improvement is attributed to two changes: replacing word counting with LLM classification (with fine-tuned FinBERT filtering price-related comments, and an ensemble of GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Flash whose outputs are merged by another LLM), and segmenting comments by the survey's domain and respondent industry before aggregation.
Load-bearing premise
The load-bearing premise is that the baseline index from the prior study was measured on the same January 2001 to June 2024 sample and with the same lag search; if the prior study used different dates, the reported improvements could come from sample differences rather than the new method.
Editorial extensions
If this is right
- For the three consumer price indices (core-core CPI, goods, services), the segmented consumer PSIs show higher time-lagged correlations than both the general PSI and the baseline, so tracking household-only comments helps nowcast consumer prices.
- The general PSI Granger-causes core-core CPI, CGPI, and SPPI at 1% significance with 12-month lags, implying the index contains predictive content beyond contemporaneous correlation.
- Corporate-goods and corporate-services PSIs have small sample sizes (about 25 comments per month), so the general PSI remains the better vehicle for tracking CGPI and SPPI.
- Because classification uses 5-shot in-context learning with API models, the same pipeline can be applied to other monthly surveys without retraining.
Reading between the lines
- A cheaper variant using a single strong classifier or majority voting across model outputs might achieve most of the integration gain; the paper only tests LLM-based integration, so a simple aggregation baseline would clarify how much the integrator adds.
- The manual mapping of 169 respondent industries to manufacturing/non-manufacturing is a hidden judgment call; an automated or sensitivity-checked mapping would show whether segmentation results are robust to that choice.
- The method implicitly assumes that comments mentioning prices are representative of actual price movements; extending the index to subnational or sectoral price indices would test whether segmentation generalizes beyond national aggregates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a pipeline for constructing price sentiment indices (PSIs) from the Cabinet Office's Economy Watchers Survey. Price-related comments are filtered with a fine-tuned FinBERT model; the direction of price movements (rise/stable/fall) is classified by multiple LLMs (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Flash) with 5-shot in-context learning, and their outputs are integrated by an LLM. The classified comments are aggregated monthly into a General PSI and five segmented PSIs (consumer general, consumer goods, consumer services, corporate goods, corporate services) by using the survey domain and respondent industry. The indices are compared with core-core CPI, CPI Goods, CPI Services, CGPI, and SPPI via maximum time-lagged correlations, and Granger causality tests are reported. The main claim is that the proposed PSIs have higher correlations than the previous word-count-based index of Nakajima et al. [7].
Significance. The paper has several strengths: the classification tasks are evaluated on labeled data with several models; the data version (2024.07.0) and API model versions are specified; the seasonal-adjustment sensitivity is examined in Appendix D; and the Granger causality tests provide a direct test of predictive content. If the comparative claim against [7] were established on a common sample, the result would be a useful demonstration that LLM-based, segmented sentiment indices add information for tracking Japanese price trends. However, as presented, the headline comparison is not yet a controlled comparison, and the statistical significance of the correlation improvements is not quantified. The contribution is therefore conditionally useful: the methodological pieces are sound, but the central evaluation needs strengthening.
major comments (3)
- [Section IV, Table VI and Appendix D] The central comparative claim that the proposed General PSI and Specific PSIs achieve higher time-lagged correlations than the baseline [7] is not established because the baseline values in Table VI are not shown to be computed on the same sample and index definitions. Appendix D states that [7] uses seasonally adjusted CPI, whereas this paper uses non-seasonally adjusted indices, and the manuscript does not state that the [7] index was recomputed on the January 2001 to June 2024 sample described in Section III-C. Since [7] is a 2021 Bank of Japan paper, its reported correlations may come from a different sample period and possibly a different survey data version. The differences are small in some cases (e.g., 0.635 vs 0.583 for core-core CPI, 0.793 vs 0.778 for CGPI), so they could be affected by sample-period, seasonal-adjustment, or index-revision differences. Please either recompute the baseline on the same sample with identical preprocessing and lag search, or clearly qualify the comparison and present the original [7] values as references rather than as a controlled benchmark.
- [Section IV, Table VI] The reported correlations are maxima over a lag grid (the values in parentheses are the lags at which the maximum is attained), and the paper gives no confidence intervals or significance tests for these correlations or for the differences between the proposed indices and the baseline. Under the null that two series are independent, the distribution of the maximum lagged correlation over a grid of candidate lags stochastically dominates that of a fixed-lag correlation, so the point estimates are likely upward-biased. In addition, the comparison across five target indices and multiple PSI variants involves multiple testing. The authors should report the full lag-correlation curves or, at minimum, provide standard errors and a test of whether the improvement over the baseline is significant, and discuss the selection of the lag and the integration model as part of the procedure.
- [Section III-B and Table IV] The selection of Gemini 1.5 Flash as the integration model (Table IV) was based on the test-set performance of the integration step, and the same test-set evaluations informed the choice of the constituent models. This selection is an additional data-dependent degree of freedom that is not accounted for when the resulting PSI is later evaluated for correlation with price indices. The paper should either use a nested or separate validation split for the integration-model choice, or report sensitivity of Table VI to alternative integration models (e.g., Claude 3.5 Sonnet), so that the reported correlation improvements cannot be attributed to overfitting the integration choice to the test set.
minor comments (5)
- [Section III-A] There is a typo: 'fune-tuned' should be 'fine-tuned'; also in Table I, 'T HE' should be 'THE', and in the NOTES section, 'the their affiliated institutions' should be 'their affiliated institutions'.
- [Section III-A and III-B] The paper does not report inter-annotator agreement for the manual labeling of the 304 comments in Price Direction Data 1 or for the price-direction labels of the 1,000 comments in Price Direction Data 2; since these labels form the gold standard for the classification evaluation, a measure of agreement (e.g., Cohen's kappa) would strengthen the results.
- [Section IV, Table VIII] The Corporate Goods PSI and Corporate Services PSI are constructed from only about 25 comments per month; the paper discusses this limitation qualitatively, but it would be useful to show the time-series volatility or confidence bands for these indices to help readers gauge their reliability.
- [Section IV, paragraph on lag correlations] The sentence 'When examining the lag correlation between CPI and SPPI, the maximum value of 0.920 is observed when CPI leads by three months' reports a correlation between two existing price indices, not between a PSI and a price index; clarify that this is a reference point from the data rather than a result about the proposed indices.
- [Equation (1), Section III-C] The definition PSI = (Rise - Fall)/(Rise + Fall + Stable) is stated, but the paper does not specify whether the baseline [7] used the same normalization; a sentence describing the baseline construction would remove ambiguity about the comparability of levels.
Circularity Check
No circularity found: the PSIs are built from survey comments with manual labels and validated against external price indices; the baseline-comparison concern is sample comparability, not a definitional reduction.
full rationale
The derivation chain is self-contained against external benchmarks. Price-related comments are filtered by FinBERT trained on 1,000 manually labeled survey comments, and price directions are classified by LLMs evaluated on 1,304 manually labeled comments; neither label set is derived from CPI, CGPI, or SPPI values. The PSI aggregation in Eq. (1) is a fixed ratio formula adopted from prior work [7] and contains no parameter fitted to the price indices, so the correlations in Table VI are external validation rather than a restatement of the inputs. The maximum-lag correlations and Granger causality tests are descriptive and inferential assessments of the constructed index against official statistics, not fitted predictions equivalent to their targets. The only substantive concern is that the Baseline [7] row in Table VI is not shown to be recomputed on the same January 2001 to June 2024 sample, and Appendix D concedes that [7] used seasonally adjusted CPI while this paper uses non-seasonally adjusted indices; this is a benchmark-comparability flaw affecting the strength of the headline claim, not a circularity reduction. Self-citations, such as [35] for the dataset and [36]/[37] for FinBERT, supply tools and data but are not load-bearing arguments that define the target result, and no uniqueness theorem or ansatz is smuggled in through citations. No equation reduces to its own input and no fitted parameter is renamed as a prediction, so the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- lag_at_max_correlation =
14 (core core CPI), 7 (goods), 16 (services), 3 (CGPI), 15 (SPPI)
- integration_model_selection =
Gemini 1.5 Flash
- industry_manufacturing_mapping =
169 industry categories manually assigned
assumptions (4)
- domain assumption Economy Watchers Survey comments are informative about actual price trends.
- domain assumption Comment domain and respondent industry map cleanly to consumer/corporate and goods/services segments.
- standard math Maximum lagged Pearson correlation and Granger causality at 12 lags are valid measures of leading information.
- domain assumption Manual labels by the authors are ground truth for filtering and direction classification.
Cite this review
Pith. "Pith review of Refined and Segmented Price Sentiment Indices from Survey Comments." pith.science (2026). https://pith.science/paper/YGP45ET3
@misc{pith2026241109937,
author = {Pith},
title = {Pith review of: Refined and Segmented Price Sentiment Indices from Survey Comments},
year = {2026},
howpublished = {\url{https://pith.science/paper/YGP45ET3}},
note = {Machine review of arXiv:2411.09937}
}
read the original abstract
We aim to enhance a price sentiment index and to more precisely understand price trends from the perspective of not only consumers but also businesses. We extract comments related to prices from the Economy Watchers Survey conducted by the Cabinet Office of Japan and classify price trends using a large language model (LLM). We classify whether the survey sample reflects the perspective of consumers or businesses, and whether the comments pertain to goods or services by utilizing information on the fields of comments and the industries of respondents included in the Economy Watchers Survey. From these classified price-related comments, we construct price sentiment indices not only for a general purpose but also for more specific objectives by combining perspectives on consumers and prices, as well as goods and services. It becomes possible to achieve a more accurate classification of price directions by employing a LLM for classification. Furthermore, integrating the outputs of multiple LLMs suggests the potential for the better performance of the classification. The use of more accurately classified comments allows for the construction of an index with a higher correlation to existing indices than previous studies. We demonstrate that the correlation of the price index for consumers, which has a larger sample size, is further enhanced by selecting comments for aggregation based on the industry of the survey respondents.
Figures
Reference graph
Works this paper leans on
-
[6]
Economic analysis using machine learning: Text mining of the Economy Watchers Survey,
K. Otaka and K. Kan, “Economic analysis using machine learning: Text mining of the Economy Watchers Survey,” Bank of Japan Working Paper Series, Tech. Rep., 2018
work page 2018
-
[7]
J. Nakajima, H. Yamagata, T. Okuda, S. Katsuki, and T. Shinohara, “Extracting firms’ short-term inflation expectations from the Economy Watchers Survey using text analysis,” Bank of Japan, Tech. Rep., 2021
work page 2021
-
[25]
Ex- tracting firms’ short-term inflation expectations from survey comments using text analysis,
J. Nakajima, T. Okuda, H. Yamagata, S. Katsuki, and T. Shinohara, “Ex- tracting firms’ short-term inflation expectations from survey comments using text analysis,” 2022
work page 2022
-
[1]
Estimation of firms’ inflation expectations using the survey DI,
J. Nakajima, “Estimation of firms’ inflation expectations using the survey DI,” Institute of Economic Research, Hitotsubashi University, Tech. Rep., 2023
work page 2023
-
[2]
Forecasting CPI inflation components with Hierarchical Recurrent Neural Networks,
O. Barkan, J. Benchimol, I. Caspi, E. Cohen, A. Hammer, and N. Koenigstein, “Forecasting CPI inflation components with Hierarchical Recurrent Neural Networks,” International Journal of Forecasting, vol. 39, no. 3, pp. 1145–1162, 2023. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0169207022000607
work page 2023
-
[3]
Forecasting UK inflation bottom up,
A. Joseph, G. Potjagailo, C. Chakraborty, and G. Kapetanios, “Forecasting UK inflation bottom up,” International Journal of Forecasting, vol. 40, no. 4, pp. 1521–1538, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0169207024000013
work page 2024
-
[4]
Internet search behavior as an economic forecasting tool: The case of inflation expectations,
G. Guzman, “Internet search behavior as an economic forecasting tool: The case of inflation expectations,” Journal of economic and social measurement, vol. 36, no. 3, pp. 119–167, 2011
work page 2011
-
[5]
Nowcasting prices using Google trends: an application to Central America,
S. Seabold and A. Coppola, “Nowcasting prices using Google trends: an application to Central America,” World Bank Policy Research Working Paper, no. 7398, 2015
work page 2015
Show all 43 references
-
[8]
Development and anal- ysis of medical instruction-tuning for Japanese large language models,
I. Sukeda, M. Suzuki, H. Sakaji, and S. Kodera, “Development and anal- ysis of medical instruction-tuning for Japanese large language models,” AIH, vol. 1, no. 2, p. 107, 2024
2024
-
[9]
Large Language Models are legal but they are not: Making the case for a powerful LegalLLM,
T. Jayakumar, F. Farooqui, and L. Farooqui, “Large Language Models are legal but they are not: Making the case for a powerful LegalLLM,” in Proceedings of the Natural Legal Language Processing Workshop 2023 , Dec. 2023, pp. 223–229. [Online]. Available: https://aclanthology.or...
2023
-
[10]
JaFIn: Japanese Financial Instruction Dataset,
K. Tanabe, M. Suzuki, H. Sakaji, and I. Noda, “JaFIn: Japanese Financial Instruction Dataset,” 2024. [Online]. Available: https: //arxiv.org/abs/2404.09260
2024 arXiv
-
[11]
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language models are few-shot learners,” in Advances in neural information processing systems, vol. 33, 2020, pp. 1877–1901
2020
-
[12]
Large language models are zero-shot reasoners,
T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa, “Large language models are zero-shot reasoners,” in Advances in neural infor- mation processing systems , vol. 35, 2022, pp. 22 199–22 213
2022
-
[13]
The causal relationship between producer price index and consumer price index: Empirical evidence from selected European countries,
S. Akcay, “The causal relationship between producer price index and consumer price index: Empirical evidence from selected European countries,” International Journal of Economics and Finance , vol. 3, no. 6, pp. 227–232, 2011
2011
-
[14]
Exchange rate pass- through to Japanese prices: Import prices, producer prices, and the core CPI,
Y . Sasaki, Y . Yoshida, and P. K. Otsubo, “Exchange rate pass- through to Japanese prices: Import prices, producer prices, and the core CPI,” Journal of International Money and Finance , vol. 123, p. 102599, 2022. [Online]. Available: https://www.sciencedirect.com/ science/ar...
2022
-
[15]
Dynamic causality between PPI and CPI in China: A rolling window bootstrap approach,
J. Sun, J. Xu, X. Cheng, J. Miao, and H. Mu, “Dynamic causality between PPI and CPI in China: A rolling window bootstrap approach,” International Journal of Finance & Economics, vol. 28, no. 2, pp. 1279– 1289, 2023
2023
-
[16]
The consumer price index prediction using machine learning approaches: Evidence from the united states,
T.-T. Nguyen, H.-G. Nguyen, J.-Y . Lee, Y .-L. Wang, and C.-S. Tsai, “The consumer price index prediction using machine learning approaches: Evidence from the united states,” Heliyon, vol. 9, no. 10, 2023
2023
-
[17]
Comparison four kernels of svr to predict consumer price index,
M. Rohmah, I. Putra, R. Hartati, and L. Ardiantoro, “Comparison four kernels of svr to predict consumer price index,” in Journal of Physics: Conference Series, vol. 1737, no. 1. IOP Publishing, 2021, p. 012018
2021
-
[18]
Predicting Socio-Economic Indicators using News Events,
S. Chakraborty, A. Venkataraman, S. Jagabathula, and L. Subramanian, “Predicting Socio-Economic Indicators using News Events,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , ser. KDD ’16, 2016, p. 1455–1464. [Online]. Av...
2016 doi
-
[19]
Measuring news sentiment,
A. H. Shapiro, M. Sudhof, and D. J. Wilson, “Measuring news sentiment,” Journal of Econometrics , vol. 228, no. 2, pp. 221– 243, 2022. [Online]. Available: https://www.sciencedirect.com/science/ article/pii/S0304407620303535
2022
-
[20]
Can we measure inflation expectations using twitter?
C. Angelico, J. Marcucci, M. Miccoli, and F. Quarta, “Can we measure inflation expectations using twitter?” Journal of Econometrics , vol. 228, no. 2, pp. 259–277, 2022. [Online]. Available: https: //www.sciencedirect.com/science/article/pii/S0304407622000227
2022
-
[21]
Sentiment Summarization of Financial Reports by LSTM RNN model with the Japan Economic Watcher Survey Data,
Y . Yamamoto and Y . Matsuo, “Sentiment Summarization of Financial Reports by LSTM RNN model with the Japan Economic Watcher Survey Data,” in Proceedings of the Annual Conference of JSAI , 2016, pp. 3L3OS16a2–3L3OS16a2, (in Japanese)
2016
-
[22]
News-based business sentiment and its properties as an economic index,
K. Seki, Y . Ikuta, and Y . Matsubayashi, “News-based business sentiment and its properties as an economic index,” Information Processing & Management, vol. 59, no. 2, p. 102795, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0306457321002739
2022
-
[23]
Stock Selection Attempt using Sentiment of Japan Company Handbook,
M. Suzuki, “Stock Selection Attempt using Sentiment of Japan Company Handbook,” in Proceedings of the Annual Conference of JSAI, 2024, pp. 2I6GS1001–2I6GS1001, (in Japanese)
2024
-
[24]
Forecasting Japanese inflation with a news-based leading indicator of economic activities,
K. Goshima, H. Ishijima, M. Shintani, and H. Yamamoto, “Forecasting Japanese inflation with a news-based leading indicator of economic activities,” Studies in Nonlinear Dynamics & Econometrics , vol. 25, no. 4, pp. 111–133, 2021
2021
-
[26]
Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language Processing,
Y . Guo, Z. Xu, and Y . Yang, “Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language Processing,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , Dec. 2023, pp. 815–821. [Online]. Available: https: //aclanthology.org...
2023
-
[27]
Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks,
X. Li, S. Chan, X. Zhu, Y . Pei, Z. Ma, X. Liu, and S. Shah, “Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Trac...
2023
-
[28]
FinBen: A Holistic Financial Benchmark for Large Language Models,
Q. Xie, W. Han, Z. Chen, R. Xiang, X. Zhang, Y . He, M. Xiao, D. Li, Y . Dai, D. Feng, Y . Xu, H. Kang, Z. Kuang, C. Yuan, K. Yang, Z. Luo, T. Zhang, Z. Liu, G. Xiong, Z. Deng, Y . Jiang, Z. Yao, H. Li, Y . Yu, G. Hu, J. Huang, X.-Y . Liu, A. Lopez-Lira, B. Wang, Y . Lai, H. W...
2024 arXiv
-
[29]
Construction of a Japanese Financial Benchmark for Large Language Models,
M. Hirano, “Construction of a Japanese Financial Benchmark for Large Language Models,” in Proceedings of the Joint Workshop of the 7th Financial Technology and Natural Language Processing, the 5th Knowledge Discovery from Unstructured Data in Financial Services, and the 4th Wo...
2024
-
[30]
Chatgpt-based investment portfolio selection,
O. Romanko, A. Narayan, and R. H. Kwon, “Chatgpt-based investment portfolio selection,” in Operations Research Forum , vol. 4, no. 4. Springer, 2023, p. 91
2023
-
[31]
Can ChatGPT improve investment decisions? From a portfolio management perspective,
H. Ko and J. Lee, “Can ChatGPT improve investment decisions? From a portfolio management perspective,” Finance Research Letters, vol. 64, p. 105433, 2024. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S154461232400463X
2024
-
[32]
Can ChatGPT assist in picking stocks?
M. Pelster and J. Val, “Can ChatGPT assist in picking stocks?” Finance Research Letters , vol. 59, p. 104786, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S1544612323011583
2024
-
[33]
LLMFactor: Extracting Profitable Factors through Prompts for Explainable Stock Movement Prediction,
M. Wang, K. Izumi, and H. Sakaji, “LLMFactor: Extracting Profitable Factors through Prompts for Explainable Stock Movement Prediction,” in Findings of the Association for Computational Linguistics ACL 2024 , Aug. 2024, pp. 3120–3131. [Online]. Available: https://aclanthology.o...
2024
-
[34]
Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models,
K. J. Koa, Y . Ma, R. Ng, and T.-S. Chua, “Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models,” in Proceedings of the ACM Web Conference 2024, ser. WWW ’24, 2024, p. 4304–4315. [Online]. Available: https://doi.org/10.1145/3589334.3645611
2024
-
[35]
Economy Watchers Survey provides Datasets and Tasks for Japanese Financial Domain,
M. Suzuki and H. Sakaji, “Economy Watchers Survey provides Datasets and Tasks for Japanese Financial Domain,” 2024. [Online]. Available: https://arxiv.org/abs/2407.14727
2024 arXiv
-
[36]
FinDeBERTaV2: Word-Segmentation-Free Pre-trained Language Model for Finance,
M. Suzuki, H. Sakaji, M. Hirano, and K. Izumi, “FinDeBERTaV2: Word-Segmentation-Free Pre-trained Language Model for Finance,” Transactions of the Japanese Society for Artificial Intelligence , vol. 39, no. 4, pp. FIN23–G 1–14, 2024, (in Japanese)
2024
-
[37]
Constructing and analyzing domain-specific language model for financial text mining,
——, “Constructing and analyzing domain-specific language model for financial text mining,” Information Processing & Management , vol. 60, no. 2, p. 103194, 2023
2023
-
[38]
Sudachi: a Japanese Tokenizer for Business,
K. Takaoka, S. Hisamoto, N. Kawahara, M. Sakamoto, Y . Uchida, and Y . Matsumoto, “Sudachi: a Japanese Tokenizer for Business,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) , 2018. [Online]. Available: https://aclanth...
2018
-
[39]
From Base to Conversational: Japanese Instruction Dataset and Tuning Large Language Models,
M. Suzuki, M. Hirano, and H. Sakaji, “From Base to Conversational: Japanese Instruction Dataset and Tuning Large Language Models,” in 2023 IEEE International Conference on Big Data (Big Data) , 2023, pp. 5684–5693
2023
-
[40]
LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs,
LLM-jp, “LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs,” 2024. [Online]. Available: https://arxiv.org/abs/2407.03963
2024 arXiv
-
[41]
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capa- bilities,
K. Fujii, T. Nakamura, M. Loem, H. Iida, M. Ohi, K. Hattori, H. Shota, S. Mizuki, R. Yokota, and N. Okazaki, “Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capa- bilities,” in Proceedings of the First Conference on Language Modeling, ser....
2024
-
[42]
Fine-tuned models: Fine-tuned parameters in FinBERT and DeBERTaV2 models are shown in Table IX
-
[43]
Yes”, and if it does not, answer “No
Prompts for LLM: The following boxes show examples of prompts used for the LLMs in the sections on price-related texts filtration (Section III-A) and price direction classification (Section III-B). We manually develop the confidence level and reasoning used for in-context lear...
2010
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.