Pith. sign in

REVIEW 13 cited by

WHEN FLUE MEETS FLANG: Benchmarks and Large Pre-trained Language Model for Financial Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.00083 v1 pith:O2D2YQE2 submitted 2022-10-31 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords financialbenchmarkslanguagedomainmodelmodelstasksdata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pre-trained language models have shown impressive performance on a variety of tasks and domains. Previous research on financial language models usually employs a generic training scheme to train standard model architectures, without completely leveraging the richness of the financial data. We propose a novel domain specific Financial LANGuage model (FLANG) which uses financial keywords and phrases for better masking, together with span boundary objective and in-filing objective. Additionally, the evaluation benchmarks in the field have been limited. To this end, we contribute the Financial Language Understanding Evaluation (FLUE), an open-source comprehensive suite of benchmarks for the financial domain. These include new benchmarks across 5 NLP tasks in financial domain as well as common benchmarks used in the previous research. Experiments on these benchmarks suggest that our model outperforms those in prior literature on a variety of NLP tasks. Our models, code and benchmark data are publicly available on Github and Huggingface.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance

    cs.CL 2025-08 reject novelty 6.0 of 10

    INSEva provides a large Chinese insurance benchmark and shows that current LLMs have basic insurance competence but lag on hard reasoning and safety compliance.

  2. CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 9,356-pair Chinese multimodal financial benchmark reveals that state-of-the-art multimodal LLMs, including GPT-4V, still score below 53% on objective and 39% on subjective financial chart tasks.

  3. VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations

    cs.MM 2025-06 conditional novelty 6.0 of 10

    VideoConviction provides the first expert-annotated multimodal benchmark of financial influencer video recommendations, showing MLLMs extract tickers better but struggle with actions and conviction, and that an invers...

  4. Dynamic Skill Adaptation for Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A training pipeline that orders generated textbook and exercise data by a skill dependency graph and dynamically updates the data during fine-tuning improves LLM performance on calculus and social studies evaluations.

  5. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  6. KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding

    cs.CL 2025-04 conditional novelty 5.0 of 10

    KFinEval-Pilot is a new Korean financial benchmark combining knowledge, legal reasoning, and toxicity tasks, and its model evaluations show clear performance and safety differences.

  7. INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

    cs.CE 2024-12 conditional novelty 5.0 of 10

    InvestorBench evaluates 13 large language models as trading agents on stock, crypto, and ETF tasks, reporting that proprietary models beat open-source ones on average.

  8. OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain

    cs.CL 2024-12 conditional novelty 5.0 of 10

    OmniEval is a finance-domain RAG benchmark that scores retrieval and generation across a 5-task by 16-topic matrix using rule-based and fine-tuned LLM metrics.

  9. SusGen-GPT: A Data-Centric LLM for Financial NLP and Sustainability Report Generation

    cs.CL 2024-12 reject novelty 5.0 of 10

    Small fine-tuned models on SusGen-30K are reported to nearly match GPT-4 on financial and ESG tasks, with a new TCFD-Bench benchmark, though the comparison is biased.

  10. Baichuan4-Finance Technical Report

    cs.CL 2024-12 reject novelty 4.0 of 10

    A finance-tuned LLM reportedly beats strong baselines on Chinese financial exams, but the training data includes exam questions and no contamination check is reported.

  11. Can ChatGPT Overcome Behavioral Biases in the Financial Sector? Classify-and-Rethink: Multi-Step Zero-Shot Reasoning in the Gold Investment

    q-fin.ST 2024-11 reject novelty 4.0 of 10

    A 'Classify-and-Rethink' prompt for ChatGPT produced higher backtested returns on gold trading than simpler prompts or buy-and-hold, though the comparison is confounded.

  12. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

  13. Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning

    cs.LG 2024-12 reject novelty 2.0 of 10

    A data-poisoning backdoor attack on audio transformers is claimed with 100 percent success on TIMIT, but the paper provides no reproducible derivation or evaluation.

Pith tools