Pith. sign in

REVIEW 2 major objections 2 minor 235 cited by

BloombergGPT: A Large Language Model for Finance

T0 review · 2 major / 2 minor · reviewed 2026-05-13 · grok-4.3

Pith's one-line read BloombergGPT, a 50 billion parameter model trained on financial plus general data, outperforms prior models on financial tasks while preserving general LLM performance.

desk verdict BloombergGPT is the first reported 50B financial LLM trained on a 363B-token domain corpus mixed with general data, and the mixed approach looks workable, but the biggest claimed gains sit on internal benchmarks that outsiders cannot audit. read the letter →

arxiv 2303.17564 v3 pith:PUSHK26R submitted 2023-03-30 cs.LG cs.AIcs.CLq-fin.GN

classification cs.LGcs.AIcs.CLq-fin.GN
keywords largelanguagemodelsfinancialNLPdomain-specifictraining50billionparametersmixeddatasetBloombergGPTbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces BloombergGPT as a 50 billion parameter language model trained on a 363 billion token financial dataset drawn from Bloomberg sources, mixed with 345 billion tokens from general datasets. This mixed training is presented as the route to strong results on financial applications such as sentiment analysis, named entity recognition, and question answering. A sympathetic reader would care because the work shows a concrete way to build a domain-specialized LLM at scale without the usual drop in broad capabilities, and it supplies training details plus internal benchmarks that match intended use cases.

What carries the argument

The mixed financial-plus-general training corpus used to pretrain the 50 billion parameter transformer model.

What would settle it

An independent evaluation on financial tasks drawn from sources outside the training corpus and the reported benchmarks would show whether the performance advantage holds.

Watch

Extended reading notes

Core claim

BloombergGPT is a 50 billion parameter model trained on a combined corpus of 363 billion financial tokens and 345 billion general tokens; the resulting model exceeds existing models by substantial margins on financial benchmarks while matching performance on standard general-purpose LLM evaluations.

Load-bearing premise

The chosen financial data sources and internal benchmarks accurately represent real financial usage and the observed gains arise from the training mix rather than from dataset artifacts or evaluation choices.

Editorial extensions

If this is right

  • Financial NLP tasks such as sentiment analysis and question answering become more accurate with the specialized model.
  • The same mixed-dataset recipe can be applied to build other domain-specific models without sacrificing general capability.
  • Releasing the training process details allows other groups to replicate or adapt the approach at similar scale.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The pattern may extend to other high-stakes domains where both specialized knowledge and general reasoning matter.
  • Collecting hundreds of billions of domain tokens appears feasible for organizations with proprietary data pipelines.
  • Public release of training logs sets a precedent for transparency that could influence future large-model projects.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript introduces BloombergGPT, a 50 billion parameter language model trained on a mixed dataset of 363 billion financial tokens drawn from Bloomberg sources and 345 billion general-purpose tokens. It claims that this training regime produces a model that outperforms prior models on financial tasks by significant margins while preserving performance on standard general LLM benchmarks. Validation is reported across standard LLM benchmarks, open financial benchmarks, and a proprietary internal benchmark suite; the authors also document modeling choices, the training process, and evaluation methodology, and release Training Chronicles in Appendix C.

Significance. If the performance claims are substantiated, the work would constitute a notable contribution as the first reported large-scale domain-specific LLM for finance. The construction of what is described as one of the largest financial token datasets and the demonstration that mixed-domain training can improve financial-task performance without degrading general capabilities would be of direct interest to both the NLP and FinTech communities. The release of training chronicles adds practical value for reproducibility.

major comments (2)
  1. [Evaluation] Evaluation section (and abstract): The headline claim that mixed training yields 'significant margins' on financial tasks rests primarily on results from the authors' internal benchmark suite, which the text states 'most accurately reflect our intended usage.' No task definitions, question sources, scoring rubrics, contamination checks, or exclusion criteria are supplied for these benchmarks. Because the largest reported gains are tied to these undisclosed evaluations, independent verification of the central empirical result is impossible and the risk of selection bias or metric-specific artifacts cannot be assessed.
  2. [Evaluation] § on open financial benchmarks: While the paper references validation on open financial benchmarks, the text supplies no numerical tables, baseline comparisons, or error bars for these results either. The absence of concrete numbers leaves the 'outperforms existing models' assertion without direct quantitative support in the manuscript.
minor comments (2)
  1. [Abstract] Abstract: The abstract asserts benchmark outperformance but supplies no numerical results, error bars, baseline details, or exclusion criteria, leaving the central claim with limited direct support from the provided text.
  2. [Appendix C] Appendix C (Training Chronicles): Confirm that the released training log includes sufficient hyper-parameter schedules, hardware details, and any observed instabilities so that the training narrative can be followed by readers.

Simulated Author's Rebuttal

2 responses · 1 unresolved

We thank the referee for their constructive and detailed review of our manuscript on BloombergGPT. The comments on the evaluation sections are well-taken, and we address each point below with clarifications and commitments to revisions where feasible while respecting necessary constraints on proprietary information.

read point-by-point responses
  1. Referee: [Evaluation] Evaluation section (and abstract): The headline claim that mixed training yields 'significant margins' on financial tasks rests primarily on results from the authors' internal benchmark suite, which the text states 'most accurately reflect our intended usage.' No task definitions, question sources, scoring rubrics, contamination checks, or exclusion criteria are supplied for these benchmarks. Because the largest reported gains are tied to these undisclosed evaluations, independent verification of the central empirical result is impossible and the risk of selection bias or metric-specific artifacts cannot be assessed.

    Authors: We appreciate the referee's emphasis on transparency for the internal benchmarks. These evaluations are constructed from Bloomberg's proprietary data and use cases to best reflect real-world financial applications, which is why full task definitions, question sources, and specific rubrics cannot be disclosed without violating confidentiality. We will revise the manuscript to provide expanded high-level descriptions of task categories (e.g., financial sentiment, report summarization, entity extraction), general scoring methodologies, and contamination mitigation steps that do not reveal sensitive details. This will better contextualize the results and address concerns about selection bias while preserving the proprietary nature of the suite. revision: partial

  2. Referee: [Evaluation] § on open financial benchmarks: While the paper references validation on open financial benchmarks, the text supplies no numerical tables, baseline comparisons, or error bars for these results either. The absence of concrete numbers leaves the 'outperforms existing models' assertion without direct quantitative support in the manuscript.

    Authors: We agree that the open financial benchmark results should be presented with explicit quantitative support in the main text. The evaluation section includes these comparisons, but to improve clarity and address the concern directly, we will add a dedicated summary table reporting numerical performance metrics on the open benchmarks (including baselines from prior models), along with error bars from multiple evaluation runs where applicable. This revision will provide the direct quantitative evidence requested. revision: yes

standing simulated objections not resolved
  • Full release of proprietary internal benchmark task definitions, question sources, and specific instances due to confidentiality and data protection requirements.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical training and benchmark evaluation

full rationale

The paper reports construction of a mixed financial+general token dataset, training of a 50B model, and empirical evaluation on standard LLM benchmarks, open financial benchmarks, and internal suites. No equations, derivations, or first-principles claims are present that reduce to self-defined quantities, fitted parameters renamed as predictions, or self-citation chains. Performance margins are reported outcomes of training and testing rather than tautological restatements of inputs. The analysis criteria for circularity (self-definitional, fitted-input-as-prediction, load-bearing self-citation, etc.) are not met.

Assumptions & free parameters 2 free parameters · 1 assumptions · 0 invented entities

The central claim depends on standard LLM training assumptions plus the unstated premise that the chosen financial data distribution and internal benchmarks are representative; no new entities are postulated.

free parameters (2)
  • Model parameter count (50 billion)
    Chosen scale for the model, likely balancing compute and performance.
  • Financial-to-general token ratio (363B:345B)
    Dataset mix proportions selected to achieve domain gains without general degradation.
assumptions (1)
  • domain assumption Standard transformer pretraining on next-token prediction transfers effectively to financial text when mixed with general data.
    Invoked to justify that mixed training preserves general capabilities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of BloombergGPT: A Large Language Model for Finance." pith.science (2026). https://pith.science/paper/PUSHK26R

@misc{pith2026230317564,
  author       = {Pith},
  title        = {Pith review of: BloombergGPT: A Large Language Model for Finance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PUSHK26R}},
  note         = {Machine review of arXiv:2303.17564}
}
read the original abstract

The use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering. Large Language Models (LLMs) have been shown to be effective on a variety of tasks; however, no LLM specialized for the financial domain has been reported in literature. In this work, we present BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data. We construct a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet, augmented with 345 billion tokens from general purpose datasets. We validate BloombergGPT on standard LLM benchmarks, open financial benchmarks, and a suite of internal benchmarks that most accurately reflect our intended usage. Our mixed dataset training leads to a model that outperforms existing models on financial tasks by significant margins without sacrificing performance on general LLM benchmarks. Additionally, we explain our modeling choices, training process, and evaluation methodology. We release Training Chronicles (Appendix C) detailing our experience in training BloombergGPT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 235 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 235 Pith citations

  1. FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

    cs.AI 2026-08 conditional novelty 8.0 of 10

    Rubrics derived from authentic financial deliverables beat prompt-only rubrics on role-specialized tasks (99.1% vs 78.0% coverage) while matching them on conventional tasks.

  2. Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers

    cs.SE 2025-06 conditional novelty 8.0 of 10

    First study of 1,899 MCP servers finds eight distinct vulnerabilities (only three traditional), 7.2% with general issues, 5.5% with tool poisoning, and 66% with code smells, urging MCP-specific security practices.

  3. TradingMoE: Routing the Right Experts in Evolving Markets

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A query-key router plus a sparse expert-selection-update mechanism lets a frozen LLM adapt its mixture-of-experts routing to trading context, beating 22 baselines in backtests.

  4. Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

    cs.CL 2026-08 conditional novelty 7.0 of 10

    A Jacobian-lens audit predicts model-level relearning recovery in LLM unlearning but cannot pick which facts return and backfires when used as a training penalty.

  5. FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

    cs.AI 2026-08 conditional novelty 7.0 of 10

    A new SEC-filing QA benchmark shows that retrieval models lose 13 to 20.5 points of ranking accuracy when plausible but wrong disclosures are used as distractors.

  6. IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

    cs.CL 2026-07 conditional novelty 7.0 of 10

    A 24,795-pair multilingual instruction dataset for Indian Knowledge Systems pedagogy brings a 7B fine-tune within 0.15 median-judge points of a strong general-purpose reference model.

  7. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    cs.CL 2026-07 conditional novelty 7.0 of 10

    On a new 615-question business-case benchmark graded by AI against instructor rubrics, frontier LLMs score 87-88% partial credit but complete only about half the questions.

  8. Meta-Benchmarks for Financial-Services LLM Evaluation

    cs.AI 2026-07 unverdicted novelty 7.0 of 10

    A meta-benchmarking framework organizes 452 LLM benchmarks into 41 O*NET Generalized Work Activities and 38 BIAN domains, using discrimination-coverage-recency weights to scale K-factors in an Elo tournament for compa...

  9. AI Trading's Alpha Singularity: Emergent Market Reasoning through Agent-to-Agent Self-Evolution

    cs.AI 2026-06 reject novelty 7.0 of 10

    Multi-agent LLM system Agora under Sealed Joint Search conditions produces +1.87 holdout Sharpe on CSI 1000 over a 91-day sealed period, exceeding the best baseline at +1.334 under favorable seed.

  10. LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    LLMs are applied in a generative pipeline for extracting, normalizing, and interpreting eligibility criteria from securities prospectuses, achieving up to 91% precision in document-level decisions with a conservative bias.

  11. InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    InvestPhilBench is a new multi-layer benchmark for LLM procedural reasoning in investment philosophy, with BASP metrics showing composite scores saturate while gate reconstruction accuracy reveals procedural deficits.

  12. MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    MacroLens is a point-in-time multi-signal benchmark dataset and seven tasks for evaluating contextual financial reasoning models under macroeconomic scenarios.

  13. The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    SEFD reconstructs SEC filings into MultiMarkdown to create a 152B-token financial pretraining corpus with low overlap to existing data and introduces EDGAR-Forecast and EDGAR-OCR benchmarks.

  14. LEDGER: A Long-Context Benchmark of Corporate Annual Reports for Grounded Financial Retrieval and Extraction

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    LEDGER provides a corpus of 4,999 annual reports with 31 labeled KPIs and three benchmarks for page-level retrieval, needle-in-haystack lookup, and full KPI extraction from long documents.

  15. It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    MUSE framework shows LLM conformity to user pushback arises from both sycophantic alignment and epistemic uncertainty, with both increasing when users appear expert or suggestions seem plausible.

  16. From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery

    cs.CE 2026-05 unverdicted novelty 7.0 of 10

    QuantEvolver applies reinforcement fine-tuning to evolve an LLM policy for generating executable alpha factor expressions, yielding higher-quality and more complementary factors than prompt-based baselines on market b...

  17. MeMo: Memory as a Model

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    MeMo encodes new knowledge into a separate memory model that integrates with frozen LLMs, showing strong performance on QA benchmarks while avoiding catastrophic forgetting and working without access to model weights.

  18. From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    A dataset-agnostic framework converts text tool-calling benchmarks to paired audio versions via TTS and noise, showing model-dependent performance with small text-to-voice gaps of 1.8-4.8 points on Confetti and When2Call.

  19. Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning

    cs.CY 2026-03 unverdicted novelty 7.0 of 10

    AWASH detects AI-washing via cross-modal inconsistency reasoning on a new trimodal benchmark of 88k corporate disclosure triplets, achieving F1 0.882 with a CMID network that grounds claims against patents and hiring data.

  20. SynBench: A Benchmark for Differentially Private Text Generation

    cs.AI 2025-09 conditional novelty 7.0 of 10

    SynBench benchmarks DP text generators across nine datasets and uses a new MIA to show that public pre-training on portions of private data overestimates synthetic text quality and breaks DP privacy bounds.

  21. CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A 5,352-pair benchmark shows thinking LLMs outperform non-thinking ones as code judges, but position bias and formatting sensitivity remain strong.

  22. RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

    cs.CL 2025-04 conditional novelty 7.0 of 10

    RAG can make language models less safe than their non-RAG equivalents, even with safe documents, and current jailbreak methods transfer poorly to RAG.

  23. Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories

    cs.SE 2024-12 conditional novelty 7.0 of 10

    An empirical study of 1,262 prompt changes across 243 GitHub repositories shows that developers mainly add and modify prompt components during feature development, rarely document the changes, and sometimes introduce ...

  24. Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Instruction tuning raises language models' confidence and reduces the variety of their reasoning rationales, without consistent accuracy gains.

  25. TIEM: Temporal Integration of Hypergraph Evidence and Skill Memory for Event-Driven Financial Forecasting

    cs.CE 2026-08 conditional novelty 6.0 of 10

    A timestamp-gated retrieval and memory framework (TIEM) with a new holdout benchmark claims consistent gains over ten baselines on five event-driven financial forecasting datasets.

  26. Context Is Not Authority: Structured Runtime Governance for Financial Market Agents

    cs.AI 2026-08 conditional novelty 6.0 of 10

    SAGE-Fin gates every financial-agent effect behind typed, state-checked receipts so that correct-sounding context cannot, by itself, authorize a response, trade, or policy.

  27. Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A hybrid analytical-ML framework predicts LLM inference latency and energy from architectural parameters, with MAPE below 5 percent on selected models and about 10 percent on a broad set.

  28. OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

    cs.CE 2026-08 conditional novelty 6.0 of 10

    OpenPM is an auditable point-in-time evaluation benchmark for LLM portfolio agents, whose 44-day case study shows analyst quality matters more than the constructor model and equal-weighting is a hard baseline.

  29. FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    FinReportBench is a fine-grained, expert-grounded benchmark for institution-grade LLM financial report generation, and its skill-evolution method improves G1 and G2 scores across model families.

  30. When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

    cs.CL 2026-08 conditional novelty 6.0 of 10

    ALiBi's linearly growing positional bias underflows floating-point attention in long contexts, zeroing out distant attention weights, with measurable but task-dependent effects on retrieval.

  31. Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory

    cs.CL 2026-08 conditional novelty 6.0 of 10

    LLM-based similarity of 10-K filings does not predict peer-firm stock reactions on acquisition announcements: pooled rank correlation +0.07, p=0.37.

  32. Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A federated fine-tuning framework that compresses foundation models on clients via SVD, aggregates adapters within groups and full-rank reconstructions across groups, then distills the result back into the full server model.

  33. Can Large Language Models Execute Parent Orders?

    cs.CE 2026-07 conditional novelty 6.0 of 10

    PACE, a planner–executor LLM framework, beats TWAP, Almgren–Chriss, and ML execution baselines by up to ~0.65 bps on Shenzhen parent orders without task-specific training.

  34. FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selective Financial Forecasting

    cs.LG 2026-07 reject novelty 6.0 of 10

    A multimodal RAG system with point-in-time retrieval and calibrated abstention is presented, with only simulated evidence that refusal reduces selective error and drawdown.

  35. Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

    cs.CL 2026-07 conditional novelty 6.0 of 10

    An LLM pipeline applied to 3,000 MIMIC-IV discharge summaries surfaced 3,460 candidate documentation inconsistencies, which the authors organize into a graded ontology of contradiction and ambiguity.

  36. Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CM-LRS is a workflow-output-layer reliability score for capital-markets LLM outputs; on five public-document workflows, frontier closed-source models score 4.09–4.31 and Llama 3.3 70B scores 3.15 under four LLM judges.

  37. FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering

    cs.IR 2026-07 conditional novelty 6.0 of 10

    FinSAgent improves financial filing QA by conditioning sub-queries on a summary of the local corpus and gating semantic reranking with a learned validity signal, beating baseline systems on five benchmarks.

  38. Planning with Transformers: Chain of Computation and Structured Context Windows

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Small transformers, trained from scratch on curated instruction traces and run inside a pointer-memory loop, solve BlocksWorld/Pancake at >99.89% and Tower of Hanoi to 20 disks.

  39. Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Among 8/8 self-consistent answers on FinQA, residual-stream probes detect wrong answers at 0.68–0.77 AUROC versus 0.55–0.63 for the best cheap output baselines across three 8–9B models.

  40. LLM-Enhanced Dynamic Financial Knowledge Graphs for Cross-Entity Signal Propagation and alpha discovery

    stat.AP 2026-07 conditional novelty 6.0 of 10

    In controlled simulations, community-aware propagation of LLM event signals on dynamic financial knowledge graphs recovers latent communities and prices incrementally beyond direct signals, though live alpha remains untested.

  41. FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance

    cs.AI 2026-07 conditional novelty 6.0 of 10

    FORCE-Bench, a 251-question enterprise-finance benchmark with rubric scoring, shows a purpose-built Microsoft finance agent outperforming general-purpose agents on ERP and public-data tasks.

  42. Grounded Event Extraction from SEC 8-K Filings with a Fine-Grained Taxonomy

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Schema-constrained, quote-grounded LLM extraction plus a second-pass quality score yields 601k auditable 8-K event tags whose precision and market reactions both improve with the score.

  43. CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    CLExEval introduces a human-annotated evaluation framework on 40 rare cases that identifies verbosity bias, hidden knowledge paradox, and 68.6% reasoning-to-output mismatch in LLMs while showing LLM-as-a-Judge overest...

  44. Fast Unlearning at Scale via Margin Self-Correction

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    MASC achieves competitive forget-retain trade-offs in language model unlearning at lower computational cost via margin self-correction and an online stopping criterion on TOFU, MUSE News, and MUSE Books.

  45. SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    SafeSteer restricts reverse KL penalty to safety tokens selected via activation steering, achieving strong safety on seven benchmarks with minimal degradation on five capability benchmarks using only 100 harmful sampl...

  46. IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    IPO-Mine releases a toolkit and large multimodal dataset for structured analysis of IPO filings and shows state-of-the-art models diverge from human judgments on chart quality and misleadingness.

  47. Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market

    q-fin.CP 2026-05 unverdicted novelty 6.0 of 10

    Using FOMC minutes to propose regime-shift candidates and a lenient text check to ratify data-detected candidates, the pipeline reaches F1=0.82 on 26 monetary-policy anchors, beating every data-only baseline.

  48. Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Distinguishable Deletion unifies knowledge erasure and refusal for LLM unlearning via an energy index that enforces boundaries during training and enables refusal at inference.

  49. Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Stateful sessions with incremental KV cache and flash queries allow O(|q|) latency in streaming transformer inference, delivering up to 5.9x speedup over conventional engines while preserving full attention.

  50. Sell Me This Stock: Unsafe Recommendation Drift in LLM Agents

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    Under manipulated tool outputs, eight LLM agents produce risk-mismatched stock recommendations in 65–99% of turns while NDCG-style quality scores stay nearly identical to clean runs.

  51. CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching

    cs.AI 2026-02 conditional novelty 6.0 of 10

    A new benchmark and training strategy show LLMs trained to internalize causal reasoning steps are less fooled by semantically similar, label-flipped questions than models using explicit chain-of-thought.

  52. All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Shapley-weighted leakage rates show that standard LLM backtests leak post-cutoff facts, and the TimeSPEC pipeline cuts measured leakage by 75-99% at the cost of accuracy on leakage-sensitive tasks.

  53. AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

    q-fin.TR 2026-02 conditional novelty 6.0 of 10

    LLMs are unreliable when asked to emit buy/sell/hold actions, so this paper benchmarks them as code-writing quantitative researchers whose generated strategies are backtested deterministically.

  54. Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation

    cs.SE 2026-02 conditional novelty 6.0 of 10

    A physics-aware two-stage SFT+GRPO training method with period/AST/sandbox rewards raises small open LLMs from ~0-2% to ~68-77% on a strict OpenSeesPy building-modeling benchmark (BMEval).

  55. Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection

    cs.LG 2025-12 unverdicted novelty 6.0 of 10

    FinFRE-RAG combines importance-guided feature reduction with label-aware retrieval-augmented generation to boost LLM performance on tabular fraud detection across four public datasets while providing human-readable ra...

  56. MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications

    cs.AI 2025-11 unverdicted novelty 6.0 of 10

    MM-Telco creates multimodal benchmarks for telecom and demonstrates that fine-tuned LLMs and VLMs achieve significant performance gains on domain-specific tasks.

  57. Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Evolution strategies can full-parameter fine-tune billion-parameter LLMs, outperforming PPO and GRPO on the Countdown task and reward robustness in a conciseness task.

  58. Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain

    cs.CL 2025-09 unverdicted novelty 6.0 of 10

    CoRT achieves 95% average attack success rate on nine LLMs by using iterative risk-concealing prompts and a controller that scores concealment levels on a new 522-instruction financial risk benchmark.

  59. From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets

    econ.GN 2025-09 conditional novelty 6.0 of 10

    LLM experts in credence goods markets reduce efficiency and consumer surplus unless liability or transparent prosocial objectives operate, and expert delegation with transparent objectives can outperform human-only markets.

  60. MM-DREX: Multimodal-Driven Dynamic Routing of LLM Experts for Financial Trading

    q-fin.TR 2025-09 reject novelty 6.0 of 10

    A multimodal LLM router dynamically weights four technical trading experts and, in cost-free backtests, outperforms 15 baselines across stocks, futures, and crypto.

See all 235 Pith citations

Reference graph

Works this paper leans on

140 extracted references · 140 canonical work pages · cited by 235 Pith papers (see all)

  1. [1]

    FinBERT: Financial Sentiment Analysis with Pre-trained Language Models

    Dogu Araci. Finbert: Financial sentiment analysis with pre-trained language models. arXiV preprint arXiV:1908.10063, 2019

  2. [2]

    PLATO - XL : Exploring the large-scale pre-training of dialogue generation

    Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhihua Wu, Zhen Guo, Hua Lu, Xinxian Huang, Xin Tian, Xinchao Xu, Yingzhan Lin, and Zheng-Yu Niu. PLATO - XL : Exploring the large-scale pre-training of dialogue generation. In Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, pages 107--118, Online only, November 2...

  3. [3]

    2019 , address =

    Iz Beltagy, Kyle Lo, and Arman Cohan. S ci BERT : A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615--3620, Hong Kong, China, November 2019. Association for Computation...

  4. [4]

    On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610--623, 2021

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610--623, 2021

  5. [5]

    The fifth PASCAL recognizing textual entailment challenge

    Luisa Bentivogli, Bernardo Magnini, Ido Dagan, Hoa Trang Dang, and Danilo Giampiccolo. The fifth PASCAL recognizing textual entailment challenge. In Proceedings of the Second Text Analysis Conference, TAC 2009, Gaithersburg, Maryland, USA, November 16-17, 2009 . NIST , 2009. URL https://tac.nist.gov/publications/2009/additional.papers/RTE5\_overview.proce...

  6. [6]

    The values encoded in machine learning research

    Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. The values encoded in machine learning research. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 173--184, 2022

  7. [7]

    PIQA: reasoning about physical commonsense in natural language

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in...

  8. [8]

    Black, G

    Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow , March 2021. URL https://doi.org/10.5281/zenodo.5297715. If you use this software, please cite it using these metadata

Show all 140 references
  1. [9]

    GPT - N eo X -20 B : An open-source autoregressive language model

    Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. GPT - N eo X -20 ...

  2. [10]

    BioMedLM

    Elliot Bolton, David Hall, Michihiro Yasunaga, Tony Lee, Chris Manning, and Percy Liang. BioMedLM . https://github.com/stanford-crfm/BioMedLM, 2023

  3. [11]

    Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, S. Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen A. Creel, ...

  4. [12]

    Byte pair encoding is suboptimal for language model pretraining

    Kaj Bostrom and Greg Durrett. Byte pair encoding is suboptimal for language model pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4617--4624, Online, November 2020. Association for Computational Linguistics. doi:10.18653/v1/2020.fin...

  5. [13]

    Popat, Peng Xu, Franz J

    Thorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och, and Jeffrey Dean. Large language models in machine translation. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning ( EMNLP - C o NLL...

  6. [14]

    Class-based n-gram models of natural language

    Peter F Brown, Vincent J Della Pietra, Peter V Desouza, Jennifer C Lai, and Robert L Mercer. Class-based n-gram models of natural language. Computational linguistics, 18 0 (4): 0 467--480, 1992

  7. [15]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jef...

  8. [16]

    Brown, Dawn Xiaodong Song, \'U lfar Erlingsson, Alina Oprea, and Colin Raffel

    Nicholas Carlini, Florian Tram \`e r, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, \'U lfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting training data from large language models. In USENIX Security...

  9. [17]

    Quantifying memorization across neural language models, 2022

    Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models, 2022. URL https://arxiv.org/abs/2202.07646

  10. [18]

    Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William H

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harrison Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder...

  11. [19]

    Training deep nets with sublinear memory cost

    Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost. arXiV preprint arXiV:1604.06174, 2016

  12. [20]

    F in QA : A dataset of numerical reasoning over financial data

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. F in QA : A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirica...

  13. [21]

    C onv F in QA : Exploring the chain of numerical reasoning in conversational finance question answering

    Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. C onv F in QA : Exploring the chain of numerical reasoning in conversational finance question answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pro...

  14. [22]

    Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, V...

  15. [23]

    B ool Q : Exploring the surprising difficulty of natural yes/no questions

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. B ool Q : Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Compu...

  16. [24]

    Think you have solved question answering? try arc, the ai2 reasoning challenge

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiV, abs/1803.05457, 2018

  17. [25]

    The pascal recognising textual entailment challenge

    Ido Dagan, Oren Glickman, and Bernardo Magnini. The pascal recognising textual entailment challenge. In Machine Learning Challenges Workshop, 2007

  18. [26]

    The commitmentbank: Investigating projection in naturally occurring discourse

    Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, pages 107--124, 2019

  19. [27]

    Bernice: A multilingual pre-trained encoder for T witter

    Alexandra DeLucia, Shijie Wu, Aaron Mueller, Carlos Aguirre, Philip Resnik, and Mark Dredze. Bernice: A multilingual pre-trained encoder for T witter. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6191--6205, Abu Dhabi, United...

  20. [28]

    8-bit optimizers via block-wise quantization

    Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 8-bit optimizers via block-wise quantization. In International Conference on Learning Representations, 2022

  21. [29]

    BERT : Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Lan...

  22. [30]

    Documenting large webtext corpora: A case study on the colossal clean crawled corpus

    Jesse Dodge, Maarten Sap, Ana Marasovi \'c , William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods i...

  23. [31]

    How twitter is changing the nature of financial news discovery

    Mark Dredze, Prabhanjan Kambadur, Gary Kazantsev, Gideon Mann, and Miles Osborne. How twitter is changing the nature of financial news discovery. In proceedings of the second international workshop on data science for macro-modeling, pages 1--5, 2016

  24. [32]

    Natural language processing in accounting, auditing and finance: A synthesis of the literature with a roadmap for future research

    Ingrid E Fisher, Margaret R Garnsey, and Mark E Hughes. Natural language processing in accounting, auditing and finance: A synthesis of the literature with a roadmap for future research. Intelligent Systems in Accounting, Finance and Management, 23 0 (3): 0 157--214, 2016

  25. [33]

    The pile: An 800gb dataset of diverse text for language modeling, 2021

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling, 2021. URL https://arxiv.org/abs/2101.00027

  26. [34]

    Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text, 2022

    Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text, 2022. URL https://arxiv.org/abs/2202.06935

  27. [35]

    The third PASCAL recognizing textual entailment challenge

    Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. The third PASCAL recognizing textual entailment challenge. In Proceedings of the ACL - PASCAL Workshop on Textual Entailment and Paraphrasing , pages 1--9, Prague, June 2007. Association for Computational Linguis...

  28. [36]

    Improving alignment of dialogue agents via targeted human judgements, 2022

    Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig...

  29. [37]

    Gordon, Zornitsa Kozareva, and Melissa Roemmele

    Andrew S. Gordon, Zornitsa Kozareva, and Melissa Roemmele. Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In International Workshop on Semantic Evaluation, 2011

  30. [38]

    News summarization and evaluation in the era of gpt-3, 2022

    Tanya Goyal, Junyi Jessy Li, and Greg Durrett. News summarization and evaluation in the era of gpt-3, 2022. URL https://arxiv.org/abs/2209.12356

  31. [39]

    Suchin Gururangan, Ana Marasovi \'c , Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. Don ' t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, page...

  32. [40]

    The second pascal recognising textual entailment challenge

    R Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. The second pascal recognising textual entailment challenge. In Proceedings of the Second PASCAL Challenges Workshop on Recognising Textual Entailment, volume 7, 2006

  33. [41]

    Gaussian error linear units (gelus)

    Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiV preprint arXiV:1606.08415, 2016

  34. [42]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ

  35. [43]

    Query-key normalization for transformers

    Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4246--4253, Online, November 2020. Association for Computational Linguistics....

  36. [44]

    Scaling laws for transfer

    Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish. Scaling laws for transfer. arXiV preprint arXiV:2102.01293, 2021

  37. [45]

    An empirical analysis of compute-optimal large language model training

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon ...

  38. [46]

    Universal language model fine-tuning for text classification

    Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328--339, Melbourne, Australia, July 2018. Association for...

  39. [47]

    Clinicalbert: Modeling clinical notes and predicting hospital readmission

    Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiV, 4 2019. URL http://arxiv.org/abs/1904.05342

  40. [48]

    Continuous speech recognition by statistical methods

    Frederick Jelinek. Continuous speech recognition by statistical methods. Proceedings of the IEEE, 64 0 (4): 0 532--556, 1976

  41. [49]

    Data governance in the age of large-scale data-driven language technology

    Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Isaac Johnson, Gerard Dupont, Jesse Dodge, Kyle Lo, Zeerak Talat, Dragomir Radev, Aaron Gokaslan, Somaieh Nikpoor, Peter Henderso...

  42. [50]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiV, 1 2020. URL http://arxiv.org/abs/2001.08361

  43. [51]

    Amazon sagemaker model parallelism: A general and flexible framework for large model training

    Can Karakus, Rahul Huilgol, Fei Wu, Anirudh Subramanian, Cade Daniel, Derya Cavdar, Teng Xu, Haohan Chen, Arash Rahnama, and Luis Quintela. Amazon sagemaker model parallelism: A general and flexible framework for large model training. arXiV preprint arXiV:2111.05972, 2021

  44. [52]

    Looking beyond the surface: A challenge set for reading comprehension over multiple sentences

    Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computati...

  45. [53]

    Reducing activation recomputation in large transformer models, 2022

    Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in large transformer models, 2022. URL https://arxiv.org/abs/2205.05198

  46. [54]

    Subword regularization: Improving neural network translation models with multiple subword candidates

    Taku Kudo. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66--75, Melbourne, Australia, July 2018. A...

  47. [55]

    S entence P iece: A simple and language independent subword tokenizer and detokenizer for neural text processing

    Taku Kudo and John Richardson. S entence P iece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66--71, Brus...

  48. [56]

    RACE : Large-scale R e A ding comprehension dataset from examinations

    Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. RACE : Large-scale R e A ding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785--794, Copenhagen, Denmark, September 20...

  49. [57]

    Teven Le Scao, Thomas Wang, Daniel Hesslow, Stas Bekman, M Saiful Bari, Stella Biderman, Hady Elsahar, Niklas Muennighoff, Jason Phang, Ofir Press, Colin Raffel, Victor Sanh, Sheng Shen, Lintang Sutawika, Jaesung Tae, Zheng Xin Yong, Julien Launay, and Iz Beltagy. What languag...

  50. [58]

    Biobert: A pre-trained biomedical language representation model for biomedical text mining

    Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36: 0 1234--1240, 2 2020. ISSN 14602059. doi:10.1093/bioinformatics/btz682

  51. [59]

    Deduplicating training data makes language models better

    Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume ...

  52. [60]

    Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael S

    Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard - Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, Rose E. Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael S. Bernstein, and Percy Liang. Ev...

  53. [61]

    Smith, Zachary Ziegler, Daniel Nadler, Peter Szolovits, Alistair Johnson, and Emily Alsentzer

    Eric Lehman, Evan Hernandez, Diwakar Mahajan, Jonas Wulff, Micah J. Smith, Zachary Ziegler, Daniel Nadler, Peter Szolovits, Alistair Johnson, and Emily Alsentzer. Do we still need clinical language models?, 2023. URL https://arxiv.org/abs/2302.08091

  54. [62]

    Levesque, Ernest Davis, and L

    Hector J. Levesque, Ernest Davis, and L. Morgenstern. The winograd schema challenge. In International Conference on Principles of Knowledge Representation and Reasoning, 2011

  55. [63]

    Limits to depth efficiencies of self-attention

    Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua. Limits to depth efficiencies of self-attention. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 22640--22651. Curra...

  56. [64]

    Pretrained language models for biomedical and clinical tasks: Understanding and extending the state-of-the-art

    Patrick Lewis, Myle Ott, Jingfei Du, and Veselin Stoyanov. Pretrained language models for biomedical and clinical tasks: Understanding and extending the state-of-the-art. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, pages 146--157, Online, November ...

  57. [65]

    Solving quantitative reasoning problems with language models, 2022

    Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models, ...

  58. [66]

    u ksekg \

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R \' e , Diana Acosta ...

  59. [67]

    Jurassic-1: Technical details and evaluation

    Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham. Jurassic-1: Technical details and evaluation. White Paper. AI21 Labs, 1, 2021

  60. [68]

    Language models of protein sequences at the scale of evolution enable accurate structure prediction

    Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, and Alexander Rives. Language models of protein sequences at the scale of evolution enable accurate structure prediction. bioRxiv, 202...

  61. [69]

    Autoregressive structured prediction with language models

    Tianyu Liu, Yuchen Eleanor Jiang, Nicholas Monath, Ryan Cotterell, and Mrinmaya Sachan. Autoregressive structured prediction with language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 993--1005, Abu Dhabi, United Arab Emirates, Decemb...

  62. [70]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7

  63. [71]

    BioGPT : generative pre-trained transformer for biomedical text generation and mining

    Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. BioGPT : generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23 0 (6), sep 2022. doi:10.1093/bib/bbac409. URL https://doi.org/10.1093

  64. [72]

    Exploring cross-sentence contexts for named entity recognition with BERT

    Jouni Luoma and Sampo Pyysalo. Exploring cross-sentence contexts for named entity recognition with BERT . In Proceedings of the 28th International Conference on Computational Linguistics, pages 904--914, Barcelona, Spain (Online), December 2020. International Committee on Comp...

  65. [73]

    Www'18 open challenge: Financial opinion mining and question answering

    Macedo Maia, Siegfried Handschuh, Andr \' e Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. Www'18 open challenge: Financial opinion mining and question answering. In Pierre - Antoine Champin, Fabien Gandon, Mounia Lalmas, and Panagiotis G. Ipeiroti...

  66. [74]

    Korhonen, Jyrki Wallenius, and Pyry Takala

    Pekka Malo, Ankur Sinha, Pekka J. Korhonen, Jyrki Wallenius, and Pyry Takala. Good debt or bad debt: Detecting semantic orientations in economic texts. J. Assoc. Inf. Sci. Technol., 65 0 (4): 0 782--796, 2014. doi:10.1002/asi.23062. URL https://doi.org/10.1002/asi.23062

  67. [75]

    Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y

    Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan. Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp, 2021. URL https...

  68. [76]

    Can a suit of armor conduct electricity? a new dataset for open book question answering

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381--2391, Brussels, Belgi...

  69. [77]

    Recurrent neural network based language model

    Tomas Mikolov, Martin Karafi \'a t, Lukas Burget, Jan Cernock \`y , and Sanjeev Khudanpur. Recurrent neural network based language model. In Interspeech, pages 1045--1048. Makuhari, 2010

  70. [78]

    A corpus and cloze evaluation for deeper understanding of commonsense stories

    Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North A merican Chapte...

  71. [79]

    BERT weet: A pre-trained language model for E nglish tweets

    Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. BERT weet: A pre-trained language model for E nglish tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 9--14, Online, October 2020. Association for Com...

  72. [80]

    Adversarial NLI : A new benchmark for natural language understanding

    Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. Adversarial NLI : A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4885--4901, Online, July...

  73. [81]

    Weinstein, Nikhil Naik, and Ali Madani

    Erik Nijkamp, Jeffrey Ruffolo, Eli N. Weinstein, Nikhil Naik, and Ali Madani. Progen2: Exploring the boundaries of protein language models. CoRR, abs/2206.13517, 2022. doi:10.48550/arXiv.2206.13517. URL https://doi.org/10.48550/arXiv.2206.13517

  74. [82]

    Train with mixed precision, 2023

    NVIDIA. Train with mixed precision, 2023. URL https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html

  75. [83]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, an...

  76. [84]

    Godel: Large-scale pre-training for goal-directed dialog

    Baolin Peng, Michel Galley, Pengcheng He, Chris Brockett, Lars Liden, Elnaz Nouri, Zhou Yu, Bill Dolan, and Jianfeng Gao. Godel: Large-scale pre-training for goal-directed dialog. arXiV preprint arXiV:2206.11309, 2022

  77. [85]

    W i C : the word-in-context dataset for evaluating context-sensitive meaning representations

    Mohammad Taher Pilehvar and Jose Camacho-Collados. W i C : the word-in-context dataset for evaluating context-sensitive meaning representations. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Languag...

  78. [86]

    Train short, test long: Attention with linear biases enables input length extrapolation

    Ofir Press, Noah Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=R8sQPpGCv0

  79. [87]

    Improving language understanding by generative pre-training, 2018

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018. URL https://gluebenchmark.com/leaderboard

  80. [88]

    Language models are unsupervised multitask learners, 2019

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019. URL https://github.com/codelucas/newspaper

  81. [89]

    Rae, Anna Potapenko, Siddhant M

    Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. Compressive transformers for long-range sequence modelling. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenRevie...

  82. [90]

    Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...

  83. [91]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21 0 (140): 0 1--67, 2020. URL...

  84. [92]

    Zero: Memory optimizations toward training trillion parameter models

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1--16. IEEE, 2020

  85. [93]

    Recipes for building an open-domain chatbot

    Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association f...

  86. [94]

    WINOGRANDE : An adversarial winograd schema challenge at scale

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WINOGRANDE : An adversarial winograd schema challenge at scale. Commun. ACM, 64: 0 99--106, 2019

  87. [95]

    Domain adaption of named entity recognition to support credit risk assessment

    Julio Cesar Salinas Alvarado, Karin Verspoor, and Timothy Baldwin. Domain adaption of named entity recognition to support credit risk assessment. In Proceedings of the Australasian Language Technology Association Workshop 2015, pages 84--90, Parramatta, Australia, December 201...

  88. [96]

    Multitask prompted training enables zero-shot task generalization

    Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyo...

  89. [97]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît S...

  90. [98]

    Japanese and korean voice search

    Mike Schuster and Kaisuke Nakajima. Japanese and korean voice search. In 2012 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5149--5152. IEEE, 2012

  91. [99]

    Neural machine translation of rare words with subword units

    Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715--1725, Berlin, Germany, August 2016. As...

  92. [100]

    When FLUE meets FLANG : Benchmarks and large pretrained language model for financial domain

    Raj Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah, Wendi Du, Sudheer Chava, Natraj Raman, Charese Smiley, Jiaao Chen, and Diyi Yang. When FLUE meets FLANG : Benchmarks and large pretrained language model for financial domain. In Proceedings of the 2022 Conference on Empirical...

  93. [101]

    GLU variants improve transformer

    Noam Shazeer. GLU variants improve transformer. arXiV preprint arXiV:2002.05202, 2020

  94. [102]

    Megatron-lm: Training multi-billion parameter language models using model parallelism

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiV preprint arXiV:1909.08053, 2019

  95. [103]

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Scharli, Aakanksha Chowdhery, Philip Mansfield, Blaise Ague...

  96. [104]

    Impact of news on the commodity market: Dataset and results

    Ankur Sinha and Tanmay Khandait. Impact of news on the commodity market: Dataset and results. CoRR, abs/2009.04202, 2020. URL https://arxiv.org/abs/2009.04202

  97. [105]

    Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model, 2022

    Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michae...

  98. [106]

    Saleh Soltan, Shankar Ananthakrishnan, Jack G. M. FitzGerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith S. Peris, Stephen Rawls, Andrew Rosenbaum, Anna Rumshisky, Chandan Prakash, Mukund Sridhar, Fabian Triefenbach, Apurv Verma, Gokhan Tur, and Premkumar Natarajan. Alexatm...

  99. [107]

    Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, ...

  100. [109]

    Roformer: Enhanced transformer with rotary position embedding

    Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. CoRR, abs/2104.09864, 2021 b . URL https://arxiv.org/abs/2104.09864

  101. [110]

    Generating text with recurrent neural networks

    Ilya Sutskever, James Martens, and Geoffrey E Hinton. Generating text with recurrent neural networks. In Proceedings of the 28th international conference on machine learning (ICML-11), pages 1017--1024, 2011

  102. [111]

    Le, Ed H

    Mirac Suzgun, Nathan Scales, Nathanael Sch \" a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. CoRR, abs/2210.09261, 2022. doi:10....

  103. [112]

    General-purpose question-answering with macaw

    Oyvind Tafjord and Peter Clark. General-purpose question-answering with macaw. arXiV preprint arXiV:2109.02593, 2021

  104. [113]

    C ommonsense QA : A question answering challenge targeting commonsense knowledge

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. C ommonsense QA : A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human La...

  105. [114]

    Scale efficiently: Insights from pre-training and fine-tuning transformers

    Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. Scale efficiently: Insights from pre-training and fine-tuning transformers. arXiV preprint arXiV:2109.10686, 2021

  106. [115]

    Scaling laws vs model architectures: How does inductive bias influence scaling? arXiV preprint arXiV:2207.10551, 2022 a

    Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q Tran, Dani Yogatama, and Donald Metzler. Scaling laws vs model architectures: How does inductive bias influence scaling? arXiV preprint arXiV:2207.10551, 2022 a

  107. [116]

    Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Siamak Shakeri, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler

    Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Siamak Shakeri, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. Ul2: Unifying language learning paradigms, 2022 b . URL https://arxiv.o...

  108. [117]

    Galactica: A large language model for science

    Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. Galactica: A large language model for science. arXiV, 11 2022. URL http://arxiv.org/abs/2211.09085

  109. [118]

    Lamda: Language models for dialog applications, 2022

    Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Q...

  110. [119]

    Tjong Kim Sang and Fien De Meulder

    Erik F. Tjong Kim Sang and Fien De Meulder. Introduction to the C o NLL -2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT - NAACL 2003 , pages 142--147, 2003. URL https://aclanthology....

  111. [120]

    LLaMA : Open and efficient foundation language models, 2023

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA : Open and efficient foundation languag...

  112. [121]

    Best practices for managing data annotation projects, 2020

    Tina Tseng, Amanda Stent, and Domenic Maida. Best practices for managing data annotation projects, 2020. URL http://rgdoi.net/10.13140/RG.2.2.34497.58727

  113. [122]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural In...

  114. [123]

    GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model

    Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model . https://github.com/kingoflolz/mesh-transformer-jax, May 2021

  115. [124]

    Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi. C ode T 5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8696--8708, O...

  116. [125]

    Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

    Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners, 2021. URL https://arxiv.org/abs/2109.01652

  117. [126]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. T...

  118. [127]

    Chi, Quoc V Le, and Denny Zhou

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in...

  119. [128]

    Ethical and social risks of harm from language models, 2021

    Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks...

  120. [129]

    Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John F. J. Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sande Minnich Brown, Zachary Kenton, William T. Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendrick...

  121. [130]

    Challenges in detoxifying language models

    Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. Challenges in detoxifying language models. In Findings of the Association for Computational Linguistics: EMNLP 20...

  122. [131]

    Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT

    Shijie Wu and Mark Dredze. Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP...

  123. [132]

    Chen, Quoc V

    Yonghui Wu, Mike Schuster, Z. Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens...

  124. [133]

    Modeling protein using large-scale pretrain language model

    Yijia Xiao, Jiezhong Qiu, Ziang Li, Chang - Yu Hsieh, and Jie Tang. Modeling protein using large-scale pretrain language model. CoRR, abs/2108.07435, 2021. URL https://arxiv.org/abs/2108.07435

  125. [134]

    Natural language based financial forecasting: a survey

    Frank Z Xing, Erik Cambria, and Roy E Welsch. Natural language based financial forecasting: a survey. Artificial Intelligence Review, 50 0 (1): 0 49--73, 2018

  126. [135]

    Detoxifying language models risks marginalizing minority voices

    Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. Detoxifying language models risks marginalizing minority voices. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human L...

  127. [136]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4800, Florence, Italy, July 2019. Associatio...

  128. [137]

    Glm-130b: An open bilingual pre-trained model

    Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. Glm-130b: An open bilingual pre-trained model. arXiV, 10 2...

  129. [138]

    Record: Bridging the gap between human and machine commonsense reading comprehension

    Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. Record: Bridging the gap between human and machine commonsense reading comprehension. arXiV, abs/1810.12885, 2018

  130. [139]

    Opt: Open pre-trained transformer language models

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. O...

  131. [140]

    DIALOGPT : Large-scale generative pre-training for conversational response generation

    Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for C...

  132. [141]

    Mics: Near-linear scaling for training gigantic model on public cloud, 2022 b

    Zhen Zhang, Shuai Zheng, Yida Wang, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li, and Xin Jin. Mics: Near-linear scaling for training gigantic model on public cloud, 2022 b . URL https://arxiv.org/abs/2205.00119

Pith tools

Reviewed May 13, 2026 · model on record in the stance chip above.