REVIEW 18 cited by
The Foundation Model Transparency Index
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Foundation models have rapidly permeated society, catalyzing a wave of generative AI applications spanning enterprise and consumer-facing contexts. While the societal impact of foundation models is growing, transparency is on the decline, mirroring the opacity that has plagued past digital technologies (e.g. social media). Reversing this trend is essential: transparency is a vital precondition for public accountability, scientific innovation, and effective governance. To assess the transparency of the foundation model ecosystem and help improve transparency over time, we introduce the Foundation Model Transparency Index. The Foundation Model Transparency Index specifies 100 fine-grained indicators that comprehensively codify transparency for foundation models, spanning the upstream resources used to build a foundation model (e.g data, labor, compute), details about the model itself (e.g. size, capabilities, risks), and the downstream use (e.g. distribution channels, usage policies, affected geographies). We score 10 major foundation model developers (e.g. OpenAI, Google, Meta) against the 100 indicators to assess their transparency. To facilitate and standardize assessment, we score developers in relation to their practices for their flagship foundation model (e.g. GPT-4 for OpenAI, PaLM 2 for Google, Llama 2 for Meta). We present 10 top-level findings about the foundation model ecosystem: for example, no developer currently discloses significant information about the downstream impact of its flagship model, such as the number of users, affected market sectors, or how users can seek redress for harm. Overall, the Foundation Model Transparency Index establishes the level of transparency today to drive progress on foundation model governance via industry standards and regulatory intervention.
Forward citations
Cited by 18 Pith papers
-
FlexOlmo: Open Language Models for Flexible Data Use
FlexOlmo merges independently trained language-model experts, trained on private data, into a single mixture-of-experts model without joint training.
-
Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?
Released BPE vocabularies can be used to estimate per-token frequency ratios of a hidden training corpus, with mean relative errors as low as 3% in controlled settings and around 6% on SmolLM.
-
Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI
A Basel-III-style two-layer system—coordinated finder-coordinator-defender reporting plus ECAR, CRTH, and ARS buffers—can detect and dampen correlated risk build-up across frontier AI labs’ internal deployments.
-
User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
All six leading U.S. AI chatbot developers, as of May 2025, appear to train their models on users' chat data by default, often without clear opt-out options.
-
MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI
MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.
-
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
Indirect data poisoning (gradient-matching prompts) makes LLMs learn secret prompt-response pairs absent from training data, detectable with certified p-values and under 0.005% contaminated tokens.
-
The AI Agent Index
The AI Agent Index catalogs 67 deployed agentic AI systems and shows that most developers publicly disclose little about safety policies and evaluations.
-
Bridging the Data Provenance Gap Across Text, Speech and Video
A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...
-
Watermarking Training Data of Music Generation Models
Injected audio watermarks can be detected in music generated by a fine-tuned MusicGen model; simple tones work best, and a neural watermark requires dozens of repeated embeddings.
-
Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
Current large language models show stigma and give clinically inappropriate responses to common mental health symptoms, so they should not be deployed as replacement therapists.
-
RedPajama: an Open Dataset for Training Large Language Models
The paper releases open pretraining corpora, RedPajama-V1 and V2, and shows that web-data quality signals can be used to filter V2 into competitive training sets.
-
AI Data Development: A Scorecard for the System Card Framework
A five-area scoring rubric for AI dataset documentation, applied to four datasets, shows consistent gaps in collection ethics and preprocessing details.
-
AI Governance through Markets
Market governance mechanisms, supported by standardized AI disclosures, can create financial incentives for responsible AI development, according to this policy paper.
-
Generative AI regulation can learn from social media regulation
Regulatory lessons from social media, especially transparency, researcher access, trust and safety, and a global perspective, can be transferred to generative AI regulation.
-
7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
The authors trained and openly released a 7B LLM, an instruction-tuned variant, a GRPO-based reasoning variant, and a VLM, claiming competitive or superior performance on zero-shot, few-shot, CoT, and VLM benchmarks.
-
Usage Governance Advisor: From Intent to AI Governance
The paper describes an IBM proof-of-concept system that combines a knowledge graph and LLM pipelines to convert a use-case description into prioritized risks, model choices, benchmarks, and mitigation actions.
-
Data-Centric Safety and Ethical Measures for Data and AI Governance
A conceptual framework that maps dataset safety practices to six stages of the AI lifecycle, synthesizing existing documentation and red-teaming recommendations.
-
Generative AI in Medicine
A stakeholder-based review of generative AI use cases in medicine and the consent, privacy, transparency, hallucination, usability, equity, evaluation, and accountability challenges that stand between prototypes and s...
Discussion (0). Continue with ORCID to comment.