pith. sign in

Jack Clark

Identifiers

  • name variant Jack Clark 0.60 · backfill

Papers (11)

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training cs.CR · 2024 · author #18
  2. Towards Measuring the Representation of Subjective Global Opinions in Language Models cs.CL · 2023 · author #17
  3. Discovering Language Model Behaviors with Model-Written Evaluations cs.CL · 2022 · author #55
  4. In-context Learning and Induction Heads cs.LG · 2022 · author #23
  5. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cs.CL · 2022 · author #36
  6. Language Models (Mostly) Know What They Know cs.CL · 2022 · author #31
  7. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cs.CL · 2022 · author #27
  8. A General Language Assistant as a Laboratory for Alignment cs.CL · 2021 · author #19
  9. Learning Transferable Visual Models From Natural Language Supervision cs.CV · 2021 · author #10
  10. Language Models are Few-Shot Learners cs.CL · 2020 · author #26
  11. Release Strategies and the Social Impacts of Language Models cs.CL · 2019 · author #3

Mentions

  • 1908.09203 #3 · arxiv_oai · confidence 0.70 Jack Clark
  • 2306.16388 #17 · arxiv_oai · confidence 0.70 Jack Clark

Frequent Coauthors