pith. sign in

Jackson Kernion

Identifiers

  • name variant Jackson Kernion 0.60 · backfill

Papers (10)

  1. Measuring Faithfulness in Chain-of-Thought Reasoning cs.AI · 2023 · author #10
  2. Discovering Language Model Behaviors with Model-Written Evaluations cs.CL · 2022 · author #25
  3. Constitutional AI: Harmlessness from AI Feedback cs.CL · 2022 · author #5
  4. Measuring Progress on Scalable Oversight for Large Language Models cs.HC · 2022 · author #20
  5. In-context Learning and Induction Heads cs.LG · 2022 · author #18
  6. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cs.CL · 2022 · author #3
  7. Language Models (Mostly) Know What They Know cs.CL · 2022 · author #23
  8. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models cs.CL · 2022 · author #169
  9. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cs.CL · 2022 · author #13
  10. A General Language Assistant as a Laboratory for Alignment cs.CL · 2021 · author #14

Mentions

  • 2211.03540 #20 · arxiv_oai · confidence 0.70 Jackson Kernion

Frequent Coauthors