pith. sign in

Chris Olah

Identifiers

  • name variant Chris Olah 0.60 · backfill

Papers (10)

  1. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet cs.AI · 2026 · author #25
  2. Emotion Concepts and their Function in a Large Language Model cs.AI · 2026 · author #15
  3. In-context Learning and Induction Heads cs.LG · 2022 · author #26
  4. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cs.CL · 2022 · author #34
  5. Language Models (Mostly) Know What They Know cs.CL · 2022 · author #35
  6. Scaling Laws and Interpretability of Learning from Repeated Data cs.LG · 2022 · author #13
  7. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cs.CL · 2022 · author #29
  8. A General Language Assistant as a Laboratory for Alignment cs.CL · 2021 · author #21
  9. Concrete Problems in AI Safety cs.AI · 2016 · author #2
  10. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems cs.DC · 2016 · author #25

Mentions

  • 2605.29358 #25 · arxiv_oai · confidence 0.70 Chris Olah
  • 2205.10487 #13 · arxiv_oai · confidence 0.70 Chris Olah

Frequent Coauthors