Chris Olah
Identifiers
- name variant Chris Olah 0.60 · backfill
Papers (10)
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet cs.AI · 2026 · author #25
- Emotion Concepts and their Function in a Large Language Model cs.AI · 2026 · author #15
- In-context Learning and Induction Heads cs.LG · 2022 · author #26
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cs.CL · 2022 · author #34
- Language Models (Mostly) Know What They Know cs.CL · 2022 · author #35
- Scaling Laws and Interpretability of Learning from Repeated Data cs.LG · 2022 · author #13
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cs.CL · 2022 · author #29
- A General Language Assistant as a Laboratory for Alignment cs.CL · 2021 · author #21
- Concrete Problems in AI Safety cs.AI · 2016 · author #2
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems cs.DC · 2016 · author #25
Mentions
- 2605.29358 #25 · arxiv_oai · confidence 0.70 Chris Olah
- 2205.10487 #13 · arxiv_oai · confidence 0.70 Chris Olah
Frequent Coauthors
- Tom Henighan 8 shared papers
- Dario Amodei 7 shared papers
- Andy Jones 6 shared papers
- Ben Mann 6 shared papers
- Catherine Olsson 6 shared papers
- Danny Hernandez 6 shared papers
- Dawn Drain 6 shared papers
- Jared Kaplan 6 shared papers
- Nelson Elhage 6 shared papers
- Nicholas Joseph 6 shared papers
- Nova DasSarma 6 shared papers
- Sam McCandlish 6 shared papers
- Tom Brown 6 shared papers
- Tom Conerly 6 shared papers
- Zac Hatfield-Dodds 6 shared papers
- Amanda Askell 5 shared papers
- Anna Chen 5 shared papers
- Deep Ganguli 5 shared papers
- Jack Clark 5 shared papers
- Jackson Kernion 5 shared papers