pith. sign in

Yuntao Bai

Identifiers

  • name variant Yuntao Bai 0.60 · backfill

Papers (14)

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training cs.CR · 2024 · author #26
  2. Discovering Language Model Behaviors with Model-Written Evaluations cs.CL · 2022 · author #53
  3. Constitutional AI: Harmlessness from AI Feedback cs.CL · 2022 · author #1
  4. Measuring Progress on Scalable Oversight for Large Language Models cs.HC · 2022 · author #43
  5. In-context Learning and Induction Heads cs.LG · 2022 · author #9
  6. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cs.CL · 2022 · author #5
  7. Language Models (Mostly) Know What They Know cs.CL · 2022 · author #17
  8. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models cs.CL · 2022 · author #444
  9. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cs.CL · 2022 · author #1
  10. A General Language Assistant as a Laboratory for Alignment cs.CL · 2021 · author #2
  11. Scattering Forms and the Positive Geometry of Kinematics, Color and the Worldsheet hep-th · 2017 · author #2
  12. Positive Geometries and Canonical Forms hep-th · 2017 · author #2
  13. The Amplituhedron and the One-loop Grassmannian Measure hep-th · 2015 · author #1
  14. The Amplituhedron from Momentum Twistor Diagrams hep-th · 2014 · author #1

Mentions

  • 1510.03553 #1 · backfill · confidence 0.70 Yuntao Bai
  • 1408.2459 #1 · backfill · confidence 0.70 Yuntao Bai
  • 2211.03540 #43 · arxiv_oai · confidence 0.70 Yuntao Bai

Frequent Coauthors