AI Safety Fundamentals: Alignment

Un pódcast de BlueDot Impact

Categorías:

83 Episodo

  1. Public by Default: How We Manage Information Visibility at Get on Board

    Publicado: 12/5/2024
  2. Writing, Briefly

    Publicado: 12/5/2024
  3. Being the (Pareto) Best in the World

    Publicado: 4/5/2024
  4. How to Succeed as an Early-Stage Researcher: The “Lean Startup” Approach

    Publicado: 23/4/2024
  5. Become a Person who Actually Does Things

    Publicado: 17/4/2024
  6. Planning a High-Impact Career: A Summary of Everything You Need to Know in 7 Points

    Publicado: 16/4/2024
  7. Working in AI Alignment

    Publicado: 14/4/2024
  8. Computing Power and the Governance of AI

    Publicado: 7/4/2024
  9. Emerging Processes for Frontier AI Safety

    Publicado: 7/4/2024
  10. Challenges in Evaluating AI Systems

    Publicado: 7/4/2024
  11. AI Control: Improving Safety Despite Intentional Subversion

    Publicado: 7/4/2024
  12. AI Watermarking Won’t Curb Disinformation

    Publicado: 7/4/2024
  13. Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small

    Publicado: 1/4/2024
  14. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

    Publicado: 31/3/2024
  15. Zoom In: An Introduction to Circuits

    Publicado: 31/3/2024
  16. Weak-To-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

    Publicado: 26/3/2024
  17. Can We Scale Human Feedback for Complex AI Tasks?

    Publicado: 26/3/2024
  18. Machine Learning for Humans: Supervised Learning

    Publicado: 13/5/2023
  19. Four Background Claims

    Publicado: 13/5/2023
  20. Biological Anchors: A Trick That Might Or Might Not Work

    Publicado: 13/5/2023

2 / 5

Listen to resources from the AI Safety Fundamentals: Alignment course!https://aisafetyfundamentals.com/alignment

Visit the podcast's native language site