Open Internet by MindsNet
AI Alignment: Superhuman AI Safety
How do we ensure AI systems with superhuman capabilities reliably pursue goals beneficial to humanity? This open research problem encompasses reward modelling, interpretability, scalable oversight, and robustness. Anthropic, DeepMind, OpenAI, and MIRI all work on this.
Computing & Technology, Computer Science, Machine Learning