AI Alignment ForumResearch
Continual learning might make your blocking monitors nearly useless
Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a…
Read at AI Alignment Forum ↗More from AI Alignment Forum on Accept All
ResearchLatent reasoning architectures would undermine CoT, our strongest oversight tool AI Alignment Forum
ResearchWhy I'm scared of RL AI Alignment Forum
ResearchWorkspaceBench: Evaluating Interpretability Methods for the Global Workspace AI Alignment Forum

ResearchShallow Beliefs: Midtraining does not inoculate against EM from reward hacking AI Alignment Forum
ResearchOp-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI AI Alignment Forum
