AI Alignment ForumResearch
Why I'm scared of RL
Summary: First, I give several different angles on how I feel about reinforcement learning: Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries…
Read at AI Alignment Forum ↗Related

ResearchCoT controllability evals seem very under-elicited AI Alignment Forum
ResearchProposal for tracking the effects of architecture on monitorability AI Alignment Forum

AsiaDeepSeek details DSec sandbox infrastructure for agent training TechNode

SiliconImproving Quantum Error Correction By Meshing Surface Code With IBM’s Heavy-Hex Architecture The Next Platform

GitHubGPT-6 Sol closes most of the alignment gap with Astra. It’s one-fifth the price. The New Stack
