AI Alignment ForumResearch
WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e.
Read at AI Alignment Forum ↗Related

ResearchCoT controllability evals seem very under-elicited AI Alignment Forum
ResearchProposal for tracking the effects of architecture on monitorability AI Alignment Forum

SiliconImproving Quantum Error Correction By Meshing Surface Code With IBM’s Heavy-Hex Architecture The Next Platform

GitHubGPT-6 Sol closes most of the alignment gap with Astra. It’s one-fifth the price. The New Stack

LabsEvaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore AWS Machine Learning
