The letter

One letter a week, in your inbox.

The signal of the week, what shipped, what to try, and the editor's note. No tracking, no ads, nothing else.

We keep your address, your language and the date you joined, nothing else. Every letter has a one-click unsubscribe link that deletes the record.

← Accept All   Archive
AI Alignment ForumResearch

Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

September 15
Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring [1] , help us do better science on current models [2] , and augment certain forms o

Read at AI Alignment Forum ↗

Related

More from AI Alignment Forum on Accept All.