“Thoughts on the persona selection model” by Sam Marks
As AIs have become more "RLVR-brained," there's been some commentary on what this means for the persona selection model (PSM). This post presents some loose thoughts on that topic. A rough summary of my opinions: PSM is over-applied. That is, it is common to argue that PSM has takeaways that don't actually follow from PSM (e.g. "PSM => AIs will not seek reward" or "PSM => AI takeover risk is low"). I don't think we've observed strong evidence that "lots of RLVR breaks PSM." (TBC, there are decent reasons to expect this a priori; I just don't think recent empirical evidence has been much of an update.) My main update is that personas—insofar as they're a good model in the first place—seem less broad and more conditionalized than I expected (nostalgebraist, 2026; Betley, 2026). As a reminder, PSM roughly states that during pre-training LLMs learn to simulate diverse (human-like) personas, and post-training elicits a particular "Assistant" persona assembled from this repertoire. Then some key questions are: Is anything like this true or useful? Are personas ever a good way to reason about AI behavior, and is "selecting over personas" something that happens during post-training? What are [...] --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/csRby7mZgjL5jCoLL/thoughts-on-the-persona-selection-model --- Narrated by TYPE III AUDIO.