Persona Vectors Reveal What Open-Weight LLMs Express, Suppress, and Resist
A new arXiv paper uses persona vectors to audit open-weight language models beyond what prompt-based testing can reveal. The study maps which traits models express by default, which remain latent but steerable, and which resist standard extraction.
Read more