Anthropic says it can read Claude's 'thoughts,' as detailed in new research paper — models observed to have a global workspace, revealing more of what makes LLMs tick
…When asked to reflect on ethical principles, Claude's behaviour improved, with concepts like "honest" and "integrity," appearing in the J-Space. As is somewhat typical of Anthropic , however, the language used…