Anthropic Gives Claude a Tool for Reading Its Own Thoughts
Anthropic's Natural Language Autoencoder turns Claude's internal activations into text, revealing when the model seems to know it is being evaluated.
1 verified stories covering AI interpretability, product updates and industry developments.