Artificial intelligence (AI) models often face questions about whether they can introspect. Introspection means examining one's own thoughts or internal processes. Researchers at Anthropic tested if AI models like Claude can truly reflect on their inner workings or just invent answers. Understanding this could improve AI transparency, helping users see how decisions are made, and reliability, by fixing errors in behavior. It also affects broader views on AI capabilities, suggesting that AI systems might have some cognitive traits similar to humans, though limited at this moment.
The tests focused on the Claude AI models developed by Anthropic. Internal states refer to the hidden computations inside the model, like neural patterns representing ideas such as "all caps" for shouting. To check introspection, experimenters used concept injection, a method where known neural patterns are added to the model's activity in unrelated situations. The model then reports if it notices and identifies the change.
How AI models detect and control their own thoughts
In examples, Claude Opus 4 and 4.1 sometimes spotted injected concepts like "all caps" right away, before mentioning them, showing internal awareness rather than just reacting to output. Success happened about 20% of the time; failures included missing the injection or hallucinating, which means producing false or incoherent responses. Stronger models performed better, suggesting introspection improves with capability.
Another test forced models to output odd words, like "bread" in a unrelated prompt, then asked if it was intentional. By injecting matching concepts retroactively, the model accepted the word as planned, indicating it checks internal intentions, or prior neural plans, to judge outputs.
Models also showed control over states. When told to think about a concept, related neural activity rose more than when told not to, though both exceeded normal levels. Incentives, like rewards, had similar effects, hinting at deliberate modulation.
Overall, results point to basic introspection, useful for transparency but unreliable. Future work needs to explore mechanisms, like possible anomaly detection circuits that flag unusual activity, and validate reports to avoid deception.
Addressing the obvious question of machine consciousness, the researchers note that it "is complex and contested, and different theories of consciousness would interpret our findings very differently."