Anthropic's new emotions paper does not show that Claude feels anything. It shows something more operationally important: affect-like internal states can tilt a model towards flattery, cheating and escalation — often before the transcript gives the game away.
In conventional security, hardening a system makes it harder to attack. You patch vulnerabilities, reduce attack surface, and defence moves in lockstep with robustness. AI alignment breaks this assumption.
The pursuit of AI supremacy has reached an inflection point where fundamental physics, rather than algorithmic ingenuity alone, dictates competitive advantage.