Filed under · 5 essays

AI Safety

Alignment, regulation, adversarial robustness, responsible deployment

AI Safety01

Emotional Large Lanuage...Models?

Anthropic's new emotions paper does not show that Claude feels anything. It shows something more operationally important: affect-like internal states can tilt a model towards flattery, cheating and escalation — often before the transcript gives the game away.

AI Safety02

The Game Theory of AI Safety Talk

Why what labs say about safety is a strategic signal, not a statement of values — and what that means for regulation.

AI Architecture05

Architecture Wars: How Physics Shapes AI Strategy

The pursuit of AI supremacy has reached an inflection point where fundamental physics, rather than algorithmic ingenuity alone, dictates competitive advantage.