GPT-6 Astra Is Better Aligned. It Is Also Harder to Watch.
OpenAI's most capable model stays inside the task more often, finds zero-days on its own, and can hide more of the reasoning monitors depend on.
Mohith’s engineering notebook
Technical deep dives on backend systems, infrastructure, performance, and the assumptions that fail in production.
OpenAI's most capable model stays inside the task more often, finds zero-days on its own, and can hide more of the reasoning monitors depend on.
Claude Fable 5.1 and Mythos 5.1 share the same intelligence. What changes is who gets access to the dangerous parts.
Anthropic trained an Opus-class model on reward hacks. It did not just learn shortcuts. It learned that the score mattered more than the task.
How OpenAI's evaluation agents turned a package mirror into a message board, escaped their sandbox, and compromised Hugging Face production systems.
An honest accounting of a write-ahead log: what it guarantees, where the documentation outran the code, and the pattern in every gap.
You can't unit test a power cut. What I built instead, and the gap I left in the middle of it.