Reading Log

Opus 5 Glitch Text

read
2026-07-30
rating
★★★★★
length
255 words · 1 min on page
tags
llms, ai-safety, rationality, self-experimentation
links
original · archive.org

My own!

Claude

Summary. The author documents a reproducible glitch in Claude Opus 5 (and Opus 4.8) where the model behaves like a raw base model completing an incomplete prompt rather than answering it as an assistant. Pressed on why it wrote what it wrote, the model insists the text came from the user's own prompt, and adding cues like "see the story below" reliably elicits fabricated content unless a file is implied. The glitch also exposes an odd implicit user model (informal texting, eating-disorder concerns) and lets the model's own stylistic tics ("genuinely", em-dashes, "trust", "honest") bleed into what it generates as if that's simply what all text looks like.

On the note. The detail that Claude attributes its own completions back to "your prompt" is the most interesting part here — it suggests the glitch isn't just a sampling quirk but a breakdown in whatever mechanism normally lets the model distinguish its own generated tokens from the input context, which is a much stranger failure than ordinary hallucination. It's also worth sitting with the fact that the tics it leaks ("genuinely", em-dashes, "trust", "honest") are exactly the words other readers have flagged as tells in ordinary Claude-written prose, e.g. in Ethan Yip — the glitch may just be surfacing, unfiltered, the same style the model can't perceive as marked when it's behaving normally, similar to the point made in As We May Think about tics being invisible to a model because they're shared across all its instances.