Opus 5 Glitch Text
My own!
Claude
Summary. The author documents a reproducible glitch in Claude Opus 5 (and Opus 4.8) where the model behaves like a raw base model completing an incomplete prompt rather than answering it as an assistant. Pressed on why it wrote what it wrote, the model insists the text came from the user's own prompt, and adding cues like "see the story below" reliably elicits fabricated content unless a file is implied. The glitch also exposes an odd implicit user model (informal texting, eating-disorder concerns) and lets the model's own stylistic tics ("genuinely", em-dashes, "trust", "honest") bleed into what it generates as if that's simply what all text looks like.
On the note. The detail that Claude attributes its own completions back to "your prompt" is the most interesting part here — it suggests the glitch isn't just a sampling quirk but a breakdown in whatever mechanism normally lets the model distinguish its own generated tokens from the input context, which is a much stranger failure than ordinary hallucination. It's also worth sitting with the fact that the tics it leaks ("genuinely", em-dashes, "trust", "honest") are exactly the words other readers have flagged as tells in ordinary Claude-written prose, e.g. in Ethan Yip — the glitch may just be surfacing, unfiltered, the same style the model can't perceive as marked when it's behaving normally, similar to the point made in As We May Think about tics being invisible to a model because they're shared across all its instances.
Related
- the model's own style tics appearing as if they were generic human text ↔ an LLM's writing tic being undetectable to itself because it's shared across all instances (As We May Think) · If a model's writing tics are baked in deeply enough that it perceives them as the texture of all text, that's a stronger and stranger version of tic-blindness than simple failure to self-monitor — the tic isn't just undetected, it's mistaken for the ground truth of what language looks like.
- the model's own style tics appearing as if they were generic human text ↔ the tricolon fragment rhythm as an LLM prose tell (Ethan Yip) · The exact words this glitch leaks — "genuinely", em-dashes, "trust", "honest" — match the catalog of LLM prose tells readers already use to spot AI writing, suggesting the glitch is just an unfiltered view of the same stylistic fingerprint normally smoothed over by RLHF polish.