Claude Opus 4.8 and The Curious "Footgun"
Opus 4.8 and 5.0 are great coding models with a confusing way of talking. A few thoughts on tone, "footguns", and why I went back to 4.6 for anything language-related.
I'm not a fan of Opus 4.8 or 5.0. They're incredible models for coding. But when I use it for explaining a topic, brainstorming, or digging into code, the tone and phrasing they use make it so confusing that it's too inefficient to actually use. Here's a list of weird phrases it uses VERY frequently, so I assume they were baked in at training:
- "footgun"
- "it's not a shape, it's a smell"
- …and plenty more in the same vein
Bear in mind it's using phrasing like this for technical tasks. "Footgun"? What the heck is a "footgun"? Have you ever heard a real person use that word? I haven't. And it talking about something, in this case a design choice in the code architecture, being a "smell" and not a "shape"? It's nonsense. For some reason the tone of 4.8 makes me think of the "girlboss" style of speaking. Maybe that's misplaced, but that's what first comes to my mind.
My background is in literature, so I've spent a fair bit of time reading and dissecting writing. I'd like to think I'm pretty good at picking out AI-written copy in LinkedIn posts or Reddit comments. It has a very distinct style and template: "it's not this… it's this," or "and what nobody saw coming… was this." Shame on anyone using AI to write your LinkedIn posts. You're better than that. But I digress.
I liked Opus 4.6 for its ability to communicate. I think it's a brilliant model for most language tasks. It has a very natural tone and a very precise understanding of English. Opus 4.7, however, was a disaster. It used emojis heavily, which Opus models never did in the past, and it lost the language intelligence that 4.6 had. My first impression was that they'd used ChatGPT output in the training data, which I assume is why 4.8 came out so quickly. It was essentially a hot patch.
I'd consider 4.8 and 5.0 (I'm ignoring 4.7 entirely) to be the two most disappointing model releases from Anthropic since I started using them. Claude Code and the coding harness still work great. Coding tasks, actually doing the work, are just fine with 4.8. But trying to understand the model's explanation of its actions or ideas is so convoluted that at this point I've given up and switched back to 4.6 for chat-based tasks.
What's your experience been like?