Ergonomics Over Benchmarks

There’s something off with the output from OpenAI’s Astra / Sol — the best way I can describe it is that the ergonomics aren’t right. This has made me realize that AI tool retention will come from process aesthetics and ergonomics, not capabilities.

Stay with me here for a bit.

I bit the bullet and gave OpenAI some money last week, after getting increasingly frustrated at what seems to be output rot in Opus.

Opus performance has been continually degrading over the last weeks — both in output speed and general quality.

Not sure if they are screwing around with feature flags, trying to cram more requests into the same hardware, or what; but Anthropic might just snatch defeat from the jaws of victory.

— Ricardo J. Méndez (@ricardo.bsky.social) September 22, 2026 at 9:56 AM

Astra performs well as long as I throw it at a project like Winesky, which has its own established structure and patterns, and have it address specific issues in it. It has also performed great at code review, finding obscure potential security issues that had escaped both Fable and myself.

But when I had it develop a Python data harness from scratch, to encode a first pass of analysis criteria, the output was… almost brutalist. And not in a clean, aesthetically pleasing way.

It was first evident in the usage documentation it generated: it was bare-bones, missed basic things, markdown formatting was inconsistent (sometimes using full lines for soft-wrapping, sometimes hard breaks), and was more of a fact dump than something meant to be read by a person.

This extended to the code itself: hard to review, somewhat hodge-podge-ish, with final output that took effort to scan — even if it did get the job done.

Then, looking back at the process, I realized the questions it asked weren’t bad, but nowhere near the level of attention to detail that Opus shows.

For example:

  • “I’m missing this mint address, paste it here”; or
  • “Describe what you mean by test and hold”

When I ran the exact same prompt through Opus 5.5 to compare, the equivalent questions were:

  • “I found these two mints using [service documented on the repository] — is it one of them or a different one?”; and
  • “I can come up with these two definitions of test-and-hold, do you want to use one, or enter your own?”

Notice that Opus picked up on the fact that I had previously used a service, included an access key in the .env, and documented it in the repository’s README — Astra either didn’t bother to see what was available, or didn’t consider it relevant.

It also felt like Opus stopped to ask for feedback and guidance at more sensible points (again, working from the same prompt). After those two questions, Astra plowed on through with what I can best describe as attempting to one-shot the task (only stopping to ask me to run the code and eyeball if the data made sense) and then vomited a bunch of assumptions at the end.

Astra is faster, but the speed is the only part of the ergonomics that feels right.

OpenAI’s models act like extremely capable code monkeys with no interest in whether their code will ever have to be extended or maintained by a human, whereas working with Opus feels closer to pair programming.

Yes, it’s possible that given enough time I could figure out a way to coax Astra into asking better questions and behaving more like Opus, but for 95% of users, the defaults are the product.

Given that models have gotten such a leap in capabilities in the last 9 months, to the point where we are approaching parity in most areas, user retention will not come from raw code quality.

That leaves two ways to hold on to users: cater to corporate buyers (like Microsoft — OpenAI’s largest backer) or provide superior developer process aesthetics and ergonomics (Apple-style).

Now, if Anthropic can stop bollocksing up whatever it is they are bollocksing up, that’d be great. They currently match my preferred approach, but there’s only so much performance decay one can put up with before the ergonomics are broken beyond repair.