astra has been unexpectedly helpful
maybe we should pace the frontier more
I have not historically been one to be spectacularly impressed by new model releases. When I first started seriously using Claude Code a few months11 It’s only been what, half a year? Jeez. Like, I guess I know it must have been less than a year from the fact that my blog last November didn’t have any little fancy embeds I coded with AI assistance, and so far this month several posts have had those (or linked to a project), but man. back, I was primarily using Opus 4.6. And I haven’t actually tried giving Opus 4.6 a coding project since upgrading, but I wouldn’t expect to find them all that much more frustrating to work with than more recent Opuses or Fable. Somewhat less efficient on the margins, certainly—I regularly notice Opus 5 making minor mistakes that Fable generally wouldn’t, and I certainly don’t expect that Opus 4.6 would be any better about this—but still more or less capable of most of the same things, at least within the limits of what I’m generally trying to do?
I’m sure my working experience has appreciably changed over the past several months. I just didn’t personally notice any particularly glaring step changes, at any point. I used Fable very heavily when it first came out, and was very excited to make the most of the Fable access I had back when it wasn’t clear how long Fable would be on subscription plans, but I’m not sure if I was really being particularly more productive than I was using Opus 4.8 or whatever. Some part of that certainly could’ve been me mostly doing relatively unproductive things with Fable at that time, or other factors like that, but nevertheless22 I think this is maybe part of why things like the model disillusionment pattern lots of people seem to go through always seemed a little weird to me. I do not typically see a new model, get super impressed, and then eventually conclude it isn’t all that crazy. I just… appreciate the marginal improvements?.
GPT-6 Astra bucked that trend a little, for me.
Aprii *️⃣@ApriiSR@TheZvi i've been using astra for about a day and it has already figured out how to resolve a math question i'd been working on with fable for like a week
The particular concrete thing there: Astra proposed a particular computation to run which ended up refuting the existence of the $9 \times 9 \times 9$ fully diagonalized Latin cube, the first one to be shown to not exist outside the corner bound. Roughly, we enumerated the collections of positions that a single symbol might take in a diagonalized Latin $9 \times 9 \times 9$, and determined that there were at most five or so compatible such collections, and to even get that high you had to choose a rather specific set of them. It is unclear to me how soon I would’ve tried that research direction relying only on Fable and Sol — it was a very recognizable sort of idea, the positions of a particular symbol within the $8 \times 8 \times 8$ is the sort of thing I remember paying attention to back in the day, I might’ve well come up with it as an avenue worth trying after giving the problem a serious think, but it would probably have been after many more days of dead ends. Astra figured out that it was promising almost immediately upon being presented the question.
If that were the only thing, it would be pretty impressive to me, but still only be one anecdote. But I’ve also found Astra to be a substantial improvement on another problem I’ve been working on. Over the past couple months, I’ve been trying to port Super Mario 63 from Flash to Godot. Now, Fable, Sol, and Opus had very little issue basically putting together the fundamentals of such a thing. The basic idea is not that difficult an ask, we had most of the movement system and stuff down within a few days. The issue is that I am incredibly picky about these things; I want TASes to sync. This means that minor edge cases in Flash’s collision detection algorithms need to be modeled correctly. There are so many fiddly things there. The behavior substantially depends upon things like what quality you have selected and how zoomed in the camera is. You need to model things like Flash’s rounding and rasterization correctly.
It did not seem to me like it ought to be all that difficult in principle to just reverse engineer most of the relevant details from looking at .swf files and Flash player binaries and so on! But for some reason this did not tend to quickly resolve all of our issues. I am not totally sure why not — possibly it was being deferred to Opus (or even Sonnet) sub-agents when it really needed a frontier model? But I think we asked Sol a few times, and Sol certainly didn’t resolve all the issues immediately.
But in my experience over the past few days, if I pose Astra a question about how Flash works and have them look at the binaries and so on, they seem to simply figure out the right answer as quickly as I had initially imagined being possible. For the past day, I have basically been pointing Astra at the next place the No Major Skips TAS desyncs, asking it to resolve the issue, and getting a fix in twenty minutes to an hour. I sort of imagine that suboptimal Claude workflow might have been creating substantial friction, but still: progress has been incredibly smooth compared to what it used to be, and the sorts of things that would have a decent chance at stumping Claude for an entire day seem to be routinely manageable.
I definitely don’t expect to switch fully to OpenAI models or anything like that, and I’m not particularly reading too much into this one set of experiences. But in terms of first impressions on me in particular, Astra has seemed like much more of a serious concrete improvement over previous models than most new releases are.
I am not sure this is a good thing
If I zoom in on only my own life — very helpful for all my random projects! That’s definitely lovely and I am very excited. But I think, on a broader level, there is quite a lot to be said for pushing the frontier forward rather gradually. I don’t think this is a very controversial position, see the calls to Pace the Frontier. As much as I may be excited to have a new super effective work partner, I do not think “how helpful does April find this model” should be a primary input in how quickly to take these things.
Now, will people doing important medical research or whatever find Astra to be helpful? Quite possibly. I think there are some pretty good reasons to try to get however much AI we can safely get, and even reasons to not do it more slowly than required. But like, I dunno man. Once we get into the seventh model or eighth model generation we might be kind of pushing it. Even if it’s not quite worthwhile to actually stop further AI development just yet, it would be nice if each new step further out into the frontier were a relatively small one!
It’s only been what, half a year? Jeez. Like, I guess I know it must have been less than a year from the fact that my blog last November didn’t have any little fancy embeds I coded with AI assistance, and so far this month several posts have had those (or linked to a project), but man.
↩I think this is maybe part of why things like the model disillusionment pattern lots of people seem to go through always seemed a little weird to me. I do not typically see a new model, get super impressed, and then eventually conclude it isn’t all that crazy. I just… appreciate the marginal improvements?
↩
