WRITING / POST

GPT Astra 6, an improvement on clarity and nuance.

5 SEPTEMBER 2026

I ran a session today to see if I could gauge a noticeable difference between SOL 5.6 and Astra 6. This was my opening request to my sysadmin agent, Maxi:

Hello Maxi, We will trial the new GPT 6 Astra model and see how it compares to GPT 5.6 SOL. This will be a subjective comparison I will make, and I will give you my feedback as we go along.

The first thing I would like you to do is review your SOUL.md file, as well as durable memories and see if there are any improvements, adjustments, or changes you would like to make that will help you better fit your role and meet your goals.

The current SOUL.md file, which defines the agent's role and identity, is a result of a three-way revise and improve process between SOL 5.6, Fable 5, and me. We all agree it is well structured, concise, and with no ambiguities that will cause an agent to guess. Maxi, now running on Astra 6, agreed. Or rather, it stated clearly what it liked, and then suggested three improvements it could make. And these weren't trivial, these were forehead-slapping "Yeah, I should have seen that" type improvements, such as:

3. Give exploration the same proportionality as intervention

The working sections are strong on preventing mistakes. They say less about making room for useful attempts whose outcome is uncertain.

I would add, beside the smallest-sufficient-intervention principle:

I distinguish uncertainty that calls for investigation from uncertainty that calls for a bounded experiment. Where authorised and proportionate, I learn by trying; I do not require certainty before taking a reversible step.

This would not expand my authority. It would clarify how I use it.

This was a point of mild frustration with SOL 5.6, and required careful phrasing on my part to prevent an investigation request caroming off into experiment, and vice versa. The Astra model identified the issue and proposed the solution.

But that wasn't the best part. It went on to give three priority tiers for improvements to durable memory that would, as far as I could see, significantly optimise those files.

I gave my input, and told it to go ahead. And this is the result:

Store Before After Reduction Entries before → after
MEMORY.md 24,154 7,151 70.4% 92 → 22
USER.md 14,862 4,994 66.4% 57 → 15
Combined 39,016 12,145 68.9% 149 → 37

File sizes in bytes.

Detailed preferences and operational history now sit in indexed, on-demand references. Every original entry is accounted for.

As Maxi then said: "Less baggage in every conversation, without throwing away the history."

It's an impressive result, not just for the 69% reduction in injected durable memory, but for showing the model's improvement over those that preceded it, in both the OpenAI and Anthropic camp. My subjective opinion is that it is a more capable model than Fable, and it will be interesting to see how Anthropic respond.