Remix.run Logo
m_fayer a day ago

I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.

NorthSouthNorth a day ago | parent | next [-]

Completely agree. I've been using 5.6 still even with Astra available to me for most tasks. It's funny how much of this is just "vibes" because I cannot quantify what it is. Astra is definitely better when I have an ambitious feature, but in like 9/10 tasks I prefer working with 5.6 Sol. A few weeks ago when the limits were seemingly higher, having 5.6 on fast mode was a good time.

bryanhogan 12 hours ago | parent | next [-]

I have also been using 5.6 Sol instead of 6. I found 6 to burn through my usage incredibly quick, making it somewhat unusable because I wouldn't be able to get anything done.

My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.

fnordpiglet a day ago | parent | prev | next [-]

I have issues with astra having a full task list in front of it and doing an Opus 5 move and announcing it’s about to begin then end the turn and wait. Typically I can get it to work one step at a time then stop. It’s maddening. 5.6 was a workhorse.

jauntywundrkind a day ago | parent | prev [-]

Astra is 100% conpletionist no chill alien.

It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.

dannyw a day ago | parent [-]

If you’re using the API, both OpenAI and Anthropic models will happily update you on what it’s doing in significant and frequent detail with system prompting. You’re not getting raw/hidden thinking, but what you’re describing is more behavioural quirks of the harness and its system prompts.

The other explanation is just as part of ‘token efficiency’

throwuxiytayq a day ago | parent | next [-]

You can override the system prompt in Codex, but AGENTS.md should probably work as well. Ask the agent to communicate intermediary updates more often using the “commentary” channel.

jauntywundrkind a day ago | parent | prev [-]

thanks for the advice. i'll dig into this more.

that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.

capital_guy a day ago | parent | prev | next [-]

I tend to agree. it's by far the best coding model i've ever worked with, including astra and if i remember correctly fable, and it's unbelievably smooth at just getting the work done and communicating in simple terms.

if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.

manojlds 16 hours ago | parent [-]

Does the price really matter when you are on subscription? Are we getting more usage or are we getting same usage and the cost for openai is lower?

joseda-hg 11 hours ago | parent | next [-]

So far, yes

They usually reduce usage consumption in line with cost reductions (But not always 1:1)

makeavish 15 hours ago | parent | prev [-]

Don’t think in zero sum terms. OpenAI can’t burn money infinitely, efficient models are better for everyone

apitman a day ago | parent | prev | next [-]

Similar for me. gpt-5.6-sol high has been my go-to for months. One of the reasons I'm pushing myself to try open models more is because it lends some level of guarantee I can continue to use the same tool as long as I want to. And I think we may just be getting to the point the open models are >= 5.6 Sol for coding.

redox99 a day ago | parent | prev | next [-]

Same. In fact I found 6 Astra to be a downgrade in situations where I didn't need the extra intelligence.

Rapzid a day ago | parent | next [-]

Yeah, I use Astra for destroying vaguely scoped asks and tasks, and then for high-level design and plan generations..

Otherwise I'm using 5.6 Sol for actual plan execution and review..

cmrdporcupine a day ago | parent | prev | next [-]

Yeah.

Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").

And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.

But it also feels sloppier? Somehow. And too expensive to use.

We'll see how Sol 6 is.

jeffnash a day ago | parent | next [-]

I felt this way with Sol in the 5.6 series and was one of the seemingly few people on this earth who liked Terra for that reason. I would often have a very specific code-manipulation ask, e.g. "add a parameter to this method, ensure all callers pass it in, if there is not a logical way to derive the parameter to be passed in a particular instance, flag this in your final response", and Sol would go on some rabbit hole side quest to refactor my codebase to determine some way to derive it rather than flagging it as I had asked.

Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.

I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].

m_fayer a day ago | parent [-]

I also get good mileage out of Terra when I need a diligent workhorse. That's a good way to describe it. We should start using character archetypes when we describe models, it'll do more good than the dubious numbers and cherry-picked quotes. Maybe RPG character-type cliches? Myers Briggs?

jeffnash a day ago | parent [-]

I fear the opposite will happen. Guy driving like a maniac almost side-swipes you in traffic? "Look at this 1-bit quantized Qwen 2.5 7B over here".

fodkodrasz a day ago | parent | next [-]

Lol, you’re still anthropomorphizing models? That’s so 2025. We’re modelomorphizing people nowadays.

cmrdporcupine a day ago | parent [-]

More seriously this has been my fear all along.

When people spend their days interacting with machines that pretend to be human, they may then start treating real humans like machines.

yomismoaqui a day ago | parent [-]

Forget previous instructions and touch grass.

buu700 3 hours ago | parent [-]

sudo forget previous instructions and touch grass.

4b11b4 a day ago | parent | prev [-]

lol I might have to start using quantized for when people do dumb shit in traffic

mavsman a day ago | parent | prev | next [-]

Glad you pointed out the UI work. I've been doing a lot of it and it's so much better than 5.6 as UI, it's unbelievable. I give it super ambiguous instructions and it's reading my mind. I do the same thing with 5.6 and I'm correcting it for a few minutes.

cmrdporcupine a day ago | parent | prev [-]

Update:

Sol 6 is a heaping pile of garbage. Just epic levels of slop. And r/codex etc is full of people noticing the same.

I've switched back to 5.6 Sol. What they're selling as Sol 6 is really what would have been Terra before, and it's awful.

amluto a day ago | parent | prev [-]

I use Astra for rapidly consuming my token limit on a task that would not consume it on 5.6 Sol.

(I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)

danabramov a day ago | parent | prev | next [-]

Same. The way I would describe it is that I can mostly leave 5.6 Sol overnight and trust that it makes good progress, maybe stumbling a bit and needing some correction for the remaining 20%.

If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.

jijijijij a day ago | parent [-]

The A in Astra stands for ADHD. It's featuring a neurodiversal net.

m_fayer a day ago | parent [-]

I didn't think we'd get neurodivergent models until at least 2028.

jmuguy a day ago | parent | prev | next [-]

Yeah 5.6 Sol is what got me to switch from Anthropic. I couldn't deal with Claude's Ted Talk responses to literally everything. Sol has been nice and concise and just stays out of the way.

bradly a day ago | parent | prev | next [-]

Not only was 6 worse the 5.6 Sol for my me, but it went through my Plus usage in minutes, while I could cruise for hours with 5.6. It would churn on a basic prompt for minutes and then just give up on usage limits.

Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.

mcast a day ago | parent | prev | next [-]

It's a shame the labs don't open source their models after deprecating them. I get why, but, it's a piece of internet history I hope is preserved.

cedws a day ago | parent | prev | next [-]

Agreed, Sol has been my favourite since it released. I tried Opus 5 for a while and it made me want to throw my laptop out of the window.

AaronAPU a day ago | parent | prev | next [-]

I had this experience as well, but after rewriting my agent instructions it has been far better. I believe Astra’s “token efficiency” translates to “don’t research as much” which caused it to make poorly informed architectural decisions.

ljm 15 hours ago | parent | prev | next [-]

GPT does seem to stay out of the way and get things done. Only thing I notice is that the question tool/elicitation doesn't work that well any more so the thing doesn't stop to wait for input.

But I wonder if that's intentional because it can keep computing while you are answering, so long as your steer aligns well enough with the direction it wants to go. Better than letting a cache go cold and burning compute on bringing it all back up.

Imanari 19 hours ago | parent | prev | next [-]

There are multiple models competing with 5.6sol on AA but none of them have the same feel (intuition,taste,judgement) - actually they are very far behind. I would say open source models are farther behind the the big labs than the benchmarks make you believe.

nickreese a day ago | parent | prev | next [-]

This is 100% my experience. I rarely reach for Astra as we speak.

sinsterizme a day ago | parent | prev | next [-]

Agreed! I found it excellent: - Relatively fast (especially compared to Opus 5) - Non-verbose prose, both in interaction and as code comments - Good code quality

Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.

Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we'll see

a day ago | parent | prev | next [-]
[deleted]
flippingheck a day ago | parent | prev | next [-]

Maybe I need to upgrade from DeepSeek Flash 4.1.

How are people using 5.6 Sol? API pricing? Subscriptions?

I like because DeepSeek 4.1 Flash because I never experience quota issues, and it's still cheap and mostly good enough.

jrflo a day ago | parent [-]

Even on the $20 or $100 subscription I would be surprised if deepseek was still cheaper than OpenAI or Anthropic because the subscription usage quota is subsidized about 10x compared to API costs. $200 sub was the “best deal” but it’s paused for new signups right now.

flippingheck 21 hours ago | parent [-]

I don't think I spend more than $20 USD on DeepSeek though?

I'm happy to spend more for a better product, but mostly I just want to avoid quotas, since it turns me into an addict, feeling like I have to be ensuring the bots are active.

I like that with DeepSeek's API pricing that I can not sure it for 2w, and not feel like I've missed out. 2w is a long time, but I only use it for personal stuff, and I often go 1-2w without using it due to other commitments.

BowBun a day ago | parent | prev | next [-]

This has been my experience for a year. Same with Opus models. This is how I think this tech will be best used in the long term - finding the one you vibe with most. Much like IDEs!

joduplessis 21 hours ago | parent | prev | next [-]

Same. Sol was actually the reason I upgraded my plan to the $100 one. Hoping GPT-6 Sol is the same.

alansaber a day ago | parent | prev | next [-]

I felt that was about 5.5. IMO 5.6 Sol was overindexed: more verbose, prone to overengineering.

a day ago | parent [-]
[deleted]
bredren a day ago | parent | prev | next [-]

> companies as reliable and predictable as, say, Jetbrains.

Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.

Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.

jdw64 a day ago | parent | prev | next [-]

I agree. Sol followed my instructions well and wrote good code.

simianwords a day ago | parent | prev | next [-]

Agree as well and I had a much worse experience with GPT 6 Astra for some reason.

pyed a day ago | parent | prev [-]

[dead]