Remix.run Logo
badatnames 8 hours ago

Using codex every day, in spite of which, I hope some day providers will just start naming their offerings small/medium/large, a bit like we eventually started doing in software testing. Trying to remember what Sol is or why it's better than the other thing is more cognitive effort than I can muster at this point. And that's a sure sign of commoditisation in itself

phoghed 8 hours ago | parent | next [-]

Sun, Earth, Moon — it’s basically L/M/S like you want but a little less boring.

Why is large better than medium to the average end user of ChatGPT though?

I don’t think there’s a way to name these things that will satisfy everyone.

mastercheif 8 hours ago | parent [-]

The naming schema actually tripped me up for a week or so.

My brain's initial conception of the concepts was earth-relative, so I mapped it as:

Sol = big, it's the sun Luna = medium, in-between sun and earth, space Terra = small, terrestrial

inexcf 7 hours ago | parent | next [-]

Pretty weird when the moon is as much between earth and sun as the earth is between the moon and the sun.

kgwgk 7 hours ago | parent [-]

And when it’s in between we cannot even see it (unless it’s exactly in line).

ModernMech 6 hours ago | parent | prev [-]

It tripped me up because I was going by distance. I thought Terra was the base model and Luna was the mid model because it’s further away.

6thbit 41 minutes ago | parent [-]

Same here. Thought Luna was for “moonshots” and sol for.. sunshots? While keeping Terra earthly.

But alas

msdz 8 hours ago | parent | prev | next [-]

Tinfoil hat time: They saw everyone referring to Mythos, and later Fable, as the new “good” models when Anthropic released those, distinguishable from the “regular” Claude (or other companies’ models) for everyone, and didn’t have that distinction for the GPT model family. That’s why the planetary names were introduced.

usef- 2 hours ago | parent [-]

A simpler explanation is that it's just a better naming system.

Calling something "small" might make it sound inferior to competitors. And S/M/L gets awkward as soon as you have more than three sizes.

This naming system can get near-infinitely bigger or smaller.

ComputerGuru 8 hours ago | parent | prev [-]

I think model naming has been atrocious in general, in part because newer "lite" models surpass the capabilities of previous "pro" models (case-in-point: Gemini Flash which now surpasses the capabilities of the latest Gemini Pro, with a newer Flash Lite vying somewhat unsuccessfully for the old Flash price/positioning), but gpt 5.6's Sol/Terra/Luna split is really not bad at all - probably easier to understand than Starbucks' cup sizing!

The problem becomes when you add in the adjustable reasoning efforts and you end up with {model, reasoning_effort} combinations that end up completely obviating particular model classes altogether for at least some percentage of queries; e.g. with GPT 5.6 the price/performance Pareto frontier is dominated by permutations of either Luna and Sol, with Terra nowhere to be seen (but then if you need "large model smells" that aren't captured by your benchmark you can't even rely on this, as a model like Luna simply isn't capable of encoding sufficient world knowledge in its weights to perform certain tasks at any reasoning level but you might be able to get away with Terra on low reasoning, but no one seems to be covering this for some reason).

s3p 8 hours ago | parent [-]

Yes but with gemini specifically they said that pro was still in training. And the comparison isn't really atrocious unless Gemini 3.5 Pro is worse than Gemini 3.5 flash