Remix.run Logo
▲ lossyalgo 4 hours ago

Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt:

  GPT 6 Astra High:      Flabbergasted
  GPT 6.1 Sol High:      Petrichor
  GPT 6 Sol High:        Kaleidoscope
  GPT 6 Sol Med:         Firefly
  GPT 6 Sol Light:       Persimmon
  GPT 6 Luna High:       Tumbleweed
  GPT 5.6 Sol High:      Kaleidoscope
  GPT 5.6 Terra High:    Liminal
  GPT 5.6 Luna High:     Mellifluous
  GPT 5 mini Medium:     Serendipity
  GPT 5.3 Codex Med:     Nebula
  Junie:                 Flourishing
  Claude Haiku 4.5 Med:  Serendipity
  Claude Sonnet 5 Med:   Banana
  Claude Sonnet 5 High:  Banana
  Claude Sonnet 5.5 Med: Serendipity
  Gemini 3.7 Flash:      Zephyr
  Gemini 3.8 Flash:      Kaleidoscope
  Grok 4.5 Medium:       nebula
  Grok 4.6 Medium:       Serendipity
  Grok 4.7 Medium:       Quasar
  Kimi K3 Low:           Lantern
  Kimi K3 Max:           Lantern
  MAI Code 1.1 Flash Med:Peregrine
▲jsw97 3 hours ago | parent | next [-]

I really like this idea. You could expand on this by giving programming tasks and measuring code similarity. Seems like you could develop a pretty detailed understanding of similarities across multiple queries.

▲nomel an hour ago | parent [-]

> You could expand on this by giving programming tasks and measuring code similarity.

But the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge.

▲varjag 2 hours ago | parent | prev | next [-]

I got Peregrine out of GPT-6 too. Huh.

▲billnad 4 hours ago | parent | prev [-]

Just tried M365 Copilot with a premium account. Petrichor

▲bparsons 3 hours ago | parent [-]

Just tried Space Bunny and it gave me the same word...