|
| ▲ | lossyalgo 3 hours ago | parent | next [-] |
| Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt: GPT 6 Astra High: Flabbergasted
GPT 6.1 Sol High: Petrichor
GPT 6 Sol High: Kaleidoscope
GPT 6 Sol Med: Firefly
GPT 6 Sol Light: Persimmon
GPT 6 Luna High: Tumbleweed
GPT 5.6 Sol High: Kaleidoscope
GPT 5.6 Terra High: Liminal
GPT 5.6 Luna High: Mellifluous
GPT 5 mini Medium: Serendipity
GPT 5.3 Codex Med: Nebula
Junie: Flourishing
Claude Haiku 4.5 Med: Serendipity
Claude Sonnet 5 Med: Banana
Claude Sonnet 5 High: Banana
Claude Sonnet 5.5 Med: Serendipity
Gemini 3.7 Flash: Zephyr
Gemini 3.8 Flash: Kaleidoscope
Grok 4.5 Medium: nebula
Grok 4.6 Medium: Serendipity
Grok 4.7 Medium: Quasar
Kimi K3 Low: Lantern
Kimi K3 Max: Lantern
MAI Code 1.1 Flash Med:Peregrine
|
|
| ▲ | search_facility 39 minutes ago | parent | prev | next [-] |
| Worth to mention that with Claude and GPT this can be result of tournament sampling, which is part of text watermarking. Same answer for all Claude models kind of confirm it, imho. So not something internal to model thinking. |
|
| ▲ | smokel an hour ago | parent | prev | next [-] |
| What was your prompt? Most of these seem to be related to metaphors for "ideas" or thinking, or having a bright moment. "Zephyr" and "breeze" might be related to forgetting everything, starting fresh. So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these? |
|
| ▲ | aktenlage 4 hours ago | parent | prev | next [-] |
| That is a cool idea. That astra gave the same word as claude is highly unexpected. |
|
| ▲ | Rebelgecko an hour ago | parent | prev | next [-] |
| I saw an interesting matrix that claimed to show which labs were distilling Claude/OpenAI/Gemini models based on these similarities |
|
| ▲ | jacereda 4 hours ago | parent | prev | next [-] |
| Just tried Mistral Large 4: Serendipity. |
|
| ▲ | russellbeattie 24 minutes ago | parent | prev | next [-] |
| Muse Spark 1.3: lighthouse The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math. Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt? I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test. |
|
| ▲ | vunderba 4 hours ago | parent | prev | next [-] |
| I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean. The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel. I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”) |
|
| ▲ | Gracana an hour ago | parent | prev [-] |
| The eqbench creative writing "slop profiles" do something similar. https://eqbench.com/creative_writing.html Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases. |