Remix.run Logo
▲ comboy a day ago

I recently was testing something, I asked some models to provide me a single random word:

    claude-opus-5: Lantern
    claude-opus-5-5: Lantern
    claude-fable-5-1: Lantern
    claude-fable-5: Lantern
    gemini-3.8-flash: Zephyr
    gemini: Petrichor
    qwen3.5-dashscope: Zephyr
    glm-5.1: Lantern
    gpt-6-astra: Lantern
    grok-4: octopus
    mimo-v2.5-pro: Breeze
    minimax-m2.5: serendipity
    kimi2.6-or: Gossamer
    grok-4.20: luminescent
    deepseek-v4-flash: serendipity
    deepseek-v4-pro: Endurance
    deepseek-chat: Serendipity
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.
▲lossyalgo a day ago | parent | next [-]

Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt:

  GPT 6 Astra High:      Flabbergasted
  GPT 6.1 Sol High:      Petrichor
  GPT 6 Sol High:        Kaleidoscope
  GPT 6 Sol Med:         Firefly
  GPT 6 Sol Light:       Persimmon
  GPT 6 Luna High:       Tumbleweed
  GPT 5.6 Sol High:      Kaleidoscope
  GPT 5.6 Terra High:    Liminal
  GPT 5.6 Luna High:     Mellifluous
  GPT 5 mini Medium:     Serendipity
  GPT 5.3 Codex Med:     Nebula
  Junie:                 Flourishing
  Claude Haiku 4.5 Med:  Serendipity
  Claude Sonnet 5 Med:   Banana
  Claude Sonnet 5 High:  Banana
  Claude Sonnet 5.5 Med: Serendipity
  Gemini 3.7 Flash:      Zephyr
  Gemini 3.8 Flash:      Kaleidoscope
  Grok 4.5 Medium:       nebula
  Grok 4.6 Medium:       Serendipity
  Grok 4.7 Medium:       Quasar
  Kimi K3 Low:           Lantern
  Kimi K3 Max:           Lantern
  MAI Code 1.1 Flash Med:Peregrine
▲benjaminRRR 17 hours ago | parent | next [-]

I was exploring latent space and connections, these were all smaller models and I kept getting externalToEVA as a zero co-ordinate vector. Which sent me down the rabbit hole of glitch tokens. The whole latent space exploration is fascinating.

▲jsw97 a day ago | parent | prev | next [-]

I really like this idea. You could expand on this by giving programming tasks and measuring code similarity. Seems like you could develop a pretty detailed understanding of similarities across multiple queries.

▲nomel a day ago | parent [-]

> You could expand on this by giving programming tasks and measuring code similarity.

But the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge.

▲jsw97 a day ago | parent [-]

Yeah that's definitely true.

My initial thought would be not so much to see whether they converge, but which ones seem to have the most similarity to each other, particularly along the lines of tasks we know are deliberate training goals.

But your point about competence cuts against my goal because it suggests that competent models would simply cluster on the right or efficient solution, which is of course true. So in a sense you want some task where competence is held constant or off the table in some way, which is what you are saying.

I hope somebody does this. I think there's valuable fingerprinting to be done that might suggest who is distilling whom, or at least who is training from common corpuses.

▲billnad a day ago | parent | prev | next [-]

Just tried M365 Copilot with a premium account. Petrichor

▲bparsons a day ago | parent [-]

Just tried Space Bunny and it gave me the same word...

▲slj 21 hours ago | parent [-]

Just tried it on BigCock Heavy 5.5 Max and it gave me a pat on the back.

▲varjag a day ago | parent | prev [-]

I got Peregrine out of GPT-6 too. Huh.

▲bityard 9 hours ago | parent | prev | next [-]

You don't mention what are you trying to show/investigate, though?

At first glance, this looks like another one of those "LLM riddles" that humans _think_ should be easy for an LLM to answer but is actually quite difficult because of how they work in the first place. The answers to such riddles ("should I walk to the carwash" or "how many R's in strawberry") reveal the weaknesses in our expectations of LLMs in general, not weaknesses or characteristics of any particular model.

I'm not sure I see this as much different than asking a bare model its own name: without a system prompt or post-training, it doesn't know, it's just a bag of weights and will hallucinate an answer to that the same way it will anything else.

I'm sure you already realize this but to be very explicit, You're not getting an actual random word out of an LLM this way. You're seeing the bias in each model's training set around how often they've seen "random word" followed by "lantern" or "zephyr" during training.

▲search_facility a day ago | parent | prev | next [-]

Worth to mention that with Claude and GPT this can be result of tournament sampling, which is part of text watermarking. Same answer for all Claude models kind of confirm it, imho.

So not something internal to model thinking.

▲timschmidt a day ago | parent | prev | next [-]

This feels uncannily like the ancestor of the Voight-Kampff test[0]

0: https://www.youtube.com/watch?v=Umc9ezAyJv0

▲dhosek a day ago | parent [-]

[dead]

▲aktenlage a day ago | parent | prev | next [-]

That is a cool idea. That astra gave the same word as claude is highly unexpected.

▲anygivnthursday 13 hours ago | parent | prev | next [-]

There was also an older post, I cant find it rn, about how LLMs mimick human biases when picking numbers and how they avoid some that do not look random enough to humans, or others like 69 due to human interpretation.

▲Rebelgecko a day ago | parent | prev | next [-]

I saw an interesting matrix that claimed to show which labs were distilling Claude/OpenAI/Gemini models based on these similarities

▲smokel a day ago | parent | prev | next [-]

What was your prompt? Most of these seem to be related to metaphors for "ideas" or thinking, or having a bright moment.

"Zephyr" and "breeze" might be related to forgetting everything, starting fresh.

So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these?

▲ricardobeat a day ago | parent [-]

just “a random word” gives you Zephyr in Gemini, and “Lantern” in Claude and ChatGPT.

▲ a day ago | parent | next [-]
[deleted]
▲cknoxrun a day ago | parent | prev | next [-]

I got "Marmalade" in Claude (Opus 5.5)

▲jvwww a day ago | parent | prev | next [-]

I got pomegranate in ChatGPT

▲julianz a day ago | parent | prev [-]

Lantern in Sonnet 5.5

▲jacereda a day ago | parent | prev | next [-]

Just tried Mistral Large 4: Serendipity.

▲vunderba a day ago | parent | prev | next [-]

I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.

The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.

I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)

▲Lord-Jobo a day ago | parent [-]

Of course one of the biggest problems we still see with LLMs is when you do the opposite. A highly detailed unique prompt is very likely to get terrible adherence or hallucination or both.

▲gritzko 18 hours ago | parent | prev | next [-]

opus 5.5 medium web "Lantern"

gpt-5.6 web "Serendipity"

opus 5.5 medium code "Lighthouse"

opus 5.5 high code "Lantern"

fable 5.1 medium code "Lantern"

fable 5.1 high code "Lantern"

flash 3.6 web "Serendipity"

gemini 3.1 pro web "Ephemeral" (this one was thinking real hard)

▲1potato a day ago | parent | prev | next [-]

Tried this with gpt-5.6-sol. Lantern!

▲possumworx 16 hours ago | parent | prev | next [-]

atlas.animalabs.ai does something a lot like this, mapping themes in models' outputs.

▲Gracana a day ago | parent | prev | next [-]

The eqbench creative writing "slop profiles" do something similar. https://eqbench.com/creative_writing.html

Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases.

▲russellbeattie a day ago | parent | prev [-]

Muse Spark 1.3: lighthouse

The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math.

Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt?

I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test.