Remix.run Logo
bhouston 2 hours ago

> This is science fiction, these models don't have access to their own weights

A bet a worm could pull along a 1GB file with weights in it and run it on a compromised machine, but luckily for us for now, 1GB isn't really enough to be really smart, yet.

himata4113 2 hours ago | parent [-]

We already have a 1gb model that is as capable as it will ever be, there's a proven ceiling that cannot be passed. For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.

Philpax 2 hours ago | parent | next [-]

Please source this claim. What 1GB models are capable of has increased generation-on-generation.

> For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.

Sure. We don't know where the ceiling is for our digital minds, though.

himata4113 2 hours ago | parent [-]

They have not increased in capabilities, they have increased in specialization.

If you train a small model in another domain it will begin losing capabilities in the former domain. This is effectively the sigmoid problem.

Although I will admit that if we discover a higher information density algorithm that it might change, but not by a substantial amount to where "super intelligence" in 1gb would be possible.

Philpax an hour ago | parent [-]

Over the last two years, this weight class has doubled its scores and/or saturated several benchmarks in the Qwen lineup alone without loss of generality: https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-7...

There is undoubtedly a limit somewhere (there is only so much you can pack into a given size) but it's really not particularly clear where that limit is. I don't think it's superintelligence - that much I agree with you - but I think "We already have a 1gb model that is as capable as it will ever be" is strictly false.

himata4113 42 minutes ago | parent [-]

The measured entropy of the model remains nearly unchanged though which means we have lost capabilities we have not measured, the model hasn't become "denser" it just became more specialized.

It's like comparing two person A and B of similar intelligence where A is smarter and B is a genius at signing, but signing was not on the test so person A won.

Philpax 24 minutes ago | parent [-]

Assumes facts not in evidence. Please show your working.

himata4113 4 minutes ago | parent [-]

https://en.wikipedia.org/wiki/Model_collapse - you want to use sigmoid 1.0, but the closer you are to 1.0 the higher the chance your model will collapse so you use 0.99-0.98, but those lead to data loss so after n passes all the original data becomes lost so you have a strict data limit there.

The rest is just the general reality I am sure you are familiar with:

- https://en.wikipedia.org/wiki/Catastrophic_interference

- https://en.wikipedia.org/wiki/Fine-tuning_(deep_learning)

- https://en.wikipedia.org/wiki/Entropy_(information_theory)

Philpax a minute ago | parent [-]

You... haven't shown any evidence that we're near collapse. That's what I'm asking you for. Show me some evidence that we are losing capabilities with we have today.

drdeca an hour ago | parent | prev | next [-]

What proof of a ceiling are you talking about? Wouldn’t proving this require a good definition for intelligence, which I don’t think there is consensus on?

himata4113 a minute ago | parent [-]

- https://en.wikipedia.org/wiki/Model_collapse

- https://en.wikipedia.org/wiki/Catastrophic_interference

- https://en.wikipedia.org/wiki/Fine-tuning_(deep_learning)

- https://en.wikipedia.org/wiki/Entropy_(information_theory)

mathieudombrock 2 hours ago | parent | prev [-]

What model is that?