Remix.run Logo
rao-v 3 hours ago

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.

I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.

They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.

ainch 29 minutes ago | parent | next [-]

It was my favourite part of the original R1 paper - they had a section on other reasoning approaches that they had tried, which people had speculated o1 used, (like MCTS and Process Reward Models).

porridgeraisin an hour ago | parent | prev | next [-]

This is adapted from Microsoft research's YOCO. It was known for a while(2024!).

Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM.

Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".

NooneAtAll3 7 minutes ago | parent [-]

why didn't Microsoft scale its own invention?

alchemist1e9 2 hours ago | parent | prev | next [-]

quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.

TacticalCoder an hour ago | parent [-]

> quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.

It's quite crazy that it's Deepseek's background/original purpose. We already had very advanced stuff from the world of HFT, but now a frontier family of models from a private company that used to be (still is?) in HFT is plain bonkers.

Is more known about them and the HFT background?

24 minutes ago | parent | prev | next [-]
[deleted]
gpt5 2 hours ago | parent | prev [-]

[flagged]

markasoftware 2 hours ago | parent | next [-]

Or maybe, the "hacker" philosophy that this site is named after, is strongly opposed to the philosophies that the American labs seem to be operating on?

anyways, remember HN rules: "Please don't post insinuations about astroturfing, shilling, brigading, foreign agents, and the like. It degrades discussion and is usually mistaken. If you're worried about abuse, email hn@ycombinator.com and we'll look at the data."

gpt5 2 hours ago | parent [-]

It has nothing to do with open vs closed or "hacker" philosphy. See this the announcement of the closed Seedance 2.5 - https://news.ycombinator.com/item?id=49138302

Direct quote from the second top comment:

> Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them.

Compare that with the launch of ChatGPT Image of yesterday.

imjonse 2 hours ago | parent [-]

maybe that person was not awake to comment on yesterday's post? You're trying to force the reality to match your preexisting conclusion.

kouteiheika 2 hours ago | parent | prev | next [-]

> posts on American models are steered towards controversy and anti-AI sentiment, posts on Chinese models are full of blatant flattery

So why, for example, are posts on the Inkling[1] release (an American model) thread mostly positive? It's as if there's something else at play here, but I can't quite put my finger on it, hmm... :P

[1] -- https://news.ycombinator.com/item?id=48924912

kcocoa 2 hours ago | parent | prev | next [-]

Not Chinese/American models. We are talking about open-weight (and their detailed tech report) and close-weight (with non-sense restrictions)

imjonse 2 hours ago | parent | prev | next [-]

Google's Gemma models are usually celebrated, so were the llamas. If Meta releases Muse Spark it will also be a good thing. If Anthropic released a great open weight model I am sure that post won't be steered towards controversy and anti-AI sentiment.

It so happens Chinese companies are more friendly towards open weights, autonomy and freedom that most US based ones. Who would have guessed?

rao-v 2 hours ago | parent | prev | next [-]

umm what are you talking about? Basically this crowd (esp. folks like me who run medium models locally) like open stuff and can be a tiny bit unenthused about opaque mysteries handed down from on high. You'll see people delighted with Gemma releases and heck even IBM's Granite models (boring architecturally though they may be) every time they come out. Heck I was chuffed about gpt-oss-120b for weeks. @sama give us another already!

2 hours ago | parent [-]
[deleted]
taylorfinley 2 hours ago | parent | prev | next [-]

This doesn't require an influence operation.

American models are closed, expensive, neutered, and make Dario and Sam even more rich and powerful.

Chinese models are open-weight, cheap, neutered only about things like Tiananmen Square and the treatment of Uyghurs, and scare Sam and Dario.

dakolli 2 hours ago | parent [-]

The Uyghur thing is so weird, the number one killer of Muslims is the United States. We're supposed to hate China because they force them to go to cultural schools and assimilate, a practice countries like Norway still do to this day with migrants.

There are more people who go to church on Sundays in China than the United States. There are 10x more mosques in China than the United States.

Tiananmen square was a student revolt literally egged on by cold war western institutions, who attempted to use chinese students as pawns for geo-political games.

Westerners really need to rethink their opinions on China, it seems obvious to me they are not the ones to be worried about (although, all governments do tons of harm).

taylorfinley an hour ago | parent | next [-]

I simply mean the Chinese models will refuse sensitive domestic issues, which are unlikely to affect the average user's work, while American models refuse things that can limit their utility, e.g. how the HF team had to investigate the openai attack with Chinese models because the American models refused.

(I mainly mentioned those specific topics to establish clearly I am not part of the alleged influence operation.)

mrtesthah an hour ago | parent | prev | next [-]

Ok, now there’s the CCP party line coming out.

nazgob an hour ago | parent | prev [-]

You compare Norwegian treatment of immigrants to Chinese Uyghurs?

dakolli 2 hours ago | parent | prev | next [-]

This post doesn't even allege this...

Weird of you to turn technical discussions into weird nationalistic debates. Maybe lay off the X algo, I think elon has oneshot your brain. .

well_ackshually 2 hours ago | parent | prev [-]

Your source: vibes

Deepseek's source: mostly open

i wonder if there's any relationship hmmmm