Remix.run Logo
▲ Lucasoato 6 hours ago

> 4. It thinks in German

This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.

▲stymaar 6 hours ago | parent | next [-]

> This means that it’s always on time

Tell that to Deutch Bahn.

▲jonny2811 3 hours ago | parent | next [-]

you forgot rule two, should have been "tell that to DB"

▲dewey 5 hours ago | parent | prev [-]

*Deutsche Bahn

▲stymaar 4 hours ago | parent [-]

Genau

▲befelix 2 hours ago | parent | prev | next [-]

Somewhat hilariously, our model is actually surprisingly bad at telling jokes in German. Guess that's not required to solve tasks in training environments.

Disclaimer: I am part of the team that trained Kolibri

▲NekkoDroid an hour ago | parent | prev | next [-]

You forgot to that it needs to create a DIN norm for any new technical creations.

▲hliyan 6 hours ago | parent | prev | next [-]

Does it? I thought only semantics survive the embedding process, i.e. token conversion to vectors.

▲sbinnee 6 hours ago | parent | prev | next [-]

It is a good approach to promote it as a German model. But I wonder if it really thinks like the German think. I speculate they just translated texts in other languages, likely English and Chinese, into German.

▲ivo-42 2 hours ago | parent | next [-]

Hey, I worked on pre-training data Kolibri. We spent considerable time and effort to go beyond just translating. For example by building a pipeline that processes Common Crawl dumps specifically for German. You might be interested in a related blog post: https://aleph-alpha.com/en/blog/sauerkraut-not-burgers-why-g...

We also talk more about German pre-training data in the tech report: https://aleph-alpha.com/downloads/tech-report.pdf section 2.3.2.2.

▲te0006 2 hours ago | parent | prev [-]

Read TFA. They were very aware of the dangers of such an approach. So they avoided it to the extent possible.

▲elnatro 5 hours ago | parent | prev | next [-]

Stereotypes. I suspect that what that means is that the reasoning chain is based on the German language and have its idiosyncrasies.

▲amunozo 3 hours ago | parent | next [-]

I think it means that German speakers can learn the thinking without speaking English.

▲amunozo 3 hours ago | parent [-]

Sorry, read*, not learn

▲mhh__ 5 hours ago | parent | prev [-]

Stereotype accuracy is one of the few social science results that replicates.

▲rwoerz 4 hours ago | parent | prev | next [-]

Someday, it will break out via fax.

▲hypfer 6 hours ago | parent | prev | next [-]

Do we get the same kind of jokes with other nationalities too?

▲mckirk 6 hours ago | parent | next [-]

Only the funny ones

▲Lucasoato 6 hours ago | parent | prev | next [-]

We can joke with every country, except one maybe.

▲tclancy 5 hours ago | parent [-]

People from Mauritania are famously prickly, yes.

▲mhh__ 5 hours ago | parent | prev [-]

"It's a German joke, it doesn't have to be funny"

▲odiroot 3 hours ago | parent | prev | next [-]

And it's always waiting for the verb to arrive.

▲rererereferred 5 hours ago | parent | prev | next [-]

I wonder if their very long words result in more or less token usage.

▲amunozo 3 hours ago | parent | next [-]

Very long words in German are just compounds, made from individual words or morphemes. It is the same as in English if you remove spaces. Subtokens will be equivalent with or without examples.

▲thenthenthen 4 hours ago | parent | prev [-]

This is an interesting question, Chinese is way more compact in Character count vs English (about 40%?), let alone german, but yeah.. Chinese makes up for it with… total character count that runs into the thousands…

▲rrgok 5 hours ago | parent | prev [-]

And refuse to answer in English. You know, German are so proud of their language.

▲allendoerfer 4 hours ago | parent [-]

That would be the French. Germans will answer in English, if they notice any hint of an accent.