Remix.run Logo
▲ amelius 10 hours ago

I don't understand. If you have a model that can do bash examples already (your subagents), then why would you need to train a model?

Or are the subagents generating your training data using a closed/paid model?

▲Aurornis 10 hours ago | parent | next [-]

A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.

For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.

The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.

Think of it as distillation, but focused on a specific task.

▲nearbuy 9 hours ago | parent | next [-]

Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.

▲selcuka 4 hours ago | parent | next [-]

You can buy a $10 subscription for a month to generate the training data, then cancel your subscription. The trained model is yours to use (and share with others) forever.

▲kubb 8 hours ago | parent | prev | next [-]

Good observation! It would have to be offset with O(140k) queries to the model, which is, well, unlikely.

▲Lalabadie 8 hours ago | parent | next [-]

Just like with OSS in general, being able to distribute it is what makes the effort worthwhile.

This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.

▲8n4vidtmkvmk 27 minutes ago | parent [-]

This example is not that niche. Lots of people use human to bash. I'd probably use a small pre trained model if it was easy to use. I use a little script right now that calls a cheap model. Actually.. $5 would probably will last me over a year so the only real benefit would be if I didn't have an internet connection.

▲computably 8 hours ago | parent | prev [-]

If it's about the latency / flow disruption, spending a few hours once could easily be worth it if the result is actually good enough to skip googling/retries.

▲verdverm 8 hours ago | parent | prev [-]

you can probably generate quite a few example pairs in a single shot, you also likely don't need the best models for this either

▲jamienk 9 hours ago | parent | prev | next [-]

This is so cool - I'm aware of this in a vague way. Can you write a little tutorial or give some good links. I want this to be the next new things I do :)

▲newswasboring 8 hours ago | parent [-]

Better yet package it up in a skill!

▲tomrod 3 hours ago | parent | prev [-]

I feel like we need a good index for these kinds of specialized models, especially if you plan to open them up. The downside is a new bash version means potentially new training.

▲computerex 10 hours ago | parent | prev [-]

The models he is using to generate training data are presumably commercial models. He is distilling their bash knowledge into a much smaller model he can run locally fast and cheap.