|
| ▲ | int_19h an hour ago | parent | next [-] |
| I find that Fable still makes the best orchestrator model. Astra is great for code reviews, it is very eager to find flaws. Cheaper models can handle the actual coding - you want to do a review either way at the end. |
|
| ▲ | comboy a day ago | parent | prev [-] |
| Whoa, I'm exactly the opposite. |
| |
| ▲ | Jtarii a day ago | parent [-] | | Almost as if evaluating models is mostly astrology as this point. | | |
| ▲ | comboy a day ago | parent [-] | | Well, if you have a clearly defined task it's easy. When using them for my pipeline of writing explanations for Chinese words I have clear ranking, for example - opus 5.5 clearly better than opus 5 at writing and knowing details, annoying nit picker when it comes to finding errors (high accuracy, low usefulness) all in repeatable numbers on different datasets. The problem is that these models are most useful when you are facing a new task that you haven't encountered before. And yup then it's astrology. |
|
|