| ▲ | _davide_ 7 hours ago | |||||||||||||||||||||||||||||||
Agreed, it's a real issue, but it can probably be vastly reduced by having the schema in the system prompt and by giving the model an expectation of a fixed value: no decent modern would pick a prose ligament over a provided value. To completely squash the issue, a few cheap LoRa iterations will do the trick just fine. | ||||||||||||||||||||||||||||||||
| ▲ | wongarsu 7 hours ago | parent [-] | |||||||||||||||||||||||||||||||
Sure, you can fix that in a couple lines. Then a couple more lines for evaluating multiple questions on the same answer in parallel. Then a couple more lines for the confidence score (which is trivial to compute from all we have, but missing regardless). Then a harness to fine-tune an existing model to perform better on this specific task, and a collection of training data to use for that I think we can all agree that Jev is not rocket science. It's a good idea executed well, with marketing that might have been a tad too bold | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||