| ▲ | bugos 5 hours ago | ||||||||||||||||
How does showing the suggested answers to the user make the conversation better for model training? They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right? | |||||||||||||||||
| ▲ | ismailmaj 5 hours ago | parent | next [-] | ||||||||||||||||
The idea is that a thread can have many reasonable follow-ups that the user would've accepted, so it is wrong to punish the model for predicting a follow up that is different from the user message, as that prediction could've been accepted by the user if it was given. | |||||||||||||||||
| ▲ | yapfrog 5 hours ago | parent | prev | next [-] | ||||||||||||||||
The user actual answer vs the user actual answer after seeing the suggested answer are different points of data | |||||||||||||||||
| ▲ | wzdd 2 hours ago | parent | prev | next [-] | ||||||||||||||||
Agreed, they already have the HF -- delta versus model prediction can be calculated at any time. If anything, showing the suggestion introduces unwanted bias. | |||||||||||||||||
| ▲ | namanyayg 5 hours ago | parent | prev | next [-] | ||||||||||||||||
Seeing the suggestion influences the decision | |||||||||||||||||
| |||||||||||||||||
| ▲ | spwa4 5 hours ago | parent | prev | next [-] | ||||||||||||||||
RL training, the second phase of LLM training, is based on "I did X, was that good/bad?" and that 1 bit of information is the training data. So you give the user a suggestion, and the user accepts -> good You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data) The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources. I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this. | |||||||||||||||||
| ▲ | wilg 4 hours ago | parent | prev [-] | ||||||||||||||||
You can press Tab+Enter to accept it. | |||||||||||||||||