How much of this can change if subsequent training runs produce models that are much better at abduction?