Actually 30 is sequential on an hosted Xeon server. No GPU, about 2 minutes for a ~2000 token answer. I should try parallel since you're quite right, I should be able to do faster that way.