Were you using the full model or a quantized version, and what harness/configuration were you using?
It sounds like you were using a quant model.