| ▲ | karmakaze a day ago | |
Personally I'm using Qwen3.8-27B (MXFP4 quant W4A8) locally hosted on a pair of AMD GPUs (with DeepSeek Harness). It starts at 250 tokens/sec down to 120 past 128k context. At work mostly Opus 4.8 (sometimes a GPT or Gemini 3.1 Pro). I find Opus 5 chatty/slower and Fable can venture into over-engineering itself into unnecessary complications. | ||