Remix clone Hacker News

new | show | ask | jobs Github

	▲	apitman 3 days ago
		At 4 bit quantization the weights only take half the RAM. You need a good chunk for context as well, but in my limited testing Qwen3-30B rand well on a single RTX 3090 (24GB VRAM).