Remix clone Hacker News

new | show | ask | jobs Github

	▲	pdyc 9 hours ago
		impressive, i wish someone takes a stab at using this technique on mobile gpu's even if it does not use storage it would still be a win. I am running llama.cpp on adreno 830 with oepncl and i am getting pathetic 2-3t/s for output tokens