| ▲ | barbacoa 3 hours ago | |
They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs. So full CPU local AI inference may become viable option in coming years. | ||
| ▲ | lallysingh an hour ago | parent [-] | |
This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage. | ||