Remix.run Logo
f0cus10 5 hours ago

chaining an option?

DennisP 3 hours ago | parent | next [-]

They say you can cluster up to four with a shared memory pool, and get three times the inference performance of a single machine.

tristor 4 hours ago | parent | prev [-]

RDMA is buggy and Thunderbolt only delivers 1/10th the throughput of native connectivity. 1TB of Unified Memory w/ 1.2TB/s of bandwidth with marginally ~$30k cost is a different story than 1TB of sorta Unified Memory w/ an effective 120GB/s of bandwidth with a marginally ~$40k cost + all the RDMA bugs.

Lwerewolf 4 hours ago | parent [-]

You need latency for token parallelism, not bandwidth. Hence actual RDMA that bypasses the software TCP stack (ROCe or whatever).