after reading the article I still have no idea how their thing performs, or if I should care how it performs since a majority of voice benchmarks still don't map to human evals