Haven't you seen the kernel optimization case study at the bottom of the page? They compare against GPT-5.6 Sol and their model is worse.