Most of the fastest inference and training code in production today is written in Python. There are no global locks on the GPU except the ones you put there