Neat!
The graphs show the "best validated result" for each model. I wonder how much variation there is between runs for a model?