What's the training corpus, just the canonical Mitchell and Webb sketch? Or ancilliary data like comments and references, or worse synthetic data generated on the primary sources?