| ▲ | theChris-in 9 hours ago | |
> should test the harness capability rather than model's knowledge/capability. Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class. | ||