| ▲ | kennywinker 5 hours ago | |
Without access to reasoning traces, we can't know that - someone inside openai/anthropic would have to run the test - and we'd have to trust their results. I would be curious to see how the open weight models do on a test like this - and then we'd be able to see the reasoning. | ||