The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.