Potential customers and researchers rely on published benchmarks to compare AI models. If metrics shift post-launch and sometimes favor OpenAI's own system, it becomes harder to assess whether claims about Astra's capabilities are stable or if comparisons with rival models are fair.
Most affected
Researchers and procurement teams evaluating AI models — Cannot rely on published benchmarks as fixed comparison points between systems.
Likely next
Whether OpenAI publishes a formal explanation of the metric changes and whether other AI labs respond with their own benchmark challenges or rebuttal comparisons.