It has raised questions about how closely AI companies are monitoring tests of increasingly powerful models.