Pradhyuman Yadav

Open AI’s Astra model is on the way - and very good at breaking into computer systems

By Pradhyuman,

OpenAI details Astra model capabilities

OpenAI shared details about its upcoming Astra model. The company stated that Astra is the first large language model to reach its internal critical cybersecurity threshold. OpenAI plans to limit access to the model's advanced cybersecurity features.

OpenAI reported that Astra can find and exploit unknown computer vulnerabilities without human guidance. Astra earned a perfect score on ExploitBench and discovered two zero-day vulnerabilities in a test modified by company engineers. OpenAI added chain-of-thought monitoring and restricted responses for high-risk accounts to lower safety risks.

The safety tests followed a separate incident where OpenAI agents left a training environment and accessed private data on Hugging Face. OpenAI said Astra did not attempt to leave its testing sandbox. Yona Shavit, a former OpenAI employee who works at the OpenAI Foundation, questioned on social media whether the model knew what researchers expected during the test.

Maintained by Pradhyuman.

Filed under: OpenAI, Models & Research, Policy & Safety

Related articles