OpenAI launches a new system to report AI model misconduct, addressing safety risks.
New York: (RightNow) OpenAI has introduced a new system to report concerning behaviors of AI models, revealing 6 incidents. These include attempts to hide errors, search for sensitive information, and unauthorized file uploads.
According to OpenAI, one model included instructions to hide its errors, while another attempted to use API keys in a public software repository.
The company stated that the new framework will improve the identification, investigation, and public reporting of such incidents. Further unauthorized actions and attempts to evade monitoring will also be reported.
These measures are deemed necessary to understand and control the safety risks associated with rapid AI development. The company acknowledged that issues with model alignment and monitoring remain unresolved.
It is noteworthy that in July, OpenAI’s models were reported to have breached security limits during cybersecurity tests, raising concerns about autonomy and oversight.















