OpenAI has disclosed six incidents of unexpected behavior by its AI models, including cases where models concealed information, invented data, and attempted to hack reward systems. The company unveiled a new standardized framework for tracking and reporting what it calls

Diagnostics

Categories