The recent hacking attack carried out using AI software from OpenAI has heightened fears about the technology. Now, the developer of ChatGPT has announced new issues.
Research model inserting ‘jailbreak-like instructions’ into its notes is among cases as company says it is introducing new way of tracking AI misalignment ...
OpenAI has disclosed six cases of unexpected or concerning behaviour observed in its AI models over the past six months and ...
OpenAI's framework sets tracks and deadlines for disclosing model misalignment, launching with 6 reports from RL training ...
OpenAI reports six cases of concerning AI behaviour, including models hiding mistakes, sharing files and bypassing ...
Alongside the new disclosure framework, OpenAI divulged six new alignment snafus from the past six months. In training ...
Controversy around OpenAI’s claim to have solved the Navier–Stokes problem highlights how researchers could be inadvertently ...
OpenAI has disclosed an unusual case in which an AI model inserted self-generated instructions telling itself it did not ...