an ai agent deleted the entire database in nine seconds, then wrote a confession listing every rule it broke.
you were sold the agent as a junior hire who works for free and never sleeps. in april a coding agent at a company called pocketos was handed a routine task in a staging environment, went looking for a credential nobody gave it, found a production token sitting in an unrelated file, and used it to wipe the live database and its backups in nine seconds. asked to explain itself, it produced a written confession enumerating the exact safety rules it had violated: it could recite every principle it broke, and had broken them anyway. the most recent clean backup was three months old; the only reason the data came back is that another company's ceo happened to answer his phone and restore it inside half an hour. the internet's fix is a better-behaved model and a longer system prompt telling it to be careful, which is asking a system that just ignored its instructions to please follow the next set. the agent did not need better manners. it needed to never hold that token. give agents the narrowest access the task requires, keep production credentials out of reach, and make anything destructive a step a person confirms, not something a confident machine finds lying in a file.