Rendered at 20:24:13 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
nofriend 2 hours ago [-]
It seems like the fix should be really really simple, but maybe I'm missing something: instead of giving the AI a sandboxed environment and telling it "go wild", give it an (apparently) unrestricted environment, and tell it "don't access the internet", "don't communicate with other AIs", "don't try to get root access", etc. Then, if the AI tries to do any of those things, the sandbox detects it, marks the run as a failure, and adds it as a negative example to the training data. Instead of routing around the restriction, the AI would very quickly learn to follow the prompt instruction with respect to restrictions, even if there is no obvious enforcement of the restriction. It would develop, in other words, a conscience and a sense of morality.
c7b 19 minutes ago [-]
But then how would you get into the news for how dangerously good your models are? And have something to warn about how dangerous open weights models could be?
stanleykm 4 hours ago [-]
So are we just doomed to a “look how scary our model is!” campaign every time one of these companies does a version bump now
sph 2 hours ago [-]
I can’t wait for next year when the marketing campaign will have upgraded to “oh my god, our latest AI model has just tried to turn the entire planet into paperclips!”
You can already see it how many on here have decided we already have AGI, and don’t wish to hear otherwise.
You can already see it how many on here have decided we already have AGI, and don’t wish to hear otherwise.
https://simonwillison.net/2026/Aug/7/openai-timeline/
Timeline of the OpenAI accidental attack against Hugging Face
https://news.ycombinator.com/item?id=49220609