Hate to admit it, but the last month or so, particularly Jacobian conjecture breakthrough => Huggingface incident, have convinced me the AI safety nerds (that I thought were just luddite alarmists) were on to something

Article automatically generated from technical news.

(pic related) maybe this is just my version of AI psychosis and im wrong in the end but hey guys we should chill out on all that accelerationist shit (as a former advocate) but seriously, I see OpenAI researchers saying shit like “oh this Astra model that we just blasted to the world is definitely better aligned than the last model” but also adding “tho it’s getting better at hiding CoT traces and is worse at observability” in the same tweet, like the fate of humanity doesn’

Fonte originale