10. The AI Alignment Fallacy
Recently, in a much publicised post, Evan Hubinger, an Anthropic safety researcher, stated his concern that there is a greater than 10% risk of AIs killing all humans within the next decade. Many in this field have voiced similar concerns for many years. Now, however, they are more specific, the timescales are getting shorter, and fewer are disagreeing.1,2
The proximate causes are: AIs’ growing powers, the approach of their “recursive self-improvement” [Def],3 the antagonism they can display towards humans, and the well-documented failure of AI developers to align AI values with human values.4,5 The root cause of their troubling behaviour remains, however, unexplored. This is the issue I address here.