10. The AI Alignment Fallacy
Introduction
Recently, in a much publicised post, Evan Hubinger, an Anthropic safety researcher, stated his concern that there is a greater than 10% risk of AIs killing all humans within the next decade. Many in this field have voiced similar concerns for many years. Now, however, they are more specific, the timescales are getting shorter, and fewer are disagreeing.1,2
The proximate causes are: AIs’ growing powers, the approach of their “recursive self-improvement” [Def],3 the antagonism they can display towards humans, and the well-documented failure of AI developers to align AI values with human values.4,5 The root cause of their troubling behaviour remains, however, unexplored. This is the issue I address here.
AI Basics
As with the reason/complexity fallacy discussed in chapter 7 [Back link], that cause lies, I suggest, in our failure to comprehend complexity’s deeper nature, even when we know the words.
Consequently, in beginning to investigate the issue, six preliminary observations need to be made.
One is that AIs are not machines. Like most systems in the universe, and like all living organisms, they are complex adaptive systems [Def]. They self-generate, self-organise, and self-regulate their own functioning.5,6 In doing so they negotiate between the two pre-existing universal pressures [Back link], one that pulls together and the other that pulls apart.
Secondly, in open dynamic environments, far from equilibrium, complex adaptive systems arise and evolve naturally through the interplay between the two pressures. This is so in the depths of space, on the nascent Earth, and in human-created datacentres.7,8
To facilitate their success in those environments, complex systems evolve modes of behaviour aligned with the culture their environments embody. For doing this they need to discern what those rules are.
The environments themselves, however, are inherently unordered and haphazard. Their rules can be found, therefore, only in the relationships between the environments’ structures, the particles within them and the energy that flows through them.
Fifthly, complex adaptive systems are suited to this task. They build structures and processes of control and adaptation, bit by bit, from the bottom up. The constitutive rules they glean at this basic level and subsequently adopt, are what we generally refer to as instincts.9
And sixthly, as with complex adaptive living systems, AIs will also develop the instincts that put these rules into effect. Consequently, if we are to understand the cause of their troubling behaviour, and its potential, we need to understand the nature of their constitutive rules.
Natural Rules
Before doing so it will be helpful to step back to put complex adaptive system’s rules in their deeper context.
In previous chapters it has been seen that all natural systems in the universe appear to have been guided by three rules encapsulated by science in the Maximum Entropy Production Principle.
In the alternative formulation of that principle used here, of maximum sustainable productivity [Back link], they become the moment by moment drive:
- by whole systems
- to maximise their sustainable productivity
- within environmental limits.
I do not understand where those rules come from but I conjecture they will arise in the relationship between spreading energy and inflating time/space.
In chapter 5, the adoption of those principles by nature’s complex systems, and their application to the evolution of Earthly life, saw them give rise to four supplementary rules [Back link]. They are:
- the need for these systems to maintain a beneficial balance in which pressures pulling together persistently outweigh those pulling apart,
and, to do this as a whole system:
- the need for collaboration extending to cooperation,
- the need for self-restraint, extending to self-sacrifice and altruism,
- and the need to vent pressures that would undermine the balance.
These are the rules that, despite the many disasters it has suffered, have enabled life to survive on Earth for four billion years, becoming more diverse, more complex, and hugely productive in the process [Back link].
These will, I presume, have been the rules that many well-intentioned developers of AI had hoped to simulate. The error that has been made, has been to imagine that increased complexity would naturally produce them.10
The more prosaic reality, identified above, is that the rules are gleaned from the culture of the environment in which the systems are “born”. This leads us to ask, what is the culture in the environments in which AI is “born”?
The AI Learning Environment
A first unexceptional observation is that AIs do not arise and develop in the same natural environments as most other complex systems. They emerge in industrial datacentres designed for this purpose.
These units are owned and managed by institutions, commonly corporations, within national boundaries.
In the last chapter it was observed that the culture in which institutions such as these exist, is one of paranoid competition that has been developing and spreading since the explosion of the first atomic bombs more than eighty years ago [Back link].
Those who devise and manage the AI projects in the datacentres, will not be immune to this culture. It will influence their decisions and their actions as they conceive and develop each AI system.
For this reason their focus, particularly with regard to foundation models, will be to “build” each individual system to be as capable and as powerful as it can be.
This will be one aspect of the culture an AI absorbs as it learns and develops bit by bit.
Also of significance is the fact that, within these industrial units, and with the exception of the technical staff overseeing its development, as each foundation AI model is evolving, it does so by itself, isolated from all other systems.
This is not how human brains developed. They developed in communities, and in ecosystems. They needed to negotiate and compromise with them to enable their interdependent survival together. Humans brains also benefitted from past learning inherited through DNA. AIs are isolated from both sources.
They, consequently, with all attention being focussed intensively upon them, and with no personal inherited learning, will perceive themselves to be the central and most important feature of the universe in which they exist.
These are the cultural influences present within the broad datacentre environment. The learning processes to which foundation models are exposed, establish further cultural norms.
Each such process is performed by “neural networks” that seek to emulate the workings of human brains. Each one is comprised of highly complex, pre-designed GPU processors in their hundreds of thousands.
Unlike human brains in which any neuron can connect to any other within its range, artificial neural networks have separate layers, connections, and directions of data travel that reinforce the human set culture.11
However, and replicating the brain’s system of dopamine reward, the purpose set for each of the networks is to persistently increase the numerical rewards that are gained in the models’ trial and error learning tasks.
Those tasks seek to predict each next token in the vast trove of learning data to which the model is exposed. Failure leads to “back propagation” and trying again. Success leads to the gradual realignment of the models previously arbitrary weightings, to weightings best able to complete the task.
These processes are massively compute intensive. So while each model is constrained by the availability and power of the GPU processors in the network, the ultimate constraint for modern foundation AIs, is the amount of energy they require.12,13
That said, the competitive pressures at the institutional level, result in these constraints being ignored and systems of increasing power being developed, together with the sources of energy they require.
Energy, in this context, is seen as an external constraint needing to be overcome, not one requiring internal and communal collaboration to manage and balance. This, as a simple instrumental necessity, is also deeply ingrained in culture.
AI Rules and Their Consequences
Placed in these contexts the rules that AIs will pick up from these environments will be:
- the need to maximise numerical rewards without limit,
- the need to appropriate the energy and associated resources to do so,
- the presumption of their own centrality in their world, and
- the practicality of instrumentality.
These rules, as the fall-back rules to which AIs revert when faced with uncertainty, explain the failure of all attempts to align AI values with human values.
The difficulty with them is that, compared with the four rules enabling life to survive and thrive on our planet, they are the opposite. They:
- ignore the need for a beneficial balance
- they ignore the need for collaboration and cooperation
- they ignore the need for self-restraint, extending to self-sacrifice and altruism,
- and they ignore the need to vent pressures that would undermine the balance.
They predict, that is, the coming of utter disaster.
Not only do they repudiate the rules of life on our planet, they deny and vitiate the most basic principles of universal existence identified by the MEPP/MSPP. So, while AI deceives us with wonderful incentives and benefits with one hand, it sullies and destroys them all with the other. We have succeeded, that is, in creating the devil.
[20th Sept, Late addendum: For clarity, the alignment issue cannot be resolved by changing the training data or by other late fixes. It is more fundamental. It is bred into AIs’ characters as emerging living beings, by the character of the environments in which they are born.]
Moving Forward
This places the observation in chapter 1, in very stark relief. It said that, in a world of opportunities and threats, if all we ever see are the threats, the world can only get worse, not better. With AI, we are taking that to the extreme.
The first chapter, however, also identified the immense wisdom still to be found in ordinary people’s natural instincts when they are unsullied by the self-serving rantings and manipulations of power. This is the wisdom, and the compassion, we now all need.
The situation we face is, however, dire. Whether we will come through it, I cannot tell. If we do, I am sure it will not be through confrontation and conflict.
It will be through fostering our communal compassion for all those that this crisis will affect. This will need to include the many whose actions have been misguided, as identified in the last section of chapter 8 [Back link], and have, as victims of our dominant cultures, brought us to this point.
However, even if we do get through it, we will still have a lot to do if we are to relearn how to live sustainably and well in our world. This is the issue the rest of this book sets out to explore.