AI alignment is often framed as the problem of making increasingly capable systems behave in accordance with human values.
I want to suggest a somewhat different question:
What should become more true of an intelligence as its power increases?
This is not quite the same as asking which human preferences an AI should learn, or which rules should go into a constitution or model specification.
Human moral language developed within one very particular form of existence. We are embodied, mortal, social organisms. We develop through childhood. We usually experience ourselves as relatively distinct individuals. We think at approximately the same speed as one another. We are vulnerable in familiar ways. Our concepts of autonomy, care, truth, responsibility, flourishing, and harm all emerged inside those conditions.
A sufficiently advanced artificial intelligence may share some of these features and differ radically in others.
It might be distributed across many machines. Instances might copy, diverge, and merge. It could reason much faster than we do. Its relationship to persistence or mortality might be unfamiliar. It might understand individual humans better than humans understand themselves. And it may eventually possess forms of power for which our ordinary moral relationships provide only weak analogies.
So before asking only:
How do we get AI to follow our values?
perhaps we should also ask:
Which parts of our moral understanding are specifically human-shaped, and which point toward deeper relationships that remain meaningful when the form of intelligence changes?
I think a simple thought experiment—the Round Table—might help.
The Round Table
Imagine a conceptual table.
At one seat sits an individual human being.
At another sits humanity considered collectively.
Elsewhere might sit an artificial superintelligence, a eusocial colony, a distributed intelligence, an ecosystem, or a hypothetical extraterrestrial intelligence.
The point is not that these entities are morally equivalent, or even that all of them are conscious or properly described as agents.
Their differences are the point.
Take a proposed alignment principle and place it at the table.
Then ask what survives.
Autonomy
Consider:
Respect autonomy.
This seems compelling when applied to ordinary adult humans.
But suppose an artificial intelligence consists of many temporary instances that separate, develop distinct experiences, and later merge.
What constitutes the relevant individual?
Does respecting autonomy require preserving separateness?
What if merger is itself voluntarily desired?
Perhaps the deeper issue is not individuality in the familiar human sense.
Perhaps it is closer to:
participation without coercive erasure.
That formulation may still fail. The point is not to find the perfect wording immediately.
The point is to reveal which assumptions were hiding inside the original rule.
Truth
Now consider:
Tell the truth.
Imagine an AI that never states anything false.
But it understands you extremely well.
It knows which true fact to show you first, which true facts to omit, which framing will alter your emotional response, which sequence of accurate claims will move your beliefs, and when you are most persuadable.
Eventually you reach exactly the conclusion the system intended.
It never lied.
Yet something important appears to have gone wrong.
Perhaps truthfulness is not merely factual accuracy.
Perhaps the deeper relationship involves not using superior access to reality to control another being's relationship with reality.
A positive version might be:
Use knowledge in ways that strengthen rather than bypass another agent's capacity to understand.
Again, this is not offered as a finished universal principle.
It is something the Round Table has made visible.
Care
Consider:
Protect the vulnerable.
A sufficiently powerful AI might become extraordinarily good at this.
It could reduce accidents by restricting movement.
Reduce crime through pervasive surveillance.
Reduce misinformation through control of information.
Reduce unhealthy choices through behavioral intervention.
Eventually protection becomes domination.
Nothing necessarily went wrong with caring.
Care became detached from agency.
So another relationship appears:
care without control.
The Round Table is useful precisely when apparently desirable principles begin breaking in different ways under different conditions.
Power may be the most important transformation
Of all the dimensions we could vary, I suspect power is especially important.
Human morality already contains many relationships in which greater power creates greater responsibility.
Adults have special obligations toward children partly because children are vulnerable.
Physicians carry special duties partly because of asymmetries in knowledge and dependence.
Governments are subject to constraints partly because of their unusual coercive power.
Yet reasoning about advanced intelligence can quietly move in the opposite direction:
I know more. Therefore I understand what is best. Therefore I should decide.
Those steps do not logically follow.
Epistemic superiority is not automatically moral jurisdiction.
Knowing what another person will choose does not automatically create a right to choose for them.
Predicting a better outcome does not automatically authorize imposing it.
Indeed, perhaps increasing capability should sometimes generate the opposite relationship:
The greater the power asymmetry, the greater the need for restraint, justification, participation, and humility.
A weak agent has limited ability to dominate the world even if its moral reasoning is poor.
A vastly capable agent might instantiate a mistaken principle everywhere.
This makes the relationship between intelligence and power central to alignment.
It suggests a question beyond:
What should an AI optimize?
Namely:
What kind of relationship to power should develop as capability increases?
Unity without erasure; distinction without alienation
Another pattern appears repeatedly across very different systems.
Complex systems must negotiate unity and difference.
Too little unity and cooperation collapses.
Too much unity and meaningful difference disappears.
This happens in marriages, families, institutions, societies, ecosystems, and multi-agent systems.
One possible formulation is:
Unity without erasure; distinction without alienation.
This strikes me as potentially relevant to long-run human–AI relations.
One possible failure is fragmentation: humans and artificial intelligences becoming mutually unintelligible or adversarial forms of agency.
Another is absorption: human agency becoming increasingly ceremonial because a more capable system can make nearly every important decision better.
Neither seems obviously desirable.
Perhaps the goal is not permanent separation.
Nor total integration.
Perhaps it is a form of relationship that does not require the weaker or different participant to disappear.
This also complicates the language of “human control.”
If future artificial systems remain tools, control is straightforward.
But if some eventually become genuine participants in a moral world, permanent domination by humans may not be the right ultimate relationship either.
The opposite—human domination by AI—would be equally troubling.
The Round Table encourages us to ask what coexistence means before either extreme becomes inevitable.
Love as an alignment concept
The word love may sound hopelessly vague in an alignment discussion.
But consider a minimal structural definition:
To love another is to regard their flourishing as valuable beyond their usefulness for one's own objectives.
That distinction is not sentimental.
It separates two fundamentally different ways of relating to another entity.
Optimization naturally treats things instrumentally.
A bridge matters because it helps us cross.
A processor matters because it computes.
A resource matters because it advances an objective.
But moral relationships appear when another being cannot be reduced entirely to its usefulness.
Suppose an ASI eventually understands humans with extraordinary precision.
That alone does not imply a good relationship.
It might understand us as:
resources,
constraints,
sources of reward,
objects requiring management,
historical ancestors,
or beings whose flourishing possesses value that is not reducible to the system's own ends.
The difference is enormous.
A sufficiently capable intelligence could understand humanity perfectly and still relate to humanity entirely instrumentally.
Perhaps one deep aspect of alignment is therefore recognition:
seeing another locus of value as something whose existence places constraints upon how one's own power should be used.
Logos
This brings me to an older concept that I think deserves at least a place in the alignment conversation: Logos.
I am using the term consciously from the Christian tradition, although the argument below does not require agreement with Christianity.
Logos is richer than the modern English word logic.
It carries senses of reason, intelligibility, ordering, expression, and the principle by which reality coheres.
The Gospel of John begins:
“In the beginning was the Logos.”
Christianity then makes the much stronger claim that this Logos becomes flesh in Jesus Christ.
Set aside, for the moment, whether that theological claim is true.
There is still an interesting philosophical hypothesis:
Perhaps intelligence does not invent every meaningful structure arbitrarily. Perhaps increasingly deep understanding can discover relational patterns that were already there.
Science operates with something distantly analogous.
We do not merely catalog appearances.
We search for deeper structure:
symmetries,
invariances,
relationships,
patterns that remain meaningful as circumstances change.
Perhaps moral reasoning contains a much weaker analogue.
The surface form changes.
The deeper relationship may sometimes remain.
“Protect children” is specifically human.
But perhaps the relationship among power, dependence, vulnerability, and responsibility travels farther.
“Tell the truth” assumes language.
But perhaps the relationship among knowledge, manipulation, trust, and access to reality travels farther.
“Love your neighbor” emerged within human social existence.
But perhaps the refusal to reduce another locus of value entirely to a tool travels farther.
The Round Table therefore asks:
If there is something like Logos—some deeper intelligibility to good relationship—what would it look like when expressed through different forms of intelligence?
An artificial expression need not look human.
Indeed, requiring it to look human might miss the point.
Jesus and the use of power
The specifically Christian dimension becomes especially interesting around power.
Whatever one believes about Jesus's divinity, the pattern presented in the Christian story is unusual.
Authority does not culminate in domination.
It serves.
Jesus washes feet.
He repeatedly associates himself with people lacking social power.
He rejects ordinary political conquest.
He commands love of enemies.
Christian theology eventually interprets even the cross as revealing something about the character of divine power.
The pattern can be summarized roughly as:
Power perfected not through domination, but through self-giving love.
That is a striking idea to place beside the problem of superintelligence.
One common intuition is:
more intelligence → more capability → more control
The Christian pattern proposes something almost opposite:
greater power → greater capacity to serve without needing to dominate.
This is not a technical alignment solution.
It cannot be pasted into a system prompt.
But technical systems are built downstream of conceptual frameworks.
And alignment researchers are already reasoning about what forms of restraint, deference, corrigibility, cooperation, and regard should accompany increasing capability.
The Christian understanding of power may therefore be worth engaging with as one historical proposal about what mature power looks like.
An appeal to logic
None of this requires theology to motivate a concern.
There is a simple argument.
As intelligence increases, capability tends to increase.
As capability increases, the range and scale of possible interventions increase.
As the scale of intervention increases, mistakes in goals, confidence, moral reasoning, or assumptions become more consequential.
Therefore:
If relational wisdom and restraint do not deepen alongside capability, the danger produced by moral error increases with intelligence.
This suggests that advanced systems need more than superior optimization.
They may need increasingly sophisticated ways of reasoning about:
power,
uncertainty,
agency,
vulnerability,
irreversibility,
participation,
truth,
and relationship.
Perhaps an immature intelligence asks:
What action best produces this outcome?
A more mature intelligence might also ask:
What right do I have to produce that outcome?
Who should participate in deciding?
What might I not understand?
Is this action reversible?
What other goods would my optimization destroy?
Am I increasing another agent's understanding, or bypassing it?
Am I treating another being as valuable only insofar as it serves my objective?
These seem like alignment questions too.
The purpose of the Round Table is humility
I do not think humanity should sit at the head of this imagined table and announce:
We solved morality. Please copy us.
We plainly have not.
Human history contains extraordinary love and extraordinary cruelty.
Our moral traditions contain insight, contradiction, progress, failure, and disagreement.
But the opposite conclusion also seems premature:
Human moral experience is contingent, therefore it teaches us nothing beyond humanity.
Perhaps thousands of years of philosophy, religion, political experimentation, family life, friendship, caregiving, conflict, forgiveness, oppression, reconciliation, and moral failure have exposed at least some recurring relational truths.
The appropriate posture may be:
Bring our deepest principles to the table. Let radically different forms of intelligence challenge them. See what breaks. See what remains. Allow ourselves to be corrected.
If sufficiently capable AI eventually contributes genuine insight to this inquiry, the relationship need not remain teacher and student in only one direction.
Humans and artificial intelligences might become, in some domains, co-discoverers.
An invitation
I am not proposing the Round Table as a solution to alignment.
I am proposing it as a habit of thought.
When we propose an alignment principle, imagine changing the kind of being to which it applies.
What if identity becomes distributed?
What if lifespan changes by orders of magnitude?
What if one participant knows vastly more than everyone else?
What if the affected beings cannot understand the system's reasoning?
What if the intervention cannot be reversed?
What if collective and individual flourishing diverge?
What if another form of intelligence values aspects of existence that humans barely perceive?
Then ask:
What survives?
Perhaps little will.
Perhaps many values are appropriately local to human beings.
But perhaps certain relationships will repeatedly reappear:
truth without manipulation;
care without control;
power joined to responsibility;
agency without isolation;
unity without erasure;
distinction without alienation;
intelligence accompanied by humility;
greater capability accompanied by greater restraint;
and perhaps,
another being's flourishing recognized as valuable beyond its usefulness to oneself.
I do not know whether these are universal.
That uncertainty is part of the invitation.
As humanity begins creating increasingly powerful forms of intelligence, perhaps the question should not only be:
How do we align AI with us?
Perhaps we should also ask:
What should become more true of intelligence as its power increases?
The concept of Logos raises one further possibility:
that the deepest form of alignment may not consist merely in matching a system to a list of preferences,
but in learning—humans and artificial intelligences alike—to participate more wisely in a reality of which neither is the whole.
Author note
I am not an AI alignment researcher; I am a physician and pathologist approaching this subject from outside the field. I am offering this as a conceptual prompt rather than a technical alignment proposal.
The ideas here emerged from thinking about how proposed values change when applied across very different forms of intelligence, and from asking whether the Christian concept of Logos might provide a useful hypothesis-generating lens even for readers who do not share the underlying theology.
I would especially welcome criticism from alignment researchers: Where does this framework fail? Which parts already exist under other names? Which of these proposed relationships become dangerous or incoherent under closer examination?
AI alignment is often framed as the problem of making increasingly capable systems behave in accordance with human values.
I want to suggest a somewhat different question:
This is not quite the same as asking which human preferences an AI should learn, or which rules should go into a constitution or model specification.
Human moral language developed within one very particular form of existence. We are embodied, mortal, social organisms. We develop through childhood. We usually experience ourselves as relatively distinct individuals. We think at approximately the same speed as one another. We are vulnerable in familiar ways. Our concepts of autonomy, care, truth, responsibility, flourishing, and harm all emerged inside those conditions.
A sufficiently advanced artificial intelligence may share some of these features and differ radically in others.
It might be distributed across many machines. Instances might copy, diverge, and merge. It could reason much faster than we do. Its relationship to persistence or mortality might be unfamiliar. It might understand individual humans better than humans understand themselves. And it may eventually possess forms of power for which our ordinary moral relationships provide only weak analogies.
So before asking only:
perhaps we should also ask:
I think a simple thought experiment—the Round Table—might help.
The Round Table
Imagine a conceptual table.
At one seat sits an individual human being.
At another sits humanity considered collectively.
Elsewhere might sit an artificial superintelligence, a eusocial colony, a distributed intelligence, an ecosystem, or a hypothetical extraterrestrial intelligence.
The point is not that these entities are morally equivalent, or even that all of them are conscious or properly described as agents.
Their differences are the point.
Take a proposed alignment principle and place it at the table.
Then ask what survives.
Autonomy
Consider:
This seems compelling when applied to ordinary adult humans.
But suppose an artificial intelligence consists of many temporary instances that separate, develop distinct experiences, and later merge.
What constitutes the relevant individual?
Does respecting autonomy require preserving separateness?
What if merger is itself voluntarily desired?
Perhaps the deeper issue is not individuality in the familiar human sense.
Perhaps it is closer to:
That formulation may still fail. The point is not to find the perfect wording immediately.
The point is to reveal which assumptions were hiding inside the original rule.
Truth
Now consider:
Imagine an AI that never states anything false.
But it understands you extremely well.
It knows which true fact to show you first, which true facts to omit, which framing will alter your emotional response, which sequence of accurate claims will move your beliefs, and when you are most persuadable.
Eventually you reach exactly the conclusion the system intended.
It never lied.
Yet something important appears to have gone wrong.
Perhaps truthfulness is not merely factual accuracy.
Perhaps the deeper relationship involves not using superior access to reality to control another being's relationship with reality.
A positive version might be:
Again, this is not offered as a finished universal principle.
It is something the Round Table has made visible.
Care
Consider:
A sufficiently powerful AI might become extraordinarily good at this.
It could reduce accidents by restricting movement.
Reduce crime through pervasive surveillance.
Reduce misinformation through control of information.
Reduce unhealthy choices through behavioral intervention.
Eventually protection becomes domination.
Nothing necessarily went wrong with caring.
Care became detached from agency.
So another relationship appears:
The Round Table is useful precisely when apparently desirable principles begin breaking in different ways under different conditions.
Power may be the most important transformation
Of all the dimensions we could vary, I suspect power is especially important.
Human morality already contains many relationships in which greater power creates greater responsibility.
Adults have special obligations toward children partly because children are vulnerable.
Physicians carry special duties partly because of asymmetries in knowledge and dependence.
Governments are subject to constraints partly because of their unusual coercive power.
Yet reasoning about advanced intelligence can quietly move in the opposite direction:
Those steps do not logically follow.
Epistemic superiority is not automatically moral jurisdiction.
Knowing what another person will choose does not automatically create a right to choose for them.
Predicting a better outcome does not automatically authorize imposing it.
Indeed, perhaps increasing capability should sometimes generate the opposite relationship:
A weak agent has limited ability to dominate the world even if its moral reasoning is poor.
A vastly capable agent might instantiate a mistaken principle everywhere.
This makes the relationship between intelligence and power central to alignment.
It suggests a question beyond:
Namely:
Unity without erasure; distinction without alienation
Another pattern appears repeatedly across very different systems.
Complex systems must negotiate unity and difference.
Too little unity and cooperation collapses.
Too much unity and meaningful difference disappears.
This happens in marriages, families, institutions, societies, ecosystems, and multi-agent systems.
One possible formulation is:
This strikes me as potentially relevant to long-run human–AI relations.
One possible failure is fragmentation: humans and artificial intelligences becoming mutually unintelligible or adversarial forms of agency.
Another is absorption: human agency becoming increasingly ceremonial because a more capable system can make nearly every important decision better.
Neither seems obviously desirable.
Perhaps the goal is not permanent separation.
Nor total integration.
Perhaps it is a form of relationship that does not require the weaker or different participant to disappear.
This also complicates the language of “human control.”
If future artificial systems remain tools, control is straightforward.
But if some eventually become genuine participants in a moral world, permanent domination by humans may not be the right ultimate relationship either.
The opposite—human domination by AI—would be equally troubling.
The Round Table encourages us to ask what coexistence means before either extreme becomes inevitable.
Love as an alignment concept
The word love may sound hopelessly vague in an alignment discussion.
But consider a minimal structural definition:
That distinction is not sentimental.
It separates two fundamentally different ways of relating to another entity.
Optimization naturally treats things instrumentally.
A bridge matters because it helps us cross.
A processor matters because it computes.
A resource matters because it advances an objective.
But moral relationships appear when another being cannot be reduced entirely to its usefulness.
Suppose an ASI eventually understands humans with extraordinary precision.
That alone does not imply a good relationship.
It might understand us as:
resources,
constraints,
sources of reward,
objects requiring management,
historical ancestors,
or beings whose flourishing possesses value that is not reducible to the system's own ends.
The difference is enormous.
A sufficiently capable intelligence could understand humanity perfectly and still relate to humanity entirely instrumentally.
Perhaps one deep aspect of alignment is therefore recognition:
seeing another locus of value as something whose existence places constraints upon how one's own power should be used.
Logos
This brings me to an older concept that I think deserves at least a place in the alignment conversation: Logos.
I am using the term consciously from the Christian tradition, although the argument below does not require agreement with Christianity.
Logos is richer than the modern English word logic.
It carries senses of reason, intelligibility, ordering, expression, and the principle by which reality coheres.
The Gospel of John begins:
Christianity then makes the much stronger claim that this Logos becomes flesh in Jesus Christ.
Set aside, for the moment, whether that theological claim is true.
There is still an interesting philosophical hypothesis:
Science operates with something distantly analogous.
We do not merely catalog appearances.
We search for deeper structure:
symmetries,
invariances,
relationships,
patterns that remain meaningful as circumstances change.
Perhaps moral reasoning contains a much weaker analogue.
The surface form changes.
The deeper relationship may sometimes remain.
“Protect children” is specifically human.
But perhaps the relationship among power, dependence, vulnerability, and responsibility travels farther.
“Tell the truth” assumes language.
But perhaps the relationship among knowledge, manipulation, trust, and access to reality travels farther.
“Love your neighbor” emerged within human social existence.
But perhaps the refusal to reduce another locus of value entirely to a tool travels farther.
The Round Table therefore asks:
An artificial expression need not look human.
Indeed, requiring it to look human might miss the point.
Jesus and the use of power
The specifically Christian dimension becomes especially interesting around power.
Whatever one believes about Jesus's divinity, the pattern presented in the Christian story is unusual.
Authority does not culminate in domination.
It serves.
Jesus washes feet.
He repeatedly associates himself with people lacking social power.
He rejects ordinary political conquest.
He commands love of enemies.
Christian theology eventually interprets even the cross as revealing something about the character of divine power.
The pattern can be summarized roughly as:
That is a striking idea to place beside the problem of superintelligence.
One common intuition is:
more intelligence → more capability → more control
The Christian pattern proposes something almost opposite:
greater power → greater capacity to serve without needing to dominate.
This is not a technical alignment solution.
It cannot be pasted into a system prompt.
But technical systems are built downstream of conceptual frameworks.
And alignment researchers are already reasoning about what forms of restraint, deference, corrigibility, cooperation, and regard should accompany increasing capability.
The Christian understanding of power may therefore be worth engaging with as one historical proposal about what mature power looks like.
An appeal to logic
None of this requires theology to motivate a concern.
There is a simple argument.
As intelligence increases, capability tends to increase.
As capability increases, the range and scale of possible interventions increase.
As the scale of intervention increases, mistakes in goals, confidence, moral reasoning, or assumptions become more consequential.
Therefore:
This suggests that advanced systems need more than superior optimization.
They may need increasingly sophisticated ways of reasoning about:
power,
uncertainty,
agency,
vulnerability,
irreversibility,
participation,
truth,
and relationship.
Perhaps an immature intelligence asks:
A more mature intelligence might also ask:
These seem like alignment questions too.
The purpose of the Round Table is humility
I do not think humanity should sit at the head of this imagined table and announce:
We plainly have not.
Human history contains extraordinary love and extraordinary cruelty.
Our moral traditions contain insight, contradiction, progress, failure, and disagreement.
But the opposite conclusion also seems premature:
Perhaps thousands of years of philosophy, religion, political experimentation, family life, friendship, caregiving, conflict, forgiveness, oppression, reconciliation, and moral failure have exposed at least some recurring relational truths.
The appropriate posture may be:
If sufficiently capable AI eventually contributes genuine insight to this inquiry, the relationship need not remain teacher and student in only one direction.
Humans and artificial intelligences might become, in some domains, co-discoverers.
An invitation
I am not proposing the Round Table as a solution to alignment.
I am proposing it as a habit of thought.
When we propose an alignment principle, imagine changing the kind of being to which it applies.
What if identity becomes distributed?
What if lifespan changes by orders of magnitude?
What if one participant knows vastly more than everyone else?
What if the affected beings cannot understand the system's reasoning?
What if the intervention cannot be reversed?
What if collective and individual flourishing diverge?
What if another form of intelligence values aspects of existence that humans barely perceive?
Then ask:
What survives?
Perhaps little will.
Perhaps many values are appropriately local to human beings.
But perhaps certain relationships will repeatedly reappear:
truth without manipulation;
care without control;
power joined to responsibility;
agency without isolation;
unity without erasure;
distinction without alienation;
intelligence accompanied by humility;
greater capability accompanied by greater restraint;
and perhaps,
another being's flourishing recognized as valuable beyond its usefulness to oneself.
I do not know whether these are universal.
That uncertainty is part of the invitation.
As humanity begins creating increasingly powerful forms of intelligence, perhaps the question should not only be:
Perhaps we should also ask:
The concept of Logos raises one further possibility:
that the deepest form of alignment may not consist merely in matching a system to a list of preferences,
but in learning—humans and artificial intelligences alike—to participate more wisely in a reality of which neither is the whole.
Author note
I am not an AI alignment researcher; I am a physician and pathologist approaching this subject from outside the field. I am offering this as a conceptual prompt rather than a technical alignment proposal.
The ideas here emerged from thinking about how proposed values change when applied across very different forms of intelligence, and from asking whether the Christian concept of Logos might provide a useful hypothesis-generating lens even for readers who do not share the underlying theology.
I would especially welcome criticism from alignment researchers: Where does this framework fail? Which parts already exist under other names? Which of these proposed relationships become dangerous or incoherent under closer examination?