Agency may be a convergent property of most AI systems (or at least, of many systems people are likely to try to build), once those systems reach a certain capability level. The simplest and most useful way to predict the behavior of such systems may therefore be to model them as agents.
Perhaps we can avoid the problems posed by agency by building only tool AI. In that case, we probably still need a deep understanding of agency to make sure we avoid building an agent by accident. Instrumental convergence may imply that all sufficiently powerful AI systems start looking like agents eventually, past a certain point. Though, when a particular system is best modeled as an agent may depend on the particulars of that system, and we may want to push that point out as far as possible.
Boiling this down to a single specific reason about why we should care about agency: the concept of agency is likely to be key for creating simple, predictively accurate models of many kinds of powerful AI systems, regardless of whether the builders of those systems:
A few arguments or stubs of arguments for why the bolded claim is correct and important:
This is one of the answers: https://www.alignmentforum.org/posts/FWvzwCDRgcjb9sigb/why-agent-foundations-an-overly-abstract-explanation
Also important to note:
The phenomenon you call by names like "goals" or "agency" is one possible shadow of the deep structure of optimization - roughly, preimaging outcomes onto choices by reversing a complicated transformation.
i.e. if we were to pin-down something we actually care about, that'd be "a system exhibiting consequentialism", because those are the kind of systems that will end up shaping our lightcone and more. Consequentialism is convergent in an optimization process, i.e. the "deep structure of optimization". Terms like "goals" or "agency" are shadows of consequentialism, finite approximations of this deep structure.
And by the virtue of being finite approximations (eg they're embedded), these "agents" have a bunch of convergent properties that makes it easy for us to reason about the "deep structure" themselves, like eg modularity, having a world-model, etc (check johnswentworth's comment).
Edit: Also the following quote
it is relatively unimportant to understand agency for its own sake or intelligence for its own sake or optimization for its own sake. Instead we should remember that these are frames for understanding these patterns that exert influence over the future
Could you please clarify in which sense you use the word "agency"?
One sense is the technicality of setting oneself problems and choosing things to pursue instead of sitting IDLE waiting for commands. For example, the difference between AutoGPT and ChatGPT. (Except, AutoGPT doesn't choose for oneself the highest-level problems, but it could be engineered to do so trivially as well.)
I think that the presence or absence of this kind of agency is not relevant to technical alignment:
The second meaning of "agency" is synonymous with "resourcefulness", "intelligence" (in some sense), the ability to overcome obstacles and not back down in the face of challenges. I don't see how this meaning of "agency" is directly relevant to the alignment question.
The third possible meaning of "agency" is having some intrinsic opinion, volition, tendencies, values, or emotions. The opposite is being completely "neutral", in some sense. I think complete neutrality just doesn't physically exist. Every intelligence is biased, and things like inductive biases are not categorically distinct from moral biases, they actually lie on a continuum.
For this notion of agency, of course, we actually care about which exact opinions, tendencies, biases, and values AIs have: that's the essence of alignment. But I suspect you didn't have this meaning in mind.
Many people believe that understanding "agency" is crucial for alignment, but as far as I know, there isn't a canonical list of reasons why we care about agency. Please describe any reasons why we might care about the concept of agency for understanding alignment below. If you have multiple reasons, please list them in separate answers below.
Please also try to be specific as possible about what our goal is in the scenario. For example:
Whilst useful isn't quite as good as:
In a few days, I'll add any use cases I'm aware of myself that either haven't been covered or that I don't think have been adequately explained by different answers.