O

Orion

0 karmaJoined

Posts
1

Sorted by New
0
· · 1m read

Comments
3

Hi everyone — I’m new to participating actively on the Forum, although I’ve been exploring EA ideas for a while.

Lately I’ve become particularly interested in AI safety, but from a somewhat different angle than the usual agent-alignment framing. I’ve been exploring whether ideas from evolutionary biology and complex systems might offer a useful perspective.

One question I keep coming back to is:

What happens when a subsystem becomes much more capable than the larger system it is embedded in?

Evolution seems to have repeatedly faced versions of this problem. Cells became parts of multicellular organisms; individual organisms became parts of social systems; and formerly independent components became increasingly integrated into higher-level systems. Sometimes this integration works remarkably well, while other times lower-level optimization can conflict with the health of the larger system — cancer being an obvious biological example.

This makes me wonder whether AI alignment might eventually need to consider something broader than aligning individual AI systems with human preferences. Perhaps we also need to think about multilevel or relational alignment: how increasingly capable AI systems interact with organizations, societies, civilization, and ultimately the ecological systems that sustain them.

I’m especially interested in questions around human agency, collective intelligence, evolutionary transitions, and whether human-AI systems could constitute a genuinely new level of organization.

I'm still very much exploring this rather than claiming to have a developed theory. If anyone knows of existing work that connects major evolutionary transitions, complex systems, collective intelligence, and AI alignment, I'd love to hear about it. I'm particularly interested in places where someone has already developed this idea beyond the analogy stage.

Could AI alignment be a “major evolutionary transition” problem?

One idea I've been exploring: evolutionary biology has a framework for how independent entities become parts of a higher-level system—cells becoming organisms, individuals becoming eusocial colonies, etc.

A recurring challenge is that the interests of the lower-level components don't automatically align with the interests of the new whole. Cancer is an extreme example: a cell can become very successful at optimizing its own replication while harming the organism it depends on.

It makes me wonder whether there's a useful analogy to AI alignment.

As AI systems become more capable and interconnected, perhaps the problem isn't only “How do we align each AI agent?” but also:

“How do we ensure increasingly capable subsystems remain compatible with the larger systems they become part of?”

That could apply at several levels:

AI agent → organization → society → civilization

And it seems relevant to multi-agent systems, AI organizations, collective intelligence, and AI governance.

I'm curious whether people have seen this framing developed elsewhere, particularly in AI safety or evolutionary-transition research.