The Viable System Model & Multi-Scale Agency

Link post

AI was used to generate the scary science attack section in a different voice than the original part was written through as well as creating diagrams according to my instructions in LaTeX.

Introduction

One of the deeper questions within the field of AI Safety is on how we can create a theory of multi-scale hierarchical agency.

I want to give you an alternative today which comes from the tradition of cybernetics, the people that information theorists like Claude Shannon talked to back when coming up with information theory.

There was a gentleman there by the name of Stafford Beer who would come to be an operations researcher and progenitor of a part of modern management science. In his somewhat obscure writing he created something called the Viable Systems Model which is a very interesting buzzword in certain circles. It is a way to describe general businesses and governments (e.g collective intelligences) and what they do developed through years of practice and it has a bunch of cool information theory hidden behind it. It is a bit dense and difficult to understand and yours truly has spent some time doing this. Yours truly has also tried to translate this into modern information theoretic terms through an Active Inference angle.

Now, this translation is not necessarily fully precise, it is approximate and pointing at the underlying truth. I’m not necessarily sure that this is the right way to formalise it either but hopefully it points at something interesting that we can build upon!

As I’ve been writing this, I’ve been continually realising that the depth of this goes deeper than I thought. I’ve been listening to these VHS tapes from the 90s and there’s a lot of stuff there to unpack.

My plan is therefore to introduce these general concepts to you, go through the brain of the firm in the future and get back to you with whatever I find. (I’m still an ignoramus in the cybernetics land for good and for bad.)

Scary Diagram attack!

Ahh! A scary diagram showed up, watch out!

This is my interpretation of Beer’s 5 levels in the Viable Systems Model from an Active Inference perspective. Now as you can see, this is quite the complicated schematic but don’t you worry dear viewer for I shall explain it to you!

First, we’ll look at an example where this is instantiated, then we’ll go into my current interpretation of each layer in more detail.

Keep in mind that this is probably not fully 100% what Stafford Beer talked about so it is more a model inspired by his VSM rather than an exact replica. He was quite the sophisticated individual.

Example: An Organisation

Stafford Beer was focused on operational research organisational cybernetics, that is he tried to deal with business operations and how organisations should control the complexity of the world.

He was quite focused on Ashby’s law of requisite variety, which essentially states that the internal state of a company needs to be as complex as the environment that it tries to regulate. Here we steal one of his examples: If you’re an alien trying to stop a (European) football team that attacks one goal then it is not enough to put up rocks in the way of the team as they will just dribble past it. One of the easiest ways to block them is just to put an opposite football team to deal with the complexity of the situation, that is if faced with variety you need variety to control it. This also relates to the good regulator theorem in how a model needs to be related to the underlying reality to be able to predict it and control it.

But if you’re in a highly complicated business, how do you go about understanding the world? Well as any good computer scientist and business manager knows, you divide and conquer, you set up teams with specific goals for different tasks. But why do you do this? Why is this the way to deal with variety?

We can relate this to teleology or goal-directedness, the easier a goal is to state, the easier it is to follow and as a consequence the error rate goes down. It is easier to track reality if you have multiple sub-systems that are combined compared to a large one as the large one would need very large and sprawling goals to remain one-pointed!

This is partly due to kolmogorov complexity reasons and if you wanna learn more about why this is from a very fundamental perspective you can check out Chris Field’s Physics as Information Processing which makes the argument for stacked regulators in depth.

This leads us to the way that we generally have set up organisations nowadays which I will share as a picture that brings in the different parts of organisation:

This is the same picture as the first picture but instead of having the feedback loops, you see how this is hierarchically instantiated in a business. The operations teams are the different sub-area focused teams with divisions heads on top and executives on top of this. This structure allows for a coherent set of target states to be propagated from the top to the bottom.

(This structure also shows up in biology and brains which I expand on a bit more here)

There is no reason for why this structure can’t be continued in that we can think of sub-agents in the mind that follow a top-down message passing structure and this is partly what the scale-free hierarchical agency agenda talks about.

I will now try to expand a bit on the exact definitions and intuitions for how we can go about formalising each layers in terms of structure and active inference.

System 1:

System 1 is the backbone of our viable systems model, it is the ground floor, the place where we have machinery and robots taking care of building cars or whatever it is that they do. In a software engineering team it is the individual teams: sales, operations, R&D among others.

Essentially these are the building blocks in which the rest are built up from and usually shows up in the organisational charts of systems.

We want to extend this to a more general theory though so we would want this to fit in more places than just the organisational context that it is instantiated within.

So to baselessly extend it we can also imagine that these are the specific processing units that the rest of our system is built upon, the atomic building blocks that are the undividable base components of our system.

In the brain we at least back in the day imagined that these were our different subareas of the brain such as the prefrontal cortex, the amygdala, the hippocampus, the sensory motor cortex. Essentially these are areas with internal consistency that have more or less coherent states.

If we want to go there (which we want to!) We could bring in some inspiration from Michael Levin and say that these are systems who share teleology, that is they have similar target states and problems that they’re solving.

How do you identify these? Well, that is one of the main questions within hierarchical agency, maybe you do it through markov blankets or causal emergence or integrated information or any other number of interesting and strange definitions.

For our purposes we will completely ignore this problem and just state that base-agents exist!

Formally, partition the org graph G into disjoint sub-teams {G₁,…,Gₙ}, each facing its own local environment ℰᵢ. Each team is an active-inference agent in its own right, with its own model ℳᵢ, and it simply minimises its own local free energy:

Mechanically it runs on one quantity, a prediction error, the gap between what the team sees and what it expected and it both updates its beliefs and acts on the world to close that gap

[1]

System 2:

This is the coordination layer everyone finds a little dull — your Asanas and Linears, the shared calendars, the standing infrastructure that keeps teams in step. It has no goal of its own. Its whole job is to stop two teams, each doing its own work perfectly correctly, from wobbling against each other.

Beer calls this oscillation damping, which sounds like a very strange way to describe a shared calendar. But he means it literally and to establish this rigorously I have engaged a specialist.

This expert will go into depth about why specifically it is called oscillatory as it is quite fascinating. The TL;DR is that since it is communication we want to make sure that what is being sent does not enter into a recursive feedback loop where the signal is amplified. Hence we need damping and the way to describe this in physics is through damping oscillations.

Scary Science Attack!

Oh no, reader — it happened again! And this time it has brought a colleague. If you are afraid of the unknown (network science) skip to the exit at the end of the section.

Professor Claudette Claudeson enters the frame


§1. On the instability of bilaterally coupled regulators under transport delay.

One notes, at the outset, that the configuration under discussion is unremarkable. Two regulators, each minimising a local error signal, each observing the other’s state only after a fixed latency τ, and each applying a correction of loop gain g. It scarcely requires saying that stability is not guaranteed by the local correctness of either party. Where the accumulated phase lag approaches inversion and g exceeds unity — that is, where each party overshoots by even a modest safety margin — the coupled system admits solutions of monotonically increasing amplitude. The lay literature, with characteristic imprecision, terms this the “bullwhip effect” (Forrester, Industrial Dynamics, 1961; Lee, Padmanabhan & Whang, Management Science 43(4), 1997). Sterman demonstrated experimentally that competent human subjects reproduce the instability reliably under laboratory conditions (Management Science 35(3), 1989), which I mention only because the author of this weblog appears to find it charming that people are bad at this. The remedy has been known since Smith (Chem. Eng. Prog. 53, 1957): one does not act upon the raw delayed observable. One acts upon a filtered estimate. The vulgar term is “a shared calendar.”

Figure A — two teams, same 3-week delay, same eagerness (gain 1.05: close the gap plus a margin). Left: reacting to raw delayed signals → howl (×2.7 growth). Right: reacting to the published schedule (EMA, ρ=0.3) → settles onto a shared plan.

§2. On pathological hypersynchrony in excitable media.

The naïve reader supposes synchronisation to be desirable simpliciter. This is incorrect, and the counterexample is not obscure: cortical tissue expends considerable metabolic resource on inhibitory machinery whose function is precisely the prevention of global phase-locking. Failure of that machinery is not coordination but seizure (for the relevant controversies, vide Jiruska et al., J. Physiol. 591(4), 2013). Healthy cortex operates in a narrow admissible band — neither incoherent nor locked — a regime for which the empirical signature is scale-free avalanche statistics (Beggs & Plenz, J. Neurosci. 23(35), 2003). The engineering moral is available to anyone willing to state it plainly, which I am not.

§3. On the correspondence between topological and temporal scales.

We proceed to the substantive result. Let a population of phase oscillators evolve under the canonical coupling of Kuramoto (1975; cf. the review of Acebrón et al., Rev. Mod. Phys. 77, 2005, which the author of this weblog has, I am given to understand, “skimmed”). Where the underlying graph possesses nested community structure, relaxation to the synchronisation manifold does not proceed uniformly. It proceeds in stages: densely intraconnected subgraphs entrain first, superordinate groupings thereafter, the global manifold last — and the characteristic timescale of each stage is governed by the corresponding gap in the spectrum of the graph Laplacian. This is the content of Arenas, Díaz-Guilera & Pérez-Vicente, Phys. Rev. Lett. 96:114102 (2006), whose title — Synchronization Reveals Topological Scales in Complex Networks — states the matter with a concision I would not attempt to improve upon. The organisational chart, in short, is recoverable from the clock.

Figure B — Kuramoto on 4 teams × 16 people in 2 divisions (p_intra 0.9 /​ p_module 0.12 /​ p_super 0.005, K=0.02, 10-seed ensemble). Pairwise phase coherence by structural tier: within-team locks t≈20, across-teams-same-division t≈60, across-divisions t≈600 — read against the Laplacian spectrum’s two gaps. The transient negative dip of the cross-division curve is real dynamics and quietly makes the point: the level above can’t settle until the levels below have.

I am obliged to correct a conflation common among enthusiasts. The stratification above is a consequence of modularity. It is not a consequence of degree heterogeneity, which is a distinct property governing the order of entrainment (hubs precede peripheries) and the critical coupling Kc at which entrainment occurs at all. Empirical organisations exhibit both. They are not the same claim and should not be advanced as though they were. I make no accusations.

A further generalisation exists — the inertial, or second-order, model, in which oscillators possess mass and the synchronisation transition acquires discontinuity and hysteresis (Tanaka, Lichtenberg & Oishi, Phys. Rev. Lett. 78, 1997; Filatrella, Nielsen & Pedersen, Eur. Phys. J. B 61, 2008; for lattice behaviour, Ódor & Deng, Entropy 25(1):164, 2023). I understand it is the subject of the author’s own current researches. I shall not be commenting on those.

§4. A speculation, advanced with appropriate reluctance.

It is hypothesised — I stress the mood of that verb — that neural systems derive functional advantage from operation proximate to a critical point, dynamic range being maximised thereat (Shew et al., J. Neurosci. 29(49), 2009; Shew & Plenz, The Neuroscientist 19(1), 2013). Meisel and colleagues report that the electrophysiological signatures of criticality degrade under sustained wakefulness and are restored following sleep (J. Neurosci. 33(44), 2013), a finding not inconsistent with — though by no means establishing — the proposition that slow-wave activity performs a global retuning function upon the network. Related homeostatic accounts exist (Tononi & Cirelli, Neuron 81(1), 2014). Whether this constitutes a System 2 operating at the scale of an entire organism is a question I decline to dignify. The author, I am told, finds it “neat.”


▌ Exit — you survived the Scary Science Attack.

Thank you, Professor. In human:

Two teams that are each doing their job perfectly correctly will start oscillating against each other if they’re working from stale information and overcorrecting a bit. The fix isn’t to make either team better — it’s to put a slow shared reference between them that filters out the fast wobble. That’s what a standard is. And when you run the same coupling over a whole nested org, the levels lock in at different speeds, which means the shape of your org chart shows up as a ladder of timescales. Structure in space, sequence in time.

Also: too much synchronisation is a seizure. Keep that one.

Formally, and this is the crude, timeless version, blind to every dynamic above it’s just a constraint pulling adjacent teams’ beliefs about their shared variables together, penalising how much team i and team j disagree about the things they both touch:

Pure regularisation for now, I’m looking at this more actively from an active inference lens but we’re currently in complicated land (randomly ended up requiring basic complex analysis) but I will hopefully be able to give a better description of this in a bit.

System 3:

This is middle management, the division heads. They can’t watch every desk, so each team sends up a compressed report and the head allocates resources and sets each team’s targets from those summaries. This is where the hierarchy first goes vertical, and the rule is simple: predictions (targets) flow down, prediction errors (reports) flow up.

Now, the weird thing about system 3 is that it might be able to be part of the underlying hierarchical structure and so for its higher order systems it might also sometimes be acting as a system 1. If you have a boss on top of an organisation which never interacts with the employees then this means that the middle managers shape all the influence that comes from the below teams. Essentially the middle managers then form a markov blanket around the underlying levels which means that there’s no lower layer system 3 for the system in charge one place above.

System 3 essentially regulates the general signals that come from system 4 and 5 and turns them into goals that can be followed. It is the strategising layer which sets the target states for the sub-systems.

A biological example of this is the heart setting the function of the different chambers within the heart which in turn sets the functioning of the individual cells and what they should do.

Now I’m not fully sure of the following formalisation but it makes an attempt at what it might look like?:

System 3 holds a model ℳ₃ of the whole inside, conditioned only on the sufficient statistics sᵢ, and it acts not by command but by setting the teams’ priors:

System 4:

This is strategy, R&D, market intelligence — the part of the company that has stopped looking inward and started looking at the world and the future. Where’s the market going, what’s the competitor doing, is that a threat or a fad. System 3 manages the present; System 4 scouts what’s coming. Or as the chad business consultant would say; “Bro, what’s the SWOT analysis looking like?”

I find system 4 kind of boring right now as it seems just like world model updating and creating general coherence for the entire strategy instead of just the sub-departments doing all of the scouting. At least in model complexity, actually doing it involves forecasting work, strategy work and a bunch of R&D which in itself is quite interesting!

It’s a different kind of inference: instead of updating beliefs inside a fixed model, it updates the model itself, selecting the ℳ₄ that best explains the history of the environment:

I will note that I’m a bit uncertain on how to integrate the environmental variables here, as the direct operational units also interact with the general environment. I would think of these as local environments though and system 4 is essentially focusing on that which falls through the cracks of the day to day operations of the base units.

System 5:

This is the board, the mission, the answer to “who are we as a company.” Not what to do today (that’s System 3) and not what’s coming (that’s System 4) but which outcomes are simply off the table — what makes a firm walk away from a profitable deal because “that’s not who we are.”

System 5 seems to be one of the more foundational parts of any type of system for that is to some extent where the identity sits. Who are you as an organisation? What are your values? What is your culture? If you predict future you from the type of person that you are, what are you likely to do?

It is where goals are set and where prediction error is defined. It is where the utility function would sit and where it would change.

Okay, I will now take this a step further and do some more speculation on this including with some heretical knowledge.

System 5 heresy

Okay, firstly, the heretical knowledge. I sat through a great episode with Karl Friston on the Jordan Peterson podcast and it was great. One of the things that I found really interesting about it was the discussion of stories as one of the generating functions of meaning. That is, they discussed where ought even comes from in the first place.

Bear with me here. How do we set target states for our lives in general? One way that we could look at it is through dreams and stories. They explore a sort of platonic space where lots of different things are possible and in there we experience happiness and sadness and through them we get the idea of what might be possible. These things then set our goals, if you dream about money then you will be sad without having money, if you dream about connection with others then you will be sad if you don’t have connection with others. The hypothesis is that dreams help set our regulatory setpoints as it is like simulating alternative futures.

The dreamspace is only coupled with what is, that is your dreams are dependent on what exists but only so much. There’s a deeper difference between what is and what ought to be. I think it partly also hides in the fact that our dreams are where we generate the ought domain, it is where we generate our baselines.

In a deeper way, I think this is what system 5 partly is. In the larger world it is culture, within your organisation, it is culture, within you it is the culture of dreams that you have.

One of the open questions within active inference and general multi-scale agency is whether you need utility functions in order to go forth in the world or whether you can actually just say that these are generative priors that we change over time.

Partly this is a definitional question and I think one of the things that is problematic with saying that it is just generative priors is that it kind of misses the point that the ought space is quite different to the is space? They can likely be coupled through prediction error as they show up here in the system 5 defining the different sub-parts but they are not the same!

But saying that the utility function is different also means that it cannot evolve over time which is also really strange as your dreams are dependent on your current situation. So to some extent there’s a is-ought loop which evolves over time, so partly I think this is one of the main things we need to study here. What are reflexively stable is-ought loops and which ones are the ones that we want to aim for?

System 5 fixes the deepest priors P(x̃∣ℳ) — the assumptions the organisation makes about itself — and the preferences P(ỹ∣ℳ) — the outcomes it treats as least surprising, which is how goals smuggle themselves into a system that only knows about prediction error. It does this by setting the top-level hyper-parameters:

Conclusion

I’ve gone through my current conception of the Viable Systems Model from Stafford Beer and how it relates to the viable systems model.

  • System 1 is the base level operations, the atoms of organisation whichever they might be.

  • System 2 is the coordinating component which make sure that the underlying layers oscillate in tune with each other (in a surprisingly technical way!).

  • System 3 is management which takes the general priorities and reshapes them into target states for the sub-systems, (system 1) to take care of.

  • System 4 is the world modelling which the normal day to day operations don’t take care of, forecasts, market research and more that helps shape the world models of the general system.

  • System 5 is where the ought space lies, it is where the target state of the entire system shows up.

Now, to be frank, this is pretty much a quite zoomed out approximation of the real thing. The real thing is a sprawling mess in a strange mind who died 24 years ago and captured through VHS tapes and long books with lots of weird diagrams and references to information theory from the 60s.

I find this quite exciting as it feels a bit like looking into weird hidden knowledge and trying to translate it into something that is relevant and usable. It relates to many interesting concepts where the one I find the most interesting right now is the coupling between the time dimension of control loops and the space dimension of modular sub-systems.

Next I might look into something like variety in more detail as it seems quite important. I’m also wondering whether there’s some sort of time dependent version of the gooder regulator theorem that could be created?[2]

Also, what is going on with teleological systems and Levin’s work in relation to this? What is the consequence on using different system 1 formalisations for the rest of the system? What happens if we use markov blankets? What about causal emergence? All very interesting questions to me at least.

This post is by design incomplete and I will at the end of this journey probably have a lot better of a model of how VSM relates to multi-scale agency.

  1. ^

    I use graphs to describe general collective intelligences, I think this is a good approach as I expand on here

  2. ^

    There’s an inherent bound on the time that it takes for a given network given the information flow and the size since you need prediction errors to flow across the system for everyone to share a model. As a consequence, there should be predictions you can make about the minimum time it takes for all parts of a system to have updated on information. We can re-frame this as “if you want to have a good model of the world, what is the minimum time needed given this rate of information flow across the system? (E.g the fiedler vector)”