Multi-agent co-design and polycentric governance of open educational commons – a living lab model
André Boeing, Sept 26th 02026
I have never really formalized the model and just published some graphical sketches with loose notes 2 yeas ago. Inspired and excited by the rich Open Course ScaDaMaLe (Scalable Data Science and Distributed Machine Learning by Dr. Raazesh Sainudiin ) I have finally received buildings blocks and “Eurekas”. This now helps to develop a descriptive language for the model in which we have been developing our programs in practical living labs over the last 3 years. What applies to human agents in the following chapters seems to apply to digital machine agents as well. This does not seem strange to me – as “AI” models originate from Human cultural, social, collective Intelligence and maybe even as a “Faculty of Agency itself!” – “Natural” and “Artificial” in Collective Intelligent Systems alike while we continue to explore, what “Intelligence” really is in essence..or could be…
Based on decades of experience in developing educational programs and formats, I found the classic Top-Down Hierarchical principle dominated model and deterministic, centralized control system to be inefficient, slow and especially preventing innovation projects in Education for Sustainable Development (ESD) from agile flourishing and rapid adaptation when conditions naturally changed during living lab developments in realtime.


Decentralized, adaptive, and feedback-driven systems helped us to move from optimization under monthlong planned certainty (which often failed, as “certainty” often showed itself as “wishful thinking” or even “hallucination”) to learning under uncertainty (which in general is a very helpful resilience skill.) Michael I. Jordans Multi-Agent Reinforcement Learning (MARL) finally gave me one key context for my insights from Practice: Stakeholders including “target groups” as Agents optimized shared objectives which resulted in temporal relational stewardship over transactional control (ongoing in our running live living lab project @ fsysgame.org)
So – Friction is not a failure; it is a signal that the commons is adapting to complex, pluralistic contexts. A conflict resolution system provides a transparent, graduated, and non-punitive pathway to resolve disagreements, align competing narratives, and adjust contributions without breaking relational trust or halting co-creation.
From what I have now learned we could treat conflict as predictive error (Jordan/MARL), institutional stress (Ostrom), and incentive misalignment (Contract Theory) and could focus on proportional response, evidence-based adjustment, and relational Healing!
A common reaction I have noticed among many multi-stakeholder agents during some projects and years is a strange believe, which says something like this: “Conflict and Friction are failures and cost lots of energy + stress, and they were also not part of our agreement plan. The way to solve this, is to not talk about it with a shared commitment to non-communication in this regard. Time eventually makes the conflicts obsolete. Latest after the project is being finished and conflict parties anyway part ways. (even before any impact goals beyond funded project time range could be achieved).
Elder Elinore Ostrom offers a gradual pathway for these shadow yet natural sides of collaboration:
🪜 The 4-Tier Graduated Pathway
|
Tier
|
Example Trigger
|
Process
|
Resolution / “Sanction”
|
Timeframe
|
|---|---|---|---|---|
|
1. Direct Peer Resolution
|
Minor PR friction, attribution confusion, wording/narrative tension, formatting disputes
|
Contributors resolve in PR comments or quick sync. Focus: clarify intent, align on ESD framing.
|
No formal sanction. Logged in PR. Merge proceeds.
|
0–7 days
|
|
2. Working Group Review
|
Unresolved pedagogy/ESD framing, competing priority claims, metric disagreements, repeated tier-1 stalls
|
Lightweight panel: 1 maintainer + 2 stakeholders (teacher/freelancer/city) + facilitator. Reviews qualitative guide data & usage metrics.
|
Graduated: Time-bound suspension of merge rights for affected branch. Priority reduced for 1–2 cycles. Mandatory re-review.
|
7–14 days
|
|
3. Community Assembly
|
Systemic friction: narrative alignment (e.g., growth vs regenerative), scaling bottlenecks, repeated tier-2 failures, cross-departmental misalignment
|
Monthly Living Review. Uses consent-not-consensus model. Public patch log records rationale. Impact dashboard reviewed.
|
Adaptive: Version branch allowed. Core commons stays intact. Divergent paths grow separately with clear provenance.
|
Monthly
|
|
4. Controlled Fork / Version Split
|
Irreconcilable values, legal/licensing ambiguity, chronic free-riding, pedagogical drift beyond calibration range
|
Fork with explicit
PROVENANCE.md & CONTRIBUTING.md reference. Commons remains open. Divergence is documented, not erased. |
Proportional: Contributor retains CC-BY-SA rights but loses co-development priority until alignment re-established via Tier 2 review.
|
Last resort
|
Note: “Sanctions” = temporary loss of influence + mandatory re-alignment, not banning. The goal is systemic convergence, not compliance.
Lets sum up with the three ingredients and components for the model: “Contract Theory in Practice”, “Elinor Ostroms 8 Principles” and “Michael I. Jordans Multi-Agent Reinforcement Learning (MARL)”
1. Contract Theory: From IP Control to Open Relational Pacts
Traditional OER projects often collapse under the tension between open licensing and closed funding/contracting. Our experimental model solves this by shifting from transactional contracts to relational contribution agreements:
-
Incentive Realignment: Creators aren’t paid for exclusivity; they’re incentivized by pedagogical impact, network access, reputation, co-development priority or just a “good time during Playshops and Fun”.
-
Monitoring via Usage & Peer Review: Instead of compliance audits, we began to track open metrics: module adoption, teacher adaptation rates, player engagement, etc. in our latest fsysgame.org project.
-
Dynamic Obligations: Contracts become living contribution charters. If – for example – a freelancer’s game module underperforms pedagogically, the network adapts (patches, co-revisions, role shifts) rather than terminating agreements.
2. Ostrom’s Principles: Governing the Educational Knowledge Commons
Our OER/ESD game suite is a commons. Ostrom’s 8 principles map directly to sustainable open education ecosystems:
| Ostrom Principle | Our Implementation |
|---|---|
|
Clearly defined boundaries |
Creative Commons licensing + project scope + contribution guidelines |
|
Rules matching local conditions |
Modular game design: core ESD mechanics + regional/contextual scenarios (e.g., local water systems, urban mobility, circular economy pilots) |
|
Collective-choice arrangements |
Co-design workshops with teachers, students, NGOs, and city depts; voting on feature priorities & adaptation guidelines |
|
Effective monitoring |
Open analytics dashboards, classroom feedback loops, peer review of pedagogical efficacy |
|
Graduated sanctions/reputation |
Contribution weighting, co-development priority, public attribution, not punitive clauses |
|
Conflict resolution |
Community patch notes, transparent design rationale, mediated scenario disputes (e.g., competing ESD narratives) |
|
Nested enterprises |
Classroom → school → city education network → national OER repository → EU ESD framework |
|
Recognized rights to organize |
Teachers/freelancers/NGOs retain IP in their contributions while embedding them in the open commons |
👉 Why it works: OER/ESD games fail when treated as “products.” Our model treat them as living institutional infrastructures that self-correct through use.
3. Prof. Michael I. Jordan’s Lens: Games as Multi-Agent Learning Systems
Jordan’s work on probabilistic graphical models, multi-agent systems, and algorithmic governance reveals why our model is somewhat computationally elegant for ESD:
| Concept | How It maps to our Projects |
|---|---|
|
Multi-Agent Reinforcement Learning (MARL) |
Each stakeholder (teacher, student, freelancer, city dept) is an agent optimizing partial objectives. The game mechanics + feedback loops approximate shared policy iteration. |
|
Online/Continual Learning |
Player responses, classroom adaptation, and teacher feedback serve as real-time reward signals. Version updates = policy adjustments. |
|
Federated Development |
OER versioning mirrors federated learning: local nodes (schools, NGOs) adapt core modules without centralizing control. |
|
Shared Reward Function |
In education, this isn’t profit—it’s systems literacy, behavioral shift, and sustainable decision-making. Our model still needs to formalize this as measurable indicators. |
|
Hierarchical/Policy Abstraction |
Analogous to hierarchical RL: low-level mechanics (game rules) → mid-level pedagogy (teacher facilitation) → high-level ESD outcomes (policy alignment, community action). |
Raaz pointed me to a chapter, from which he thinks it might realllly interest me. And holy, has he been right. Remember my increasing discomfort with “classic Top-Down Hierarchical principle dominated models and deterministic, centralized control systems” because I found them to be from practice experience inefficient, slow and especially preventing innovation projects in Education for Sustainable Development (ESD) from agile flourishing and rapid adaptation when conditions naturally changed during living lab developments in realtime?
I let this post end with a beautiful, both scientific and poetic question and quest Raaz invites us into within prACTical group projects. Next week I can tell the story of happened next…in real, physical, rapid micro living labs way up the Northern lands…
7. The IAS ladder: Part V hook
This section is clearly separated from the course exposition above and is intended for students
continuing into Part V (advanced seminars / group projects).The IAS (Identityless Agent Substrate) architecture models the operator–agent pairing as a
repeated principal-agent relationship. In each session:
• The principal is the human operator who cannot directly inspect the agent’s internal state
— only the observable outputs (commits, tool calls, design documents).
• The agent is the IAS-governed agent incarnation, whose behaviour is a function of its
specification and lineage.
• The type is the agent’s behavioural distribution over tasks — how reliably it follows the
specification under the operator’s prompting style.Across sessions, the principal is learning the agent’s type. Classical contract theory says the
principal’s posterior over type is the sufficient statistic for optimal contract selection once it has
been computed. The IAS question is: How do you compute and propagate this posterior when
agent incarnations are ephemeral and stateless?