Skip to main content

Multi-agent co-design and polycentric governance of open educational commons – a living lab model

Contract Theory in Practice, Elinor Ostroms 8 Principles and Michael I. Jordans  Multi-Agent Reinforcement Learning (MARL) remixed in practical living labs for educational game practice design.

André Boeing, Sept 26th 02026

Since 2022 I have been experimenting with an alternative model for developing Open Educational Resources and programs in multi-stakeholder projects with municipal, governmental, non-profit, non-governmental, freelancers, people with disabilities, youth, elders and student agents. 

I have never really formalized the model and just published some graphical sketches with loose notes 2 yeas ago. Inspired and excited by the rich Open Course ScaDaMaLe (Scalable Data Science and Distributed Machine Learning by Dr. Raazesh Sainudiin ) I have finally received buildings blocks and “Eurekas”. This now helps to develop a descriptive language for the model in which we have been developing our programs in practical living labs over the last 3 years. What applies to human agents in the following chapters seems to apply to digital machine agents as well. This does not seem strange to me – as “AI” models originate from Human cultural, social, collective Intelligence and maybe even as a “Faculty of Agency itself!” –  “Natural” and “Artificial” in Collective Intelligent Systems alike while we continue to explore, what “Intelligence” really is in essence..or could be…

Based on decades of experience in developing educational programs and formats, I found the classic Top-Down Hierarchical principle dominated model and deterministic, centralized control system to be inefficient, slow and  especially preventing innovation projects in Education for Sustainable Development (ESD) from agile flourishing and rapid adaptation when conditions naturally changed during living lab developments in realtime. 

Decentralized, adaptive, and feedback-driven systems helped us to move from optimization under monthlong planned certainty (which often failed, as “certainty” often showed itself as “wishful thinking” or even “hallucination”)  to learning under uncertainty (which in general is a very helpful resilience skill.) Michael I. Jordans Multi-Agent Reinforcement Learning (MARL) finally gave me one key context for my insights from Practice: Stakeholders including “target groups” as Agents optimized shared objectives which resulted in temporal relational stewardship over transactional control (ongoing in our running live living lab project @ fsysgame.org)

And then Elinor Ostrom’s Work pushed my buttons in a good way and opened me up for the “missing in plain sight puzzle piece”.  In my usual euphoric, positive and sometimes a bit too euphemistic side, my model assumed harmony – just because the model is beautiful. But Ostrom also emphasizes how disputes are handled! And disputes, conflicts, frictions DO come up, when collaborating over many months with diverse, multi-stakeholder agents in intense production sprints! Multi-stakeholder systems DO imply epistemic diversity and methodological divergence as a Feature! Our own home planet with it’s marvelous, homeostatic, biodiverse system is the perfect prove for a regenerative incredible long living system. We do good, to relearn from it. Real Living labs with real impact value ain´t created behind desks or in sterile labs!

So – Friction is not a failure; it is a signal that the commons is adapting to complex, pluralistic contexts. A conflict resolution system provides a transparent, graduated, and non-punitive pathway to resolve disagreements, align competing narratives, and adjust contributions without breaking relational trust or halting co-creation.

From what I have now learned we could treat conflict as predictive error (Jordan/MARL), institutional stress (Ostrom), and incentive misalignment (Contract Theory) and could focus on proportional response, evidence-based adjustment, and relational Healing!

A common reaction I have noticed among many multi-stakeholder agents during some projects and years is a strange believe, which says something like this:  “Conflict and Friction are failures and cost lots of energy + stress, and they were also not part of our agreement plan. The way to solve this, is to not talk about it with a shared commitment to non-communication in this regard. Time eventually makes the conflicts obsolete. Latest after the project is being finished and conflict parties anyway part ways. (even before any impact goals beyond funded project time range could be achieved).

Elder Elinore Ostrom offers a gradual pathway for these shadow yet natural sides of collaboration:

🪜 The 4-Tier Graduated Pathway

Tier
Example Trigger
Process
Resolution / “Sanction”
Timeframe
1. Direct Peer Resolution
Minor PR friction, attribution confusion, wording/narrative tension, formatting disputes
Contributors resolve in PR comments or quick sync. Focus: clarify intent, align on ESD framing.
No formal sanction. Logged in PR. Merge proceeds.
0–7 days
2. Working Group Review
Unresolved pedagogy/ESD framing, competing priority claims, metric disagreements, repeated tier-1 stalls
Lightweight panel: 1 maintainer + 2 stakeholders (teacher/freelancer/city) + facilitator. Reviews qualitative guide data & usage metrics.
Graduated: Time-bound suspension of merge rights for affected branch. Priority reduced for 1–2 cycles. Mandatory re-review.
7–14 days
3. Community Assembly
Systemic friction: narrative alignment (e.g., growth vs regenerative), scaling bottlenecks, repeated tier-2 failures, cross-departmental misalignment
Monthly Living Review. Uses consent-not-consensus model. Public patch log records rationale. Impact dashboard reviewed.
Adaptive: Version branch allowed. Core commons stays intact. Divergent paths grow separately with clear provenance.
Monthly
4. Controlled Fork / Version Split
Irreconcilable values, legal/licensing ambiguity, chronic free-riding, pedagogical drift beyond calibration range
Fork with explicit PROVENANCE.md & CONTRIBUTING.md reference. Commons remains open. Divergence is documented, not erased.
Proportional: Contributor retains CC-BY-SA rights but loses co-development priority until alignment re-established via Tier 2 review.
Last resort

Note: “Sanctions” = temporary loss of influence + mandatory re-alignment, not banning. The goal is systemic convergence, not compliance.

Lets sum up with the three ingredients and components for the model: “Contract Theory in Practice”, “Elinor Ostroms 8 Principles” and “Michael I. Jordans Multi-Agent Reinforcement Learning (MARL)”

1. Contract Theory: From IP Control to Open Relational Pacts

Traditional OER projects often collapse under the tension between open licensing and closed funding/contracting. Our experimental model solves this by shifting from transactional contracts to relational contribution agreements:

  • Incentive Realignment: Creators aren’t paid for exclusivity; they’re incentivized by pedagogical impact, network access, reputation, co-development priority or just a “good time during Playshops and Fun”.

  • Monitoring via Usage & Peer Review: Instead of compliance audits, we began to track open metrics: module adoption, teacher adaptation rates, player engagement, etc. in our latest fsysgame.org project.

  • Dynamic Obligations: Contracts become living contribution charters. If  – for example – a freelancer’s game module underperforms pedagogically, the network adapts (patches, co-revisions, role shifts) rather than terminating agreements.

2. Ostrom’s Principles: Governing the Educational Knowledge Commons

Our OER/ESD game suite is a commons. Ostrom’s 8 principles map directly to sustainable open education ecosystems:

Ostrom Principle Our Implementation

Clearly defined boundaries

Creative Commons licensing + project scope + contribution guidelines

Rules matching local conditions

Modular game design: core ESD mechanics + regional/contextual scenarios (e.g., local water systems, urban mobility, circular economy pilots)

Collective-choice arrangements

Co-design workshops with teachers, students, NGOs, and city depts; voting on feature priorities & adaptation guidelines

Effective monitoring

Open analytics dashboards, classroom feedback loops, peer review of pedagogical efficacy

Graduated sanctions/reputation

Contribution weighting, co-development priority, public attribution, not punitive clauses

Conflict resolution

Community patch notes, transparent design rationale, mediated scenario disputes (e.g., competing ESD narratives)

Nested enterprises

Classroom → school → city education network → national OER repository → EU ESD framework

Recognized rights to organize

Teachers/freelancers/NGOs retain IP in their contributions while embedding them in the open commons

👉 Why it works: OER/ESD games fail when treated as “products.” Our model treat them as living institutional infrastructures that self-correct through use.

3. Prof. Michael I. Jordan’s Lens: Games as Multi-Agent Learning Systems

Jordan’s work on probabilistic graphical models, multi-agent systems, and algorithmic governance reveals why our model is somewhat computationally elegant for ESD:

Concept How It maps to our Projects

Multi-Agent Reinforcement Learning (MARL)

Each stakeholder (teacher, student, freelancer, city dept) is an agent optimizing partial objectives. The game mechanics + feedback loops approximate shared policy iteration.

Online/Continual Learning

Player responses, classroom adaptation, and teacher feedback serve as real-time reward signals. Version updates = policy adjustments.

Federated Development

OER versioning mirrors federated learning: local nodes (schools, NGOs) adapt core modules without centralizing control.

Shared Reward Function

In education, this isn’t profit—it’s systems literacy, behavioral shift, and sustainable decision-making. Our model still needs to formalize this as measurable indicators.

Hierarchical/Policy Abstraction

Analogous to hierarchical RL: low-level mechanics (game rules) → mid-level pedagogy (teacher facilitation) → high-level ESD outcomes (policy alignment, community action).

Raaz pointed me to a chapter, from which he thinks it might realllly interest me. And holy, has he been right. Remember my increasing discomfort with “classic Top-Down Hierarchical principle dominated models and deterministic, centralized control systems”  because I found them to be from practice experience inefficient, slow and  especially preventing innovation projects in Education for Sustainable Development (ESD) from agile flourishing and rapid adaptation when conditions naturally changed during living lab developments in realtime?

I let this post end with a beautiful, both scientific and poetic question and quest Raaz invites us into within prACTical group projects. Next week I can tell the story of happened next…in real, physical, rapid micro living labs way up the Northern lands…


7. The IAS ladder: Part V hook 

This section is clearly separated from the course exposition above and is intended for students
continuing into Part V (advanced seminars / group projects).

The IAS (Identityless Agent Substrate) architecture models the operator–agent pairing as a
repeated principal-agent relationship. In each session:
• The principal is the human operator who cannot directly inspect the agent’s internal state
— only the observable outputs (commits, tool calls, design documents).
• The agent is the IAS-governed agent incarnation, whose behaviour is a function of its
specification and lineage.
• The type is the agent’s behavioural distribution over tasks — how reliably it follows the
specification under the operator’s prompting style.

Across sessions, the principal is learning the agent’s type. Classical contract theory says the
principal’s posterior over type is the sufficient statistic for optimal contract selection once it has
been computed. The IAS question is: How do you compute and propagate this posterior when
agent incarnations are ephemeral and stateless?