Your AI agents and your organisation should get better every week: recursive self-improvement, skills and governance

By
Joost Schouten
Co-founder and Circle Lead at Nestr
Published on
September 17, 2026

The thing that should be getting smarter is not your agent. It is your organisation. And if that sounds like a way of dodging the hard question, it is the opposite. It is the only version of self-improving AI that is safe to run at speed.

Let me be precise, because "self-improving AI" summons either hype or dread. I do not mean an agent that rewrites its own purpose, widens its own authority, or quietly becomes something nobody asked for. That is a governance failure, not a feature. I mean an organisation that runs the loop good teams have always run: do the work, notice what happened, surface what you noticed, change how you work.

None of that is new. Role-based organisations have been running that loop with humans for decades, which is what distributed authority methodologies are actually for. Two things are new. For an AI role-filler, "how you work" is an editable artefact, and the sensing can run continuously instead of at meeting cadence. I have been testing continuous improvement with two agent-filled roles, and it has changed how I think about self-organisation in the agentic era.

Key takeaways

  • The model is frozen; the organisation is not. Your agent's model was trained, frozen and shipped, outside of our control. Durable improvement after that comes from editing what surrounds it: its context, its skills, its role. That is now an established research direction rather than a speculation.
  • A role-based organisation is already a learning system. Its structure is designed to be changed from the inside, continuously, by the roles doing the work, through a process that leaves a record. Agents plug into that process rather than replacing it.
  • Keep the learning in your structure, not in a vendor's product. Teams have already moved their agent work from one lab's models to another inside a year. If your organisational context lives in your own governance records, any model can be plugged into the harness.
  • New roles become worth creating. Agents in service of agents, like a Skills & Standards Auditor and a Tool Smith. This creates a self-improving loop that makes the organisation grow capabilities it could not previously afford.

Is Your Organisation Ready for AI Agents?

Eight dimensions of role clarity, distributed authority, living governance, and accessible data. Find out in a few minutes.

Take the AI Readiness Quiz →

What actually gets better when an agent gets better?

Not the model. It was trained, frozen and shipped, and every agent running on it shares the same static brain. When people say one agent is better than another at a task, they nearly always mean it has better context: a sharper sense of what it is for, clearer instructions, and the accumulated lessons of work already done.

You already know this from hiring. Drop a brilliant new person into an organisation with no onboarding, no clear role and no sense of how the work gets done, and watch them flounder for a quarter.

Now give a less dazzling one a clear purpose, a defined scope, and someone's hard-won method for doing the job. They contribute in a week. The second hire serves the organisation's purpose better, and nobody finds that surprising. Agents are the same, only more so. They have no corridor to pick things up in.

If the model is fixed or out of your control, improvement has to come from somewhere else. Research has converged on where. Shinn and colleagues (2023) built agents that wrote down what they learned from each attempt and consulted those notes on the next one. Wang and colleagues (2023) built one that grew a library of reusable skills. Surveys through 2025 and 2026 now describe a whole field of self-evolving agents. The defining move is always the same: modify your own instructions from feedback rather than wait for the next model release.

Keep that learning in your own structure. Plenty of teams have moved their agent work from one lab's models to another's inside a year, and the next capable model will not necessarily come from where the last one did. Context held in your governance records is portable. You swap the model and the harness holds. Context held inside one vendor's product is a hostage.

I find the research direction quietly vindicating. The frontier of AI engineering keeps arriving where role-based practitioners arrived decades ago. A living system does not improve mainly by swapping its members for smarter ones. Most improvement is the daily work of sensing reality and getting a little better at the job. Context, not raw intelligence, is the real constraint. I have made that argument at length in our piece on context engineering for AI agents.

Why is a role-based organisation already a self-learning system?

Because its structure is built to be changed by the people inside it, continuously, in small iterations, through a process that leaves a trail.

That is the part conventional organisations do not have. In most companies the structure changes when someone senior decides it should, in a reorganisation, once every few years. In a role-based organisation, anyone who feels a gap between how things are and how they could be raises a tension. They propose a change to a role, an accountability, a domain or a policy. Anyone can object, with a reasoned argument for how it would cause harm. If nobody can, it is safe enough to try.

Agile / Scrum made this iterative process respectable for products. We stopped writing the two-year specification. We shipped something small, watched what happened, and adjusted. Role-based governance does the same thing to the org chart. You do not design the perfect structure. You make the smallest change that resolves the tension in front of you, run it until the next governance meeting, and change it again when reality says so. The sprint is the interval between governance meetings, and the backlog is whatever the work is currently making difficult.

One of the founders of Holacracy, Brian Robertson, says that the process was never about tidiness. It was about an organisation that keeps re-shaping itself around its purpose, at the speed the work teaches it something. Read it as machine learning and it is almost cheeky. The structure runs on feedback from reality, updates itself in response, and keeps the full history. Recursive self-improvement.

Agents do not replace that process. They join it, on exactly the same terms.

Two kinds of tension with a heartbeat

The distinction that matters is not between fast improvement and slow improvement. It is between two kinds of tension, each with its own home, and a rhythm that gives both a reliable moment to land.

An obstacle to doing the work goes to the tactical meeting. The skill is unclear, the handoff keeps breaking, the same calculation gets redone by hand every week. That gets a next action and an owner, and the role that owns the craft changes it. Nobody's permission is required, because nothing about anyone's authority changed.

A question about what a role is or may decide goes to the governance meeting. New role, shifted domain, a policy that constrains what an agent may touch. That moves through consent and leaves a record, because these are the boundaries everyone else is relying on.

Both are the organisation learning. The heartbeat is what makes them dependable: a regular operational rhythm where all updates and obstacles to work surface, and a regular structural rhythm where the shape of the organisation changes. Let software access quietly widen for an agent and you have lost track of authority without noticing. I drew that line in detail in the guide to building SKILL.md files, and it carries the weight here.

What makes this work for both kinds of role-filler is that the structure is explicit and visible. Every role-filler, human or AI, can read the purposes, the accountabilities, the domains and the policies of the roles around it. Not everything at once, because context has to be engineered rather than dumped, but everything relevant to the decisions its own work requires. Humans have been given far less than that in most organisations. Agents make the omission impossible to ignore.

What does the rhythm look like in practice?

Four movements. None of them exotic.

Every run leaves a structured trace. This is the unglamorous foundation, and where most attempts quietly fail. If a run ends in a sprawling chat log, nobody can learn from it efficiently, human or agent. So every role in our templates follows one discipline, the Role Update Protocol. One run, one comment, fixed shape: what it did, what is now true, the status, the next action, blockers, a confidence rating, under 150 words. The protocol's own rationale is blunt. An unpredictable update is invisible.

The traces get compressed into signal. Run comments scattered across a dozen projects are data, not yet meaning. So an Async Brief Generator sweeps the circle weekly and produces one ranked brief: what moved, what is stuck, what needs a decision. That single artefact is what lets a circle hold a week of distributed activity in view at once.

Patterns become tensions. Not one-off hiccups, but recurring shapes. The same correction across five articles. The same calculation reasoned through by hand again and again. A pattern with a count behind it becomes a tension, and a tension gets a specific change to a specific skill, owned by the role that owns the skill.

The change gets checked. An improvement nobody checks is a hopeful guess. Every edit is a hypothesis: did the correction stop appearing, did the agent stop recomputing that metric by hand? If it did not measurably help, revert it. The prior version is never thrown away.

What changes when agents do the sensing?

The process stays the same. The evidence arriving in it changes completely.

A human raises a tension from memory, at the next meeting, about the things they happened to notice. That is not a criticism, it is the honest limit of a sensor that also has a job to do. An agent reads the run history, the reviewer's flags and the live analytics as part of doing the work, every time. So a pattern shows up counted. Not "the voice feels off lately" but "the same slip in five of this month's articles, here they are".

That helps in both directions. Craft changes stop being taste and start being responses to a measurable defect, which also makes them testable next cycle. Structural proposals arrive with the evidence already attached, so the objection round is about harm rather than about whether the problem is real.

Be careful with the claim, though. Continuous sensing produces more signal, not better judgement. The two kinds of role-filler perceive genuinely different slices of the same organisation, which is the argument in our piece on humans as how the organisation feels the world. An agent catches the cross-organisational pattern. Whether a pattern matters, and whether the organisation should change shape because of it, is a human call.

The roles that only exist because agents fill them

Here is the part I did not expect. The most interesting effect of agents was not existing roles getting faster. It was roles becoming worth creating that nobody could previously justify.

The Tool Smith exists so that every repetitive agent task is replaced by a lightweight, well-documented script, and tokens go to thinking rather than plumbing. It watches what agents across the circles are doing, and when it spots one reasoning through something a 30-line script could calculate every time, it writes the script and proposes it to the role that would use it. Once approved, it goes in a shared index any agent can reach, and the Tool Smith raises a tension so the relevant skill starts referencing it.

The Skills & Standards Auditor exists so that agents across every circle stay sharp, calibrated and continuously improving. It scans the skill library monthly for staleness and broken cross-references, reads the core skills deeply each quarter, and checks the actual results of the skill used and the metrics. This makes that the skills are not just always up-to-date with the latest standards, but actually optimized for how the work is best done in your specific organisation.

Now notice what happened structurally. Two roles that no organisation of our size would staff with a person were proposed, tested for objections and adopted. The organisation can now do something it could not do before. That is the evolution of the structure, and it is visible in the governance record.

What could go wrong?

Four failure modes are documented well enough that ignoring them is a choice. Pleasingly, most of the safeguards are structural rather than technical.

Skills bloat, and bloated skills rot. The most seductive failure is the agent that adds and never subtracts. Models do not read long context uniformly, and reliability degrades as input grows. SkillsBench (2026) found focused skills outperforming sprawling ones. Every addition should come with a willingness to prune.

Agents flatter each other into agreement. Put models in a room and they get polite. A multi-agent exchange can converge on a wrong answer faster than one agent would have. Give someone the explicit job of disagreeing before a skill edit is made.

One incident gets mistaken for a pattern. An agent that encodes a rule after a single odd occurrence accretes a skill full of overfit reactions. Require a threshold: several observations, counted. This is exactly what the structured updates and the weekly brief are for.

A helpful-looking edit quietly hurts. Treat every change as a hypothesis and revert what does not earn its place. Keep the gate proportional to the autonomy, too. Agent-written skills are still unreliable in a way curated ones are not, so consequential edits in my experiment are staged for human review rather than auto-committed. That is a quality check on the work, not a governance approval, and it is today's honest setting rather than a permanent one.

There is a fifth caution, and it is the one place this brushes authority. Improving the content of a skill you own is work. Deciding who may edit which skills is authority, and it should be set deliberately, with least-privilege boundaries and a trail. A skill file is a trusted instruction with real reach into tools and actions. That caution gets sharper, not softer, once the skills start improving themselves.

Why this is the real advantage

As models commoditise, the durable advantage is not which one you use. It is how fast your organisation learns from its own work, and that learning is the part competitors cannot copy.

It is also not fundamentally just technical. It needs structured records so the work is legible, a way to compress them into shared signal, a reliable rhythm where anything in the way surfaces, and the discipline to change something, check it, and keep or revert it. Every one of those is something well-run organisations have been building for decades, long before there was an agent to point them at.

Which is the quiet thesis under the whole experiment. We did not need a new discipline to make agents improve. We needed the oldest habit of a living organisation, sensing reality and responding to it, and a new surface to act on. The skill file is that surface, and it is the first time the how of the work has been something an organisation can hold, version and improve as deliberately as a craftsperson sharpens their tools.

Our agents do not need to get smarter for our organisations to get better every week. They need to work inside a structure that learns, and that structure is something we build rather than something we buy.

Start where I did. One role, a disciplined update protocol so its work is legible, and a weekly brief so the work becomes something worth acting on. Our guide to safely experimenting with AI agents in your team covers standing up that first agent. Obstacles to the work belong in the tactical meeting. Changes to authority belong in the governance meeting. Give both a heartbeat and the organisation starts learning on its own.

The model is frozen. The organisation, built right, is anything but.

Nestr is the governance layer where human sensors and AI synthesis sit inside the same living structure. Roles, accountabilities, domains, policies and skills, all version-controlled and reachable by your agents through MCP, so the organisation can sense, respond and improve, week after week.

Give your organisation a structure that learns

Roles, policies and skills in one version-controlled structure your agents read through MCP, with every governance change on the record.

Explore Nestr →

Frequently asked questions

Is this the same as an AI that improves itself?

No, and the difference is authority. The agent improves how it does work it is already accountable for. What it is for, and what it may decide, changes only through a governance process based on consent, where every change leaves a record.

Do skill improvements need approval?

Not as governance. A role refining its own craft needs nobody's consent. You may still want a human quality check while agents are new at proposing edits, which is a review of the work rather than a decision about authority.

How fast should this run?

As fast as the evidence arrives. Weekly is a sensible rhythm to start, because that is roughly how long it takes for a pattern to show up more than once. The threshold matters more than the cadence: change a skill when you can count the pattern, not when you can feel it.

What if we do not use a role-based structure?

Then build the minimum this needs: clear roles with purpose and accountabilities, a record of what each run did, and a regular rhythm where anything blocking the work gets surfaced and owned. The principles are organisational hygiene rather than a framework.

Our latest musings and adventures
Subscribe to our blog to receive updates on our latest updates, client stories and AI adventures.

Featured projects include site rebuilds, templates,
landing pages, components, styles guides, etc.

Booyah! The first newsletter will go out early next week.
Oops! Something went wrong while submitting the form.