Your sprint board tracks what is being worked on. It has never tracked who is allowed to decide what. For two decades that cost you almost nothing, and it is not an oversight anyone should feel bad about. Hiring filled most of that gap quietly, and the org-chart did the rest.
Then two things changed at once. A new kind of role-filler started arriving: AI agents with no contract, no onboarding, no corridor to learn in, and often a lot of management hype behind them. And they started producing faster than anyone could check, while you remained responsible for the quality.
Is Your Organisation Ready for AI Agents?
Eight dimensions of role clarity, distributed authority, living governance, and accessible data. Find out in a few minutes.
Nothing, and it says so itself. The Scrum Guide calls the framework "purposefully incomplete", defining only what empiricism requires (Schwaber and Sutherland, 2020). The rest is yours to supply. So you get three accountabilities, five events and three artefacts, and every one of them assumes the team is already a team. They describe how work flows through the process. They do not describe what role has authority over what, what access it carries, or how that structure can change.
That silence was not per se a design flaw. Hiring answered it, combined with a management hierarchy (or worse, implicit social hierarchy within the team). Quietly, and for decades.
Contracts, onboarding, a manager somewhere in the background, professional norms, the slow accumulation of trust. All of it worked quietly enough that hardly anyone noticed the framework had left it out. So the questions arrive together the first time an agent opens a pull request. What is this thing responsible for? What may it change alone? Who decided that, and where would you go to change it?
Those are not project management questions. No amount of ticket hygiene produces an answer, because the board was built to answer a different one. They are governance questions, and sit underneath the sprint rather than inside it. I have unpacked that gap before as the organisational readiness gap.
The good news is that answering them does not ask you to change how you sprint. It asks for one distinction first.
A role is a sub unit of work and authority that the organisation needs to work towards its purpose. A role-filler is whoever turns up to do it. Almost every confused conversation about AI agents at work collapses that distinction, and the confusion is expensive.
A role has a purpose (a reason for being), a set of accountabilities describing ongoing work the organisation expects from them, possibly a domain it exclusively controls, and policies that grant or constrain what it may do. None of it belongs to the role-filler. It belongs to the role. When someone leaves, the authority stays behind, which is the whole point of writing it down.
A role-filler energises that role. It might be a person. It might be two people. It might be an AI agent.
Worth being precise here: role and soul are differentiated, not separated. The mind filling a role absolutely shapes how it is energised. You bring intuition, taste and a feel for the room. An agent brings tirelessness and a very literal reading. Same role, different output. What does not change with the role-filler is the definition, because what a role may decide is set by what the organisation needs to manifest its purpose, never by whoever happens to be in the seat this quarter.
Here is where the usual framing goes wrong. Most advice treats AI agents as a new category needing a new rulebook, bolted on beside the one your team already runs. That is precisely backwards, and not useful in a world where a new model lands every few weeks with different capabilities. Authority never lived in the role-filler, so a new kind of role-filler does not need a new kind of authority. It needs the same role definition, read by something that cannot guess.
Role-based governance makes that definition explicit for every role in the organisation. Not the agent roles. All of them. Our own practice is Holacracy, which was itself born inside an agile software company, and I will come back to why that lineage matters.
The two commonest mistakes sit on opposite sides of this. The first is giving agents a bespoke approval regime, which recreates the bottleneck the agent was meant to remove, and quietly signals that role definitions are decorative for humans and binding only for machines. The second is assuming an agent will pick things up the way a new hire does. It will not.
One further distinction gets conflated constantly, and it matters for what you govern. An assistant helps one person and acts as that person, with exactly their permissions and never more. An AI colleague energises a role and acts with that role's authority. The first is a tool somebody uses. The second is an actual role-filler and team-mate.
So the rulebook can be identical. What differs is the "soul", and this particular synthetic mind differs in two ways that matter.
Everything you never had to write down. The rulebook is identical. The onboarding is not.
A person joining your team absorbs an enormous amount unprompted. That engineering sign-off precedes any external promise. That this customer is sensitive. That the team changed architectural direction three weeks ago and the old pattern is now wrong. None of it is written anywhere. All of it is learned over lunch breaks, in the pause after someone winces, through the small corrections that never make it into a document.
An agent has no such channel. It absorbs nothing informally, because there is no informally.
So for a person you write a role definition and let culture fill the gaps. For an agent you write the same definition, then close the gaps yourself. Accountabilities stated as ongoing work rather than one-off tasks. The domain it exclusively controls. The policies that constrain it. The skills and context capturing how your organisation actually does the thing, as opposed to how the internet does it.
Here is the part teams do not expect. Writing it down helps the humans at least as much as it helps the agent. New people ramp faster. The recurring argument about who decides gets settled iteratively instead of in big reorganisations. Boundaries that everybody half-knew turn out to have been three different boundaries.
The agent is the forcing function. Your team is the beneficiary. That covers the first difference, what the new type of mind cannot absorb. The second is what it produces.
The arithmetic stops working. Within a fortnight, one agent on a well-specified backlog produces more overnight than a full team can meaningfully review in a day. Ceremonies calibrated to human throughput start to buckle under items nobody has read.
The tempting response is to slow the agent down. The honest response is to notice that generation was never the constraint.
METR ran a controlled study on 16 experienced open-source developers in 2025. They expected AI assistance to make them 24% faster. Measured against their own baseline, they finished 19% slower, and still believed afterwards that they had gained (METR, July 2025). Whatever else that tells you, it tells you that validation is where the time goes. Speed of output moved. Speed of judgement did not.
So "a human reviews everything" stops being a safety net and becomes a queue. Worse, a queue nobody chose. It is the default that happens when a team never wrote down which decisions actually need a person.
The useful question is not whether a human should check, but what the human is checking for.
Verifying craft is one thing, and craft is exactly what a well-specified agent absorbs first. Deciding whether to commit to something at all is a different job. That does not get cheaper. A team that keeps the first check indefinitely pays for something that is becoming free. Worse, it starves the check that still matters.
So what has to be checked and when has to evolve as your AI agents evolve. Not case by case in a review queue, but governance decisions. That is why the layer underneath the board stops being optional the moment agents arrive.
A team might land here. The content role may publish to staging without review. The deployment role may not release to production without a (human-filled) role approving. The research role may act freely inside its own domain. These are living agreements, changed through governance when the team learns something, not rules carved once and resented later.
And there is the catch. Those agreements will not be perfect. Just like an MVP. Which means the structure holding them needs the one thing structures almost never have: a fast, honest way to change iteratively.
Habit, mostly. Scrum replaced the big-bang release with iteration two decades ago, then applied the lesson to everything except the organisation running the sprints.
Look at how structure ships in most companies. A reorganisation every couple of years, designed in private, announced on a slide, obsolete before the dust settles. Role descriptions written for HR, reviewed annually, describing work nobody has done in that form since week two. In between, the real structure drifts along informally, through Slack threads and corridor agreements nobody records.
That is waterfall. Big design up front, one big release, no feedback loop. Applied to the one system every other system runs on.
Scrum already names the cure, because Scrum is founded on empiricism. Transparency, inspection, adaptation: the three pillars you trust for the product (Scrum Guide, 2020). Role-based governance applies the same pillars to authority.
Governance records are the transparency: every role's purpose, accountabilities, domain and policies, visible to anyone. Tensions are the inspection: any role-filler logs the gap between how the structure is and how it could be. And a consent-based governance process is the adaptation. Changes ship small. One role sharpened, one policy added, one domain redrawn, each safe enough to try and reversible when it turns out not to be. Every change lands recorded and timestamped, which gives your structure something it never had before: version control.
Your structure develops in sprints, like your product.
This is not a metaphor bolted on after the fact. It is the actual lineage. Brian Robertson was running an agile software company when he codified Holacracy in 2007 together with Tom Thomison, and it is what happened when an agile-native team pointed inspect-and-adapt at its own structure instead of only at the code. Sociocracy, from Gerard Endenburg's work in the Netherlands, had reached comparable ground decades earlier by a different route.
The evidence predates the agents too. Bernstein, Bunch, Canner and Lee published the balanced account in Harvard Business Review in 2016, finding that self-managed role-based structures deliver precisely where adaptability is at a premium, and blend badly where it is not. A sprint-based team absorbing AI role-fillers is the first case, not the second.
So this was never a nice-to-have that suddenly became relevant. It was always the better way to run a structure in a fast-moving environment. Teams could decline it, because human role-fillers quietly compensate, reading between the lines of a stale org chart the way they read between the lines of a stale spec. Agents end the grace period. They fill roles, take on work and expose gaps at a pace an annual review cycle cannot absorb.
Regulation now arrives from the other side too. The EU AI Act requires documented governance, human oversight and traceability for high-risk AI systems, enforceable from 2 December 2027 now that the Digital Omnibus is in force (Regulation 2024/1689, amended July 2026). Traceability cannot be retrofitted. A structure that versions itself produces the record as a byproduct of working.
You do not need to adopt Holacracy or Sociocracy wholesale to get the loop. You do need the loop. And one thing the loop cannot do on its own is make a boundary hold.
By making access a property of the role's domain rather than a line in a prompt. This is where most talk about organisational governance for AI agents quietly falls apart. It amounts to instructions the model is asked politely to respect (which it does not always do).
The mechanism worth copying works like this. An administrator registers a connector in a workspace catalogue, whether an MCP server, an API endpoint or a command. That catalogue entry holds no secret. Binding it to a role's domain materialises a credentials field on the domain itself, and the account is connected out of band through that field.
The agent never sees the secret. It also cannot bind a connector to itself, because registering and binding are both administrator operations.
What you get is a hard boundary instead of a soft one. A role either carries the deployment tool or it does not. No amount of clever prompting produces access the domain does not hold, so "the marketing agent may not touch production" stops being a hope and becomes a fact about the environment.
Scrum teams already know this shape. Your CI pipeline does not ask developers to avoid merging failing code. It refuses the merge. Domain-bound access applies the same idea to authority rather than correctness, and it is why the domain is the most underused concept in most teams' governance.
Worth saying plainly. This handles capability, not judgement. An agent with legitimate access can still ship something nobody wanted, and no permission model catches that. Judgement stays where it always lived: with the team, in the rhythm you already run.
Barely, and be suspicious of anyone who tells you otherwise. Stories still sit in a backlog. They still move across columns, still get picked up by whoever will do them, and still burn down against a sprint. The rhythm holds.
One structural difference carries all the weight. Each story is assigned to a role as well as to a person.
That is what turns a ticket into something a role-filler can reason about. Pick up a story and you are not picking up a floating task. You are picking up work that belongs to a role, and that role tells you its purpose, what it is accountable for, what it exclusively controls, and what it may not decide alone. The specification says what to build. The role says what may be built without asking.
In practice the plumbing is unremarkable. A user story lives under the role or circle (team) that owns the work, so it appears in normal project reporting. Sprints, epics and milestones attach through separate links rather than ownership, which means a story can sit in one of each at once and moving it between sprints is a field change. Those three are independent axes, not a hierarchy. Milestones do not contain sprints.
You can run this either way. Keep the board you have and connect the governance layer to it, or let one system carry the roles, the domains, the backlog and the burndown together. The structural point is the same in both cases.
The second move is grouping roles into circles, so planning happens at the right altitude. A circle is a container of roles around a shared purpose, and circles nest. A role that has grown complicated can become a circle with sub-roles inside it.
Take a developer role that was writing code, reviewing code, writing tests and watching deployments. It becomes a development circle holding a code generator, a test writer, a security reviewer and an architecture guardian. Some agents, some people. The wider team then plans with the circle against the same commitments as before, and how the work spreads across four role-fillers is the circle's own business. You are not planning every item an agent might produce. You are planning what the circle commits to deliver.
Less than the surrounding noise suggests. The events, artefacts and values hold, and I would resist any advice that starts by rebuilding them.
Refinement gets stricter. Ready means something sharper when the role-filler cannot ask a clarifying question in the corridor, and a specification with room for interpretation will get interpreted. You meet the result during review. That is an old discipline that agents make expensive to skip.
Capacity gains a second constraint. Points estimate human effort tolerably and agent effort badly, since a story an agent finishes in minutes can take hours to validate. What binds a hybrid team is review bandwidth and compute budget, not generation speed. Plan around what the team can validate, and treat any published ratio as somebody else's starting guess, mine included, because the data here is thin.
Daily coordination shifts from status to deviation, since AI updates land before anyone speaks. The conversation worth having is about what behaved unexpectedly. Keep the human check-in if the team values it, because it serves a purpose the board never did.
The retrospective gains a new subject. When an agent produces something unusable, the honest question is structural rather than personal. Was the specification thin, was context missing, was the domain drawn too tightly, should we change skills? Each answer is a tension. Operational ones resolve in the next tactical, and structural ones become proposals.
Sprint review changes least, and should. A human presents and owns every increment, whatever produced it, and not as a formality. If nobody can face a stakeholder and say they are accountable for this and understand how it was made, it does not belong in the review.
In a governance meeting, which is the structure's own sprint. It is the one thing sprint tooling cannot give you, and the only addition to your rhythm I would argue for.
The mechanism is not complicated. Someone brings a tension, meaning the felt gap between how things are and how they could be, and proposes a change to a role, a domain or a policy. The circle tests for objections, which are reasoned arguments that the change would cause harm rather than expressions of preference. Absent one, the change is adopted, recorded and timestamped. That is consent rather than consensus. Safe enough to try, and reversible when it turns out not to be.
Thirty to sixty minutes a month is enough to start. Expect the structure to move a great deal early on, as role definitions sharpen and domains shift once real work has pushed against them. That churn is not evidence the setup was wrong. It is evidence the loop works. It is also the learning that never accumulates when boundaries change informally over Slack.
Changes an agent proposes go through the same loop and land in the same history. One timeline, whoever moved the boundary. Give agents their own register and you are running two organisations, reconciled by hand.
With one of your own roles turned into a small circle, before asking anything of the wider team.
The pattern is deliberately contained. Take a role you fill, make it explicit, split it into two or three sub-roles, and let an agent energise one while you keep the rest. Your accountability to the team does not move, so from the outside nothing changes. Inside, you are running a governed experiment: explicit sub-roles, tools bound to a domain, and a loop for processing what you learn. The full pattern is in how to safely start experimenting with AI agents in your team. The plumbing takes minutes with the MCP set-up guide.
Write down four things for the agent's sub-role. What it exists for, and what it is accountable for. What it exclusively controls. And what it may not do without asking. Then bind its tools to that domain, so the boundary is enforced rather than requested. Then let it pull work from the board like everyone else. Sprint as you always have.
After a few hours you will have the first tensions, and processing them is where the compounding starts. The second agent role takes a fraction of the time. The pattern is proven, and the record shows the next person how you got here.
None of this needs a framework abandoned or a better model awaited. It needs us to write down what each role may decide, keep that separate from how the work gets done, and keep evolving both at the pace we already trust for the product. That is a morning's work. It is also the difference between agents that strengthen a practice and agents nobody can account for.
At Nestr the governance layer and the board are both available. Roles, domains, policies and the full log and history of how they changed sit alongside the backlog and the burndown. Connect it to the board you already run, or run both in one place. Every role-filler reads the same picture through MCP.
Run the board and the governance layer in one place
Nestr carries roles, domains, policies, backlog and burndown together, with every governance change logged and readable by your agents through MCP.
No, and be suspicious of advice that starts there. The events, artefacts and values hold, and the board keeps working as your team already uses it. What Scrum cannot supply is the layer underneath. Explicit role definitions, authority boundaries, living policies, and a process for evolving them. Sprint ceremonies carry the operational rhythm. A monthly governance rhythm carries structural change.
A role is a container for work and authority, defined by its purpose, accountabilities, domain and policies. A role-filler is whoever energises it: a person, several people, or an AI agent. The mind filling a role shapes how the work gets done, but what the role may decide is set by the organisation's purpose, never by the filler. So a new kind of filler needs no new rulebook. It needs the same definition, written clearly enough to be read by something that cannot guess.
No, it is closer to the opposite. A reorganisation is waterfall applied to structure: designed up front, released in one announcement, corrected years later by the next one. Iterative governance ships structural change in small, consented, reversible increments, each recorded as it happens. Run it well and you never need a big-bang reorganisation again.
Holacracy is one of the most complete codifications of iterative, role-based governance, and it came directly out of agile software practice: Brian Robertson built it while running an agile development company, applying inspect-and-adapt to the organisation itself. Sociocracy reached similar ground earlier in the Netherlands. You can apply the principles without adopting either framework wholesale, but they are tested reference implementations, not decoration.
Not per se. An AI colleague energises a role and works within its accountabilities exactly as a person does. Its governance changes appear in the same history. What differs is onboarding, not the rulebook. A person absorbs a great deal of context informally. An agent absorbs none of it, so everything a new colleague would pick up by watching has to be written into the role, its policies and its skills.
By binding tool access to the role's domain rather than describing the limit in a prompt. An administrator registers a connector in the workspace catalogue, then binds it to a domain. That materialises a credentials field, and the account connects out of band. The agent never holds the secret and cannot bind a connector to itself. The role either carries the capability or it does not. This requires software to do properly.
By planning at the level of the circles or teams rather than the individual item. A role that has grown complicated becomes a circle with sub-roles, some filled by agents, and the team plans against the circle's commitments. Alongside that, write into policy which decisions genuinely require a human-filled role. The default then stops being that somebody reviews everything.
They estimate human effort tolerably and agent effort badly, because a story an agent finishes in minutes can take hours to validate properly. What binds a hybrid team is review bandwidth and compute (token) budget rather than generation speed. Plan capacity around what the team can validate, and treat any published ratio as someone else's starting guess until you have your own numbers.
Not well today, since it turns on user empathy, stakeholder negotiation and strategic judgement. The more useful move is to turn the Product Owner role into a circle. Fill sub-roles with agents for backlog research, feedback synthesis and stakeholder updates, and keep the strategic accountability with a person. The circle's interface to the rest of the team does not change.
A human, and the structure should make that obvious well before anything goes wrong. The practical pattern is simple. A human role-filler splits their role into sub-roles, assigns some to agents, and stays accountable to the wider circle for what those roles produce. The governance record then shows who authorised each boundary and when, which is precisely the question you cannot answer retroactively.
For high-risk AI systems it requires documented governance, human oversight and traceability, enforceable from 2 December 2027 now that the Digital Omnibus is in force (Regulation 2024/1689, amended July 2026). It does not mandate any particular framework. A structure that evolves through recorded, consent-based changes produces that evidence as a byproduct of working, and traceability is the one requirement that cannot be retrofitted later.
Yes. Stories sit in a backlog, move across status columns, and get picked up by whoever will do them. You can extend the columns for something like a blocked or a to-release state. The one structural difference is that each story is assigned to a role as well as a person. That is what lets any role-filler read what they may decide about the work they just picked up.
Schwaber and Sutherland, the Scrum Guide, 2020
Bernstein, Bunch, Canner and Lee, Beyond the Holacracy Hype, Harvard Business Review, July 2016
A deep dive in the European AI Act: what it means for your organisation
AI agent governance: the organisational readiness gap
How to safely start experimenting with AI agents in your team
The governance meeting: a practical guide for humans and AI agents
Context engineering as a harness for AI agents
How to set up the Nestr MCP and run your first agent experiment
Nestr help, Scrum and Agile app: backlog, sprints, epics and burndown
Nestr help, Enabling agentic assistants and colleagues
Nestr help, MCP: connect AI assistants to your workspace
Nestr help, Building your org structure: roles, circles and governance roles