MSc thesis · AI & Governance, Vrije Universiteit Amsterdam · Submitted August 2026

The Agentic Artificial Bureaucrat

Agentic AI is increasingly described as capable of planning, coordinating and acting across multi-step workflows. In public administration, this means AI may do more than support an individual decision. It may organise parts of the process through which decisions are prepared, routed, reviewed and carried out.

My thesis examines how public-sector-facing advisory documents imagine this emerging administrative actor. It asks what roles and capabilities are attributed to agentic AI, what work and authority may be delegated to it, and how control, oversight and accountability are expected to operate.

The analysis shows that agentic AI cannot be understood simply as a technical tool placed inside an otherwise unchanged organisation. Its behaviour depends on a wider arrangement of models, tools, information, permissions, workflows and human responsibilities.

Human oversight is part of that arrangement, but the phrase conceals materially different forms of involvement. Approval, monitoring, exception handling, takeover, audit and nominal human responsibility do not provide the same kind of control.

Public administration depends on delegation. Tasks and authority are distributed between institutions, officials, contractors and technical systems. Agentic AI introduces a new possible participant in these arrangements: a system that may pursue objectives, select actions, use tools, coordinate tasks and continue operating with varying degrees of human involvement.

This creates difficulties for familiar ideas about oversight and responsibility.

A human may remain formally responsible without having the information, authority, time or practical means needed to govern what the system does. A final decision may be reviewed by a person even though earlier agentic actions have already shaped the available evidence, options or route through the process. Controls described as "human in the loop" may therefore represent very different distributions of judgement and authority.

The relevant question is not simply whether a human remains involved. It is what the system is configured to do, which consequential choices it can make, and whether the surrounding human and organisational arrangement can meaningfully direct, inspect and correct its behaviour.

How do public-sector-facing advisory documents characterise agentic AI as an administrative actor, and connect its configuration to delegation, discretion, control and accountability?

The analysis examined four related questions:

  • How are the roles, capabilities and architectures of agentic AI characterised?
  • What tasks, authority and discretion are delegated to these systems?
  • How are configuration and anticipated failures connected to control, oversight and accountability?
  • How do these accounts differ between commercial, public-interest and official producers?

The thesis uses qualitative documentary analysis to examine 17 public-facing advisory documents about agentic AI in or relevant to the public sector.

The corpus contains material produced by:

  • commercial consultancies and technology organisations;
  • public-interest and intermediary organisations;
  • governments and other official bodies.

More than 1,000 passages were coded across six main analytical domains:

  • the administrative actor being imagined;
  • system configuration;
  • delegation and discretion;
  • risks and failure pathways;
  • control and oversight;
  • responsible actors.

The analysis considered both the prevalence of different ideas and the relationships constructed between them. The aim was not to treat the documents as neutral descriptions of a settled technology. They were analysed as accounts that help define what agentic AI is understood to be, what it may legitimately be asked to do and how responsibility for its actions is allocated.

The documents do not describe one stable form of agentic AI. They describe different configured arrangements with different combinations of capabilities, access, authority and human involvement.

Across the corpus, several broad patterns emerged.

Consequential choices may occur before the final decision.

An agent may select information, decompose a task, choose tools, coordinate other systems, determine whether an exception exists or decide what should be escalated. A human approving the final output does not necessarily review or understand all the choices that shaped it.

Human oversight is highly elastic.

The term is used for arrangements ranging from direct approval to occasional monitoring, exception handling, retrospective audit and general human responsibility. These arrangements give people very different opportunities to influence system behaviour.

Control is often described more abstractly than capability.

Documents may explain in considerable detail what an agent can do while describing the corresponding human or organisational controls through broad terms such as guardrails, monitoring or human oversight.

Formal responsibility does not guarantee governing capacity.

Responsibility is frequently retained by human officials or organisations. Less attention is given to whether those actors have sufficient visibility, competence, authority and practical ability to intervene.

Configuration, risk and control are connected unevenly.

Documents identify many relevant controls, but they do not consistently show which system property or failure pathway each control is meant to address. This makes it difficult to assess whether a proposed control is adequate for the particular arrangement being described.

Producer groups emphasise different concerns.

Commercial documents tend to give greater attention to capabilities, implementation and the reduction of routine human involvement. Official and public-interest documents place more emphasis on accountability, legitimacy, transparency and the continued role of human judgement. Considerable variation also exists within each group.

Terms such as autonomy, memory, tools and human oversight are often used as if each described a single property. In practice, they compress distinctions that matter.

For example:

  • retaining information is not the same as adapting future behaviour through experience;
  • recording an action is not the same as giving a responsible person usable information;
  • providing an intervention mechanism is not the same as giving someone authority and practical ability to use it;
  • an agent recognising the need for escalation is not the same as having a functioning escalation route;
  • assigning responsibility is not the same as ensuring that the responsible actor can exercise control.

These distinctions affect what the system can do, how problems can arise and which controls might work. Governance therefore needs to engage with the configured properties of the system rather than relying only on broad labels.

The thesis exposed a wider problem: technical and governance discussions often describe agentic systems at different levels.

Technical accounts provide detailed explanations of architectures, memory systems, planners, tools and orchestration. Governance accounts discuss autonomy, risk, oversight, accountability and guardrails. What is often missing is a shared middle level for describing the system properties that matter for behaviour and consequences.

My post-thesis research develops this idea into a design grammar for agentic behaviour.

The aim is to identify the distinct constructs needed to:

  • describe an agentic arrangement precisely;
  • design the conditions under which intended behaviour should occur;
  • diagnose why observed or anticipated behaviour occurs;
  • identify how unwanted behaviour can be prevented, constrained, redirected, detected or reversed.

The central governance question then becomes: What does this configuration make possible, and what are the consequences?

This provides a possible meeting point for engineering, governance, behavioural science, law and operations. Each discipline can contribute its own expertise while working from a shared description of the system being designed or governed.