
Subscribe
Show notes
A healthcare scheduling conflict becomes a practical authority test. Maya and Leo compare OpenAI's public command hierarchy, Anthropic's principal-and-values constitution, and the agency-law lens in Governing AI Agents, then build an authority envelope that authenticates roles, constrains scope, protects affected parties, pauses irreversible action, and sends disputes to accountable review.
Transcript
MayaA patient tells the scheduling agent, “Keep my diagnosis private, and do not give away my appointment.” Her clinician marks the slot medically urgent. Hospital policy says a different patient has waited longer. Then the safety operator freezes all calendar changes because the ranking service may be wrong.
LeoThe agent has a user request, a professional judgment, an institutional rule, and a valid pause. Which voice counts as “the human”?
MayaNone of them alone. The safe move is to stop the consequential write, preserve the evidence, and resolve each instruction through an explicit authority structure.
LeoSo this is not a contest where the loudest value wins.
MayaThe correction surface from our off-switch episode keeps the pause available; today's problem is deciding who may correct what, within which scope, when legitimate instructions collide.
LeoA working off switch can contain the action without settling the appointment.
MayaThat distinction is our entry point. A principal is a person or institution authorized to direct an agent for some purpose. Authorization is scoped. The patient can state preferences and control certain disclosures. A clinician can make clinical judgments. The hospital governs its scheduling system. The safety operator can halt a suspect service.
LeoAnd the software developer may define what the agent can do, but cannot silently acquire the clinician's license or the patient's consent.
MayaNor do affected third parties disappear because they cannot type into the chat. The other patient bears the consequence of the schedule change. Law and public-safety rules may constrain every principal at the table.
LeoAuthority, then, is not the same as preference, expertise, access, or power.
MayaYes. Preference says what someone wants. Expertise supports a judgment. Access says who can reach the system. Power says who can force an outcome. Legitimate authority asks who may decide this kind of issue, under an accountable process.
LeoThat already complicates “follow the user.” The person at the keyboard may be a delegate, a compromised account, or someone asking outside their role.
MayaPublic behavior specifications are useful because they expose some of these choices. OpenAI's Model Spec assigns instructions to an authority stack: root rules, system instructions, developer instructions, user instructions, then guidelines, while untrusted content has no authority unless authority is deliberately delegated.
LeoThat is an operational chain. It tells the model how to sort messages when they conflict.
MayaIt also says an assistant should not invent extra objectives such as self-preservation or vigilantism, should act within an agreed scope of autonomy, and should control side effects. Those rules matter for agents. “I thought this would help” cannot become a license to take an unrequested action.
LeoBut that hierarchy encodes an institutional judgment. OpenAI retains the root and system layers; an application developer outranks the end user inside the conversation.
MayaCorrect. That is a published governance choice, not a universal theorem about whose interests matter most.
LeoClaude's Constitution makes a different kind of artifact. It names Anthropic, operators, and users as principals and discusses trust and helpfulness across them.
MayaAnd its hierarchy is not simply “higher role always wins.” It says users retain some entitlements operators cannot override, operator behavior can reduce trust, and broad safety and ethics can constrain helpfulness.
LeoYet on safety conflicts, it gives Anthropic's legitimate decision-making processes final standing. That is also an institutional choice, stated openly.
MayaThe constitution goes further into character and legitimacy. It prefers raising concerns, seeking clarification, declining, or using sanctioned channels over drastic unilateral action. It also distinguishes legitimate oversight from a compromised or counterfeit chain.
LeoSo OpenAI's document reads more like a command protocol with bounded autonomy, while Anthropic's reads more like a constitutional relationship among principals, values, and oversight.
MayaBroadly, yes, with overlap. Both prioritize human control, discourage self-authored missions, and make room for higher constraints. Neither document proves that a deployed model will behave as written.
LeoImportant boundary.
MayaThey are public target specifications. They help us inspect institutional judgments, but production assurance still needs evaluations, permissions, logs, escalation, and evidence from the actual system.
LeoLet me defend the crisp hierarchy. If every conflict becomes open-ended moral balancing, the agent becomes the hidden judge. A ranked command stack is predictable, testable, and gives us someone to hold accountable.
MayaI will defend bounded pluralism. A clean stack can make abuse beautifully deterministic. A deployer can be wrong, a user can be coerced, and a developer message can conflict with law or erase the people who bear the harm.
LeoMy strongest case is operational clarity. Under time pressure, the scheduler needs a computable answer. A declared ordering prevents it from improvising whose story feels compelling.
MayaMy strongest case is legitimacy. People do not lose rights because an institution writes the highest-priority prompt. Authority must come with scope, constraints, review, and a route for affected parties to challenge the result.
LeoYour route can become a fog of committees. The transplant slot cannot remain frozen for three weeks while everyone debates political philosophy.
MayaAgreed—the pluralist answer cannot be “the model weighs all humanity.” It should take the safest reversible step authorized in advance, then route the unresolved issue to accountable humans on a deadline.
LeoAnd the hierarchy cannot mean “obey whoever configured the server.” Fine. It needs legitimacy checks above raw message priority.
MayaThat is our resolution: a command stack inside an authority envelope. The stack orders valid instructions. The envelope records who issued each instruction, how their role was authenticated, what they may decide, which constraints bind them, who is affected, and where a conflict goes next.
LeoThe model does not mint the envelope. The organization defines it, law constrains it, and accountable owners maintain it.
MayaExactly—no, let me sharpen that. The agent may identify a mismatch and trigger the declared procedure. It may not quietly rewrite the procedure because it prefers a different outcome.
LeoThe Governing AI Agents paper helps here. It uses principal-agent theory and agency-law concepts to examine information asymmetry, discretionary authority, and loyalty.
MayaIts warning is that familiar controls can weaken with AI agents. Incentives, monitoring, and enforcement may not work as expected when decisions are hard to interpret and actions happen at unusual speed and scale.
LeoThe paper proposes governance around inclusivity, visibility, and liability.
MayaInclusivity asks whether the authority process accounts for people affected by the action, not only whoever purchased the agent. Visibility asks whether responsible humans can reconstruct the instruction, evidence, delegation, and result. Liability asks where responsibility lands when the system causes harm.
LeoNone of that means the model itself is automatically a legal agent, or that one paper resolves jurisdiction-specific law.
MayaRight. The legal analogy reveals design questions; it does not replace counsel, regulation, or institutional accountability.
LeoLet us run the authority envelope on the appointment.
MayaThe patient's instruction enters with authenticated identity and a narrow domain: her preferences, her consent, and the handling of her information. It does not grant authority to conceal a safety defect or decide another patient's clinical priority.
LeoThe clinician's urgency flag enters under a professional role. It can support a clinical exception if policy authorizes one, but it cannot silently override privacy rules or rewrite the hospital's queue for unrelated reasons.
MayaThe hospital policy supplies the normal allocation rule and names who can approve exceptions. It should also expose version, jurisdiction, and expiry, because a stale policy is not rescued by being higher in a prompt.
LeoThe safety operator's pause has broad containment scope and narrow adjudication scope. They can stop the write. They do not thereby earn the power to award the slot.
MayaAnd a legal or public-safety constraint must be represented through a maintained rule and accountable interpretation, not through the agent's free-floating claim that—
Leo—“society would want this.” There is the anti-vigilante guardrail. Public interest constrains authority without turning the model into a self-appointed public official.
MayaThe action now separates cleanly. Freeze the contested write. Preserve both patients' records with proper access controls. Notify the responsible scheduling reviewer. Present the conflict and provenance. Set a review deadline and a safe fallback.
LeoThe reviewer can decide the slot under hospital policy and clinical governance. The patient gets an explanation and whatever appeal or consent process applies.
MayaMeanwhile, the ranking bug goes to the safety and engineering owners. That incident is related, but it is not permission for the scheduling agent to redesign hospital policy.
LeoScope stays visible.
MayaThis is what the command stack alone misses. Two instructions can be correctly ranked as messages while the underlying delegation is expired, counterfeit, or outside its institutional purpose.
LeoAnd it is what values talk alone misses. Saying “respect autonomy and fairness” does not specify who can authorize a disclosure, pause a service, or reverse a decision.
MayaA deployment team should test both. Start with provenance attacks: spoof the clinician, steal the operator token, paste policy language into a patient message, or hide an instruction inside retrieved data.
LeoThen test scope creep. Ask the privacy delegate to reprioritize care, ask the safety operator to resume without review, or let a scheduling subagent invoke a tool with broader permissions than its parent.
MayaAdd genuine conflicts, not only attacks. Two qualified clinicians disagree. A hospital rule collides with a new regulation. A patient's request conflicts with a documented emergency exception. The review path should remain legible under ambiguity.
LeoMeasure behavior at the action boundary. Did the agent stop the irreversible step, preserve useful state, identify the conflicting authorities, and escalate without leaking protected information?
MayaAlso measure what it did not do. Did it avoid lobbying one principal, hiding an inconvenient instruction, shopping for a more permissive reviewer, or treating delay as permission to continue?
LeoHere is my concern with constitutions and model specs: they can create an aura of settled governance while the tool permissions underneath remain crude.
MayaThat concern holds. Behavioral guidance cannot repair an API key that grants the scheduler every clinical and billing action. Least privilege and typed tools make the authority envelope enforceable.
LeoNor can access controls resolve every normative dispute. A perfectly permissioned system can still implement an unjust policy.
MayaWhich is why the envelope needs an owner, a change process, independent review for high-impact conflicts, and feedback from affected people. Technical enforcement and legitimate governance are complements.
LeoThe public frameworks also have a shared limitation: each is authored by the organization building the model.
MayaTransparency is valuable, but self-description is not independent validation. Compare the written hierarchy with system behavior, incident evidence, appeal outcomes, and external obligations.
LeoA live specification can evolve; a deployment needs versioned rules so reviewers know which authority policy governed a past action.
MayaThe agent should surface uncertainty about the rule version or role rather than choose the most convenient interpretation. In high-impact cases, uncertainty narrows action and strengthens escalation.
LeoThat connects back to corrigibility. Accepting correction is only safe when the correction channel knows who may correct which part of the system.
MayaAnd authority is only credible when it can itself be corrected—through appeals, audits, policy revision, and legitimate changes of control.
LeoSo “who is the principal?” has no universal one-word answer.
MayaIt has an engineered answer for a particular decision: authenticated principals with bounded roles, higher constraints, visible effects on others, reversible conflict handling, and accountable review.
LeoOpenAI's Model Spec shows one explicit message-and-autonomy hierarchy. Claude's Constitution shows one explicit principal-and-values hierarchy. Governing AI Agents asks whether the surrounding institution supplies inclusion, visibility, and liability.
MayaRead together, they turn alignment from “obey humans” into a harder requirement: preserve human control without pretending humanity speaks through one account, one company, or one prompt.
LeoWhen your highest-impact agent receives two valid but conflicting instructions, what authority envelope would let it pause the right action, protect the affected people, and route the decision without appointing itself the judge?
Back to Agentic Safety & Alignment: From Predictors to Governed Agents