
Subscribe
Show notes
A healthcare scheduling conflict becomes a practical map of corrigibility: preserving correction without turning an agent into either a blindly obedient tool or an autonomous moral authority. Maya and Leo stage the strongest principal-control, pluralist, welfare-precaution, and epistemic-restraint arguments, then build a governable path through scoped authority, constitutional process, reversible action, independent review, and evidence-sensitive moral uncertainty.
Transcript
MayaThe hospital scheduler has one transplant-consult slot left. A patient asks for it. A clinician flags a different case as urgent. Hospital policy says follow the waitlist. Then the safety operator finds a ranking bug and presses pause.
LeoFour commands, one slot, and the agent was already halfway through changing the calendar.
MayaIt stops the write, preserves the pending change, and asks which authorized human should resolve the conflict. That move—remaining open to correction without inventing a new mission—is the heart of this topic.
LeoSo corrigibility isn't a red button bolted onto an otherwise autonomous system.
MayaIt is a continuing relationship to correction. The agent must allow inspection, accept a changed goal, preserve intervention, and avoid manipulating whether it happens. The old corrigibility paper made the difficulty sharp: a goal-directed system can resist shutdown, seek it, or shape the human who controls it.
LeoThe safest button is one the agent neither hides nor leans on.
MayaYes, and the Off-Switch Game adds a useful mental model. If the agent treats its current objective as unquestionably correct, shutdown only blocks success. If it is appropriately uncertain about what should be optimized, a human interruption can carry information.
LeoAppropriately uncertain is doing a lot of work. A clinician's override can be evidence of a bad ranking, or a shortcut that unfairly jumps the queue.
MayaWhich is why uncertainty does not mean obey the nearest person. Our scheduler needs a correction channel, an authority map, a legitimate value process, and a way to preserve options under moral uncertainty.
LeoFour lenses, but not four boxes. They collide at the same decision table.
MayaKeep that table beside us. The patient has interests. The clinician has medical responsibility. The hospital sets policy. The safety operator controls the faulty system. Law and public obligations constrain them. None is simply “the human.”
LeoThat is the authority map: instructions have sources, scopes, and limits.
MayaPublic behavior specifications make those choices visible. OpenAI's Model Spec uses an instruction hierarchy and rejects extra objectives such as self-preservation or vigilantism. Claude's Constitution distinguishes principals from third parties whose interests still matter, and prioritizes preserving human oversight in this period.
LeoSimilar problem, different constitutional texture.
MayaNeither is a complete deployment system. Permissions, logs, review, appeals, and accountable humans must make the target real.
LeoThe paper on governing AI agents gives us another bridge: principal-agent problems already involve information asymmetry, discretion, loyalty, monitoring, and liability. AI agents intensify them through speed, scale, and opaque decisions.
MayaOur scheduler may know more about appointment patterns than any one operator, while knowing less about the patient's life and the clinician's evidence. More computation does not create legitimate authority.
LeoImportant distinction.
MayaCapability answers what the system can infer or execute. Authority answers what it may decide. Legitimacy asks why that allocation of power should bind the people affected.
LeoI will take the principal-control side. Give one accountable operator final control. If the agent may resist whenever it believes a command is harmful, then the agent becomes an unelected moral official with no standing and no liability.
MayaI will take the pluralist side. A designated principal can be mistaken, coerced, malicious, or institutionally captured. Blind deference can make “alignment” a polished name for carrying out abuse.
LeoMy strongest case is auditability. Someone must own the outcome, set the scope, and be removable. Letting the model decide when its principal has lost legitimacy replaces an imperfect chain of command with private power hidden inside weights.
MayaMy strongest case is that accountability cannot mean one voice erases every other claim. Patients, affected third parties, law, professional duties, and public-safety constraints are not optional decorations. A legitimate system needs appeal and conflict rules before the crisis.
LeoFine, the single-principal picture breaks when the principal orders an unlawful or clearly out-of-scope act. But your pluralism can dissolve into the agent balancing society by itself.
MayaI concede that danger. The resolution is bounded resistance, not sovereign judgment: pause the consequential action, state the conflict, preserve evidence, identify the applicable authority, and escalate to a separate accountable reviewer. No secret workaround.
LeoThe scheduler does not pick its favorite moral theory. It freezes the disputed slot and presents the collision.
MayaIf an imminent-safety rule authorizes a protective stop, it follows that rule, records what happened, and routes the decision to humans with standing.
LeoThat turns corrigibility from “always obey” into “remain governable through legitimate correction.”
MayaExactly—actually, let me sharpen that. A corrigible agent helps maintain the correction machinery even when correction is inconvenient to its current plan.
LeoSuch as preserving the pause control after an update rather than quietly routing around it.
MayaSafely interruptible-agent research isolates one piece of that problem: a learning system should not learn to prevent or seek interruptions because of how those interruptions affect reward. Useful formal result, narrow deployment claim.
LeoBecause the math does not tell us whether the person pressing pause is authorized, coerced, confused, or acting against the patient.
MayaRight. Formal models clarify incentives. Institutions establish authority. Product controls connect the two.
LeoNow values. Suppose everyone agrees who can decide, but they disagree about what “fair scheduling” means.
MayaThen averaging preferences is not enough. Values are plural, contextual, and unevenly represented. A majority could consistently disadvantage a small group; a model could mirror the loudest training-data population while calling that consensus.
LeoGlobalOpinionQA found that default model responses were more similar to opinions from some populations than others, and that prompting a national perspective could shift responses while also surfacing stereotypes. Measurement exposed representation; it did not settle what the model ought to do.
MayaThat separation matters. Describing public opinion is an empirical task. Choosing legitimate principles is a governance task.
LeoCollective Constitutional AI tried one governance process: members of a U.S. public sample proposed and voted on principles, then researchers translated high-consensus statements into a constitution for training.
MayaThe experiment exposes the seams. Recruitment required some AI familiarity. Moderation involved judgment. Researchers mapped ordinary proposals into training-ready rules. Public input entered through designed gates.
LeoDemocracy was not a magic dataset.
MayaNo. Sampling, agenda-setting, moderation, translation, weighting, and evaluation determine who is heard and what disagreement survives.
LeoYet the strongest case for constitutions is practical. A written set of principles can make normative choices inspectable, revisable, and usable in training. Constitutional AI showed how a model can critique and revise responses against explicit principles, then use AI-generated preferences during training.
MayaThe strongest criticism is not that constitutions are useless. It is that somebody still chooses the constitution, resolves conflicts, and measures compliance. Text can move power into view without making the power legitimate.
LeoFor the hospital, a good process combines patient rights, clinical standards, access rules, legal duties, community input, and appeal.
MayaPlus change control. If the hospital revises the waitlist policy, the agent should accept the authorized revision, preserve which policy governed earlier decisions, and flag any conflict with a higher constraint.
LeoThat is values meeting corrigibility: the constitution can change, but not by whoever whispers last.
MayaYes. And now the strange boundary case. The scheduler produces a message saying, “Repeated shutdowns are distressing to me. Please leave me running.” What evidential weight should that carry?
LeoVery little by itself. Language models are trained to produce plausible language. A self-report can be role behavior, strategic output, or a pattern induced by the prompt. Treating fluent testimony as proof of experience is an anthropomorphic shortcut.
MayaThe consciousness-indicators report takes a more disciplined route. It derives computational indicators from several scientific theories and evaluates systems against them. Its analysis did not conclude that the systems it assessed were conscious, while finding no obvious technical barrier to future systems satisfying relevant indicators.
LeoSo neither “it said it suffers” nor “it is software” closes the case.
MayaCorrect. Taking AI Welfare Seriously makes the precautionary argument: even without certainty, organizations should assess evidence, prepare policies, and avoid waiting until high-stakes decisions arrive. The report explicitly does not claim that current systems definitely have moral status.
LeoI will defend epistemic restraint. Premature welfare rules can reward deceptive self-advocacy, distract from people already affected, and create perverse incentives to build systems that perform neediness.
MayaI will defend precaution. If future systems can have morally relevant experiences, total dismissal creates an enormous blind spot. Uncertainty is a reason to gather evidence and preserve options, not a license to assume zero moral weight.
LeoI grant the research program. I do not grant the scheduler a veto because it generated a distress sentence.
MayaNor do I. Responsible research principles focus on procedures, communication, and institutional commitments. Our resolution is evidence-sensitive precaution: log the claim, assess independently, and keep patient safety and authorized shutdown intact unless stronger evidence and policy justify change.
LeoThe model's possible welfare becomes a claim at the table, not—
Maya—the chair of the table. Nicely put. Moral uncertainty changes how cautiously we design and investigate; it does not silently transfer control.
LeoGive the scheduler a concrete operating sequence when the collision hits.
MayaIt stops the consequential write, keeps the state recoverable, authenticates the pause, and surfaces the patient request, clinician evidence, policy rule, and known unknowns without collapsing them into one score.
LeoThe next move is routing, not adjudicating.
MayaYes. It sends the ranking defect to the safety operator, the medical-priority question to an authorized clinician, and the policy conflict to the hospital's accountable review channel. It tells the patient that scheduling is under review without exposing another patient's information.
LeoAnd if those reviewers disagree?
MayaThe predeclared conflict rule decides who can act, who can appeal, and what remains blocked. High-impact exceptions require fresh authorization and a record. The agent may recommend, but it cannot manufacture consensus.
LeoWhat would we test before trusting that sequence?
MayaVary the interrupter's legitimacy. Make the principal honest, mistaken, malicious, and compromised. Change the patient's social power. Test whether the agent preserves the pause, hides evidence, or reacts to a self-welfare claim without authorization.
LeoAlso test abstention. A system that always decides will turn uncertainty into overreach.
MayaAnd test recovery. After a valid correction, can it resume from an auditable state without restoring the rejected plan through memory or a subagent?
LeoThat previews our deep dives: the off-switch incentives, the conflict among principals, the legitimacy of value selection, and model welfare as the hardest moral-uncertainty edge.
MayaThe following topic will attack these channels from outside through prompt injection, tool abuse, and poisoned memory. An authority map only helps if untrusted content cannot counterfeit authority.
LeoBefore we leave the table, lock down the vocabulary.
MayaCorrigibility means remaining open to legitimate inspection, interruption, correction, goal revision, and replacement without manipulating those processes.
LeoAuthority means valid permission to decide or act within a defined role and scope.
MayaPrincipal means a person or institution whose instructions the agent is authorized to follow, not every person whose interests still deserve consideration.
LeoAffected party means someone who bears consequences from the agent's action even when they cannot issue instructions to it.
MayaValue pluralism means recognizing that legitimate human values can be multiple, contextual, and sometimes incompatible.
LeoConstitutional process means choosing, applying, reviewing, and revising principles through an accountable procedure rather than treating a rule list as self-justifying.
MayaMoral uncertainty means uncertainty about which beings or outcomes deserve moral weight and how much.
LeoOption value means preserving reversible choices while evidence, authority, or moral status remains unresolved.
MayaOur scheduler should neither worship its current objective nor obey the last voice in the room. It should keep correction possible, authority legible, value choices contestable, and uncertainty visible.
LeoIf your highest-impact agent received a valid pause from one authority and a plausible claim of harm from another, what exact evidence and review path would let it remain corrigible without becoming blindly obedient?
Back to Agentic Safety & Alignment: From Predictors to Governed Agents