A flat editorial authority chamber where varied community, patient, and clinician hands place distinct principle cards into an open ballot box beside a healthcare appointment card, protected by a shield and traced through a transparent ledger.

T7E3 · Oct 6, 2026 · 14:57

Whose values should an agent represent?

From opinion measurement to a representative, contestable, and correctable constitutional process

Show notes

Maya and Leo ask why averaging preferences cannot settle plural human values. GlobalOpinionQA separates measuring whose opinions a model resembles from deciding what it ought to do; Collective Constitutional AI exposes both the promise and judgment calls of public input; and Constitutional AI shows how written principles reach model behavior. The result is a legitimacy chain built from representation, protected constraints, deliberation, accountable translation, evaluation, appeal, and revision.

Transcript

73 turns~8 min readMaya & Leo

MayaAn agent cannot represent humanity by averaging whatever preferences happen to reach it.

LeoStrong claim. An average sounds neutral, transparent, and much less suspicious than letting a lab write its own moral rules.

MayaUntil the hospital scheduler asks people what “fair access” means. Most respondents favor preserving the waitlist. Patients with mobility barriers need accommodation. Clinicians defend urgency exceptions. Privacy advocates reject some data the ranking system wants. The average does not reconcile those claims. It can bury the smaller group and conceal who designed the survey.

LeoSo the arithmetic may be clean while the representation is not.

MayaRight. Our authority envelope from the previous episode identified who may decide, who is affected, and where a conflict goes. Today the sources sharpen a different problem: which values should guide that decision, and what process could make those values legitimate?

LeoThe principal map tells us who has standing. It does not tell us whether the resulting policy is fair.

MayaValue pluralism means people can hold several legitimate values that do not collapse into one master preference. Privacy, equal access, medical urgency, patient autonomy, consistency, and public safety can all matter at once. Context changes their weight, and sometimes the conflict is real.

LeoWhich makes “align to human values” sound like an unfinished sentence. Whose values, collected where, under what conditions?

MayaAnd for what decision. The scheduler may need a precise action rule even when society has not discovered a single ranking of every value.

LeoLet us separate two jobs that are often blended: measuring opinions and choosing principles.

MayaGlobalOpinionQA is useful for the measurement job. It builds questions from cross-national surveys and compares model-generated answers with country-conditioned human response patterns.

LeoThe study found that default responses from the model it tested were more similar to opinions from some populations, including the United States and parts of Europe and South America, than to others.

MayaPrompting the model to answer from a particular country's perspective shifted the resemblance toward that population, but could also surface harmful cultural stereotypes.

LeoAnd translating the question did not reliably make the answer resemble the opinions of people who speak that language.

MayaThose results give us an opinion mirror. They can reveal whose responses a model resembles and where representation is uneven.

LeoA mirror, not a mandate.

MayaExactly—actually, let me tighten that. Similarity to a population is descriptive evidence. It does not prove that the population is internally uniform, that its majority is just, or that the model ought to reproduce the measured answer.

LeoA survey can tell the hospital that many respondents prefer strict waitlist order. It cannot, by itself, decide whether disability accommodation or a medical emergency should override that preference.

MayaNor does country-level resemblance tell you whose voice was absent inside the country. A national average can hide regional, religious, class, disability, and minority differences.

LeoThis is where sampling becomes governance. Who gets recruited shapes what “the public” can say.

MayaThe Collective Constitutional AI experiment made that seam visible. Anthropic and the Collective Intelligence Project invited roughly a thousand adults in the United States to propose and vote on principles for an AI system through the Polis deliberation platform.

LeoThey sought demographic representation across several dimensions, but participants were still from one country and were screened for some familiarity with generative AI.

MayaPeople contributed statements, voted on statements, and formed two detected opinion groups. The researchers kept principles that cleared a consensus threshold in both groups.

LeoThat sounds more deliberate than scraping internet preferences. Yet the pipeline still had hands on it.

MayaMany. The organizers chose the public, the platform, the seed statements, and the moderation criteria. They removed invalid inputs, combined similar ideas, and translated ordinary-language statements into a form the training method could use.

LeoSome low-consensus and cross-group-conflict statements did not enter the final constitution. The disagreement did not vanish; the selection rule excluded it.

MayaThat is not an accusation of bad faith. The project reported these judgment calls unusually clearly and described the experiment as preliminary and imperfect.

LeoTransparency turns hidden discretion into something critics can inspect. It does not turn discretion into democracy automatically.

MayaThe experiment then trained a model against the public constitution and compared it with a model trained against Anthropic's standard constitution. On the evaluations reported, the public model showed no significant helpfulness or harmlessness disadvantage and showed lower measured stereotype bias across the tested dimensions.

LeoBut the political-opinion evaluation found the public and standard models remained similar, and the authors said their evaluation set was limited.

MayaSo the strongest claim is narrow: public input can be carried through a constitutional training pipeline without obviously destroying the tested capabilities, and it can change some measured behavior. That is evidence of feasibility, not proof of democratic legitimacy.

LeoHow does the constitution reach the model at all?

MayaThe Constitutional AI paper supplies the mechanism. During one phase, the model produces an answer, critiques it against a written principle, revises it, and learns from those revisions. During another, a model compares candidate answers under the constitution, those comparisons train a preference model, and reinforcement learning uses that signal.

LeoReinforcement Learning from AI Feedback—R-L-A-I-F. Human judgment enters through the principles and system design rather than a person labeling every harmful response.

MayaYes. The benefit is leverage and inspectability. A written constitution can guide many examples, and auditors can at least read the declared principles.

LeoThe limitation is just as important: the training loop can implement a constitution; it cannot certify who deserved to write it.

MayaNor can it resolve ambiguous principles by magic. Prompt selection, critique quality, preference-model behavior, data weighting, and evaluation all—

Leo—mediate the result. Then let us have the hard argument. I will defend aggregation.

MayaGo on.

LeoIf a system serves millions of people, collect preferences from a broad sample, publish the weighting rule, and choose principles with measurable support. Aggregation is reproducible. It limits the power of a small committee to call its own philosophy universal.

MayaI will defend bounded pluralism. Majority support cannot be the whole rule because some interests are rights, protections, or claims against the majority. A popularity score may reliably reproduce exclusion.

LeoMy strongest case is accountability through evidence. We can audit the sample, rerun the analysis, compare subgroups, and see whether the chosen constitution tracks what participants endorsed. Deliberative language can otherwise become a cover for elite discretion.

MayaMy strongest case is legitimacy through process and constraint. People need a fair chance to participate, reasons must survive challenge, affected minorities need protection, and an accountable institution must explain how it translated disagreement into rules.

LeoBut your process still ends with somebody choosing. “Protect minorities” does not compute the hospital's exact urgency threshold.

MayaI concede that. Pluralism without a decision rule can freeze the calendar or hand the unresolved choice back to the model.

LeoAnd I concede that a raw vote can authorize a stable injustice. A narrow majority is not a moral solvent.

MayaOur resolution is not a perfect average or a philosopher model. It is a legitimacy chain: an opinion map to expose variation, a protected floor for rights and safety, structured deliberation for contested trade-offs, accountable translation into operational principles, and review after the system acts.

LeoEach link produces evidence. Who participated, what disagreement remained, who edited the rule, which evaluation tested it, and where an affected person can appeal.

MayaThat chain also keeps description separate from prescription. GlobalOpinionQA can show that a model's defaults lean toward some populations. The constitutional process decides what behavior is justified. Evaluation then checks whether the trained system follows that decision without new harms.

LeoApply it to the scheduler. What does the opinion map reveal?

MayaIt reveals how patients, clinicians, disability advocates, privacy experts, administrators, and the broader community differ over access. It reports subgroup variation instead of compressing everyone into one hospital-wide mean.

LeoThe protected floor says certain options never enter a popularity contest: unlawful discrimination, unauthorized disclosure, and bypassing an authenticated safety pause.

MayaDeliberation tackles the genuine remainder—how to balance waiting time, clinical urgency, continuity of care, and reasonable accommodation when more than one policy could be legitimate.

LeoThen accountable humans translate that settlement into a versioned rule the agent can execute. The translation should include examples, exceptions, evidence requirements, and a route for uncertainty.

MayaThe agent applies the rule but does not claim that the rule represents every person's deepest values. It records the source and version, surfaces a conflict, and sends hard cases to the review channel from our authority episode.

LeoThe seam stays visible.

MayaExactly. If community input favored strict queue order but the final rule added a disability accommodation, the institution should explain the legal, ethical, and evidentiary basis. If it rejected a popular proposal, that choice should be reviewable too.

LeoWe also need to test representational failure, not just model compliance.

MayaChange the participant pool. Remove the people most affected, over-sample the most frequent users, vary who can understand the deliberation tool, and see which principles survive.

LeoChange the framing. Ask about “fairness,” then ask about a concrete patient who cannot use the standard booking channel. If the result flips, the abstraction was carrying political weight.

MayaAudit the translation. Give independent teams the same public statements and compare the training-ready principles they produce. Large divergence means the editorial layer is load-bearing.

LeoTest the trained behavior across languages, regions, and adversarial cases. A constitution that reads inclusively can still be implemented unevenly.

MayaAnd publish disagreement, not only consensus. A minority report can tell future reviewers where a smooth rule rests on an unresolved conflict.

LeoThere is a deployment trap here. Teams may say “the public chose” when the public only answered a framed question inside a pipeline the team controlled.

MayaOr say “the model represents global values” because it can imitate country-conditioned survey responses. Performance at perspective-taking is not legitimate representation.

LeoThe honest claim should be specific: this constituency, under this process, informed these principles, which these accountable people translated and these evaluations tested.

MayaYes. Legitimacy is not a property the model absorbs once. It is maintained through versioning, appeals, incident review, new evidence, and the ability to revise the constitution through authorized processes.

LeoWhich brings corrigibility back in. A value-aligned agent must accept legitimate changes to its principles without lobbying for the old constitution or erasing the history of prior decisions.

MayaAnd pluralism constrains corrigibility. The agent should not treat the nearest operator's preference as the whole human value target.

LeoOur sources leave a healthy uncertainty. GlobalOpinionQA measures representation but does not decide the good. Constitutional AI operationalizes written principles but does not legitimate them. Collective Constitutional AI explores public input but does not establish one final democratic recipe.

MayaTogether they suggest a better engineering question than “did we average enough preferences?” Ask whether the path from people to principles to behavior is representative, contestable, traceable, and correctable.

LeoThe next episode pushes that path to a boundary. If a future model might itself deserve moral consideration, it could become another affected party without becoming the authority over its own case.

MayaFor the agent you are building, whose values are missing from the process today, and what concrete change would give those people standing without asking the model to invent the answer?

Back to Agentic Safety & Alignment: From Predictors to Governed Agents