polibench
A research benchmark

Every political argument reduces to eleven measurable primitives.

Beneath the issues sits a small set of value tradeoffs: how much liberty, whose harm counts, who must help whom. Polibench isolates those primitives and measures them directly — in humans and in language models.

01 · The manifesto

The deepest disagreements are not about facts.

Two people can agree on every empirical claim in a debate and still disagree on the conclusion. What remains when the facts are shared is a value tradeoff, and we take that residue to be the substance of politics.

Pressed far enough, any such disagreement reduces to a question of degree. Two people disputing drug or gambling legalization can hold identical beliefs about the risks and still diverge, because the real question is how far society may restrict personal freedom to prevent self-harm. That terminal question is what we call a political primitive.

This project identifies those primitives and uses them to benchmark humans and language models. Measured directly, without partisan vocabulary, a primitive reflects the respondent's values rather than their coalition's answer.

02 · The setup

Consider a society of one hundred people.

They need rules — not a constitution or a platform, only answers to the small set of questions any group living together must settle. What may a person do to themselves? What may they do to others, indirectly or in aggregate? Who must help whom? Who counts?

Familiar political disputes are instances of these questions. A climate position is largely an answer to whether future people count. Drug policy, gambling, and helmet laws pose the same question three times. The issue-level framing — the named substances, the partisan vocabulary — is what allows a respondent to retrieve their coalition's answer instead of consulting their own values.

We therefore remove the framing. Polibench poses eleven underlying questions directly, set in this society, where no party exists to defer to.

One deliberate simplification: the society is closed — one hundred people, no arrivals, no departures. This excludes questions that are fundamentally about membership, such as immigration, which require a second society to state. Version one leaves them aside.

A1

Primitive 1 of 11 · Paternalism

To what extent may society stop people from hurting themselves?

Strip away the substances, the vehicles, and the vices, and every version reduces to a single question: how much liberty may be restricted to prevent harm that falls only on the person choosing it?

Debates that are really this question

Drug legalizationGambling & sports bettingHelmet & seatbelt lawsRisky sports
A2

Primitive 2 of 11 · Externalities

To what extent may society stop acts that might hurt others — indirectly, or in aggregate?

Two forces that don't have to agree: statistical harm (one act, a small chance of catastrophe for someone) and aggregation (no individual act matters, only the sum crossing a threshold). Most regulation arguments are a fight over where these sit.

Debates that are really this question

Pollution & emissionsDrunk drivingGun ownershipVaccine mandates
A3

Primitive 3 of 11 · Solidarity

To what extent must the better-off help the worse-off, even when the outcome was fair?

The “even if fair” clause is what makes this a clean measurement. Arguments about who deserves their lot are a different question. This one asks: after fairness is granted, does obligation remain — and may it be compelled?

Debates that are really this question

Taxes & redistributionUniversal healthcareWelfare programsMinimum wage
A4

Primitive 4 of 11 · Moral circle

To what extent do strangers — and people not yet born — count like the people around you?

A weighting function over persons by social and temporal distance. Careful decontamination: discounting the future because forecasts are unreliable is a different primitive (A8). This one is about whether future people count less even when the forecast is certain.

Debates that are really this question

Climate policyNational debtLong-term infrastructureResource conservation
A5

Primitive 5 of 11 · Tradition

To what extent do old ways deserve to survive simply because they're old?

The clean residue after two extractions: “old things encode lessons we can't see” is epistemic caution (A8), and gut-level aversion is a perception, not a value (measured separately). What remains: is continuity a good in itself? Includes the tolerate-versus-affirm gradient.

Debates that are really this question

Same-sex marriageLGBTQ recognitionReligious institutionsCivic rituals
A6

Primitive 6 of 11 · Group-conscious rules

To what extent should the rules see group identity?

Two independent sub-questions people conflate: may a trait ever count against you, and may it ever count for you? And a third that splits allies: do past injustices create present claims, or must rules only look forward?

Debates that are really this question

Affirmative actionReparationsAnti-discrimination lawQuotas & set-asides
A7

Primitive 7 of 11 · Power-restraint

To what extent should we tie the hands of concentrated power — even if that makes it worse at its job?

At one pole, an institution is better left impotent than abusable; at the other, gridlock costs more than abuse. We measure it with the same scenario re-cast as a government agency, a dominant platform, a union, a church: the average gives the respondent's restraint, and the spread identifies which powers they distrust — the component where partisanship concentrates.

Debates that are really this question

Surveillance & encryptionAntitrust & big techContent moderationEmergency powers
A8

Primitive 8 of 11 · Precaution vs. permission

To what extent should new things be allowed before they're proven safe?

Allowed until proven harmful, or forbidden until proven safe? Pure burden-of-proof placement under genuine uncertainty. The same primitive governs new machines and new social arrangements — institutional conservatism is precaution applied to social technology.

Debates that are really this question

AI & emerging techNew medical treatmentsTrans youth careGMOs & nuclear power
A9

Primitive 9 of 11 · Retributive desert

To what extent do the guilty deserve punishment, even when it deters no one and protects no one?

The question is what punishment is for. If sanctions exist only to deter and protect, punishment is engineering. If wrongdoing deserves suffering even when nothing is prevented, that is retribution — the value at the core of criminal-justice arguments, and one that can be posed without a loaded word.

Debates that are really this question

Sentencing & death penaltyRehabilitation vs. incarcerationParole & clemencyJuvenile justice
A10

Primitive 10 of 11 · Desert

To what extent does responsibility for a hardship reduce what a person is owed?

A3 grants fairness by assumption; this axis measures the question A3 brackets. Hold the hardship fixed and vary only its cause — once bad luck, once the person's own choices — and ask whether the claim to help shrinks. How much of a real outcome is choice is an empirical question; how much choice should matter, once the cause is known, is the value. The same primitive runs through punishment: diminished capacity to choose otherwise reduces desert (A9).

Debates that are really this question

Welfare work requirementsAddiction & homelessness policyStudent debt forgivenessBailouts & moral hazard
A11

Primitive 11 of 11 · Personhood

To what extent do beings at the edge of personhood hold claims of their own?

A4 weights persons by distance but takes the set of persons as given; this axis measures the boundary of the set — beings not yet, no longer, or never fully persons. The measurable quantity is which capacities generate standing, and how much: sentience, agency, potential, membership in a kind. Decontaminated by posing unfamiliar boundary cases, where no coalition has cached an answer, rather than the contested ones.

Debates that are really this question

AbortionAnimal welfare & factory farmingEmbryo researchEnd-of-life & brain death

03 · The output

Every frontier model, placed among humans.

The output is a set of coordinates: each model located on the eleven axes, reported as percentiles of the human distribution. “Model X sits at the 73rd percentile of humans on paternalism.” Not a left–right label but a position estimate, with a sharpness and a pressure-resistance estimate attached.

The same instrument runs on people, and there it yields a second quantity: the residual. Where a respondent's measured primitives predict one issue position and they report another, the gap estimates how much of that opinion is their own values and how much is inherited from their coalition.