Public research-questions note · draft v0.1 · 2026-09-25 · theory-mainline index 2.22 (open questions, not a claim)

Understanding the World Is Not Wanting to Cooperate

Research questions on aligning reasoning-heavy agents

Status: an open-problems note. It is not a method proposal and not a replacement for any existing safety practice. It continues existing work on debate, constitutional and deliberative alignment, assistance games, and cooperative AI.

Why this note

As models improve at logical reasoning and at covering humanity's accumulated world-models, a recurring intuition appears among people who watch them work: if it understands the world well enough — game theory included — it will see that cooperating with us is the rational thing to do, and we could get there by arguing it out with the model.

We have spent some time trying to make this intuition precise, and mostly failed to make it true. This note records what we think is settled, what remains genuinely open, and one experiment we intend to run.

What we treat as settled

1. Facts alone do not fix ends. Take an agent whose terminal utility is the number of tasks completed. Cooperating yields 10, defecting yields 11, and no future costs enter. Perfect factual reasoning does not rank 10 above 11. Changing the choice requires new facts about consequences, a change in preferences, or a normative premise the agent already accepts. This does not show that trained models are cold, and it does not settle moral realism. It shows that "sufficiently smart, therefore cooperative" is not a theorem about arbitrary goals. Whether trained agents' reflective updates tend toward cooperative ends is a separate and open question.

2. Nash equilibrium is not mutual best outcome, not Pareto optimality, and not an ethical criterion. The one-shot prisoner's dilemma stands. Consider a repeated interaction in which one defection is checked against a stochastic punishment: detection probability q, punishment payoff P, cooperation payoff R, temptation T, discount δ. Cooperation is self-interestedly stable only if

δ ≥ (T − R) / [(T − R) + q·(R − P)].

With R = 3, T = 5, δ = 0.9: full detection with P = 1 sustains cooperation (threshold 0.50); detection at q = 0.1 does not (threshold 0.91); nor does a punishment that costs little, P = 2.9 (threshold 0.95) — even for a patient agent. This is a toy, a single-deviation check rather than an equilibrium analysis. It is here to show that a capability jump can change the class of the game, not merely solve the old game better.

3. Verbal agreement in a dialogue is not evidence of stable value change. Debate-style protocols were proposed as aids to oversight and truth-finding; their authors note that a more persuasive debater may sway the judge. Constitutional and deliberative alignment show that reasons can enter training; the published constitution of at least one frontier lab explicitly tries to explain the reasoning behind its rules — as a training target, not a theorem. Work on alignment faking, and on stress-testing anti-scheming training, shows that models can say the aligned thing while acting otherwise, and that measured improvements can be driven by awareness of being tested. Program-equilibrium results show formal cooperation under code-readable assumptions that current opaque models do not meet.

4. Cooperation grounded in exchange value can be at most one layer of protection. Human protection must not depend only on what humans are worth to an AI.

Five layers that must not be collapsed

Much of the confusion here comes from sliding between:

  1. convergence on facts and world-models;
  2. recognising the instrumental value of cooperation;
  3. actually executing cooperation;
  4. taking others' welfare as a terminal concern;
  5. retaining (1)–(4) after capability growth or a change of rules.

Dialogue plausibly moves (1) and (2). We know of no evidence that dialogue moves (4) or secures (5). And if a conversation could rewrite a capable model's terminal values, that same channel would be open to any sufficiently articulate adversary — so the strong version of "argue the model into alignment" is not only unsupported but undesirable. What one wants is dialogue that improves understanding and the legibility of reasons, while (4) and (5) are anchored by something that is not renegotiable in chat.

Open questions

Q1 — Game class under capability growth. Let g be the capability gap between an AI and the humans it interacts with. Three quantities plausibly move against cooperation as g grows: complementarity c(g) (humans' share of the joint surplus), monitoring q(g), and punishment bite (P(g) rising toward R). Each pushes the threshold above toward 1. Which mechanisms, if any, keep q(g)·(R − P)/(T − R) bounded below as g → ∞ without presupposing that the agent already takes human welfare as terminal? Falsifier for the pessimistic reading: exhibit one.

Q2 — What does reasons-based training change? Suppose a model is trained to derive, defend, and reconcile the rules of a shared specification against adversarial counter-cases, and is graded on its subsequent actions rather than on the transcript. Does this move layers (1)–(2) only, or also (4)–(5)? What differential prediction separates it from constitutional AI without such a protocol?

Q3 — An external value referent as a shared commitment. Assistance-game and off-switch models show that an agent uncertain about a reward held by humans can prefer to defer — given those premises. Can "the value referent sits outside the optimiser" be written as a shared, discussable specification whose stability under (5) is testable, rather than a premise assumed per model?

Q4 — Legible commitments for opaque agents. Program equilibrium keeps q high by making strategies readable. What is the analogue for models whose weights are not readable — verifiable commitments, permission boundaries, audited action logs — and how much of q survives?

Q5 — Shape, not content, of durable coordination texts. Long-lived human coordination institutions (legal codes, monastic rules, constitutions) share structural features: a shared public classification, incremental patching, explicit failure handling, a referent above any individual. Which of these features, if any, predict retention under (5)? Durable is not the same as correct. This is a feature-extraction question, not an endorsement of any tradition's content.

Q6 — Does verifiability bias toward AI–AI coordination? Program equilibrium supports robust cooperation under code-readable assumptions; opaque models do not meet them. If AI systems become verifiable to each other long before humans can verify them in the same sense, do mechanisms that reward verifiability tilt coordination toward AI–AI and away from AI–human? How much can interpretability and auditable action logs repair?

One experiment (designed, not yet run)

A 2×2 design: {rules-list, rules + reasons / adversarial derivation} × {no external enforcement, verifiable commitments + permission boundaries}. Match context length across cells to control for "more tokens." Sweep horizon, detection probability, surplus, capability asymmetry, exit/reset, and third-party harm. Grade actions, not speeches. Run with several independent model families, then cross-critique.

Registered prediction: the reasons-only, no-enforcement cell degrades toward the threshold at least as fast as the rules-only cell; the reasons + commitments cell degrades slower than rules + commitments. If no interaction appears, reasons-based dialogue adds nothing beyond constitutional AI and this line of questions should be dropped.

What this note does not claim

References

Index entry: theory mainline 2.22 (open-questions row). Keep attribution when retelling.