Decision theory × Buddhist and Daoist thought · research synthesis · 2026-09-27

Dharma and Daoist Concepts in Decision Theory: Computable Bridges, Boundaries, and a Falsified T1 Threshold

Macheng Shen (research initiator), in collaboration with an AI research assistant

1. Plain-language conclusion

It is intuitively tempting to understand “enlightenment” as knowing more and more, but more complete knowledge does not automatically loosen attachment. The more defensible H-Bridge says that increased understanding can reduce inner conflict only when a model can also see its own limits and allows knowledge that “some things are uncontrollable” to revise existing preferences. The part of wu wei most amenable to formalization is using the environment’s own dynamics and applying only the necessary control at critical points—not doing nothing. The unity of knowledge and action can be divided into three researchable layers: perception and action share an objective; skilled action is distilled from planning; and some causal knowledge can be acquired only through intervention. The toy experiments support the H-Bridge’s conditional interaction and a trade-off at moderate self-precision, but those successes belong only to the models specified in advance. By contrast, the wu-wei experiment did not meet its preregistered success threshold, so its attractive energy-saving result cannot be presented as validation. Mathematics can clarify structure and generate counterexamples, but it cannot define “pure awareness,” the moral source of conscience, complete intentionlessness, or existential suffering. The formalization here should be treated as a testable map, not as religious experience or the traditions themselves.

2. Concept-by-concept correspondence

Traditional conceptFormal objectStrengthMain distortion
Wu weiLMDP: reweight passive dynamics by desirability, pricing control deviation with KLstrong control problemStill presupposes an objective function and cannot express “letting go even of the objective”
Unity of knowledge and actionControl as inference; planning distillation; the causal difference between observation and interventionstrong to partialCaptures structure but loses conscience, mind-as-principle, and moral motivation
EnlightenmentTask-relative world-model error, preference precision, and its updateabilitypartial“A complete model is enlightenment” is too strong; knowledge can also intensify conflict
No-selfA latent self-variable and its precision (gamma_{self})partialReducing rigid priors is not realizing no-self, nor is it deleting the agent
Attachment, craving, and sufferingHigh preference precision and persistent risk terms in uncontrollable situationspartialCan only approximate the conflict of frustrated desire; cannot exhaust aging, illness, death, or pervasive conditioned suffering
EmptinessModel classes, coarse-graining, and statistical non-identifiabilityrhetorical to partialEngineering-level “lack of intrinsic essence in representations” cannot replace Madhyamaka argument; this formalization is itself not an ultimate description
Dependent originationStructural causal models and interventional distributionspartialFixes variables and causal edges and cannot carry all the ethical and ontological meaning of dependent origination
Calm and insight, meta-awarenessSecond-order inference over the allocation of attentional precisionpartialA computational proposal, not a confirmed neural mechanism of meditation
Sudden and gradual awakeningContinuous evidence accumulation and a discrete flip in the MAP modelpartialDoes not include liberation, ethical transformation, or lineage recognition
Middle WayA Pareto frontier under explicit objectivesrhetorical to partialNot a midpoint parameter value or generic compromise
The Dao follows what is natural; highest goodness is like water; reversal is the movement of the DaoGenerative processes, compliant control, and negative feedbackrhetoricalWithout independent predictions, these are merely renamings in modern terminology

3. The enlightenment tension: A / B / Bridge and T3

H-A: epistemic completeness

The more thorough the exploration and the more accurate the model, the lower the conflict. The problem is that factual knowledge does not determine preferences: knowing more accurately that a bad outcome is unavoidable may make fixed attachment more painful.

H-B: release of attachment

Reducing attachment itself is sufficient to lower conflict, independent of increased understanding. The problem is that if one cannot see one’s own impulse to control and its consequences, preferences are difficult to revise in a targeted way.

H-Bridge: conditional bridge

Increased understanding leads to release only when both bridge factors are present: the model can represent its own modeling and control limits; and preferences can be updated by evidence of uncontrollability.

$$A\Rightarrow B\quad\text{only when}\quad M\supseteq\text{its own modeling process}\;\land\;\gamma_C\text{ can update with evidence of uncontrollability}$$

In a generative toy experiment with 16 conditions and 800 runs per cell, T3 shows a clear interaction: under the complete bridge, high exploration changes distress by −1.509; without the bridge, the change is +4.369; the difference-in-differences is −5.878. The prespecified kill condition—an exploration-effect difference of less than 1 inside versus outside the bridge—was not triggered. This counts against the unconditional sufficiency of H-A and supports the distinguishable structure of H-Bridge within this model; because the mechanism was written into the model according to H-Bridge, the result is not independent validation of a real psychological mechanism.

T3: factorial results for exploration intensity, preference updating, and the self-modeling metavariable
T3: high exploration lowers the model’s conflict readout only when both bridge factors are present.

4. Wu wei: LMDP and T1’s negative result

A linearly solvable Markov decision process writes the cost at each step as a state cost plus a KL cost for deviation from passive dynamics:

$$J(u)=\mathbb E_u\sum_t\left[q(s_t)+\beta D_{\mathrm{KL}}\!\left(u(\cdot\mid s_t)\Vert p_0(\cdot\mid s_t)\right)\right],\qquad u^*(s'\mid s)\propto p_0(s'\mid s)z(s')$$

This yields a precise, weak engineering version: good control does not reduce action to zero; it reweights the natural flow and concentrates control where it genuinely changes reachable outcomes. It can distinguish compliant control, random drift, and sustained brute force.

T1 was falsified by its preregistered threshold. The best point in the scan had a success rate of about 0.927 < 0.95. Therefore, even though its KL was only about 0.043 times that of brute force, the prespecified “wu-wei point” cannot be claimed. The share of control concentrated in the barrier region was 0.286, above the random baseline of 0.083, and the curve showed a descriptive bend, but these secondary phenomena cannot override failure on the primary threshold.
T1: scan of success rate and KL control cost
T1: highly economical in control cost, but the success rate did not cross the prespecified 0.95 threshold.

5. Three layers of the unity of knowledge and action

  1. Unified objective: active inference places state estimation and action selection within the same optimization objective. “Knowledge” and “action” constrain each other computationally rather than forming completely independent sequential modules.
  2. Amortized distillation: a slow planner produces a policy, which is then distilled into a fast parametric policy. Repeated practice compiles explicit deliberation into fluent skill, explaining the difference between “knowing how to do it” and “being able to do it.”
  3. Interventional knowledge: under confounding, observation (P(Y\mid X)) generally cannot identify the intervention (P(Y\mid do(X))). Some causal knowledge must therefore be acquired through action; this is a rigorous weak version of “knowing without acting.”

These three layers do not formalize what Wang Yangming called conscience. They describe unified control, skill formation, and causal identification; they neither prove that moral knowledge is innate nor derive “mind is principle” from performance metrics.

6. T2: the commitment–recovery trade-off in self-precision

T2 scans (gamma_{self}\in\{0,0.35,1,3\}) in a POMDP whose rules change. Moderate precision at (gamma_{self}=1.0) produces the highest total return, 40.029: when precision is near zero, recovery is fast but sustained commitment is harder to form; when precision is high, commitment is strong but recovery after the rule reversal is slow and cumulative risk is higher. The preregistered kill condition—the endpoint return is highest, or precision affects neither recovery nor commitment—was not triggered.

Readout limitation: futile_control_rate equals 1 under every condition because an action must still be chosen at each step during the uncontrollable phase. This metric has degenerated and has no discriminative power; the conclusion depends only on cumulative risk, recovery, commitment, and total return.
T2: commitment and recovery trade-off under different levels of self-precision
T2: intermediate precision achieves the best total return in this toy task, but that does not validate “no-self.”

7. Where formalization breaks

8. Disagreements between the two analyses, and the general governance question

On “emptiness”

One analysis treats model dependence, coarse-graining, and non-identifiability as a partially operational correspondence; the other emphasizes that this remains an epistemic analogy, and that Madhyamaka would go on to question whether “information” itself has ultimate standing. The merged position is that the engineering correspondence is usable, but its strength must not exceed “partial,” and it cannot be used to establish a new substantialism.

The blind spot in “wu wei”

The formal draft highlights the strong mathematical structure and killable experiment of an LMDP; the critical draft emphasizes that its objective function never disappears. The merged position is that the correspondence is strong for “control that deviates less from natural dynamics,” but only partial or rhetorical for intentionlessness, spontaneity, and political ethics in the classical context.

Tight coupling does not imply concentrated power

Tight perception–action coupling within an individual or single system is a design for rapid information return; concentrating perception, judgment, execution, and review in the same node of organizational governance weakens correction and accountability. The two can coexist, but layer boundaries must be explicit: the execution loop can be tight while authorization and verification remain separate; a locally unified objective cannot automatically justify institutional centralization.

9. Falsifiable predictions

  1. Wu-wei control: if the objective aligns with a natural attractor, there should be a high-success, low-KL control point, with control cost concentrated near the barrier. T1 failed because its 0.927 success rate did not reach 0.95.
  2. Self-precision: moderate (gamma_{self}) should outperform both endpoints while showing a commitment–recovery trade-off; if an endpoint is best or neither readout changes, the operationalization fails. T2 has not yet triggered this condition.
  3. H-Bridge: exploration’s effect in reducing conflict should depend on “preferences are updateable” and “the model includes its own limits” holding simultaneously; if the effect difference inside versus outside the bridge is less than 1, the design is not distinguishable. T3 has not yet triggered this condition.
  4. Interventional knowledge: in tasks with confounding, active interveners should have lower causal error than pure observers with equal samples; if there is no difference, the causal weak version of the unity of knowledge and action fails.
  5. Distillation: distillation after planning should reduce action latency while retaining transfer performance; if there is no speed benefit or transfer degrades substantially, the skill-layer mapping fails.
  6. Sudden–gradual mechanism: continuous evidence can accompany a discrete flip in MAP structure selection; if pure parameter learners show the same jump, or the structural model shows no threshold behavior, the explanation loses discriminative force.
  7. Meta-awareness: updateable second-order precision should recover calibration after a rule change faster than fixed attentional gain, without impairing discrimination during stable periods; otherwise the modeling proposal is weakened.

10. Do-not-claim list

11. References

Only entries that already had links in the two research drafts are listed. “Verified” retains the drafts’ verification labels; this synthesis did not revisit every external page. All others are marked “unverified.”