Move, response, position

Your move changes their options. Their response changes yours.

Your movea
changes their incentives
Their responseBR(a)
creates the position you keep
New positionplay again, redesign, or exit

Most important decisions involve other people. They respond to your move, and their response changes what your move was worth.

Ask whether a choice still works after everyone responds. A good first move can produce a bad final position.

I use this as a working note. A script computes every equilibrium because Chicken and Stag Hunt are easy to remember backwards. Bad maths quickly becomes bad advice.

maxa   EV(a)  −  BR(a)

a
one action available to you
maxa
choose the action with the highest resulting value
EV(a)
the expected value of action a before anyone responds
BR(a)
the cost created by the other players' best response to a
subtract the response cost from the action's expected value

01 Recognise the game

Recurring problems often share a structure. Once you recognise it, your options become clearer.

Read a matrix like this. You pick a row, the other player picks a column, and the cell shows what each of you gets as (uyou, uthem). A cell is a Nash equilibrium when neither of you improves by changing alone:

u1(s1*, s2*)  ≥  u1(s1, s2*)   for all s1

u1
player one's payoff function
s1*
player one's equilibrium strategy
s2*
player two's equilibrium strategy, held fixed in this condition
s1
any alternative strategy available to player one
switching alone cannot give player one a higher payoff
for all
the condition must hold for every s1 in player one's strategy set
player 2
the same condition also applies with the two players reversed

A Nash equilibrium is stable. It does not have to be good for either player.

Solver

Pick a game, then click a cell
NE marked by the solver, not by hand

All eight, by shape

Two things worth noticing as you click through. The Prisoner's Dilemma and Public Goods are the games where both players have a dominant strategy. In both cases that strategy leads somewhere neither wanted. Matching Pennies is genuinely zero-sum and has no stable pure strategy.

Salary talks, arguments with a partner and supplier negotiations rarely have a fixed pot. How you play can create or destroy value. You can win the argument and leave both sides worse off.

02 See all five boards

A major decision affects five games at once. Winning on one board can hide losses on the others.

Layers

Same decision, five boards
Depth increases downward

The intrapersonal layer is easy to overlook. Today you can make commitments that your future self cannot reverse. Debt, specialisation, lifestyle and reputation all move on this board.

03 Lengthen the horizon

Repeat the Prisoner's Dilemma and cooperation can become rational. The people stay the same; the prospect of another round changes their incentives.

Suppose you cooperate as long as they do, and stop for good if they betray you. Cooperating forever is worth R + δR + δ2R + …, and betraying once pays T now and P forever after. Cooperation survives when:

R1 − δ T + δP1 − δ δ T − RT − P

δ
how much the next round is worth relative to this one, so how much the future matters
T
temptation, what betraying a cooperator pays
R
reward for mutual cooperation
P
punishment when both defect
1 − δ
the denominator of the infinite discounted stream of future payoffs
the expression on the right is the same condition after rearranging the terms

Model

Move the horizon
Values read from the matrix above
δ* keep cooperating betray once, then nothing δ = 0  no future δ → 1  future dominates

Trustworthiness depends partly on the situation. A contractor who will never see you again, a colleague in their notice period, a counterparty in a one-off deal: each is in a low δ game whatever their character. If you want cooperation, lengthen the relationship or add a stake that outlives it. Changing the incentive may be easier than replacing the person.

04 Change the game itself

A better move can still leave you stuck in a bad game. Sometimes you need to change the rules, players, or exit.

Seven things define a game. When you feel stuck, you are usually trying to improve your move while treating the other six as fixed:

G = (N, S, u, I,  >, R, E)

G
the complete game being described
N
the players involved
S
the strategies available to those players
u
the payoff assigned to each outcome
I
who knows what when a decision is made
>
the order in which the players move
R
the rules that constrain the players and actions
E
the outcome available when you exit or refuse to play

Diagnostic

Seven things you can change
Tap one

I use the last lever most often. Your bargaining power is largely set before you walk in by how survivable no-deal is. In the Nash bargaining solution the split depends on the disagreement point (d1, d2):

max  (u1 − d1)(u2 − d2)

max
find the agreement that makes the product as large as possible
u1
player one's payoff under the proposed agreement
u2
player two's payoff under the proposed agreement
d1
player one's payoff if no agreement is reached
d2
player two's payoff if no agreement is reached
ui − di
player i's gain from reaching an agreement
( )( )
the product balances both players' gains over their no-deal outcomes

A better outside option raises d1 before either side speaks. Another offer, a second customer or more runway can do this.

05 Size the bet before the upside

Some risks deserve a chance. A loss that ends the game needs a hard limit.

Sort any decision on two axes: can you undo it, and how bad is the bad case. The four quadrants want genuinely different behaviour.

Sorter

Reversibility against downside
Hover a quadrant
Act now reversible, small downside Slow down irreversible, survivable Pilot it reversible, real cost Refuse irreversible and ruinous Harder to undo Smaller downside

Positive expected value is not enough

A bet can have a genuine edge and still ruin you. You never experience the average across every possible future. You experience one path, in order, and a path that touches zero stops.

For a repeated favourable bet, the growth-optimal stake is the Kelly fraction:

f* = bp − qb

f*
the fraction of current wealth that maximises long-run growth
p
the probability of winning one round
q
the probability of losing, equal to 1 − p
b
the net amount won for each unit staked
bp − q
the edge after weighting the win and loss outcomes by their probabilities

Real edges are uncertain, so treat f* as a ceiling. More uncertainty calls for a smaller stake.

Simulation

Sixty paths through the same favourable bet
Red paths hit the floor
ruin round 0 round 60
Edge per unit
0
Kelly fraction
0%
Paths ruined
0%
Median outcome

Set the edge positive and the stake high and watch what happens. A favourable game, played too large, still kills a good fraction of the paths. Ruin needs a hard constraint:

P(ruin) < ε   and   integrity intact

P(ruin)
the probability that a loss ends your ability to keep playing
ε
the maximum ruin probability you are willing to accept
<
ruin risk must stay below that chosen limit
integrity
a hard constraint that the strategy is not allowed to violate
and
both constraints must hold at the same time

Optimisation happens inside these limits. A high expected value does not cancel ruin risk or an integrity failure.

06 Before you decide

These nine questions reveal the assumptions behind a decision.

Checklist

Name the game first
Mark each one honestly
1 Who are the players, including future you?Present you and future you have different incentives and rarely negotiate.
2 What does each player want?Compare what they say with the choices they repeatedly make.
3 Who moves first?Order determines who commits and who responds, and often decides the outcome.
4 Does this game repeat?One-shot and repeated versions of the same game have opposite correct strategies.
5 Is the available value fixed?Zero-sum tactics can destroy value when both sides could have gained.
6 What does each side know?Hidden information is where most bad deals are made.
7 What happens if I refuse to play?Your no-deal outcome sets your bargaining power before you say anything.
8 Which consequences are irreversible?Reversible decisions reward speed. Irreversible ones reward patience.
9 What response am I ignoring?The move that fails is usually the one that assumed nobody would react.

07 Seventeen domains

Each part of life has its own players, incentives and risks. All seventeen draw from the same limited pool of time, money, energy and attention.

What winning looks like

Capability that is hard to replace, with results other people can verify.

The characteristic failure

Becoming indispensable in a role that creates no bargaining power. Praise rises while outside options stay flat.

Skill → Scarcity → Evidence → Bargaining Power → Choice

Does my future bargaining power grow with this responsibility?

SignallingTournamentRepeated bargaining
What winning looks like

Pay that grows with judgement, scarcity or ownership. Hours worked become less important.

The characteristic failure

One employer, one client, one credential. Dependency turns a good income into a fragile one.

Income ∝ Hours → Income ∝ Judgement, Scarcity, Ownership

What mechanism decides what this can eventually pay?

BargainingInformation asymmetryMarket competition
What winning looks like

Staying in the game long enough for returns to compound.

The characteristic failure

Position sizes that carry real ruin risk. One bad sequence and the compounding stops.

Wt+1 = Wt(1 + r) − C  subject to  P(ruin) < ε

If the downside lands at the worst moment, am I still playing?

Repeated stochastic gameRuin problem
What winning looks like

Holding a claim on value you helped create, so the upside is not entirely someone else's.

The characteristic failure

Entering because ownership sounds like freedom, without a customer or an unfair advantage.

What happens if this works and the incumbent responds rationally?

Market entryRepeated competitionSignalling
What winning looks like

A relationship in which conflict can be repaired reliably.

The characteristic failure

Treating it as zero-sum. Winning the argument while lowering the value of the relationship.

δ ≥ (T − R) / (T − P)

Does repeated interaction here increase trust and capability, or steadily reduce them?

Repeated Prisoner's DilemmaStag HuntTrust game
What winning looks like

Contributions and authority made explicit, so the load does not silently fall on one person.

The characteristic failure

Free-riding on invisible labour. Nobody intends it, and the structure produces it anyway.

Is this a character problem or an incentive problem?

Public goodsCommon-pool resourceCoordination
What winning looks like

A smaller set of trusted relationships can be worth more than a large contact list. These people would still take your call without the job titles.

The characteristic failure

Treating a network as something to extract from. People remember that behaviour.

Am I contributing here, or only withdrawing?

Coalition formationRepeated trustPublic goods
What winning looks like

Signals that are expensive to fake. Shipped work, kept promises, consistent judgement.

The characteristic failure

Spending a decade of trust on one transaction. It rebuilds far slower than it breaks.

Tt+1 = (1 − ρ)Tt + aRt − bDt,   b > a

Could someone without the underlying capability afford to send this signal?

Costly signalling
What winning looks like

Responsibility, authority and accountability roughly in balance.

The characteristic failure

Responsibility far exceeding authority. You absorb the blame for outcomes you cannot control.

Responsibility ≈ Authority ≈ Accountability

Am I being given power, or only exposure?

Mechanism designPrincipal-agent
What winning looks like

Knowledge that transfers, compounds, and eventually changes what you do.

The characteristic failure

Consuming information because it feels productive, without using it to change a decision or action.

UCBi = μ̂i + c√(ln t / ni)

Can I use this yet, or have I only read about it?

Explore versus exploit
What winning looks like

Work at the intersection of genuine curiosity and non-generic knowledge.

The characteristic failure

Optimising novelty without usefulness, or usefulness without distinctiveness.

Value = Novelty × Relevance × Execution × Distribution

Which term in the product is closest to zero?

DifferentiationAttention competition
What winning looks like

Confidence built through evidence, repetition and recovery.

The characteristic failure

Outsourcing your sense of self to recognition. Praise then sets your direction, and so does criticism.

Confidence = f(Evidence, Repetition, Recovery)

Does repeating this make me more capable, or only more approved of?

Self-signallingSocial comparison
What winning looks like

Deciding from equilibrium. Recognising which internal player currently holds the pen.

The characteristic failure

Local relief that raises the future cost. Avoidance and overwork both do this.

Which version of me is making this call, and on what horizon?

Present self versus future self
What winning looks like

Sources of meaning that do not depend on outranking anyone.

The characteristic failure

Rank-based satisfaction. The reference group climbs with you and the target keeps moving.

Satisfaction = f(rank) ⇒ no terminal state

Would I value this with no spectators?

Finite life, unbounded appetite
What winning looks like

An environment that amplifies you. No location is good in the abstract.

The characteristic failure

Committing permanently on holiday data. A visit and a life are different samples.

What version of me does this place encourage?

MatchingNetwork effectsOption value
What winning looks like

Recovery scheduled as part of the work, including when the work remains unfinished.

The characteristic failure

Every domain withdrawing from one biological account with nobody governing the withdrawals.

Demand > Regeneration ⇒ fragility

Is this replenishing energy or borrowing it from tomorrow?

Common-pool resource
What winning looks like

Output that remains sustainable for years.

The characteristic failure

Converting a temporary biological capacity into a permanent expectation.

Sustainable = Intensity × Consistency × Years

Is the current pace something I could hold for a decade?

Long-horizon resource preservation

A strong result in one domain may do little for a failing one. Money cannot replace recovery, and status cannot supply meaning. Scoring the domains separately makes those gaps harder to ignore.

Explore against exploit

Learning, career and creativity share a choice: use what already works or test an option whose value is still unknown. The upper-confidence-bound formula gives that choice a shape.

UCBi = μ̂i + c · ln t√ni

UCBi
the upper-confidence score assigned to option i
i
one option or arm being compared
μ̂i
the average result observed so far for option i
c
how much weight you give to exploring uncertain options
t
the total number of trials across all options
ni
the number of times option i has been tried
ln
the natural logarithm, which makes the exploration bonus grow slowly over time
the square root keeps the uncertainty bonus from growing too quickly

The first term rewards observed performance. The second adds a temporary bonus when evidence for an option is limited.

Model

Three options, one exploration dial
Grey is evidence, blue is uncertainty

c =

Explore early, after the environment changes, or when returns flatten. Use the proven option when evidence is strong and the advantage is compounding. Revisit the balance as the evidence changes.

Trust as an asset with asymmetric accounting

Tt+1 = (1 − ρ)Tt + aRtbDt,    b > a

Tt
trust held at the start of period t
Tt+1
trust carried into the next period
ρ
the fraction of trust that fades during one quiet period
Rt
reliable behaviour observed during period t
Dt
defection or betrayal observed during period t
a
how much one unit of reliability adds to trust
b
how much one unit of defection removes from trust
b > a
one defection costs more trust than one reliable act adds

Model

Forty rounds of reliability, then one defection
Move the asymmetry
round 0 round 60

b / a =

08 Find the floor that binds

One collapsed part of life can cap everything else, however good the rest looks.

U  ≈  min(x1, x2, …, xn)   compared with   wi xi

U
overall life utility in this simplified model
an approximation, not an exact equality or measured law
xi
the score for life domain i
n
the number of domains included
min
the lowest domain score, treated as the binding constraint
Σ
the alternative model that adds all domain scores
wi
the importance weight assigned to domain i

If the minimum model is useful, the next hour may matter most in the weakest domain. A career at nine does little for a relationship at three.

Dashboard

Score at least five, honestly
Stays in this browser
Career Is my capability growing along with my responsibility?
Moneyfloor Is economic resilience improving?
Ownership Do I capture upside from the value I create?
Relationshipsfloor Is trust increasing?
Homefloor Does home restore energy or consume it?
Emotionalfloor Am I deciding from equilibrium or from strain?
Learning Am I gaining useful future options?
Network Are more of these relationships built on trust?
Reputation Are my claims backed by evidence?
Authority Does my authority match my responsibility?
Meaningfloor Would I value this life with no spectators?
Recoveryfloor Is energy genuinely being replenished?
Optionality Do I have more credible moves than a year ago?

Nothing else substitutes for a domain marked floor. Money cannot buy back recovery, and status cannot supply meaning. A broken floor deserves attention even when everything else looks good.

09 Which regime are you in

Different periods call for different strategies. Expansion rewards scale; recovery protects capacity.

Periods

Six regimes, six different correct answers
Tap one

Readiness against opportunity

Two questions help identify the regime. How ready are you, and how good is the opportunity?

Sorter

Where you are this year
Hover a quadrant
Position Expand Recover Participate More opportunity outside More ready inside

High readiness with little opportunity can feel like failure. It is useful preparation time. Openings arrive on their own schedule, and preparation determines what you can do with them.

10 Fifteen laws

These are the rules I take from the models above.

Laws I actually keep 0 of 15
1
Survival precedes optimisation

Nothing compounds after ruin. Expected value stops being the right measure once a loss ends your ability to keep playing.

2
Most human games repeat

The same people come back. That is what makes reputation an asset and betrayal expensive.

3
Most valuable games are positive-sum

Treating a joint problem as a contest destroys the surplus you were both there to create.

4
Some games really are zero-sum

One promotion, one contract, one auction. Misreading these as cooperative is its own kind of loss.

5
Include the response

Evaluate a decision after everyone else has replied to it.

6
Change the game before playing harder

Effort inside a badly structured game is the most expensive way to lose.

7
Information has option value

Under uncertainty, a cheap experiment is usually worth more than a confident decision.

8
Escalate commitment with evidence

Small bet, then evidence, then large bet. Reversing that order is how people lose big.

9
Sunk cost has no vote

What you already spent is a fact about the past. Only future costs and benefits are decidable.

10
Improve your outside option first

Bargaining power is built before the conversation, by making no-deal survivable.

11
Reputation compounds and decays asymmetrically

It takes years to build and one transaction to spend. Price it accordingly.

12
Incentives beat intentions

When the stated goal and the reward structure disagree, the reward structure wins.

13
Avoid success that costs you agency

A well-paid role you cannot leave is a trap that looks like an achievement.

14
Optimise the trajectory

Ask whether today's choice improves the sequence of states that follows.

15
Choose games carefully

Strong players spend less of their life in games where winning means nothing.

11 Two loops

Both loops compound. One opens future choices. The other makes leaving harder, even while the rewards grow.

The loop that opens doors
  1. Mastery produces something people will pay for
  2. Economic value buys optionality
  3. Optionality improves the choices available
  4. Better choices put you in better environments
  5. Better environments improve judgement
  6. Better judgement deepens mastery

Each round widens the set of games you can enter.

The loop that closes them
  1. Status arrives, and with it commitments
  2. Commitments raise fixed costs
  3. Higher costs create dependency
  4. Dependency makes leaving expensive
  5. Expensive exit produces conformity
  6. Conformity is rewarded with more status

Each round narrows it, and the narrowing is what the reward pays for.

The tell is the direction of your exit costs. If leaving got harder this year while the reward got larger, you are in the second loop regardless of how the first one felt.

arg maxa LTV × Optionality × Alignment TailRisk × Irreversibility × Depletion

a
one action available to you
arg maxa
the action that gives the highest value of the full ratio
LTV
the action's expected long-term value
Optionality
the useful future choices the action preserves or creates
Alignment
how well the action fits your values and intended direction
TailRisk
damage in unusually bad outcomes
Irreversibility
the cost or difficulty of undoing the action
Depletion
the energy and capacity the action consumes over time
×
a weak factor can sharply reduce the whole numerator or enlarge the denominator

The action must also keep P(ruin) < ε, preserve integrity, and leave every foundational domain above its minimum.

12 Why I keep this

I work on calibration: the gap between what a system believes and what is true. Game theory adds an environment that watches the system and responds.

This is also how I think about safety work. A model may optimise its stated objective and ignore how people respond. That is the same error as evaluating a decision without the BR term. More intelligence does not repair a badly specified game.

I do not use the formalism to predict outcomes. Real payoffs are unknown, people are inconsistent, and the model assumes you can rank outcomes you have never experienced. I use it to ask three questions: who else moves, what happens next, and does the bad case leave me able to keep playing?

Treating relationships as games has a cost. Run this analysis on someone you love and you will get an answer. You may dislike who you became to get it. I use game theory on structures and incentives, and try to leave people out of it.

Protect the downside. Keep the ability to walk. Test cheaply before committing heavily. Leave games whose rewards cost you your agency or integrity.

On the maths. Every equilibrium shown on this page is computed by scripts/solve_games.py from the payoff matrix displayed beside it, then checked against the standard properties: dominance, Pareto efficiency, and indifference at any mixed equilibrium. The data file is generated, so every stated equilibrium follows from the displayed numbers. Change a payoff and re-run the script.

The payoffs are ordinal and chosen so each game has its canonical structure. They are not measurements. The ruin simulator uses a fixed seed, which keeps the paths stable as you move a slider.

Nothing you type here is transmitted anywhere. The checklist and dashboard live in your browser's local storage and clearing browser data removes them.

The same treatment is applied to Marcus Aurelius and Epictetus, George Mack's High Agency, and my own operating manual.