Move, response, position
Your move changes their options. Their response changes yours.
Most important decisions involve other people. They respond to your move, and their response changes what your move was worth.
Ask whether a choice still works after everyone responds. A good first move can produce a bad final position.
I use this as a working note. A script computes every equilibrium because Chicken and Stag Hunt are easy to remember backwards. Bad maths quickly becomes bad advice.
maxa EV(a) − BR(a)
- a
- one action available to you
- maxa
- choose the action with the highest resulting value
- EV(a)
- the expected value of action a before anyone responds
- BR(a)
- the cost created by the other players' best response to a
- −
- subtract the response cost from the action's expected value
01 Recognise the game
Recurring problems often share a structure. Once you recognise it, your options become clearer.
Read a matrix like this. You pick a row, the other player picks a column, and the cell shows what each of you gets as (uyou, uthem). A cell is a Nash equilibrium when neither of you improves by changing alone:
u1(s1*, s2*) ≥ u1(s1, s2*) for all s1
- u1
- player one's payoff function
- s1*
- player one's equilibrium strategy
- s2*
- player two's equilibrium strategy, held fixed in this condition
- s1
- any alternative strategy available to player one
- ≥
- switching alone cannot give player one a higher payoff
- for all
- the condition must hold for every s1 in player one's strategy set
- player 2
- the same condition also applies with the two players reversed
A Nash equilibrium is stable. It does not have to be good for either player.
Solver
Pick a game, then click a cell
All eight, by shape
Two things worth noticing as you click through. The Prisoner's Dilemma and Public Goods are the games where both players have a dominant strategy. In both cases that strategy leads somewhere neither wanted. Matching Pennies is genuinely zero-sum and has no stable pure strategy.
Salary talks, arguments with a partner and supplier negotiations rarely have a fixed pot. How you play can create or destroy value. You can win the argument and leave both sides worse off.
02 See all five boards
A major decision affects five games at once. Winning on one board can hide losses on the others.
Layers
Same decision, five boards
The intrapersonal layer is easy to overlook. Today you can make commitments that your future self cannot reverse. Debt, specialisation, lifestyle and reputation all move on this board.
03 Lengthen the horizon
Repeat the Prisoner's Dilemma and cooperation can become rational. The people stay the same; the prospect of another round changes their incentives.
Suppose you cooperate as long as they do, and stop for good if they betray you. Cooperating forever is worth R + δR + δ2R + …, and betraying once pays T now and P forever after. Cooperation survives when:
R1 − δ ≥ T + δP1 − δ ⇔ δ ≥ T − RT − P
- δ
- how much the next round is worth relative to this one, so how much the future matters
- T
- temptation, what betraying a cooperator pays
- R
- reward for mutual cooperation
- P
- punishment when both defect
- 1 − δ
- the denominator of the infinite discounted stream of future payoffs
- ⇔
- the expression on the right is the same condition after rearranging the terms
Model
Move the horizon
Trustworthiness depends partly on the situation. A contractor who will never see you again, a colleague in their notice period, a counterparty in a one-off deal: each is in a low δ game whatever their character. If you want cooperation, lengthen the relationship or add a stake that outlives it. Changing the incentive may be easier than replacing the person.
04 Change the game itself
A better move can still leave you stuck in a bad game. Sometimes you need to change the rules, players, or exit.
Seven things define a game. When you feel stuck, you are usually trying to improve your move while treating the other six as fixed:
G = (N, S, u, I, >, R, E)
- G
- the complete game being described
- N
- the players involved
- S
- the strategies available to those players
- u
- the payoff assigned to each outcome
- I
- who knows what when a decision is made
- >
- the order in which the players move
- R
- the rules that constrain the players and actions
- E
- the outcome available when you exit or refuse to play
Diagnostic
Seven things you can change
I use the last lever most often. Your bargaining power is largely set before you walk in by how survivable no-deal is. In the Nash bargaining solution the split depends on the disagreement point (d1, d2):
max (u1 − d1)(u2 − d2)
- max
- find the agreement that makes the product as large as possible
- u1
- player one's payoff under the proposed agreement
- u2
- player two's payoff under the proposed agreement
- d1
- player one's payoff if no agreement is reached
- d2
- player two's payoff if no agreement is reached
- ui − di
- player i's gain from reaching an agreement
- ( )( )
- the product balances both players' gains over their no-deal outcomes
A better outside option raises d1 before either side speaks. Another offer, a second customer or more runway can do this.
05 Size the bet before the upside
Some risks deserve a chance. A loss that ends the game needs a hard limit.
Sort any decision on two axes: can you undo it, and how bad is the bad case. The four quadrants want genuinely different behaviour.
Sorter
Reversibility against downside
Positive expected value is not enough
A bet can have a genuine edge and still ruin you. You never experience the average across every possible future. You experience one path, in order, and a path that touches zero stops.
For a repeated favourable bet, the growth-optimal stake is the Kelly fraction:
f* = bp − qb
- f*
- the fraction of current wealth that maximises long-run growth
- p
- the probability of winning one round
- q
- the probability of losing, equal to 1 − p
- b
- the net amount won for each unit staked
- bp − q
- the edge after weighting the win and loss outcomes by their probabilities
Real edges are uncertain, so treat f* as a ceiling. More uncertainty calls for a smaller stake.
Simulation
Sixty paths through the same favourable bet
Set the edge positive and the stake high and watch what happens. A favourable game, played too large, still kills a good fraction of the paths. Ruin needs a hard constraint:
P(ruin) < ε and integrity intact
- P(ruin)
- the probability that a loss ends your ability to keep playing
- ε
- the maximum ruin probability you are willing to accept
- <
- ruin risk must stay below that chosen limit
- integrity
- a hard constraint that the strategy is not allowed to violate
- and
- both constraints must hold at the same time
Optimisation happens inside these limits. A high expected value does not cancel ruin risk or an integrity failure.
06 Before you decide
These nine questions reveal the assumptions behind a decision.
Checklist
Name the game first
07 Seventeen domains
Each part of life has its own players, incentives and risks. All seventeen draw from the same limited pool of time, money, energy and attention.
Capability that is hard to replace, with results other people can verify.
Becoming indispensable in a role that creates no bargaining power. Praise rises while outside options stay flat.
Does my future bargaining power grow with this responsibility?
Pay that grows with judgement, scarcity or ownership. Hours worked become less important.
One employer, one client, one credential. Dependency turns a good income into a fragile one.
What mechanism decides what this can eventually pay?
Staying in the game long enough for returns to compound.
Position sizes that carry real ruin risk. One bad sequence and the compounding stops.
If the downside lands at the worst moment, am I still playing?
Holding a claim on value you helped create, so the upside is not entirely someone else's.
Entering because ownership sounds like freedom, without a customer or an unfair advantage.
What happens if this works and the incumbent responds rationally?
A relationship in which conflict can be repaired reliably.
Treating it as zero-sum. Winning the argument while lowering the value of the relationship.
Does repeated interaction here increase trust and capability, or steadily reduce them?
Contributions and authority made explicit, so the load does not silently fall on one person.
Free-riding on invisible labour. Nobody intends it, and the structure produces it anyway.
Is this a character problem or an incentive problem?
A smaller set of trusted relationships can be worth more than a large contact list. These people would still take your call without the job titles.
Treating a network as something to extract from. People remember that behaviour.
Am I contributing here, or only withdrawing?
Signals that are expensive to fake. Shipped work, kept promises, consistent judgement.
Spending a decade of trust on one transaction. It rebuilds far slower than it breaks.
Could someone without the underlying capability afford to send this signal?
Responsibility, authority and accountability roughly in balance.
Responsibility far exceeding authority. You absorb the blame for outcomes you cannot control.
Am I being given power, or only exposure?
Knowledge that transfers, compounds, and eventually changes what you do.
Consuming information because it feels productive, without using it to change a decision or action.
Can I use this yet, or have I only read about it?
Work at the intersection of genuine curiosity and non-generic knowledge.
Optimising novelty without usefulness, or usefulness without distinctiveness.
Which term in the product is closest to zero?
Confidence built through evidence, repetition and recovery.
Outsourcing your sense of self to recognition. Praise then sets your direction, and so does criticism.
Does repeating this make me more capable, or only more approved of?
Deciding from equilibrium. Recognising which internal player currently holds the pen.
Local relief that raises the future cost. Avoidance and overwork both do this.
Which version of me is making this call, and on what horizon?
Sources of meaning that do not depend on outranking anyone.
Rank-based satisfaction. The reference group climbs with you and the target keeps moving.
Would I value this with no spectators?
An environment that amplifies you. No location is good in the abstract.
Committing permanently on holiday data. A visit and a life are different samples.
What version of me does this place encourage?
Recovery scheduled as part of the work, including when the work remains unfinished.
Every domain withdrawing from one biological account with nobody governing the withdrawals.
Is this replenishing energy or borrowing it from tomorrow?
Output that remains sustainable for years.
Converting a temporary biological capacity into a permanent expectation.
Is the current pace something I could hold for a decade?
A strong result in one domain may do little for a failing one. Money cannot replace recovery, and status cannot supply meaning. Scoring the domains separately makes those gaps harder to ignore.
Explore against exploit
Learning, career and creativity share a choice: use what already works or test an option whose value is still unknown. The upper-confidence-bound formula gives that choice a shape.
UCBi = μ̂i + c · √ln t√ni
- UCBi
- the upper-confidence score assigned to option i
- i
- one option or arm being compared
- μ̂i
- the average result observed so far for option i
- c
- how much weight you give to exploring uncertain options
- t
- the total number of trials across all options
- ni
- the number of times option i has been tried
- ln
- the natural logarithm, which makes the exploration bonus grow slowly over time
- √
- the square root keeps the uncertainty bonus from growing too quickly
The first term rewards observed performance. The second adds a temporary bonus when evidence for an option is limited.
Model
Three options, one exploration dial
c =
Explore early, after the environment changes, or when returns flatten. Use the proven option when evidence is strong and the advantage is compounding. Revisit the balance as the evidence changes.
Trust as an asset with asymmetric accounting
Tt+1 = (1 − ρ)Tt + aRt − bDt, b > a
- Tt
- trust held at the start of period t
- Tt+1
- trust carried into the next period
- ρ
- the fraction of trust that fades during one quiet period
- Rt
- reliable behaviour observed during period t
- Dt
- defection or betrayal observed during period t
- a
- how much one unit of reliability adds to trust
- b
- how much one unit of defection removes from trust
- b > a
- one defection costs more trust than one reliable act adds
Model
Forty rounds of reliability, then one defection
b / a =
08 Find the floor that binds
One collapsed part of life can cap everything else, however good the rest looks.
U ≈ min(x1, x2, …, xn) compared with ∑ wi xi
- U
- overall life utility in this simplified model
- ≈
- an approximation, not an exact equality or measured law
- xi
- the score for life domain i
- n
- the number of domains included
- min
- the lowest domain score, treated as the binding constraint
- Σ
- the alternative model that adds all domain scores
- wi
- the importance weight assigned to domain i
If the minimum model is useful, the next hour may matter most in the weakest domain. A career at nine does little for a relationship at three.
Dashboard
Score at least five, honestly
Nothing else substitutes for a domain marked floor. Money cannot buy back recovery, and status cannot supply meaning. A broken floor deserves attention even when everything else looks good.
09 Which regime are you in
Different periods call for different strategies. Expansion rewards scale; recovery protects capacity.
Periods
Six regimes, six different correct answers
Readiness against opportunity
Two questions help identify the regime. How ready are you, and how good is the opportunity?
Sorter
Where you are this year
High readiness with little opportunity can feel like failure. It is useful preparation time. Openings arrive on their own schedule, and preparation determines what you can do with them.
10 Fifteen laws
These are the rules I take from the models above.
Survival precedes optimisation
Nothing compounds after ruin. Expected value stops being the right measure once a loss ends your ability to keep playing.
Most human games repeat
The same people come back. That is what makes reputation an asset and betrayal expensive.
Most valuable games are positive-sum
Treating a joint problem as a contest destroys the surplus you were both there to create.
Some games really are zero-sum
One promotion, one contract, one auction. Misreading these as cooperative is its own kind of loss.
Include the response
Evaluate a decision after everyone else has replied to it.
Change the game before playing harder
Effort inside a badly structured game is the most expensive way to lose.
Information has option value
Under uncertainty, a cheap experiment is usually worth more than a confident decision.
Escalate commitment with evidence
Small bet, then evidence, then large bet. Reversing that order is how people lose big.
Sunk cost has no vote
What you already spent is a fact about the past. Only future costs and benefits are decidable.
Improve your outside option first
Bargaining power is built before the conversation, by making no-deal survivable.
Reputation compounds and decays asymmetrically
It takes years to build and one transaction to spend. Price it accordingly.
Incentives beat intentions
When the stated goal and the reward structure disagree, the reward structure wins.
Avoid success that costs you agency
A well-paid role you cannot leave is a trap that looks like an achievement.
Optimise the trajectory
Ask whether today's choice improves the sequence of states that follows.
Choose games carefully
Strong players spend less of their life in games where winning means nothing.
11 Two loops
Both loops compound. One opens future choices. The other makes leaving harder, even while the rewards grow.
The loop that opens doors
- Mastery produces something people will pay for
- Economic value buys optionality
- Optionality improves the choices available
- Better choices put you in better environments
- Better environments improve judgement
- Better judgement deepens mastery
Each round widens the set of games you can enter.
The loop that closes them
- Status arrives, and with it commitments
- Commitments raise fixed costs
- Higher costs create dependency
- Dependency makes leaving expensive
- Expensive exit produces conformity
- Conformity is rewarded with more status
Each round narrows it, and the narrowing is what the reward pays for.
The tell is the direction of your exit costs. If leaving got harder this year while the reward got larger, you are in the second loop regardless of how the first one felt.
arg maxa LTV × Optionality × Alignment TailRisk × Irreversibility × Depletion
- a
- one action available to you
- arg maxa
- the action that gives the highest value of the full ratio
- LTV
- the action's expected long-term value
- Optionality
- the useful future choices the action preserves or creates
- Alignment
- how well the action fits your values and intended direction
- TailRisk
- damage in unusually bad outcomes
- Irreversibility
- the cost or difficulty of undoing the action
- Depletion
- the energy and capacity the action consumes over time
- ×
- a weak factor can sharply reduce the whole numerator or enlarge the denominator
The action must also keep P(ruin) < ε, preserve integrity, and leave every foundational domain above its minimum.
12 Why I keep this
I work on calibration: the gap between what a system believes and what is true. Game theory adds an environment that watches the system and responds.
This is also how I think about safety work. A model may optimise its stated objective and ignore how people respond. That is the same error as evaluating a decision without the BR term. More intelligence does not repair a badly specified game.
I do not use the formalism to predict outcomes. Real payoffs are unknown, people are inconsistent, and the model assumes you can rank outcomes you have never experienced. I use it to ask three questions: who else moves, what happens next, and does the bad case leave me able to keep playing?
Treating relationships as games has a cost. Run this analysis on someone you love and you will get an answer. You may dislike who you became to get it. I use game theory on structures and incentives, and try to leave people out of it.
Protect the downside. Keep the ability to walk. Test cheaply before committing heavily. Leave games whose rewards cost you your agency or integrity.
On the maths. Every equilibrium shown on this page is computed by
scripts/solve_games.py from the payoff matrix displayed beside it, then checked
against the standard properties: dominance, Pareto efficiency, and indifference at any mixed
equilibrium. The data file is generated, so every stated equilibrium follows from the displayed
numbers. Change a payoff and re-run the script.
The payoffs are ordinal and chosen so each game has its canonical structure. They are not measurements. The ruin simulator uses a fixed seed, which keeps the paths stable as you move a slider.
Nothing you type here is transmitted anywhere. The checklist and dashboard live in your browser's local storage and clearing browser data removes them.
The same treatment is applied to Marcus Aurelius and Epictetus, George Mack's High Agency, and my own operating manual.