A distribution fixes the total

After thiswhat you will be able to doCheck whether finite shares form a distribution, combine them into event weights, and renormalise the survivors after removing outcomes.

Questionwhat this lesson answersWhen several outcomes share one fixed whole, what does a distribution commit you to, what does that fixed total let you calculate, and what can no single share say?

Not coveredwhat this lesson leaves outWe stay with finitely many separate outcomes. We do not cover continuously varying quantities, averages or named families of distributions, and we name but do not develop how assigned shares are tested against what happened.

A rolled die must settle on exactly one face. A forecast for tomorrow must eventually meet one recorded weather outcome. A machine choosing the next word must place one word after the last. The objects differ, but the problem has the same shape: one result will occur out of several possible results, and it has not occurred yet.

Before that choice, all you can attach to each possible result is some amount of weight. You may treat all six die faces alike, put more weight on rain than snow, or let a procedure produce a different weight for every possible word. A possible result in a list where exactly one will occur is called an outcome.

Make one whole on purpose

Attach a number pi{p_i} to outcome i{i} and call it that outcome’s share. The complete list of shares is called a distribution only when it obeys two rules:

  1. No share is negative.
  2. All the shares add to exactly one.

Both rules are definitions. They were chosen because they make the list behave like parts of one whole. Nothing in a die, a cloud or a machine forces us to use them.

Allow a negative share and adding an outcome to a collection could make the collection weigh less. The number would no longer describe a part of a whole. Let the total float and the scale becomes arbitrary: the lists (2,3,5){(2, 3, 5)} and (20,30,50){(20, 30, 50)} have the same proportions but different totals. You could not find the weight of everything outside a collection by subtracting from one. Dividing every non-negative weight by their total fixes that scale. This operation is called normalising.

The fixed total does work

Take several outcomes and ask whether any one of them occurs. A collection used as one combined question is called an event. Its weight is the sum of its outcomes’ shares. Give every face of a six-faced die the same share, then the event “an even number” contains three faces:

W(even)=16+16+16=12.W(\text{even}) = \dfrac{1}{6} + \dfrac{1}{6} + \dfrac{1}{6} = \dfrac{1}{2}.

No extra rule for even numbers was needed. The rule for a collection follows from the shares. Everything outside an event has weight 1W(A){1 - W(A)}. Every individual share is also at most one, because it is one non-negative part of a total that is exactly one. That upper bound is a consequence, not a third rule.

The same constraint has a cost. Raise one share by some amount and the other shares together must fall by exactly that amount. They are locked together by the total. When shares express confidence, there is no way to become more confident about one outcome without becoming less confident about the rest. That is a real restriction on what a distribution can represent.

A share does not report what happened

A share of 0.9{0.9} is not a statement that its outcome will occur. If the other outcome occurs, that one case has not falsified the share. A share is also not a recommendation, a quality score or a claim that its outcome is good. A 0.9{0.9} share on rain says nothing about whether rain is welcome.

When a share is intended as a claim about how often an outcome occurs, it is usually called a probability. That name does not turn it into a promise. It makes a claim about repeated assignments: among all comparable cases given a share of 0.9{0.9}, the outcome should occur about nine times in ten. Agreement between assigned shares and long-run occurrence rates is called calibration. One case can never establish or refute it. A long run can.

The list does not reveal its source

Suppose one hundred recorded days contain seventy dry days, twenty rainy days and ten snowy days. Normalising those counts gives

(70,20,10)(0.7,0.2,0.1).(70, 20, 10) \longmapsto (0.7, 0.2, 0.1).

Those shares summarise that set of observations. Whether they carry over to tomorrow depends on whether the recorded days are relevant and representative, which has to be checked against more weather.

A forecaster, a mathematical model or a person could instead assert the identical list without counting those days. The rules guarantee that the result is a distribution. They do not guarantee that it matches anything that happens. Nothing in (0.7,0.2,0.1){(0.7, 0.2, 0.1)} records which origin it had. Treating an asserted list as if it were an observed frequency is the practical trap.

Compare how spread out the shares are

The flattest possible list gives every outcome the same share. This even arrangement is called a uniform distribution. With n{n} outcomes, each share is 1n{\dfrac{1}{n}}. At the other end, nearly all the weight sits on one outcome.

One number puts both ends on the same scale. Define the effective number of choices by

Neff=2ipilog2pi.N_{\mathrm{eff}} = 2^{-\sum_i p_i \log_2 p_i}.

A zero share contributes zero to the sum. This lesson uses the logarithm in that construction without deriving it. A uniform list over n{n} outcomes has exactly n{n} effective choices. As a list concentrates on one outcome, its effective number approaches one.

Removing outcomes makes a different distribution

Start with shares (0.50,0.30,0.15,0.05){(0.50, 0.30, 0.15, 0.05)} and keep only the two largest. The survivors hold S=0.80{S = 0.80} and the discarded outcomes held 0.20{0.20}. To make a distribution again, divide each survivor by the amount left. Making the survivors add to one again is called renormalising:

p1=0.500.80=0.625,p2=0.300.80=0.375.\begin{aligned} p'_1 &= \dfrac{0.50}{0.80} = 0.625, \\ p'_2 &= \dfrac{0.30}{0.80} = 0.375. \end{aligned}

The first survivor gains 0.125{0.125} and the second gains 0.075{0.075}, exactly the discarded 0.20{0.20} between them. Their ratio stays fixed:

0.6250.375=0.500.30=53.\dfrac{0.625}{0.375} = \dfrac{0.50}{0.30} = \dfrac{5}{3}.

Move one weight, move every share

One whole has no spare room

Treat these as four forecasts with exactly one recorded for tomorrow. Move any starting weight. Every share is worked out again, so the whole remains fixed while the other outcomes give up or take back room.

  • weight 5 × 101
    Before cut5 × 10-1
    After cut5 × 10-1

    Gains 0

  • weight 3 × 101
    Before cut3 × 10-1
    After cut3 × 10-1

    Gains 0

  • weight 1.5 × 101
    Before cut1.5 × 10-1
    After cut1.5 × 10-1

    Gains 0

  • weight 5
    Before cut5 × 10-2
    After cut5 × 10-2

    Gains 0

Remove outcomes, then make one whole again
Total before cut
1
Total after cut
1
Effective choices before
3.133
Effective choices after
3.133
Weight discarded
0
Total gained by survivors
0

The shares before cutting total 1. The shares after cutting total 1. The survivors gained 0, exactly the weight removed from the discarded outcomes.

Renormalising has not put the discarded weight somewhere neutral. It has moved that weight onto the survivors in proportion to what they already held. Every surviving share now makes a different claim. The new list is a different distribution.

Doorswhat to read next, and why

Symbolswhat each one means, and whether we defined it, measured it, or just started there

p_iStatus: defined
The number p_i is defined as the share assigned to outcome i in this finite list.
iStatus: defined
The label i is defined as the position of one outcome in the finite list.
p_i >= 0Status: defined
Shares are defined never to be negative, so adding an outcome to a collection cannot lower that collection's weight.
sum of p_i = 1Status: defined
The shares are defined to add to one whole, which fixes their scale and makes every complement available by subtraction.
W(A)Status: defined
The weight W(A) of a collection A is defined as the sum of the shares of the outcomes in that collection.
AStatus: defined
The label A is defined as one chosen collection of outcomes, also called an event.
calibrationStatus: empirical
Whether outcomes assigned a given share occur at that rate across many cases is measured against what happened and can be false.
counted or assertedStatus: empirical
Whether a list came from counted cases or was asserted by a person or procedure is a fact about its history that cannot be read from its shape.
u_i = 1/nStatus: defined
The flattest list is defined to give each of its n outcomes the same share of one over n.
N_effStatus: defined
The effective number of choices is defined as the size of a flat list with the same spread as the shares being examined.
log_2Status: door
The base-two logarithm is used inside the effective-choice definition without being derived here, and is earned in the linked logarithms lesson.
SStatus: defined
The number S is defined as the total share held by the outcomes that survive a cut.
p_i' = p_i / SStatus: defined
Renormalising is defined to divide every surviving share by the surviving total S, which preserves survivor ratios while making their new total one.
What these classifications mean
defined
circular by construction, true because we chose it
empirical
a measured claim about the world that could have come out otherwise
bottoms out
a primitive of the model, with nothing under it here
door
used here, explained elsewhere