A distribution fixes the total
After thiswhat you will be able to doCheck whether finite shares form a distribution, combine them into event weights, and renormalise the survivors after removing outcomes.
Questionwhat this lesson answersWhen several outcomes share one fixed whole, what does a distribution commit you to, what does that fixed total let you calculate, and what can no single share say?
Not coveredwhat this lesson leaves outWe stay with finitely many separate outcomes. We do not cover continuously varying quantities, averages or named families of distributions, and we name but do not develop how assigned shares are tested against what happened.
A rolled die must settle on exactly one face. A forecast for tomorrow must eventually meet one recorded weather outcome. A machine choosing the next word must place one word after the last. The objects differ, but the problem has the same shape: one result will occur out of several possible results, and it has not occurred yet.
Before that choice, all you can attach to each possible result is some amount of weight. You may treat all six die faces alike, put more weight on rain than snow, or let a procedure produce a different weight for every possible word. A possible result in a list where exactly one will occur is called an outcome.
Make one whole on purpose
Attach a number to outcome and call it that outcome’s share. The complete list of shares is called a distribution only when it obeys two rules:
- No share is negative.
- All the shares add to exactly one.
Both rules are definitions. They were chosen because they make the list behave like parts of one whole. Nothing in a die, a cloud or a machine forces us to use them.
Allow a negative share and adding an outcome to a collection could make the collection weigh less. The number would no longer describe a part of a whole. Let the total float and the scale becomes arbitrary: the lists and have the same proportions but different totals. You could not find the weight of everything outside a collection by subtracting from one. Dividing every non-negative weight by their total fixes that scale. This operation is called normalising.
The fixed total does work
Take several outcomes and ask whether any one of them occurs. A collection used as one combined question is called an event. Its weight is the sum of its outcomes’ shares. Give every face of a six-faced die the same share, then the event “an even number” contains three faces:
No extra rule for even numbers was needed. The rule for a collection follows from the shares. Everything outside an event has weight . Every individual share is also at most one, because it is one non-negative part of a total that is exactly one. That upper bound is a consequence, not a third rule.
The same constraint has a cost. Raise one share by some amount and the other shares together must fall by exactly that amount. They are locked together by the total. When shares express confidence, there is no way to become more confident about one outcome without becoming less confident about the rest. That is a real restriction on what a distribution can represent.
A share does not report what happened
A share of is not a statement that its outcome will occur. If the other outcome occurs, that one case has not falsified the share. A share is also not a recommendation, a quality score or a claim that its outcome is good. A share on rain says nothing about whether rain is welcome.
When a share is intended as a claim about how often an outcome occurs, it is usually called a probability. That name does not turn it into a promise. It makes a claim about repeated assignments: among all comparable cases given a share of , the outcome should occur about nine times in ten. Agreement between assigned shares and long-run occurrence rates is called calibration. One case can never establish or refute it. A long run can.
The list does not reveal its source
Suppose one hundred recorded days contain seventy dry days, twenty rainy days and ten snowy days. Normalising those counts gives
Those shares summarise that set of observations. Whether they carry over to tomorrow depends on whether the recorded days are relevant and representative, which has to be checked against more weather.
A forecaster, a mathematical model or a person could instead assert the identical list without counting those days. The rules guarantee that the result is a distribution. They do not guarantee that it matches anything that happens. Nothing in records which origin it had. Treating an asserted list as if it were an observed frequency is the practical trap.
Compare how spread out the shares are
The flattest possible list gives every outcome the same share. This even arrangement is called a uniform distribution. With outcomes, each share is . At the other end, nearly all the weight sits on one outcome.
One number puts both ends on the same scale. Define the effective number of choices by
A zero share contributes zero to the sum. This lesson uses the logarithm in that construction without deriving it. A uniform list over outcomes has exactly effective choices. As a list concentrates on one outcome, its effective number approaches one.
Removing outcomes makes a different distribution
Start with shares and keep only the two largest. The survivors hold and the discarded outcomes held . To make a distribution again, divide each survivor by the amount left. Making the survivors add to one again is called renormalising:
The first survivor gains and the second gains , exactly the discarded between them. Their ratio stays fixed:
Move one weight, move every share
One whole has no spare room
Treat these as four forecasts with exactly one recorded for tomorrow. Move any starting weight. Every share is worked out again, so the whole remains fixed while the other outcomes give up or take back room.
- Before cut5 × 10-1After cut5 × 10-1
Gains 0
- Before cut3 × 10-1After cut3 × 10-1
Gains 0
- Before cut1.5 × 10-1After cut1.5 × 10-1
Gains 0
- Before cut5 × 10-2After cut5 × 10-2
Gains 0
- Total before cut
- 1
- Total after cut
- 1
- Effective choices before
- 3.133
- Effective choices after
- 3.133
- Weight discarded
- 0
- Total gained by survivors
- 0
Renormalising has not put the discarded weight somewhere neutral. It has moved that weight onto the survivors in proportion to what they already held. Every surviving share now makes a different claim. The new list is a different distribution.
Doorswhat to read next, and why
- LogarithmsThe effective number of choices is built from a logarithm. This lesson uses that construction without deriving why logarithms turn multiplication into addition.
- ExpectationThis lesson gives outcomes shares but no values. Expectation is what becomes available when each outcome carries both.
- Sampling and decodingnot written yetThis lesson defines shares but stops before selecting an outcome, and it does not compare the decoding rules that turn a model's shares into one vocabulary entry.
- Failure modesnot written yetThis lesson says assigned shares can fail calibration, but it does not explain why a model can produce a false answer with the same fluency as a true one.
- Probability density for continuous quantitiesA quantity that can vary continuously has infinitely many possible values, so assigning a separate share to every value no longer builds the right object.
Symbolswhat each one means, and whether we defined it, measured it, or just started there
- p_iStatus: defined
- The number p_i is defined as the share assigned to outcome i in this finite list.
- iStatus: defined
- The label i is defined as the position of one outcome in the finite list.
- p_i >= 0Status: defined
- Shares are defined never to be negative, so adding an outcome to a collection cannot lower that collection's weight.
- sum of p_i = 1Status: defined
- The shares are defined to add to one whole, which fixes their scale and makes every complement available by subtraction.
- W(A)Status: defined
- The weight W(A) of a collection A is defined as the sum of the shares of the outcomes in that collection.
- AStatus: defined
- The label A is defined as one chosen collection of outcomes, also called an event.
- calibrationStatus: empirical
- Whether outcomes assigned a given share occur at that rate across many cases is measured against what happened and can be false.
- counted or assertedStatus: empirical
- Whether a list came from counted cases or was asserted by a person or procedure is a fact about its history that cannot be read from its shape.
- u_i = 1/nStatus: defined
- The flattest list is defined to give each of its n outcomes the same share of one over n.
- N_effStatus: defined
- The effective number of choices is defined as the size of a flat list with the same spread as the shares being examined.
- log_2Status: door
- The base-two logarithm is used inside the effective-choice definition without being derived here, and is earned in the linked logarithms lesson.
- SStatus: defined
- The number S is defined as the total share held by the outcomes that survive a cut.
- p_i' = p_i / SStatus: defined
- Renormalising is defined to divide every surviving share by the surviving total S, which preserves survivor ratios while making their new total one.
What these classifications mean
- defined
- circular by construction, true because we chose it
- empirical
- a measured claim about the world that could have come out otherwise
- bottoms out
- a primitive of the model, with nothing under it here
- door
- used here, explained elsewhere