The archive against exact theory.

Three questions where mathematics gives an exact answer before anyone looks at the data — how many numbers should still be waiting for a 1st prize, what shape the digit sums should take, and how long a number should sit out between wins. Then 40 years of real boards, checked against all three.

The archive matches the maths. 20 tests in this batch, 0 below the flag line of p < 2.5e-3. That line is 0.05 divided by the number of tests, because running this many comparisons guarantees a couple of lucky-looking results on perfectly fair data.

Digit sums vs the exact curve

Do winning numbers really cluster around a digit sum of 18?

Matches theory

Add the four digits of a number and you get something between 0 (0000) and 36 (9999). Those sums are not equally likely: exactly one number sums to 0, one sums to 36, and 670 sum to 18. The full curve is computed exactly by convolution over all 10,000 numbers — no simulation — and it is perfectly symmetric about 18. So winners DO cluster around 18, but only because there are more numbers there to draw.

Exact theoryObserved
0369121518212427303336
Digit sum 0-36 along the bottom. Grey is what the exact combinatorics predicts for this many prize numbers; blue is what actually came out.
Result
612,168 published prize numbers. Chi-square 38.08 on 36 degrees of freedom, p = 0.375. Mean digit sum 17.995 against a theoretical 18.

Binning rule: no expected cell below 5 (Cochran). The smallest expected cell here holds 61, so 0 cells needed merging — the tails are tested at full single-sum resolution.

What the profile stream's sum band should be

Our winner-profile prediction stream filters to numbers whose four digits all differ, then keeps a hand-chosen digit-sum band of 13-23. Against the exact distribution of those 5,040 numbers, 13-23 keeps 74.3% of them — barely a filter at all. The narrowest band centred on 18 that holds half the mass is 15-21: 52.4% of the pool, 2,640 numbers. Over all 10,000 numbers instead, the equivalent band is 14-22 (55.2%).

We report this number; we have not changed the stream. Its picks are locked before each draw and scored in public, so moving the pool would reset the comparison that makes the scoreboard worth reading.

Numbers within one board are drawn without replacement, so the ~23 numbers of a single draw are very slightly dependent. At 23 in 10,000 the effect on the test is negligible and it errs conservative.

The "never drawn" club

How many 4-digit numbers have never taken a 1st prize?

Matches theory

This is the coupon-collector problem. After n independent draws from 10,000 numbers, the expected number still uncollected is 10,000 x (1 - 1/10,000) to the power n. Magnum has run more 1st prizes than any other operator and still should have thousands of numbers waiting — that is arithmetic, not bad luck.

Operator1st prizesExpectedObservedzp
Magnum 4D6,8055063.55,0911.000.316
Da Ma Cai 1+3D5,9765501.15,501-0.010.996
SportsToto 4D5,6745669.85,667-0.110.911
Grand Dragon2,2417992.37,9960.280.776
Special CashSweep1,5268584.78,5850.040.972
Sandakan 4D1,5218589.08,6031.480.138
Singapore 4D1,4688634.68,6420.810.421
Sabah 88 4D1,3858706.68,7221.770.077
All operators together

1st prize only: across 26,608 1st prizes from every operator, 668 numbers have never taken one anywhere. Theory says 699 (z = -1.35, p = 0.176).

Any prize tier: across 612,168 published prize slots, 0 numbers have never appeared at any operator in any tier. These are different questions and they have different answers.

When an operator shows many more never-drawn numbers than another, the reason is almost always a shorter archive, not a shy machine — the expectation column already accounts for it, which is why every row lands close to its own prediction. A count well ABOVE its own expectation would point at missing draws in our archive first.

How long a number waits before it returns

Is a number that hasn't shown up for 300 draws unusual?

Matches theory

Gaps are counted in DRAWS, never in days, and always within one operator: Magnum draws about three nights a week and Grand Dragon draws seven, so a gap in days would be measuring the schedule instead of the numbers. Each board publishes about 23 of the 10,000 numbers, so under fairness a number returns with probability about 23/10,000 each draw and the gap follows a geometric distribution. We take that probability from each operator's OWN average board size rather than assuming 23.

OperatorDrawsTheory gapObserved gapGaps countedp
Magnum 4D22.98 numbers per board6,800435406median 281146,2600.400χ²=13.6/13
Da Ma Cai 1+3D22.97 numbers per board5,966435402median 280127,0630.472χ²=12.7/13
SportsToto 4D23.05 numbers per board5,674434399median 278120,7680.195χ²=17.1/13
Grand Dragon22.97 numbers per board2,241435333median 24041,5490.880χ²=7.4/13
Special CashSweep23.00 numbers per board1,526435287median 21325,4010.995χ²=3.6/13
Sandakan 4D23.00 numbers per board1,521435283median 20825,2800.131χ²=18.8/13
Singapore 4D22.98 numbers per board1,468435284median 21224,0770.317χ²=14.8/13
Sabah 88 4D22.98 numbers per board1,385435275median 20522,1970.315χ²=14.9/13
Magnum 4D — gap histogram, observed against theory
1340 / 359
2371 / 358
3363 / 357
4345 / 356
5373 / 356
6–101,771 / 1,765
11–203,451 / 3,465
21–406,786 / 6,681
41–8012,159 / 12,419
81–16021,657 / 21,469
161–32032,260 / 32,148
321–64036,131 / 36,352
641–128024,069 / 24,028
1281–67996,184 / 6,148
All operators together
The per-operator fits are independent, so their statistics and degrees of freedom add: chi-square 102.9 on 104 degrees of freedom, p = 0.512.
Why the observed gaps look shorter than 1/p

A number's wait since its last appearance has not ended, so it is not a completed gap and is excluded — 78,617 of them are still running right now. That exclusion biases the sample toward SHORTER gaps, because a long wait is more likely to still be running when the archive ends. The theory side is corrected the same way: in an archive of N draws there are only N - g places a completed gap of length g can sit, so long gaps are rarer than the textbook geometric says.

Skip that correction and every operator fails spectacularly on data that is perfectly fair: Sandakan 4D alone gives chi-square 2249 on 13 degrees of freedom. That is the single most common way to get this analysis wrong, so the wrong answer is published next to the right one.

13 pairs of consecutive draw dates in the archive carry byte-identical boards (Magnum 4D, Da Ma Cai 1+3D). That is one board filed twice, not a night when all 23 numbers repeated, so those boards are dropped from the sequence. Left in, they put a 71% excess on the one-draw gap and nothing anywhere else.

Baseline for a single number's page

Across every operator's board pooled together, a number waits a median of 136 days between wins, 235 days on average, with the middle half falling between 53 and 302 days (599,135 completed waits). This is the day-based, any-operator figure printed beside each number's own gap on its page — a different question from the per-operator draw counts above.

A long gap is not a signal. Every draw is independent: a number absent for 800 draws has exactly the same chance tonight as one that won yesterday. The histogram tells you how common such an absence is, not what happens next.

How to read this page

A p-value answers one question: if the draws were fair and the theory correct, how often would the archive look at least this far off by luck alone? A small p is surprising under fairness. It is not the probability that anything is wrong.

We ran 20 tests in a single batch. About one in twenty lands under p = 0.05 even on perfectly fair data, so with this many tests a couple of "significant" results are guaranteed noise. The flag line is therefore p < 2.5e-3 — 0.05 divided by the number of tests.

Both chi-square tests use the same binning rule: no expected cell below 5, with the sparse tails merged into their neighbours until that holds. The number of merged cells is published beside each test so you can see how much of the tail was folded away.

The expected outcome for all of these is agreement, and agreement is a publishable result. Exact theory that survives 40 years of real draws is what lets the rest of this site say "that pattern is just the maths" instead of merely asserting it.