The archive against exact theory.
Three questions where mathematics gives an exact answer before anyone looks at the data — how many numbers should still be waiting for a 1st prize, what shape the digit sums should take, and how long a number should sit out between wins. Then 40 years of real boards, checked against all three.
Digit sums vs the exact curve
Do winning numbers really cluster around a digit sum of 18?
Add the four digits of a number and you get something between 0 (0000) and 36 (9999). Those sums are not equally likely: exactly one number sums to 0, one sums to 36, and 670 sum to 18. The full curve is computed exactly by convolution over all 10,000 numbers — no simulation — and it is perfectly symmetric about 18. So winners DO cluster around 18, but only because there are more numbers there to draw.
Binning rule: no expected cell below 5 (Cochran). The smallest expected cell here holds 61, so 0 cells needed merging — the tails are tested at full single-sum resolution.
Our winner-profile prediction stream filters to numbers whose four digits all differ, then keeps a hand-chosen digit-sum band of 13-23. Against the exact distribution of those 5,040 numbers, 13-23 keeps 74.3% of them — barely a filter at all. The narrowest band centred on 18 that holds half the mass is 15-21: 52.4% of the pool, 2,640 numbers. Over all 10,000 numbers instead, the equivalent band is 14-22 (55.2%).
We report this number; we have not changed the stream. Its picks are locked before each draw and scored in public, so moving the pool would reset the comparison that makes the scoreboard worth reading.
Numbers within one board are drawn without replacement, so the ~23 numbers of a single draw are very slightly dependent. At 23 in 10,000 the effect on the test is negligible and it errs conservative.
The "never drawn" club
How many 4-digit numbers have never taken a 1st prize?
This is the coupon-collector problem. After n independent draws from 10,000 numbers, the expected number still uncollected is 10,000 x (1 - 1/10,000) to the power n. Magnum has run more 1st prizes than any other operator and still should have thousands of numbers waiting — that is arithmetic, not bad luck.
| Operator | 1st prizes | Expected | Observed | z | p |
|---|---|---|---|---|---|
| Magnum 4D | 6,805 | 5063.5 | 5,091 | 1.00 | 0.316 |
| Da Ma Cai 1+3D | 5,976 | 5501.1 | 5,501 | -0.01 | 0.996 |
| SportsToto 4D | 5,674 | 5669.8 | 5,667 | -0.11 | 0.911 |
| Grand Dragon | 2,241 | 7992.3 | 7,996 | 0.28 | 0.776 |
| Special CashSweep | 1,526 | 8584.7 | 8,585 | 0.04 | 0.972 |
| Sandakan 4D | 1,521 | 8589.0 | 8,603 | 1.48 | 0.138 |
| Singapore 4D | 1,468 | 8634.6 | 8,642 | 0.81 | 0.421 |
| Sabah 88 4D | 1,385 | 8706.6 | 8,722 | 1.77 | 0.077 |
1st prize only: across 26,608 1st prizes from every operator, 668 numbers have never taken one anywhere. Theory says 699 (z = -1.35, p = 0.176).
Any prize tier: across 612,168 published prize slots, 0 numbers have never appeared at any operator in any tier. These are different questions and they have different answers.
When an operator shows many more never-drawn numbers than another, the reason is almost always a shorter archive, not a shy machine — the expectation column already accounts for it, which is why every row lands close to its own prediction. A count well ABOVE its own expectation would point at missing draws in our archive first.
How long a number waits before it returns
Is a number that hasn't shown up for 300 draws unusual?
Gaps are counted in DRAWS, never in days, and always within one operator: Magnum draws about three nights a week and Grand Dragon draws seven, so a gap in days would be measuring the schedule instead of the numbers. Each board publishes about 23 of the 10,000 numbers, so under fairness a number returns with probability about 23/10,000 each draw and the gap follows a geometric distribution. We take that probability from each operator's OWN average board size rather than assuming 23.
| Operator | Draws | Theory gap | Observed gap | Gaps counted | p |
|---|---|---|---|---|---|
| Magnum 4D22.98 numbers per board | 6,800 | 435 | 406median 281 | 146,260 | 0.400χ²=13.6/13 |
| Da Ma Cai 1+3D22.97 numbers per board | 5,966 | 435 | 402median 280 | 127,063 | 0.472χ²=12.7/13 |
| SportsToto 4D23.05 numbers per board | 5,674 | 434 | 399median 278 | 120,768 | 0.195χ²=17.1/13 |
| Grand Dragon22.97 numbers per board | 2,241 | 435 | 333median 240 | 41,549 | 0.880χ²=7.4/13 |
| Special CashSweep23.00 numbers per board | 1,526 | 435 | 287median 213 | 25,401 | 0.995χ²=3.6/13 |
| Sandakan 4D23.00 numbers per board | 1,521 | 435 | 283median 208 | 25,280 | 0.131χ²=18.8/13 |
| Singapore 4D22.98 numbers per board | 1,468 | 435 | 284median 212 | 24,077 | 0.317χ²=14.8/13 |
| Sabah 88 4D22.98 numbers per board | 1,385 | 435 | 275median 205 | 22,197 | 0.315χ²=14.9/13 |
A number's wait since its last appearance has not ended, so it is not a completed gap and is excluded — 78,617 of them are still running right now. That exclusion biases the sample toward SHORTER gaps, because a long wait is more likely to still be running when the archive ends. The theory side is corrected the same way: in an archive of N draws there are only N - g places a completed gap of length g can sit, so long gaps are rarer than the textbook geometric says.
Skip that correction and every operator fails spectacularly on data that is perfectly fair: Sandakan 4D alone gives chi-square 2249 on 13 degrees of freedom. That is the single most common way to get this analysis wrong, so the wrong answer is published next to the right one.
13 pairs of consecutive draw dates in the archive carry byte-identical boards (Magnum 4D, Da Ma Cai 1+3D). That is one board filed twice, not a night when all 23 numbers repeated, so those boards are dropped from the sequence. Left in, they put a 71% excess on the one-draw gap and nothing anywhere else.
Across every operator's board pooled together, a number waits a median of 136 days between wins, 235 days on average, with the middle half falling between 53 and 302 days (599,135 completed waits). This is the day-based, any-operator figure printed beside each number's own gap on its page — a different question from the per-operator draw counts above.
A long gap is not a signal. Every draw is independent: a number absent for 800 draws has exactly the same chance tonight as one that won yesterday. The histogram tells you how common such an absence is, not what happens next.
How to read this page
A p-value answers one question: if the draws were fair and the theory correct, how often would the archive look at least this far off by luck alone? A small p is surprising under fairness. It is not the probability that anything is wrong.
We ran 20 tests in a single batch. About one in twenty lands under p = 0.05 even on perfectly fair data, so with this many tests a couple of "significant" results are guaranteed noise. The flag line is therefore p < 2.5e-3 — 0.05 divided by the number of tests.
Both chi-square tests use the same binning rule: no expected cell below 5, with the sparse tails merged into their neighbours until that holds. The number of merged cells is published beside each test so you can see how much of the tail was folded away.
The expected outcome for all of these is agreement, and agreement is a publishable result. Exact theory that survives 40 years of real draws is what lets the rest of this site say "that pattern is just the maths" instead of merely asserting it.