Sukoro
Put a number from 1 to 4 in some of the cells and leave the rest empty. A number says how many of its four neighbours also hold numbers. Two numbers that touch must differ, and all the numbered cells form one connected group. Click a cell to cycle it through number → struck out → undecided. You never type a digit: the number is the count of numbered neighbours you have just drawn, so the board writes it for you.
the board
The numbers are not written on the answer — they are the answer
Read the rule again with graph glasses on. A number says how many of its neighbours hold numbers, which is to say: a number is the degree of its own cell in the subgraph induced by the numbered cells. It is not data laid on top of a shape; it is a function of the shape. And the second rule — neighbours must differ — then says no edge of that subgraph joins two vertices of equal degree, which is the textbook definition of a locally irregular graph.
answer = connected, locally irregular induced subgraph of the grid, minimum degree ≥ 1
Everything else on this page falls out of that sentence. Three things
straight away. A lone cell is illegal, because its degree would be 0
and 0 is not one of the numbers. A domino is illegal, because both
cells read 1 and they touch — K₂ is the standard example
of a graph that is not locally irregular, and here it is,
printed on paper. So the smallest legal group is three cells reading
1–2–1. And a solid block is illegal once it has two interior cells:
both read 4, and they touch. Answers can be neither thin nor solid,
which caps how much of a board one can cover — measured below, and
the ceiling is a wall rather than a tendency.
The blank board
Before any clue at all: how many answers does an empty grid have? That is the same as counting the objects themselves.
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1×n | 0 | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 |
| 2×n | 0 | 4 | 12 | 26 | 48 | 80 | 126 | 190 | 278 | 398 | 560 | 778 | 1,070 | 1,460 |
| 3×n | 1 | 12 | 54 | 162 | 428 | 1,012 | 2,322 | 5,134 | 11,338 | 24,580 | ||||
| 4×n | 2 | 26 | 162 | 702 | 2,810 | 10,514 | 38,928 | 143,254 | ||||||
| 5×n | 3 | 48 | 428 | 2,810 | 18,230 | 110,138 | 666,757 | |||||||
| 6×n | 4 | 80 | 1,012 | 10,514 | 110,138 | 1,074,750 | ||||||||
| 7×n | 5 | 126 | 2,322 | 38,928 | 666,757 |
Every figure is exact. The small ones are re-derived from scratch on each test run, and the blank boards up to 4×4 are counted again by the second engine.
The first row is the rule in miniature. A 1×n strip has exactly n − 2 answers. In one row a group is a run of consecutive cells, and a run reads 1, 2, 2, …, 2, 1 — so a run of one has a 0 in it, a run of two has two 1s touching, a run of four or more has two 2s touching, and only a run of exactly three survives. There are n − 2 places to put it. The blank 1×3 is then the one board in this puzzle that is a legal puzzle with nothing printed on it at all: one answer, zero clues.
None of the other rows — 2×n, 3×n, 4×n, 5×n, 6×n, 7×n — return anything from OEIS. The 2×n row looks polynomial for a while and is not: the cubic through 2×2 … 2×5 predicts 80 for 2×6, which is right, and then 124 for 2×7, where the true value is 126.
A clue can only ever say yes
Here is the asymmetry that makes this puzzle's clue language strange. A printed number asserts there is a number in this cell, and it is this one. There is no notation for this cell is empty. Every clue is positive; a setter can only ever add cells to the answer, never take them away. Two things follow, and they pull in opposite directions.
The first is a lemma. Print every number in an answer and the board has exactly one answer. Any rival answer must contain the whole clue set, and if it were strictly bigger, connectivity would put one of its extra cells next to a clued cell and raise that cell's degree above the number printed on it. So the rival is the same set — and the numbers, being the degrees, follow. Checked on 100 of 100 randomly drawn answers, which is all of them.
The second is that almost nothing a setter could print means anything. A clue set is a partial assignment: choose some cells, write 1–4 in each. On a small board that alphabet is finite, so every clue set a setter could possibly print can be enumerated and solved — 1,254,128 of them here.
| board | clues | clue sets | impossible | ambiguous | exactly one |
|---|---|---|---|---|---|
| 3×4 | 1 | 48 | 29.2% | 70.8% | 0.0% |
| 3×4 | 2 | 1,056 | 65.8% | 27.0% | 7.2% |
| 3×4 | 3 | 14,080 | 91.2% | 5.0% | 3.8% |
| 3×4 | 4 | 126,720 | 98.5% | 0.6% | 0.9% |
| 3×5 | 1 | 60 | 26.7% | 73.3% | 0.0% |
| 3×5 | 2 | 1,680 | 60.8% | 34.6% | 4.6% |
| 3×5 | 3 | 29,120 | 87.6% | 9.3% | 3.2% |
| 3×5 | 4 | 349,440 | 97.3% | 1.6% | 1.1% |
| 4×4 | 1 | 64 | 25.0% | 75.0% | 0.0% |
| 4×4 | 2 | 1,920 | 58.6% | 40.1% | 1.3% |
| 4×4 | 3 | 35,840 | 85.6% | 12.0% | 2.4% |
| 4×4 | 4 | 465,920 | 96.6% | 2.3% | 1.1% |
| 4×5 | 1 | 80 | 22.5% | 77.5% | 0.0% |
| 4×5 | 2 | 3,040 | 49.0% | 50.2% | 0.8% |
| 4×5 | 3 | 72,960 | 77.6% | 20.4% | 2.0% |
| 5×5 | 1 | 100 | 20.0% | 80.0% | 0.0% |
| 5×5 | 2 | 4,800 | 42.9% | 57.0% | 0.2% |
| 5×5 | 3 | 147,200 | 68.5% | 30.8% | 0.6% |
91.5% of everything printable is impossible: no answer satisfies it. Only 1.2% is a puzzle. And the dead space grows with the clue count rather than shrinking — at 3×4 with 4 clues it reaches 98.5%, because each extra number is another chance to assert something the grid cannot do. A single clue is never enough anywhere: 0 of the 352 one-clue boards has one answer.
The same census, restricted to the truth
Strip out the lies and the question changes. Take every answer of a blank board, then every subset of that answer's numbers as a clue set. These can never be unsolvable — the answer they came from still satisfies them — so the only failure left is ambiguity, which is exactly the setter's real problem.
| board | answers | true clue sets walked | fewest clues that ever work | median answer needs | worst answer needs |
|---|---|---|---|---|---|
| 3×4 | 162 | 13,582 | 2 | 3 | 4 |
| 3×5 | 428 | 131,428 | 2 | 3 | 6 |
| 4×4 | 702 | 261,346 | 2 | 3 | 6 |
Two numbers can already pin an answer, and the median answer needs 3. That is a very different picture from the raw census above, and the gap between them is the clue language: a setter is not searching a space of clue sets, they are searching a space of answers and then throwing numbers away.
Scaling up
| board | clues | random clue sets: impossible | random: exactly one | true clue sets: exactly one |
|---|---|---|---|---|
| 6×6 | 3 | 53% | 0.2% | 4.8% |
| 6×6 | 5 | 87.4% | 0.2% | 3.1% |
| 6×6 | 8 | 99.8% | 0% | 6.3% |
| 6×6 | 12 | 100% | 0% | 21.1% |
| 8×8 | 3 | 39.3% | 0% | 1.1% |
| 8×8 | 5 | 64.3% | 0% | 0% |
| 8×8 | 8 | 95% | 0% | 0% |
| 8×8 | 12 | 99% | 0% | 2.3% |
| 8×8 | 16 | 100% | 0% | 3.7% |
| 8×8 | 20 | 100% | 0% | 9% |
| 8×8 | 25 | 100% | 0% | 16.9% |
| 10×10 | 3 | 36.7% | 0% | 0% |
| 10×10 | 5 | 62.5% | 0% | 0% |
| 10×10 | 8 | 85% | 0% | 0% |
| 10×10 | 12 | 96.7% | 0% | 0% |
| 10×10 | 16 | 99.2% | 0% | 0% |
| 10×10 | 20 | 100% | 0% | 0% |
| 10×10 | 25 | 100% | 0% | 0% |
Scattering numbers at a board is hopeless well before it gets interesting: 2 of 4,940 random clue sets across every size and count sampled here produced a puzzle, and from 8×8 upwards not one of them did. Reading the clues off a real answer is better but not a method either: at 10×10 a random subset of the answer's own numbers came out unique 0 times in 66 draws, spread over every clue count from 3 to 25. That is a small sample — drawing a random 10×10 answer is itself expensive — but it never once worked. The shipped boards at that size print 16–22 numbers and are unique every time — so which numbers you keep matters as much as how many. The generator does not sample: it draws an answer, prints all of it, and then erases one number at a time for as long as the board stays unique.
What the number is worth, and what the position is worth
A printed number does two jobs at once: it says a cell is numbered, and it says which number. Those can be separated. Replace every clue with a bare mark — "a number lives here, but not which one" — and the first job survives while the second is thrown away.
| board | the shipped clues, numbers and all | the same cells, numbers erased | every answer cell marked, no numbers | every number printed |
|---|---|---|---|---|
| 6×6 | 24 / 24 | 0 / 24 | 0 / 24 | 24 / 24 |
| 8×8 | 24 / 24 | 0 / 24 | 0 / 24 | 24 / 24 |
| 10×10 | 24 / 24 | 0 / 24 | 0 / 24 | 24 / 24 |
The number matters, and it matters more than I expected. Erasing the values off the shipped clue positions leaves 0 of 72 boards unique. Marking every cell of the answer — telling the solver the exact shape and withholding only the digits — also leaves 0 of 72. Writing the numbers on those same cells pins 72 of 72, which is the lemma above. So even a complete map of where the numbers are leaves 72 boards ambiguous. Knowing the shape of the answer everywhere is not the same as knowing the answer.
Erasing a clue cannot break a board
One more consequence of the clues being purely positive. Erase one and the original answer still satisfies what remains, so the board can gain answers but never lose its last one. Move a clue and that protection is gone.
| board | clues erased one at a time | answers left, median | boards made unsolvable | clues moved instead | still unique | made unsolvable |
|---|---|---|---|---|---|---|
| 6×6 | 54 | 6 | 0 | 162 | 3.1% | 80.2% |
| 8×8 | 90 | 8+ | 0 | 270 | 2.2% | 73.7% |
| 10×10 | 153 | 8+ | 0 | 459 | 2.4% | 82.8% |
Counting stops at 8 answers, so 8+ means "at least 8".
0 of 297 erasures made a board unsolvable, which is what the argument predicts and is worth checking anyway. Every one of them cost uniqueness, though — that is what makes the shipped set minimal. Moving a clue somewhere else has no such protection: it keeps uniqueness 3.1%, 2.2% and 2.4% of the time and kills the board outright 80.2%, 73.7% and 82.8% of the time, for the same reason 98.5% of the exhaustive census is impossible. Erasing and moving are the same size of edit and they are not the same kind of edit at all.
What each rule is worth
| rule set | boards that stop being unique | of those, boards left with no answer |
|---|---|---|
| nothing removed | 0 / 72 | 0 |
| adjacent numbers may repeat | 72 / 72 | 0 |
| the numbers need not be connected | 70 / 72 | 0 |
| both removed | 72 / 72 | 0 |
Relaxing a rule can only ever add answers — the intended one still satisfies the weaker rule set — so unlike most ablations in this series nothing here becomes unsolvable, and the column of zeroes on the right is a prediction rather than a surprise. What is left is ambiguity, and there is plenty. Drop the local irregularity rule and all 72 boards lose uniqueness; drop connectivity and 70 of 72 do, the survivors being boards whose clues happen to leave no room for a second group anywhere.
Connectivity also carries the lemma this generator is built on. Printing every number pins the answer on 100/100 answers with the rule in place and on only 39/100 (39%) without it — because the argument for the lemma is exactly "an extra cell would have to touch the clued ones", and that is connectivity talking.
What the ladder costs
Branch points needed to prove the shipped boards unique, at each rung, summed over the set.
deg | neq | conn | probe | |
|---|---|---|---|---|
| 6×6, 24 boards | 3,840,046* | 48,958 | 932 | 2 |
| 8×8, 24 boards | 9,497,265* | 1,862,473* | 5,280 | 0 |
| 10×10, 24 boards | 9,600,024* | 7,618,877* | 19,296 | 64 |
* a lower bound. Each board was cut off at 400,000 branch points, and 71 of the 288 board-and-rung runs hit that ceiling — every one of them on the bottom two rungs. On deg alone the count never finished on 53 of the 72 boards, including every 10×10.
And what each rung settles from an untouched board, before any guess — which is exactly what the checkbox above the board displays:
| unclued cells settled before any guess | deg | neq | conn | probe |
|---|---|---|---|---|
| 6×6 | 11.5% | 23.9% | 27.2% | 97.1% |
| 8×8 | 10.2% | 18.7% | 21.3% | 100% |
| 10×10 | 10.7% | 17% | 19.1% | 94.7% |
deg is pure arithmetic: a cell holding v needs
exactly v numbered neighbours, so v has to sit
between the neighbours already settled as numbered and the neighbours
that could still become numbered. neq is the local
irregularity rule as an elimination. conn is the only
rung that looks at the whole board at once: it strikes out everything
that cannot reach the settled cells, and forces a number into any cell
whose removal would cut the settled cells in two — an articulation
point of the "could still be numbered" graph. Together they settle
27.2%, 21.3% and 19.1% of the
unclued cells and stop. One level of lookahead finishes
68 of 72 boards outright
with no search at all, leaving
2 at 6×6, 0 at 8×8 and 64 at 10×10
branch points to spend on the rest.
Which numbers actually appear
| 1 | 2 | 3 | 4 | |
|---|---|---|---|---|
| the 72 shipped answers | 24.5% | 38.0% | 27.3% | 10.2% |
| the numbers those boards print | 42.8% | 24.0% | 26.6% | 6.6% |
| every answer of a blank 3×5 | 46.5% | 29.5% | 21.9% | 2.2% |
| every answer of a blank 4×4 | 45.2% | 30.0% | 21.7% | 3.0% |
| every answer of a blank 4×6 | 37.6% | 34.3% | 23.2% | 4.9% |
| every answer of a blank 5×5 | 37.0% | 34.8% | 23.1% | 5.2% |
The commonest number in an answer is a 2 (38.0%), and 4s are the rarest everywhere (10.2% in the shipped answers, 5.2% across every answer of a blank 5×5). The geometry says why: a 4 needs all four neighbours numbered and none of them a 4, so it wants a cross of smaller numbers around it, and the local irregularity rule keeps taking that away. On the blank boards the commonest number is a 1 instead. Those are the boards small enough to enumerate — 5×5 and under, where most cells are on an edge and cannot reach 4 at all — so the gap is at least as much about board size as about the rule, and it is recorded rather than explained.
The second row is the one I did not expect. The numbers a setter ends up printing are not drawn in the proportions they occur in: 42.8% of the printed numbers are 1s against 24.5% of the numbers in the answers, and 6.6% are 4s against 10.2%. Nothing in the minimiser knows what a digit is — it erases whatever it can and keeps whatever it cannot. A 1 says its cell has exactly one numbered neighbour, which strikes out three cells at once; a 4 says all four are numbered, which strikes out nothing. The clue that survives erasure is the clue that forbids the most.
How full an answer is allowed to get
| blank board | answers | fullest answer | as a share of the board | average answer |
|---|---|---|---|---|
| 2×4 | 26 | 6 / 8 | 75% | 3.7 |
| 3×3 | 54 | 9 / 9 | 100% | 4.3 |
| 3×4 | 162 | 9 / 12 | 75% | 5.4 |
| 3×5 | 428 | 12 / 15 | 80% | 6.8 |
| 4×4 | 702 | 12 / 16 | 75% | 7.2 |
| 4×5 | 2,810 | 15 / 20 | 75% | 9.4 |
| 4×6 | 10,514 | 18 / 24 | 75% | 11.5 |
| 3×7 | 2,322 | 17 / 21 | 81% | 9.8 |
| 5×5 | 18,230 | 20 / 25 | 80% | 12.2 |
Roughly a fifth to a quarter of the board has to stay empty, and it is a wall rather than a tendency: no thin arms, because two 2s cannot touch, and no solid interior, because two 4s cannot touch. There are exactly 2 exceptions in the whole family — over every rectangle up to 8×12, only 1×3 and 3×3 can be filled completely. The solid 3×3 reads 2 3 2 / 3 4 3 / 2 3 2 and is locally irregular by luck: one interior cell, so there is no second 4 for the 4 to sit beside. On the boards shipped here:
| board | boards | numbers in the answer | numbers printed | printed as a share of the answer |
|---|---|---|---|---|
| 6×6 | 24 | 19 (52.8%) | 7 (19.4%) | 36.8% |
| 8×8 | 24 | 35.5 (55.5%) | 11.7 (18.3%) | 33% |
| 10×10 | 24 | 53.9 (53.9%) | 18.8 (18.8%) | 34.9% |
Every figure above is generated from src/stats.json and
src/ledger.json. All 72 shipped boards are
cross-checked by the second engine, with
0 disagreements.