The pattern
Local Moran’s I
Global Moran’s I
—
Moran’s I compares each square with the squares around it. Near +1 the pattern is clustered: like sits beside like. Near −1 it is dispersed: each square sits beside its opposite, as on a chessboard. Near 0, neighbours are no more alike than unlike.
Dispersed does not mean random. A random scatter scores near zero, not below it. A score below zero means the squares take turns, which is a strong pattern of its own.
And zero does not prove random. It means this measure cannot tell the pattern apart from a random one, which is not the same thing. Load Stripes: it scores exactly zero and is obviously not random. The matching neighbours above and below cancel the differing ones left and right.
Zero does not mean there is no pattern. Load Stripes and look: the score is exactly zero because the matching neighbours above and below cancel the differing ones left and right.
When every square is one colour the number cannot be worked out. There is no variation to compare. The same is true if no square counts as a neighbour.
The second grid gives every square its own score. In the cluster view each square is sorted into one of four kinds: black sitting among black (high–high, red), white among white (low–low, blue), and the two odd ones out, black among white (high–low) and white among black (low–high), in paler shades. Switch to Local I to see the underlying numbers instead: their average is exactly the single number above.
Which squares count as neighbours is a choice you make, not something the data tells you. The small grid sets it. The dot in the middle is the square being measured; the numbers around it say how much each nearby square counts. Only the ratios matter: doubling every number changes nothing.
Testing against chance. Could a square’s neighbourhood have come out this way by luck? To find out, the grid is shuffled 999 times. Each square keeps its own colour while the squares around it are dealt again at random, and we count how often luck alone produces something as striking as what is really there. If that almost never happens, the square is worth noticing.
Why the two buttons differ. Uncorrected judges each square as if it were the only one you asked about. But you are asking about every square at once, and rare things stop being rare when you look often enough — roll a die enough times and a six is not surprising. Corrected takes that into account: it keeps the share of mistakes among the squares it marks down to about one in twenty.
See it happen. Load Random 2, choose a wide kernel, and switch between the two. Uncorrected marks a scatter of squares in a pattern with no structure whatsoever. Corrected removes every one of them.
Why nothing is ever significant with Rook. The strongest result four neighbours can give is all four matching, and that happens by luck about 6 times in 100 — just above the usual 5 in 100 mark. So no square can pass, whatever the pattern and however big the grid. Sides and corners gives 3 in 1,000, and a wider kernel far less. How many neighbours you count decides what you are able to detect at all.
One limit worth knowing: Moran’s I cannot see direction. Load Half queen, which keeps only the four neighbours on one side, and compare it with Queen. The number does not move at all. Pointing the weights one way changes nothing.
Click a square to change it, and drag to keep painting the same colour. With a keyboard, move with the arrow keys and press space. Press p for the big screen version.
For more, see: Moran, P.A.P. (1950) “Notes on Continuous Stochastic Phenomena.” Biometrika 37(1–2): 17–23, where the measure comes from · Tobler, W.R. (1970) “A Computer Movie Simulating Urban Growth in the Detroit Region.” Economic Geography 46(sup1): 234–240, for the idea that near things are more alike than distant ones.
Made by Luke Bergmann with Claude. Code MIT, text CC BY 4.0. Source.
The number above is one figure for the whole map. The second grid gives every square its own. A square scores high when it and the squares around it are alike, whichever colour they are, and low when it is the odd one out among its neighbours.
Clusters sorts every square into one of four kinds: black sitting among black, white among white, and the two ways of being the odd one out. This is the map most software draws, and the four names — high–high, low–low, high–low, low–high — are what you will meet there. Here high simply means black, because a black square is above the average and a white one below it.
Local I shows the numbers behind those colours. Their average is exactly the single number above: the global figure is not a separate measurement, it is the average of these.
For more, see: Anselin, L. (1995) “Local Indicators of Spatial Association — LISA.” Geographical Analysis 27(2): 93–115.
Which squares count as neighbours is a choice you make, not something the data hands you. The small grid is that choice, drawn out. The dot at the centre is the square being measured; the numbers around it say how much each nearby square counts towards it. Only the ratios matter, so doubling every number changes nothing.
Rook counts the four squares sharing an edge, queen adds the four corners — named, as you would guess, after the way the chess pieces move — and the decay kernels let distant squares count for less rather than not at all. The same map can give very different answers under different choices — the checkerboard scores −1.00 under rook and about −0.03 under queen.
Counting more neighbours is not just a nicer answer, it is the difference between having one and not. With only four neighbours the strongest possible result — all four matching — happens by luck about six times in a hundred, which is commoner than the usual one-in-twenty bar. So under rook nothing can ever be significant, whatever the pattern. Eight neighbours get you to about three in a thousand, which can.
One limit worth knowing: Moran’s I cannot see direction. Load Half queen, which keeps only the four neighbours on one side, and compare it with Queen. The number does not move at all.
For more, see: Cliff, A.D. & Ord, J.K. (1981) Spatial Processes: Models and Applications. London: Pion — the standard treatment of spatial weights, rook’s-case and queen’s-case contiguity among them, and of the distribution theory behind Moran’s I.
The question is whether a square’s neighbourhood could have come out this way by luck. Picture dealing the grid again: this square keeps its own colour, every other square is shuffled at random, and you look at the neighbourhood luck produced. How often does chance do as well as reality? If almost never, the square is worth noticing. That fraction is the p-value.
Sometimes we can work it out, and sometimes we have to try it. When every neighbour counts the same — Rook, Queen, Half queen, Even — the answer has an exact formula, because the neighbours are just so many squares drawn at random from the rest of the grid. No shuffling needed. When neighbours count for different amounts, as in the decay kernels, there is no such formula, so the widget really does shuffle, 999 times.
Shuffling is an estimate, and estimates have their own error. Under Rook, 999 shuffles will report a smallest p-value of about 0.04, when the smallest that is genuinely possible is 0.059 — the simulation inventing extremes that cannot occur. That is why the exact answer is used wherever it exists.
Uncorrected judges each square as though it were the only one you asked about. But you are asking about every square at once, and rare things stop being rare when you look often enough — roll a die twice and a six is mildly interesting, roll it two hundred times and it is not.
Corrected takes that into account, holding the share of mistakes among the squares it marks to about one in twenty. Load Random 2, choose a wide kernel, and switch between the two: uncorrected marks a scatter of squares in a pattern with no structure whatsoever, and correction removes every one.
The cost is real in the other direction too. Correction throws away the weakest genuine findings along with the flukes, which is the trade it exists to make.
The number below the buttons sets the bar, and it is a choice as well. At 0.10 you accept more mistakes in order to miss less; at 0.01 the reverse. Nothing in the data tells you which to pick, and one in twenty is a convention rather than a fact about the world.
It is not the same quantity in the two modes, which is why its name changes. Uncorrected, it is the significance level α: the chance of calling one square remarkable when it is not. Corrected, it is q, the false discovery rate: the share of the squares actually marked that you should expect to be mistakes. The first is a promise about each test taken alone, the second a promise about the map as a whole.
A further nicety: because these p-values come from shuffling rather than from a formula, they are pseudo p-values, and the smallest one 999 shuffles can produce is 0.001.
For more, see: Anselin, L. (1995) “Local Indicators of Spatial Association — LISA.” Geographical Analysis 27(2): 93–115. · Benjamini, Y. & Hochberg, Y. (1995) “Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing.” Journal of the Royal Statistical Society B 57(1): 289–300. · Caldas de Castro, M. & Singer, B.H. (2006) “Controlling the False Discovery Rate: A New Application to Account for Multiple and Dependent Tests in Local Statistics of Spatial Association.” Geographical Analysis 38(2): 180–208.
The squares are the same size in both. What changes is how much of the map you are looking at: the smaller grid is a window cut out of the middle of the larger one, not a different pattern.
Where you draw the boundary changes the answer. Load Patches and switch between the two. The window scores higher than the whole, because it happens to sit over the middle of a large patch. Move the boundary and you would get a different number from the same world — a real difficulty in spatial analysis, not an artefact of this widget.
There is a second effect pulling the other way. More squares means more questions asked at once, which correction has to allow for — but it also means more genuine signal, and that usually wins. Here the two effects fight and the boundary wins: 73 in every 100 squares survive correction on the window against 67 on the whole.