Circle area is the size of the number. In each term, red adds to the prediction and blue takes away.
This is R2, written here as a percentage. It is the share of the differences between areas that the two predictors account for. At 0 per cent the model tells you nothing. At 100 per cent it would predict every area exactly.
A higher R2 does not mean a better description of the world. Grouping areas together usually raises it. Fotheringham and Wong grouped 871 Buffalo block groups into fewer and fewer zones: fit averaged about 0.4 at 800 zones and about 0.85 at 100. Their conclusion was that you can reach any level of accuracy you like by aggregating enough.
Where the lines go matters as much as how many there are. Openshaw took the 99 counties of Iowa and searched for the 6-region grouping that would give him the answer he wanted. He could drive the correlation anywhere from −0.99 to +0.99. He called it applied gerrymandering. Even without searching, ordinary six-region groupings of the same data ranged from 0.26 to 0.86.
For more, see: Openshaw, S. (1984) The Modifiable Areal Unit Problem. Concepts and Techniques in Modern Geography 38. Norwich: Geo Books. · Fotheringham, A. S. and Wong, D. W. S. (1991) The modifiable areal unit problem in multivariate statistical analysis. Environment and Planning A 23(7), 1025–1044.
One equation for the whole city. It predicts how many incidents of the chosen kind were reported in an area, from two things: how many private households the area holds, and the median household income there, in thousands of dollars.
Each number in front of a predictor says how much the prediction changes when that predictor goes up by one, holding the other still. A term is crossed out when it could easily have come out that size by chance (p is 0.05 or more), which is the usual reason to leave it out of the equation you write down.
The equation describes areas, not people. Nobody in this data is a person. Reading a coefficient as a fact about residents is the ecological fallacy, and it is one of the things this page is built to make visible.
The income term is where that goes wrong most easily. A negative coefficient does not say that people with lower incomes steal or damage more. It says areas whose residents report lower median incomes hold more records. Those areas also differ in what the land is used for, how many people pass through in a day, how much of the space is public, and how closely it is policed.
The count is of who lives here. Most of these incidents happen where people are, not where they sleep. Work on Vancouver found that using the population actually present rather than the resident population changed which places looked like hot spots, and for several kinds of incident they nearly disappeared.
For more, see: Andresen, M. A. and Brantingham, P. J. (2007) Hot Spots of Crime in Vancouver and their Relationship with Population Characteristics. Ottawa: Department of Justice Canada.
The residual errors are what the model failed to account for, and they should be scattered. If areas where it guesses too low sit next to other areas where it guesses too low, something with a geography to it is missing from the model.
Moran's I measures that clustering: near 0 means scattered, positive means neighbours resemble each other. Neighbours here are areas that touch, each counted equally.
The usual reading of clustered errors is that a variable with a geography is missing from the model. Recording and policing are two such variables: where officers are sent, how readily an incident is reported, and how it is written up all vary across a city. Part of what clusters here is the recording rather than the thing recorded, and no amount of care with the statistics separates them.
The p-value comes from shuffling the residual errors between areas 999 times and asking how often chance alone beats what we see. Treat it as a rough guide: they are left over from a model already fitted to this data, so the shuffling is not quite testing what it appears to.
For more, see: Cliff, A. D. and Ord, J. K. (1972) Testing for spatial autocorrelation among regression residuals. Geographical Analysis 4(3), 267–284. · Cliff, A. D. and Ord, J. K. (1973) Spatial Autocorrelation. London: Pion.