The river-basin commission wants a world choropleth of river kilometres per country from the index it commissioned. The first draft used quintiles, the safe default, and came back with three of its five classes labelled '0 – 0'. The cartographer wants the failure quantified before choosing a scheme, because the same thing will happen on every map of a variable that most countries do not have.
Compute river kilometres per country exactly as the river-kilometres challenge does — cut each river at every border, measure the pieces geodesically with pyproj.Geod, sum per country, zero where there is nothing — and then classify the 241 values:
countries_at_zero — how many countries have exactly zeroq40 — the second quintile break (the 40th percentile, linear interpolation), whole kmq60 — the third, whole kmq80 — the fourth, whole kmequal_interval_bottom_count — how many countries fall in the lowest of five equal-interval classes over the observed rangeequal_interval_top_count — how many fall in the highestThe quintile breaks are the argument: when more than half the values are the same number, several breaks land on it and the classes they bound are empty in every sense but the legend's. Equal intervals fail differently, and the two counts show how.
These are where the teaching is. Read them twice.
Zero-heavy variables — facilities per district, incidents per ward, rivers per country — are the commonest kind of choropleth data and the one every default classification handles worst.
Work the problem in whatever tool you like, then enter the answers here.
countries_at_zeroq40q60q80equal_interval_bottom_countequal_interval_top_count6 scored, 0 informational
Countries with exactly zero river kilometres
15%Second quintile break
15%Third quintile break
15%Fourth quintile break
15%Countries in the lowest equal-interval class
23%Countries in the highest equal-interval class
15%