A shipping-data vendor is merging this port registry into its own. Its matching rule is 'same name', which merges every port called after a saint into one and leaves two entries for one harbour spelled the same way at either end of the quay. The data lead wants the registry sized for de-duplication before anyone writes the merge: how many names repeat, how many of those repeats are close enough to be one place, and how much of the registry lacks the website that would have settled it.
Report:
names_used_more_than_once — how many distinct name values appear on two or more portsduplicate_pairs_within_20km — how many pairs of ports share a name and lie within 20 km of each other (great-circle)no_website — ports whose website is nullbusiest_name — the name used by the most portsA pair is two ports; three ports with one name and all within 20 km of each other are three pairs. A shared name 400 km apart is a homonym, not a duplicate, and is not counted.
These are where the teaching is. Read them twice.
Name-plus-distance is the standard first pass of every facility-registry merge, from ports to pharmacies, and the pair count is the number the merge is budgeted on.
Work the problem in whatever tool you like, then enter the answers here.
names_used_more_than_onceduplicate_pairs_within_20kmno_websitebusiest_name4 scored, 0 informational
Distinct names used more than once
20%Same-name pairs within 20 km
40%Ports with no website
20%The most-used name
20%