I faced a similar problem years ago (the program was written in C) when processing millions of addresses to eliminate duplicates for a mailing house. It wasn't feasible to compare every address with every other address. What I ended up doing was partitioning the addresses into "buckets" by postcode and comparing all addresses with all other addresses in the same bucket (i.e. with the same postcode) only. To ensure reasonable reliability of postcodes, I had another program to compare suburb names with postcodes, looking for errors.
In reply to Re: Efficient Fuzzy Matching Of An Address
by eyepopslikeamosquito
in thread Efficient Fuzzy Matching Of An Address
by Limbic~Region
| For: | Use: | ||
| & | & | ||
| < | < | ||
| > | > | ||
| [ | [ | ||
| ] | ] |