I would be interested in other users’ comments on RM’s duplicate search function. With a large tree, finding duplicates is an important task. RM helps with this - I have certainly found some duplicates using its duplicate search - but it seems to me that it’s duplicate search is not as good as it could be.
I think that there are two types of duplicates
- ‘Technical duplicates’ where I (or Ancestry’s wonderful software) entered a child, spouse or parent twice.
- ‘Organic duplicates’ where a person who already features in my tree, usually as a child or a spouse, then crops up in a different part of the tree as a child or a spouse. Most frequently, a child in one part of the tree is a spouse in another, although sometimes a spouse in one part of the tree is also a spouse in the other.
The criteria for finding people in the two groups are different. For the first group, I would like to find people who are absolutely or partially duplicated. They should have at least one relative in common - a spouse or a parent - and their names and dates should not be inconsistent.
When I run the tool on my 67,000 person database (criteria: names spelled the same, no blank names; compare birth and death places including blanks) I get more than 6,000 pairs. The results come back surprisingly quickly, and I would happily wait much longer for a better result. The first person on the list qualifies as one that I think RM should find (although it is in fact legitimate) but none of the others that I have browsed through do; they all have inconsistent information - different parents, inconsistent dates (usually one child of the same name to the same parents died before another was born) etc. There are far too many pairs on the list for me to browse through them all or to contemplate marking them as not duplicates. It obviously helps that RM sorts them in some kind of priority order; when it has helped me find real duplicates in the past, they have been at the top.
The most common occurrence in the second type, is a person who is a child in one part of the tree and a spouse in another. The characteristics here are often that the names are the same (although I may only know a name and initial in one or both instances), the child has parents and birth details and the spouse has marriage details and perhaps a death, with parents and birth unknown. The time between the child’s birth and the spouse’s marriage should be reasonable. In the other case, the person may appear as spouse twice, perhaps with no parents or birth details in either case, but certainly without inconsistent details and with two marriages close enough to be within the same lifetime.
The criteria here should be that the names are identical or at least consistent, that there is no inconsistency in parents or birth details and that the date of birth in the one instance is consistent with the date of marriage in the other; birth and marriage physically close to each other (same county, state, country) should probably add points. Similarly, an exact match on names should get more points than a match on initials, and matches on rare names more points than on common names.
This morning, I spotted an obvious duplicate in my tree; two results popped up when I searched Ancestry for ‘Elizabeth Joyce Trapnell’. In one instance she was born in England in 1908; her father was English but her mother Irish. This person had no marriage event. In another instance, she married in Ireland in 1946. (Her husband was born in 1903 and his first wife died in 1945.) This person had no birth event. Neither entry had a death event. The Irish marriage register confirms her father’s name, but that was only for confirmation; it was already obvious that the two entries were duplicates.
Try as I may, I can’t get this pair to appear in RM’s duplicate search. I would have thought that if I de-selected the options to compare birth and death dates and places, (just retaining a tick by names spelled the same) and told the tool to start at the name ‘Trapnell’ that she would appear at the top, but she didn’t feature at all, even though her name was spelled identically in her two instances. I can’t work out why not.
Although I can alter the criteria for the searches, I can’t really match what I want for either type of duplicate. It seems to me that the functionality could generally be better, but I am not sure exactly how I would design it if I were building it myself. I would be interested in other users’ comments.





