Thoughts on duplicate search

I would be interested in other users’ comments on RM’s duplicate search function. With a large tree, finding duplicates is an important task. RM helps with this - I have certainly found some duplicates using its duplicate search - but it seems to me that it’s duplicate search is not as good as it could be.

I think that there are two types of duplicates

  1. ‘Technical duplicates’ where I (or Ancestry’s wonderful software) entered a child, spouse or parent twice.
  2. ‘Organic duplicates’ where a person who already features in my tree, usually as a child or a spouse, then crops up in a different part of the tree as a child or a spouse. Most frequently, a child in one part of the tree is a spouse in another, although sometimes a spouse in one part of the tree is also a spouse in the other.

The criteria for finding people in the two groups are different. For the first group, I would like to find people who are absolutely or partially duplicated. They should have at least one relative in common - a spouse or a parent - and their names and dates should not be inconsistent.

When I run the tool on my 67,000 person database (criteria: names spelled the same, no blank names; compare birth and death places including blanks) I get more than 6,000 pairs. The results come back surprisingly quickly, and I would happily wait much longer for a better result. The first person on the list qualifies as one that I think RM should find (although it is in fact legitimate) but none of the others that I have browsed through do; they all have inconsistent information - different parents, inconsistent dates (usually one child of the same name to the same parents died before another was born) etc. There are far too many pairs on the list for me to browse through them all or to contemplate marking them as not duplicates. It obviously helps that RM sorts them in some kind of priority order; when it has helped me find real duplicates in the past, they have been at the top.

The most common occurrence in the second type, is a person who is a child in one part of the tree and a spouse in another. The characteristics here are often that the names are the same (although I may only know a name and initial in one or both instances), the child has parents and birth details and the spouse has marriage details and perhaps a death, with parents and birth unknown. The time between the child’s birth and the spouse’s marriage should be reasonable. In the other case, the person may appear as spouse twice, perhaps with no parents or birth details in either case, but certainly without inconsistent details and with two marriages close enough to be within the same lifetime.

The criteria here should be that the names are identical or at least consistent, that there is no inconsistency in parents or birth details and that the date of birth in the one instance is consistent with the date of marriage in the other; birth and marriage physically close to each other (same county, state, country) should probably add points. Similarly, an exact match on names should get more points than a match on initials, and matches on rare names more points than on common names.

This morning, I spotted an obvious duplicate in my tree; two results popped up when I searched Ancestry for ‘Elizabeth Joyce Trapnell’. In one instance she was born in England in 1908; her father was English but her mother Irish. This person had no marriage event. In another instance, she married in Ireland in 1946. (Her husband was born in 1903 and his first wife died in 1945.) This person had no birth event. Neither entry had a death event. The Irish marriage register confirms her father’s name, but that was only for confirmation; it was already obvious that the two entries were duplicates.

Try as I may, I can’t get this pair to appear in RM’s duplicate search. I would have thought that if I de-selected the options to compare birth and death dates and places, (just retaining a tick by names spelled the same) and told the tool to start at the name ‘Trapnell’ that she would appear at the top, but she didn’t feature at all, even though her name was spelled identically in her two instances. I can’t work out why not.

Although I can alter the criteria for the searches, I can’t really match what I want for either type of duplicate. It seems to me that the functionality could generally be better, but I am not sure exactly how I would design it if I were building it myself. I would be interested in other users’ comments.

agreed that identifying many duplicate scenarios could/should be better and that how that could happen would be an easy task (or what it would look like) – that does not add much other than that it sounds like we agree.

Check that Elizabeth Joyce Trapnell surname is actually in the surname field on both of them.

2 Likes

Actually I have had that happen with Fam Search or others— the Surname was not in the correct field – using CAPS on surnames might help identify

Since you are able to find both, open each edit screen and see if you can identify any differences. Hope you have RIN turned on :slight_smile:

I have since merged the two records, so I can no longer check them individually, but I have checked the merged record which did not create an alternate name, so I am pretty sure that the two were the same.

To double check, I have added her again in a different family and run the duplicate search check again. I still can’t see her in the results.

I had to browse through the search results. Even starting at ‘Trapnell’ I get more than 6,000 when I don’t check birth or death dates or places, but I can’t see her in the list.

A backup would still have them both.

Run the database tools to rule that out.

1 Like

Alan-- I copied your info for Elizabeth-- just both sets of parents, birthdate/place and hubby–ran the same search and she showed up at the top of my list as she was the only one in the database that needs merged-- her score was 25.0 because beside hubby/ no hubby, you had different parents

Thanks to everyone for their contributions. As @nkess pointed out, experiment trying to repeat the first circumstances did not properly do so, as I ended up with inconsistent parents.

And as @rzamor1 pointe out, it is a good idea to run the database tools. (I had already rebuilt the indexes, as I know this is always required after a manual merge.)

So I set about, first to run the tools, then to create a new Elizabeth Joyce Trapnell without the inconsistent parents, and ran the duplicate search again, starting with Trapnell and only choosing to match on names spelled the same as before. I can’t entirely rule out the possibility that she might be hidden in the list somewhere, but I still can’t see her.

So I set up a new test database. This contains five people, Elizabeth Joyce with her correct birth details and parents and the second Elizabeth Joyce with her correct husband and marriage details. I ran the duplicate search twice, once including birth dates as criteria (no results) and once only matching on name. This did of course produce the desired result. (I never imagined that the code contained a filter to prevent showing people called Elizabeth Joyce Trapnell.)

However, she only got a score of 12.0.

If I run the duplicate search on my main database, with no start surname and only matching on name spelled the same, I get 30,000 entries. The highest score is 50.96 and the lowest 11.53.

I can’t see Elizabeth Joyce in there, but whether she is or is not is not the point. The point is that I would like to be able to find duplicates like her, with the same name, a birth and parents on one record and a marriage and spouse on the other, whether the birth and marriage dates and places are consistent. I don’t see how I can use the current tool to do this.

Instead, I get loads of entries like this

Both people have the same name and parents, but their birth dates, although close, are clearly inconsistent.

As an aside, and to the RM tool’s credit, top of the list of duplicates, one above Doughty Wormell, was a real duplicate who also had a duplicate wife. This was a third category, separate from the two in my first post, where I had duplicated a child in the wrong family.

1 Like

It’s just a “legacy” auxilliary command with some functionality from back in earlier versions… likely why it’s not under Search or Tools.

The points are dependent on the amount of information the tool is working with. There’s a weighted-value assigned to each element of potentially matching information. Some things, like the Middle name being Joyce may not even matter …or if it does, maybe it doesn’t add much in terms of “points”, because many folks could have that middle name and thusly it’s help in differentiating or confirming has little value. A Given and Surname match (ONLY) is also likely low-scoring in and of itself. (ie. there could be a hundred David Jones).

A subsequent comparative test of those same 5, but with a more complete picture of facts (events with sex/marriages/dates/places/other matching fields filled in), and same (or other) parents plus same (or other) children (themselves accompanied by a more complete picture of events with dates/places/other matching fields filled in) and the scores go up or down based on potential duplicateness. The lower(est) point scores are the ones with a whole lotta info on one side and very little on the other OR the ones with very many differences.

Have you tried selecting just Names Sound alike (slower)? Does it maybe bring up a match? Maybe try Start with surname S instead of Trapnell.

One-off issues like this are best shared with Support, before any subsequent change that disappears it from occurring any longer… unless it’s reproducible in some consistent fashion.

1 Like

Duplicate search appears under person tools, perhaps illogically headed Merge Duplicates.

I am more interested in the general issue of how to improve the search functionality than whether there was a specifical problem in my search for Elizabeth Joyce Trapnell which is why I prefer to discuss it here rather than go to support.

In general, I think that an approach using scores (like FamilySearch’s search functions) can work much better than alternative approaches (like Ancestry’s search for example.) Much depends on what criteria are considered and how the scores are awarded, although one of the advantages of score systems is that you can play with them until you get the best results. However there is also an important question of how candidates are selected before calculating scores. In a database like mine with more than 67,000 people there are 4.5 billion possible pairs. The system needs an initial filter and I think that tick boxes in RM’s duplicate search are used in this stage.

For example, if I select to match on names spelled the same, then there is no match at all for Elizabeth Joyce Trapnell and Elizabeth J Trapnell. (Even choosing to match blank given names doesn’t help; they only match if you choose ‘sounds like.’) I would have been inclined to have a wider net, matching those with the same first and surnames, and then adjusting points for middle name matches, mismatches or various combinations with initials.

When it comes to scores, two 'David Smith’s get the same score as two 'Elizabeth Joyce Trapnell’s. I would have given lots of extra points for the matched middle name and the rare surname. When matching birth dates, things work sort of as you would have expected, although with less difference in points than I would have used; an exact match of dates get more points than a match of a detailed date to the same year with no month or day; two exact but different dates a few months apart get fewer points, but not much fewer.

If I put two David Smiths under the same parents, then they get extra points for having specific birth and death dates both of which are fairly close to each other, even if they are specific and specifically different. To my mind points should be deducted for having different dates. And no points are deducted for one having a death date before the other’s birth; I think that lots of points should be deducted as these strongly look like different people.

If I set up my two David Smiths, one of whom died before the other was born, and my two Elizabeth Joyce Trapnells, one of whom has a birth and the other one who has a marriage 38 years later, and allow matches on names only (ie no requirement to match birth or death dates) then I get two matches, the Davids and the Elizabeth Joyces. The Davids, whose data are to my mind inconsistent enough to rule them out as matches, get 32 points, and the two Elizabeth Joyces, whose data to my mind makes them an obvious duplicate, get 12. This buries matches like this at the bottom of the many thousands in my tree. I can’t think of any way to adjust the available criteria to bring them to the top, which is why I think that the tool could be significantly improved. However, I recognise that it is not an easy problem.

maybe a result grid view with sortable cols could be one approach not sure – I really have not used most of the duplicate tools in RM (for some of the reason described in this thread) its not ideal for many scenarios it seems

Yes, I misspoke there, apologies.
And I have, generally, always preceded using those facilities by publishing a Duplicate List report to eyeball first… where the separated groupings seem easily human-readable for a starting point.

I have been wondering about writing some comments to conclude this thread and have been prompted to do so by discovering another duplicate in my tree; William H Simpson who married a cousin of mine in Cheshire in 1935 was the cousin of mine William Haughton Simpson born in Cheshire in 1908. The duplicate is obvious when you see it, but it is outside the scope of RM’s current duplicate search.

Anyway, my comments

  • The duplicate search function (person, tools, merge, merge duplicates) works in two stages. The first stage uses firm criteria to choose the pairs of potential matches which are then scored in the second stage.
  • The duplicate list (publish, duplicate list) uses the same criteria to produce the same pairs of potential matches but does not add the scores. The advantage of the duplicate list is that you can export the results, although unfortunately it does not have much detail (eg it does not contain the spouses or parents of the people concerned which are included in the search function)and is not in the most useful format.
  • For both functions, you are first presented screens which, although different in format, contain the same options for you to tick and date ranges to choose; it is worth understanding how these work to get the best results. For example, my duplicate pair (William H Simpson and William Haughton Simpson) would only have appeared if I had picked the option ‘sounds like’, which would also have picked many other pairs whose names were more significantly different. Note that the options to compare blank names, dates or places will include pairs whose other details match but for whom the relevant value is present in one record and absent in the other; even if you choose not to compare blanks you will get cases in which both values are blanks, and in my tree there are lots of these.
  • Although I have found duplicates with the tool in the past, and think that it could be improved significantly, I don’t think that I have made the best possible use of it; it is worth running it several times with different criteria. For example you could choose to match death places, match death dates to the same year but also to match blank birth dates and places, and vice versa. As noted above, matching on sounds like can add useful results, and it must also be true that matching on blank surnames could help match duplicates of married women with their single selves. I have not previously used these options.

I think that there are some easy ways in which the current criteria (for selecting candidate matches) and scoring (for ranking them) could be improved. On the criteria

  • I would include options only to select those whose birth or death dates match exactly, not just to the same year. This can be a very useful way to narrow a search down, and it can also be useful in finding duplicated events on otherwise different people, for example if the same death record has been attached to two different Alan Watsons.
  • I would include another option for names, in which first name, family name and initials match, so that my example of William H Simpson and William Haughton Simpson would match. There might be merit in allowing more options than this; for example it is always useful to match Ann with Anne or similar, but the current ‘sounds like’ also matches surnames ‘Bell’ with ‘Bewley,’ ‘Stewart’ with ‘Stordy’ and other such things which seem clearly different to me. Something in the middle would be good.
  • I would allow an option to exclude blank/blank matches rather than treating them like value/value matches, or even do this by default. When I run the tool, for example matching on birth places and with birth dates matching to the same year but not matching death details, the vast majority of my matches are those for whom there is no birth event present. This is not terribly useful.
  • I might allow matching blank surnames only for women, although perhaps this would be over-engineering. However, without some limitation (notably excluding blank/blank matches) the option to match blank surnames is unusable. If I put everything else to match as tightly as possible (birth places, death places, birth dates same year, death dates same year, names spelled the same but blank surnames allowed, then I get 35,000 potentially duplicated pairs.

On the scoring

  • I would add extra points for matching relatively rare names. (I assume that this data is available somewhere.)
  • I would subtract a lot of points for inconsistent details, notably if one person in the pair had died before the other was born, even if both events had been in the same year, but also for people whose births or deaths were close to each other but different. For example, I would give most points to two exactly matching deaths, eg both 31 Oct 1960 and quite a few to one as 31 Oct 1960 with the other as abt Oct 1960 or abt 1960, but many fewer to one as 31 Oct 1960 and the other as 5 Nov 1960. I would also subtract points for inconsistent parents, giving more points for parents present in one and missing in the other.
  • If blank/blank matches are still included, I would give them very few points.
  • I would compare birth, death and marriage dates between those in a pair. For example, if one has a birth date but not a death date and the other the other way around, are the two compatible or did one die before the other was born or was one born 300 years before the other died? Similarly for marriage dates. Ideally, one would do something for places too, but this would obviously be much more difficult to achieve in practice.

PS. I have not checked how place matching or place scoring works. Do the two places need to be identical? For example would Hume, Manchester, Lancashire, England match with Hume, Lancashire, England or Hume, Manchester, Lancashire, England, United Kingdom? I don’t know.