Sign Up

Sign Up to our social questions and Answers Engine to ask questions, answer people’s questions, and connect with other people.

Have an account? Sign In

Have an account? Sign In Now

Sign In

Login to our social questions & Answers Engine to ask questions answer people’s questions & connect with other people.

Sign Up Here

Forgot Password?

Don't have account, Sign Up Here

Forgot Password

Lost your password? Please enter your email address. You will receive a link and will create a new password via email.

Have an account? Sign In Now

You must login to ask a question.

Forgot Password?

Need An Account, Sign Up Here

Please briefly explain why you feel this question should be reported.

Please briefly explain why you feel this answer should be reported.

Please briefly explain why you feel this user should be reported.

Sign InSign Up

The Archive Base

The Archive Base Logo The Archive Base Logo

The Archive Base Navigation

  • SEARCH
  • Home
  • About Us
  • Blog
  • Contact Us
Search
Ask A Question

Mobile menu

Close
Ask a Question
  • Home
  • Add group
  • Groups page
  • Feed
  • User Profile
  • Communities
  • Questions
    • New Questions
    • Trending Questions
    • Must read Questions
    • Hot Questions
  • Polls
  • Tags
  • Badges
  • Buy Points
  • Users
  • Help
  • Buy Theme
  • SEARCH
Home/ Questions/Q 1080743
In Process

The Archive Base Latest Questions

Editorial Team
  • 0
Editorial Team
Asked: May 16, 20262026-05-16T22:04:41+00:00 2026-05-16T22:04:41+00:00

I’m developing a shopping comparison website, and the project is in a very advanced

  • 0

I’m developing a shopping comparison website, and the project is in a very advanced stage. We index 50 million products daily using merchant feeds from various affiliate networks. Most of the problems I had is already solved, including the majority of the performance bottlenecks.

What is my problem: Please, first of all, we are using apache solr with drupal BUT, this problem IS NOT specific to drupal or solr, if you do not have knowledge of them, it doesn’t matter.

We receive product feeds from over 2000 different merchants, and those feeds are a mess. They have no specific pattern, each merchant send the feeds the way they want. We already solved many problems regarding this, but one remains. Normalizing the taxonomy terms for the faceted browsing functionality.

Suppose that I have a “Narrow by Brands” browsing facet on my website. Now suppose that 100 merchants offer products from Microsoft. Now comes the problem. Some merchants put in the “Brands” column of the data feed “Microsoft”, others “Microsoft, Inc.”, others “Microsoft Corporation” others “Products from Microsoft”, etc… there is no specific pattern between merchants and worst, some individual merchants are so sloppy that they have different strings for the same brand IN THE SAME DATA FEED.

We do not want all those different brands appearing in the navigation. We have a manual solution to the problem where we manually map the imported brands to the “good” brands table (“Microsoft Corporation” -> “Microsoft”, “Products from Microsoft” -> “Microsoft”, etc..). We have something like 10,000 brands in the database and this is doable. The problem is when it comes with bigger things like “Authors”. When we import books into the system, there are over 800,000 authors and we have the same problem and this is not doable by hand mapping. The problem is the same: “Tom Mike Apostol”, “Tom M. Apostol”, “Apostol, Tom M.”, etc…

Does anybody know a good way to automatically solve this problem with an acceptable degree of accuracy (85%-95% accuracy)?

Thanks you for the help!

  • 1 1 Answer
  • 0 Views
  • 0 Followers
  • 0
Share
  • Facebook
  • Report

Leave an answer
Cancel reply

You must login to add an answer.

Forgot Password?

Need An Account, Sign Up Here

1 Answer

  • Voted
  • Oldest
  • Recent
  • Random
  1. Editorial Team
    Editorial Team
    2026-05-16T22:04:41+00:00Added an answer on May 16, 2026 at 10:04 pm

    Some idea that comes to my mind, altough it’s just a loose thought:

    1. Convert names to initials (in your example: TMA). Treat ‘-‘ as spaces, so fe. Antoine de Saint-Exupéry would be ADSE. Problem here is how to treat “,”, altough, it’s common usage is to have surname before forename, so just swapping positions should work (so A,TM would be TM,A, get rid of comma – TMA).
    2. Filters authors in database by those initials
    3. For each intitial, if you have whole name (Tom, Apostol) check if it match, otherwise (M.) consider it a match automatically.
    4. If you want some tolerance, you can compare names with Levenshtein distance and tolerate some differences (here you have Oracle implementation)
    5. Names that match you treat as the same authors, to find the whole name, for each initial (T, M, A) you look up your filtered authors (after step 2) and try to find one without just initial (M.) but with whole name (Mike), if you can’t find one, use initial. Therefore, each of examples you gave would be converted to the same value, which would be full name (Tom Mike Apostol).

    Things that are worth to think about:
    Include mappings for name synonyms (would be more likely maximally hundred of records, like Thomas <-> Tom
    This way is crucial to have valid initials (no M instead of N etc.).

    edit: I’ve coded such thing some time ago, when I had to identify a person by it’s signature, ignoring scanning problems, people sometimes sign by Name S. Surname, or N.S. or just by Name Surname (which is another thing maybe you should consider in the solution, to allow the algorithm to ignore second name, altough in your situation it would be rather rare to ommit someone’s second name I guess).

    • 0
    • Reply
    • Share
      Share
      • Share on Facebook
      • Share on Twitter
      • Share on LinkedIn
      • Share on WhatsApp
      • Report

Sidebar

Related Questions

No related questions found

Explore

  • Home
  • Add group
  • Groups page
  • Communities
  • Questions
    • New Questions
    • Trending Questions
    • Must read Questions
    • Hot Questions
  • Polls
  • Tags
  • Badges
  • Users
  • Help
  • SEARCH

Footer

© 2021 The Archive Base. All Rights Reserved
With Love by The Archive Base

Insert/edit link

Enter the destination URL

Or link to existing content

    No search term specified. Showing recent items. Search or use up and down arrow keys to select an item.