Is Latent Semantic Indexing (LSI) a Statistical Classification algorithm? Why or why not?
Basically, I’m trying to figure out why the Wikipedia page for Statistical Classification does not mention LSI. I’m just getting into this stuff and I’m trying to see how all the different approaches for classifying something relate to one another.
No, they’re not quite the same. Statistical classification is intended to separate items into categories as cleanly as possible — to make a clean decision about whether item X is more like the items in group A or group B, for example.
LSI is intended to show the degree to which items are similar or different and, primarily, find items that show a degree of similarity to an specified item. While this is similar, it’s not quite the same.