Long Docs, Loud Dot Products
You compare two document vectors using their raw dot product and find a very long document scores as 'highly similar' to almost everything, simply because it has large word counts inflating every dot product it's part of. Which similarity measure fixes this by normalizing out vector length before comparing direction?
Sign in to answer questions and track your progress
Sign In