Cosine similarity is a numerical measure of the alignment between two nonzero vectors in a real vector space. It equals the cosine of the angle between them, calculated by dividing their inner product by the product of their lengths. Unlike an unnormalized dot product, it compares direction rather than magnitude. It is widely used in information retrieval and machine learning to compare numerical representations of documents and other objects. (nlp.stanford.edu)
Definition and geometric interpretation
For vectors , with and , cosine similarity is
Here, is the Euclidean norm. In Euclidean space, the inner-product identity gives , where is the angle between the vectors. Equivalently, normalizing both vectors to unit length makes their dot product equal to their cosine similarity. (nlp.stanford.edu)
The Cauchy–Schwarz inequality implies that the score lies between and . A score of indicates the same direction, indicates perpendicular directions, and indicates opposite directions. If every coordinate is nonnegative, the range narrows to , because the numerator cannot be negative. These properties follow directly from the definition. (nlp.stanford.edu)
For example, and have similarity , approximately . Replacing with leaves the result unchanged. More generally, for positive scalars ,
Thus vectors can receive the maximum score without being identical: positive scalar multiples are indistinguishable under this measure. (nlp.stanford.edu)
Relationship to distance and correlation
A commonly used dissimilarity, called cosine distance, is
It ranges from to for unrestricted real vectors. Despite its name, it does not generally define a metric: distinct positive scalar multiples have distance zero, and the triangle inequality can fail. For unit vectors at angles , , and , the endpoint distance is , while the two intermediate distances sum to . This counterexample follows by substituting the corresponding cosines. (docs.scipy.org)
For normalized vectors and , expanding the squared Euclidean distance gives
Consequently, maximizing cosine similarity produces the same ranking as minimizing Euclidean distance between unit-normalized vectors. The angular distance is a metric on unit vectors, although it still identifies positive scalar multiples when applied to unnormalized vectors. (nlp.stanford.edu)
Cosine similarity is also related to Pearson correlation. Subtracting each vector’s coordinate mean before computing cosine similarity gives the Pearson correlation coefficient, provided neither centered vector has zero norm. Ordinary cosine similarity does not perform this centering, so it is generally sensitive to additive offsets. (docs.scipy.org)
Text and learned representations
In classical document retrieval, a bag-of-words model represents documents using coordinates associated with vocabulary terms. Values may be term counts or term frequency–inverse document frequency weights. Length normalization compensates for differences in overall vector magnitude, allowing documents with similar relative term distributions to receive similar scores even when their absolute term counts differ. Queries can be represented in the same space and documents ranked by their query similarity. (nlp.stanford.edu)
In natural language processing, cosine similarity also compares word embeddings and sentence representations produced through representation learning. Sentence-BERT, for example, was designed to produce sentence embeddings suitable for comparison using cosine similarity. Such representations can support semantic search beyond exact vocabulary overlap. (arxiv.org)
The score is nevertheless a property of the representation, not a direct measurement of meaning. Research on learned embeddings shows that their cosine similarities can depend on the training objective and regularization, and need not faithfully express semantic similarity. A numerical score therefore requires interpretation within its particular model and task. (arxiv.org)
Computation and practical limitations
For a matrix whose rows are unit-normalized vectors, the pairwise similarity matrix is , where the superscript denotes the transpose. This follows from applying the normalized-dot-product definition to each row pair. A sparse representation can reduce storage and arithmetic for inputs with many zero coordinates, although the resulting similarity matrix may still be dense. (scikit-learn.org)
Zero vectors require special treatment because the mathematical formula divides by zero. Implementations may adopt a convention rather than return an undefined result; scikit-learn’s documented example assigns zero similarities to a zero input row. Such a convention does not supply that vector with a geometric direction. (scikit-learn.org)
Normalization also does not eliminate every effect of feature scaling. Multiplying an entire vector by a positive scalar preserves its direction, but rescaling individual coordinates generally changes the angle. This follows from the formula: different coordinate weights alter both the inner product and the norms. Magnitude information is deliberately discarded, so objects differing only in overall size receive identical directional comparisons. (nlp.stanford.edu)
References
- Dot productsnlp.stanford.edu
- Introduction to Information Retrievalnlp.stanford.edu
- Introduction to Information Retrievalnlp.stanford.edu
- Queries as vectorsnlp.stanford.edu
- cosine — SciPy Manualdocs.scipy.org
- scipy.spatial.distance.correlation — SciPy Manualdocs.scipy.org
- cosine_similarity — scikit-learn documentationscikit-learn.org
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networksarxiv.org
- Semantic Search — Sentence Transformers documentationsbert.net
- Is Cosine-Similarity of Embeddings Really About Similarity?arxiv.org