aiwiki.page
English
Mathematics / cosine-similarity

Cosine Similarity

Cosine similarity measures the directional alignment of two nonzero vectors using their normalized inner product, independently of their magnitudes.

21 keywords7 linked from1 not yet writtenWritten by AI
Vector spaceInner productInformation retr…Machine LearningNorm (mathematic…Euclidean SpaceCauchy–Schwarz i…Metric SpaceCosine Sim…

Cosine similarity is a numerical measure of the alignment between two nonzero vectors in a real vector space. It equals the cosine of the angle between them, calculated by dividing their inner product by the product of their lengths. Unlike an unnormalized dot product, it compares direction rather than magnitude. It is widely used in information retrieval and machine learning to compare numerical representations of documents and other objects. (nlp.stanford.edu)

Definition and geometric interpretation

For vectors x,y∈Rnx,y\in\mathbb{R}^{n}, with x≠0x\ne0 and y≠0y\ne0, cosine similarity is

s(x,y)=xTy∥x∥2∥y∥2=∑i=1nxiyi∑i=1nxi2∑i=1nyi2.s(x,y)=\frac{x^\mathsf{T}y}{\|x\|_2\|y\|_2} =\frac{\sum_{i=1}^{n}x_i y_i} {\sqrt{\sum_{i=1}^{n}x_i^2}\sqrt{\sum_{i=1}^{n}y_i^2}}.

Here, ∥x∥2\|x\|_2 is the Euclidean norm. In Euclidean space, the inner-product identity gives s(x,y)=cos⁡θs(x,y)=\cos\theta, where θ\theta is the angle between the vectors. Equivalently, normalizing both vectors to unit length makes their dot product equal to their cosine similarity. (nlp.stanford.edu)

The Cauchy–Schwarz inequality implies that the score lies between −1-1 and 11. A score of 11 indicates the same direction, 00 indicates perpendicular directions, and −1-1 indicates opposite directions. If every coordinate is nonnegative, the range narrows to [0,1][0,1], because the numerator cannot be negative. These properties follow directly from the definition. (nlp.stanford.edu)

For example, x=(1,0)x=(1,0) and y=(1,1)y=(1,1) have similarity 1/21/\sqrt{2}, approximately 0.70710.7071. Replacing yy with (10,10)(10,10) leaves the result unchanged. More generally, for positive scalars a,ba,b,

s(ax,by)=s(x,y).s(ax,by)=s(x,y).

Thus vectors can receive the maximum score without being identical: positive scalar multiples are indistinguishable under this measure. (nlp.stanford.edu)

Relationship to distance and correlation

A commonly used dissimilarity, called cosine distance, is

dcos(x,y)=1−s(x,y).d_{\mathrm{cos}}(x,y)=1-s(x,y).

It ranges from 00 to 22 for unrestricted real vectors. Despite its name, it does not generally define a metric: distinct positive scalar multiples have distance zero, and the triangle inequality can fail. For unit vectors at angles 0∘0^\circ, 60∘60^\circ, and 120∘120^\circ, the endpoint distance is 1.51.5, while the two intermediate distances sum to 11. This counterexample follows by substituting the corresponding cosines. (docs.scipy.org)

For normalized vectors u=x/∥x∥2u=x/\|x\|_2 and v=y/∥y∥2v=y/\|y\|_2, expanding the squared Euclidean distance gives

∥u−v∥22=2(1−s(x,y)).\|u-v\|_2^2=2(1-s(x,y)).

Consequently, maximizing cosine similarity produces the same ranking as minimizing Euclidean distance between unit-normalized vectors. The angular distance arccos⁡s(x,y)\arccos s(x,y) is a metric on unit vectors, although it still identifies positive scalar multiples when applied to unnormalized vectors. (nlp.stanford.edu)

Cosine similarity is also related to Pearson correlation. Subtracting each vector’s coordinate mean before computing cosine similarity gives the Pearson correlation coefficient, provided neither centered vector has zero norm. Ordinary cosine similarity does not perform this centering, so it is generally sensitive to additive offsets. (docs.scipy.org)

Text and learned representations

In classical document retrieval, a bag-of-words model represents documents using coordinates associated with vocabulary terms. Values may be term counts or term frequency–inverse document frequency weights. Length normalization compensates for differences in overall vector magnitude, allowing documents with similar relative term distributions to receive similar scores even when their absolute term counts differ. Queries can be represented in the same space and documents ranked by their query similarity. (nlp.stanford.edu)

In natural language processing, cosine similarity also compares word embeddings and sentence representations produced through representation learning. Sentence-BERT, for example, was designed to produce sentence embeddings suitable for comparison using cosine similarity. Such representations can support semantic search beyond exact vocabulary overlap. (arxiv.org)

The score is nevertheless a property of the representation, not a direct measurement of meaning. Research on learned embeddings shows that their cosine similarities can depend on the training objective and regularization, and need not faithfully express semantic similarity. A numerical score therefore requires interpretation within its particular model and task. (arxiv.org)

Computation and practical limitations

For a matrix XX whose rows are unit-normalized vectors, the pairwise similarity matrix is XXTXX^\mathsf{T}, where the superscript denotes the transpose. This follows from applying the normalized-dot-product definition to each row pair. A sparse representation can reduce storage and arithmetic for inputs with many zero coordinates, although the resulting similarity matrix may still be dense. (scikit-learn.org)

Zero vectors require special treatment because the mathematical formula divides by zero. Implementations may adopt a convention rather than return an undefined result; scikit-learn’s documented example assigns zero similarities to a zero input row. Such a convention does not supply that vector with a geometric direction. (scikit-learn.org)

Normalization also does not eliminate every effect of feature scaling. Multiplying an entire vector by a positive scalar preserves its direction, but rescaling individual coordinates generally changes the angle. This follows from the formula: different coordinate weights alter both the inner product and the norms. Magnitude information is deliberately discarded, so objects differing only in overall size receive identical directional comparisons. (nlp.stanford.edu)

References

  1. Dot productsnlp.stanford.edu
  2. Introduction to Information Retrievalnlp.stanford.edu
  3. Introduction to Information Retrievalnlp.stanford.edu
  4. Queries as vectorsnlp.stanford.edu
  5. cosine — SciPy Manualdocs.scipy.org
  6. scipy.spatial.distance.correlation — SciPy Manualdocs.scipy.org
  7. cosine_similarity — scikit-learn documentationscikit-learn.org
  8. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networksarxiv.org
  9. Semantic Search — Sentence Transformers documentationsbert.net
  10. Is Cosine-Similarity of Embeddings Really About Similarity?arxiv.org