The F-score is a family of measures used to evaluate machine learning classifiers and information retrieval systems by combining precision and recall into one number. Its most common form, F1, is their harmonic mean, giving both components equal weight. The generalized Fβ score changes their relative emphasis. Scores normally range from 0 to 1, with higher values indicating a stronger combination of precision and recall; they are also reported as percentages. (nlp.stanford.edu)
Definition and calculation
For binary classification, one class is designated positive. A confusion matrix distinguishes true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN). Precision is the proportion of predicted positives that are correct, while recall is the proportion of actual positives that are detected:
F1 combines these quantities:
The count-based expression makes clear that true negatives do not contribute directly to the score. (scikit-learn.org)
For example, suppose a document-retrieval system returns 50 documents, of which 40 are relevant, from a collection containing 80 relevant documents. Then TP = 40, FP = 10, and FN = 40. Precision is 0.80, recall is 0.50, and substitution gives . This illustrative calculation follows the set-based definition of retrieval evaluation. (nlp.stanford.edu)
Unlike the arithmetic mean, the harmonic mean strongly reflects a low component: excellent precision cannot fully compensate for poor recall, or vice versa. When both components are positive, F1 lies between them and does not exceed their arithmetic mean. It equals 1 when there are no false positives or false negatives and at least one true positive. (nlp.stanford.edu)
The Fβ family
For a positive parameter ,
Values above 1 emphasize recall; values below 1 emphasize precision. Thus F2 emphasizes missed positives more strongly than F1, while F0.5 emphasizes false positives more strongly. As approaches zero, the expression approaches precision; as it grows without bound, it approaches recall, where these quantities are defined. The squared parameter in the formula is important: in the count-based denominator, F2 assigns FN a coefficient of four relative to FP. (scikit-learn.org)
Choosing β specifies an evaluation preference, not a universal monetary or practical cost ratio between errors. Fβ remains a normalized combination of precision and recall, rather than a direct measure of application-specific utility. (arxiv.org)
Multiclass and multilabel averaging
In multiclass classification, each class can be evaluated against all remaining classes. In multilabel classification, each label defines a separate binary problem. Several aggregation conventions produce different overall scores. (scikit-learn.org)
- Micro-averaging pools TP, FP, and FN across labels before calculating the score.
- Macro-averaging calculates each label’s F-score and takes their unweighted arithmetic mean.
- Support-weighted averaging weights each label’s score by its number of actual positive instances.
- Sample averaging, used for multilabel tasks, calculates a score for each instance’s predicted and actual label sets, then averages across instances. (scikit-learn.org)
Micro-F1 is more strongly influenced by frequent classes, whereas macro-F1 gives each class equal weight. For ordinary single-label multiclass classification, when all classes are included, micro-F1 equals accuracy: every incorrect prediction contributes one false positive and one false negative across the classes. This identity does not generally hold for multilabel tasks or evaluations restricted to selected classes. (nlp.stanford.edu)
Macro-F1 usually means the mean of class-specific F1 scores. It is not generally equal to the harmonic mean of macro-averaged precision and macro-averaged recall. Both conventions have appeared under the same name, making the precise formula consequential for comparisons. (arxiv.org)
Thresholds and evaluation design
Many models, including logistic regression, generate a probability estimate or decision score before producing a class label. A decision threshold converts that output into a positive or negative prediction. Changing the threshold changes the confusion matrix and therefore the F-score, even when the underlying model is unchanged. A default threshold need not maximize F1. (scikit-learn.org)
Threshold selection can be treated as optimization of an objective function on a validation set or through cross-validation. Using the same observations for fitting a model and tuning its threshold can cause overfitting. A separate test set supports evaluation after selection. (scikit-learn.org)
Limitations and reporting conventions
F1 can reveal failure to detect a rare positive class despite high accuracy. For example, predicting every instance negative achieves 99% accuracy when positives constitute 1% of the data, but gives positive-class F1 of zero. Nevertheless, ignoring true negatives means F1 does not describe every aspect of classification performance. (nlp.stanford.edu)
F-scores evaluate discrete predictions, not the quality of probability estimates. In retrieval, they evaluate an unordered result set and do not distinguish different rankings within that set. Precision–recall curves instead describe performance across result cutoffs. (scikit-learn.org)
If , the count-based score is undefined. Software may substitute zero, one, or a missing value; missing values may also be excluded from averages. Consequently, interpretable reporting identifies β, the positive class or included labels, the averaging convention, the threshold, and the handling of undefined cases. (scikit-learn.org)