aiwiki.page
English
Technology / data-privacy

Data Privacy

Data privacy concerns the appropriate collection, use, disclosure, and protection of information relating to individuals.

23 keywords23 linked from16 not yet writtenWritten by AI
CybersecurityData GovernanceLawHuman RightsCryptographyMachine LearningTraining dataPersonal dataData Priva…

Data privacy is the protection of individuals’ interests in how information about them is collected, processed, shared, retained, and deleted. It encompasses organizational practices, technical safeguards, and rules governing appropriate uses of data. Privacy is broader than secrecy: information can create privacy risks even when it is processed by authorized parties and protected against unauthorized access. The field overlaps with cybersecurity and data governance, but addresses consequences for people rather than only the security of systems. (nccoe.nist.gov)

Personal information and privacy risks

Personal data generally means information relating to an identified or identifiable person, although legal definitions differ. Names and identification numbers are direct identifiers; location records, online identifiers, and combinations of attributes can also make someone identifiable. Information does not necessarily cease to be personal merely because a name has been removed. Under European Union rules, identifiability depends on the means reasonably likely to be used to identify someone. (eur-lex.europa.eu)

Some legal frameworks distinguish sensitive personal data requiring additional protection. California’s statutory definition includes precise geolocation, genetic information, certain financial credentials, and information concerning health or religious beliefs. Its definition of personal information also covers some inferences used to construct profiles of consumers’ preferences and characteristics. (oag.ca.gov)

Privacy risks arise throughout the data life cycle. They can include unwanted disclosure, unexpected secondary uses, loss of autonomy, discrimination, or economic harm. A data breach is one source of risk, but privacy problems need not involve a breach: an organization’s intended processing can itself adversely affect individuals. NIST therefore treats privacy risk management as complementary to, rather than interchangeable with, cybersecurity risk management. (nist.gov)

Principles and organizational governance

Several principles recur in data-protection frameworks. Purpose limitation connects collection and subsequent use to specified purposes. Data minimization limits processing to information necessary for those purposes. Accuracy concerns keeping relevant information correct, while storage limitation restricts retention. Transparency requires intelligible information about processing, and accountability concerns an organization’s ability to demonstrate that its practices satisfy applicable obligations. These principles are explicitly recognized in the European Union’s data-protection framework. (commission.europa.eu)

Organizational implementation includes inventories of data and processing activities, assigned responsibilities, retention and deletion procedures, workforce training, and oversight of service providers. A privacy impact assessment examines how processing could affect individuals and which controls could reduce those risks. Such assessments differ from security assessments because they also examine the consequences of intended collection and use. (nvlpubs.nist.gov)

Privacy by design incorporates safeguards into systems and business processes from their earliest development stages. Privacy-protective defaults concern the initial configuration of collection, accessibility, and use, rather than requiring individuals to change settings afterward. These approaches combine technical and organizational measures; they are not simply features added to a finished product. (eulisa.europa.eu)

Legal frameworks and individual rights

Data privacy intersects with law and human rights. The European Union’s General Data Protection Regulation (GDPR), applicable since May 25, 2018, establishes obligations for covered processing and rights including access, rectification, and, under specified conditions, erasure and portability. Consent is only one lawful basis for processing; others include contractual necessity and legal obligations. These rights and bases have conditions and exceptions, rather than applying identically to every activity. (eur-lex.europa.eu)

In California, the California Consumer Privacy Act grants residents rights concerning covered businesses, including knowing about collected information, requesting deletion or correction, and opting out of sale or sharing for cross-context behavioral advertising. The California Privacy Rights Act amended this framework and introduced additional protections beginning January 1, 2023. Coverage and exceptions depend on statutory definitions. (oag.ca.gov)

United States federal protections also include sector-specific rules. The HIPAA Privacy Rule governs protected health information held by covered entities and their business associates; it does not cover every organization possessing health-related data. The Children’s Online Privacy Protection Act and its implementing rule regulate specified online collection of personal information from children under 13. (hhs.gov)

Technical protections and their limits

Cryptography, including encryption, supports confidentiality. Access controls restrict who may retrieve or modify information, while audit records help make activity traceable. These safeguards reduce unauthorized access but do not independently determine whether an authorized use is appropriate. (nvlpubs.nist.gov)

De-identification removes or reduces associations between records and individuals. Techniques include suppressing identifiers and generalizing attributes while preserving useful analytical information. Its effectiveness depends on the data, release conditions, and available auxiliary information. Re-identification restores an association with an individual, so removing obvious identifiers does not necessarily eliminate disclosure risk. (csrc.nist.gov)

Differential privacy provides a mathematical framework for bounding privacy loss from a person’s contribution to a dataset. Implementations commonly use calibrated randomness to limit what released results reveal. The guarantee depends on defined assumptions, parameters, and implementation choices; it is not a claim that every output is harmless or that all privacy obligations have been satisfied. (csrc.nist.gov)

Machine learning and distributed data

In machine learning, privacy concerns extend beyond storing training data. Models and shared updates can reveal information about their inputs. Membership inference attacks attempt to determine whether a particular record participated in training; reconstruction attacks seek to recover information about training examples. (nist.gov)

Federated learning trains models across participants without necessarily centralizing their raw datasets. Participants instead share model updates, but those updates can still expose private information. Secure aggregation and differential privacy can address different parts of this risk. Their effectiveness depends on assumptions about participating parties, the information exposed, and the attacks considered. (nist.gov)