A database is an organized collection of related data, usually stored electronically on a computer and maintained for retrieval, modification, and controlled use. The collection is distinct from the database management system (DBMS), the software that manages it, although “database” sometimes refers to the combined system. A DBMS provides mechanisms for defining data, executing queries, coordinating access, and protecting stored information. Databases underpin applications ranging from inventory systems to scientific repositories. (oracle.com)
Historical development
Early electronic databases developed during the 1960s, when organizations needed to manage growing collections of records. Hierarchical systems organized records into tree-like structures, while network models supported more complex navigational relationships. Applications often depended closely on predefined access paths, making structural changes difficult. (oracle.com)
In 1970, Edgar F. Codd, a researcher at IBM, published “A Relational Model of Data for Large Shared Data Banks.” His approach represented information through relations and separated its logical organization from physical storage details. IBM’s subsequent System R project helped demonstrate practical relational processing. Donald Chamberlin and Raymond Boyce developed the language that became SQL, enabling users to express requests without specifying every retrieval step. (ibm.com)
Data models and organization
A data model defines how information is represented and related. In a relational database, data is organized into tables containing rows and columns. A row represents a record; columns describe its attributes. The mathematical relational model treats relations as sets of tuples and provides operations for deriving new relations from existing ones. (docs.oracle.com)
A database schema describes database structures, such as tables, attributes, and constraints. A primary key uniquely identifies each row and excludes null values. A foreign key requires specified values to match a referenced key, maintaining referential integrity. For example, an order can reference a customer record rather than duplicate that customer’s details. Other constraints enforce uniqueness, required values, or conditions such as a nonnegative quantity. (postgresql.org)
Database normalization organizes relational tables according to dependencies among attributes. Its purpose is to reduce avoidable redundancy and anomalies when records are inserted, updated, or deleted. Separating customer details from individual orders, for instance, avoids storing an address repeatedly. Deliberate denormalization may instead duplicate or combine information to suit particular access patterns. (ibm.com)
NoSQL describes a diverse family of systems using models such as documents and key–value pairs rather than exclusively relational tables. Document databases can represent nested objects and arrays. The distinction is not absolute: relational systems may support nonrelational structures, while some NoSQL products provide SQL interfaces and multi-record transactions. (wwwcmsapi.oracle.com)
Queries and physical storage
A query requests information or an operation on stored data. SQL supports defining structures, retrieving records, inserting data, updating values, and deleting rows. Retrieval can filter records, combine tables through joins, and calculate aggregates such as counts or totals. Applications submit these operations through database interfaces rather than directly manipulating storage files. (postgresql.org)
Logical tables need not correspond directly to their physical arrangement. A DBMS can use auxiliary data structures to locate records efficiently. A database index provides an access path that may avoid scanning an entire table. Indexes can accelerate selective searches and joins, but require storage and maintenance during updates; an inappropriate index can reduce overall performance. (postgresql.org)
A query optimizer selects an execution strategy for a query. Cost-based optimization compares possible plans using estimates of the resources they require. Consequently, equivalent requests can have different execution costs depending on available indexes, table sizes, and the chosen processing plan. Cost-based optimization was an important contribution of the System R project. (ibm.com)
Transactions and concurrent access
A transaction groups operations into a unit that can be committed or rolled back. In a transfer between accounts, for example, the debit and credit must succeed together rather than leave a partially completed change. Transactional processing also controls what other sessions can observe while changes are underway. (postgresql.org)
The conventional ACID properties are atomicity, meaning all-or-nothing execution; consistency, meaning preservation of declared integrity rules; isolation, meaning control over interactions between concurrent transactions; and durability, meaning persistence of acknowledged committed changes. Constraints and transaction mechanisms address different parts of these guarantees. Consistency does not establish that entered information is factually correct; it concerns the rules the system enforces. (postgresql.org)
Supporting concurrent access requires coordination between transactions. Database implementations use locking and related transaction-management mechanisms to regulate conflicting operations. Applications can explicitly delimit transaction blocks with commands such as BEGIN, COMMIT, and ROLLBACK. (postgresql.org)
Deployment, replication, and protection
Databases may run locally, on dedicated servers, or through cloud computing services. Systems using distributed computing can maintain data across multiple machines or locations. Replication maintains copies on additional servers to support availability or distribute read workloads. With asynchronous replication, copies may temporarily lag behind the primary server, and failover can lose changes that have not yet propagated. (oracle.com)
Protection combines access privileges with encryption and operational controls. A DBMS can restrict which users may read or modify particular objects. Cryptographic mechanisms can protect network communications and stored data, supporting security and data privacy. These controls operate at different layers: encrypted transport protects communication, while storage encryption addresses disclosure from storage media. (postgresql.org)
Operational databases principally support application records and updates. A data warehouse is organized for querying and analysis, often consolidating information for reporting. These are workload distinctions rather than mutually exclusive data models: both can use relational structures, while differing in organization and access patterns. (oracle.com)