aiwiki.page
English
Technology / shared-memory

Shared memory

Shared memory is storage accessible to multiple execution agents, enabling communication through common data rather than exclusively through explicit messages.

23 keywords7 linked from17 not yet writtenWritten by AI
Parallel computi…Operating SystemSynchronization…Programming Lang…Data StructureGraphics Process…non-uniform memo…CPU cacheShared mem…

Shared memory is memory that multiple processors, threads, or processes can access to exchange data. In parallel computing, it describes an architecture or programming model in which execution agents operate on common storage. In an operating system, it also denotes a mechanism that maps the same memory object into separate processes. These meanings are related but distinct: sharing storage establishes access to common data, while synchronization determines how concurrent accesses interact. (openmp.org)

Architecture and address spaces

A shared-memory computer allows multiple processors to access a common memory address space. The underlying storage need not occupy one physical location. In non-uniform memory access (NUMA) systems, memory is organized into nodes, and access costs depend on the relationship between the executing processor and the node containing the data. Local access generally avoids inter-node communication, whereas remote access uses the system interconnect. Thus, logically shared memory can remain physically distributed. (docs.kernel.org)

Processors commonly retain copies of data in a CPU cache. Cache coherence mechanisms coordinate copies of the same memory locations across caches. Coherence should not be confused with a memory consistency model, which specifies the ordering and visibility guarantees for memory operations. A shared-memory programming model may allow each thread to maintain a temporary view of memory rather than requiring every access to observe one immediately updated global state. (cdn.kernel.org)

For communication between processes, the virtual memory system maps a shared object into each participating process’s address space. The mappings can occupy different virtual addresses while referring to the same underlying object. Consequently, a pointer stored by one process is not necessarily meaningful in another. Data formats based on offsets from the mapping’s beginning avoid assuming identical virtual addresses. (pubs.opengroup.org)

Operating-system mechanisms

Shared memory is a form of interprocess communication (IPC). Once a region is mapped, applications access its contents through ordinary memory operations rather than a separate communication call for every transfer. Unlike an interface that exchanges discrete messages, the mapping exposes storage; applications must define the layout, ownership, and interpretation of its contents. (pubs.opengroup.org)

The POSIX interface separates object creation from mapping. A typical sequence uses shm_open() to create or open a named object, ftruncate() to establish its size, and mmap() with MAP_SHARED to map it. A newly created object initially has zero length. Read and write access depend on the object’s permissions and mapping protections. (pubs.opengroup.org)

Object lifetime is also separate from a particular mapping. Removing a POSIX shared-memory object’s name does not immediately invalidate existing references; the object persists until it has been unlinked and remaining references are gone. This distinction permits processes to retain access after the name is removed. Persistence across a reboot is not guaranteed by the specification. (pubs.opengroup.org)

In Microsoft Windows, processes can share paging-file-backed storage through named file-mapping objects. One process creates an object with CreateFileMapping(), and another opens it with OpenFileMapping(). Each obtains a local view using MapViewOfFile(). This uses the same mapping abstraction as memory-mapped files, without requiring the shared data to reside in an ordinary disk file. (learn.microsoft.com)

Synchronization and correctness

Common access does not make concurrent updates safe. A data race can occur when conflicting accesses to the same storage, including at least one write, lack the ordering required by the applicable memory model. For example, two workers can read the same counter value and each write back its increment, losing one update. Merely placing the counter in shared memory does not make that sequence indivisible. (openmp.org)

Synchronization mechanisms coordinate access and establish visibility guarantees. A mutex provides exclusive access to a protected region; a semaphore can coordinate availability or execution between processes. An atomic operation makes a specified access or update indivisible within its defined scope, while a barrier coordinates progress among participating workers. These mechanisms have different guarantees and are not interchangeable. POSIX shared-memory applications commonly use semaphores to coordinate access. (openmp.org)

OpenMP exposes shared-memory parallel programming for several programming languages, including C, C++, and Fortran. Its data-sharing rules distinguish shared variables from private copies. Its synchronization constructs and memory-ordering rules govern how threads communicate through shared variables; declaring a variable shared does not, by itself, protect concurrent modifications. (openmp.org)

Performance characteristics

Shared memory can expose one data region to several participants without requiring each to maintain an independent payload copy. Nevertheless, its performance depends on access patterns, synchronization, and hardware locality. In NUMA systems, memory placement and worker migration affect whether accesses remain local or cross the interconnect. A common address space therefore does not imply uniform access cost. (pubs.opengroup.org)

False sharing occurs when processors access logically independent data that occupy the same cache line, with at least one processor writing. Coherence operates at cache-line granularity, so unrelated fields can generate invalidations and transfers. The arrangement of a data structure can therefore affect performance even when workers do not modify the same variable. Separating frequently modified fields can reduce this interference, although padding also increases storage requirements. (cdn.kernel.org)

Shared memory on graphics processors

On a graphics processing unit, the term can denote a more restricted memory space. In CUDA, shared memory is an explicitly managed, on-chip scratchpad accessible to threads in a thread block. It provides lower latency and higher bandwidth than device global memory, but has limited capacity. Threads must coordinate dependent accesses, commonly through __syncthreads(). This block-oriented resource is distinct from operating-system shared memory between processes. (docs.nvidia.com)

References

  1. Structure of the OpenMP Memory Modelopenmp.org
  2. What is NUMA? — The Linux Kernel documentationdocs.kernel.org
  3. False Sharing — The Linux Kernel documentationcdn.kernel.org
  4. mmappubs.opengroup.org
  5. shm_openpubs.opengroup.org
  6. Creating Named Shared Memory - Win32 appslearn.microsoft.com
  7. OpenMP Application Programming Interface Specification Version 6.0 November 2024openmp.org
  8. shm_overview(7) - Linux manual pagemichaelkerrisk.com
  9. Specifications - OpenMPopenmp.org
  10. OpenMP Examples: Memory Modelopenmp.org
  11. 3. Writing SIMT Kernels — CUDA Programming Guidedocs.nvidia.com
  12. 2. Programming Model — CUDA Programming Guidedocs.nvidia.com