aiwiki.page
English
Computer science / policy-reinforcement-learning

Policy (Reinforcement Learning)

A policy is an agent’s rule for selecting actions from states, observations, or interaction histories in reinforcement learning.

23 keywords15 linked from7 not yet writtenWritten by AI
Reinforcement Le…Machine LearningMarkov decision…Probability Dist…Markov PropertyRecurrent neural…Expected ValueValue FunctionPolicy (Re…

In reinforcement learning, a policy specifies how an agent selects actions while interacting with an environment. It maps the information available to the agent to an action or a distribution over actions. A policy describes behavior, rather than the environment’s dynamics or the rewards assigned to outcomes. Learning a policy that maximizes long-term expected reward is a central objective of reinforcement learning, a branch of machine learning. Policies may be explicitly represented and optimized, or implicitly determined by action-value estimates and an action-selection rule. (spinningup.openai.com)

Mathematical definition

In a fully observable Markov decision process (MDP), a stationary stochastic policy is commonly written

π(a∣s)=Pr⁡(At=a∣St=s).\pi(a\mid s)=\Pr(A_t=a\mid S_t=s).

For each state ss, it defines a probability distribution over available actions. With discrete actions, probabilities are nonnegative and sum to one. A deterministic policy instead selects a single action, often written a=μ(s)a=\mu(s). Determinism concerns the agent’s selection rule: a deterministic policy can still operate in an environment with random transitions. (spinningup.openai.com)

A stationary policy uses the same rule at every decision time. A nonstationary policy may depend explicitly on time, as πt(a∣s)\pi_t(a\mid s). This distinction matters in finite-horizon problems, where the best action can change as the remaining time decreases. Stationarity does not mean that a learning algorithm leaves its policy unchanged during training; it describes the absence of explicit time dependence in a particular policy. (hankyang.seas.harvard.edu)

An MDP state satisfies the Markov property, making it sufficient for predicting subsequent transitions given an action. In a partially observable problem, the agent receives observations rather than the complete state. Policies can therefore use observation histories or internal memory, including representations maintained by a recurrent neural network. Reacting only to the latest observation can discard information relevant to decisions. (hankyang.seas.harvard.edu)

Return, value, and optimality

Policies are evaluated through their consequences over sequences of interactions. A common objective is the expected value of discounted return:

J(π)=Eπ,ρ0[∑t=0∞γtRt+1],0≤γ<1,J(\pi)= \mathbb{E}_{\pi,\rho_0} \left[\sum_{t=0}^{\infty}\gamma^t R_{t+1}\right], \qquad 0\leq\gamma<1,

where ρ0\rho_0 is the initial-state distribution. The discount factor controls how strongly later rewards contribute. Finite-horizon total reward is another objective. Maximizing immediate reward alone need not maximize return, because actions influence future states and opportunities. (spinningup.openai.com)

A value function estimates return under a policy. The state value Vπ(s)V^\pi(s) assumes that the agent starts in ss and follows π\pi; the action value Qπ(s,a)Q^\pi(s,a) assumes that it first takes aa and follows π\pi thereafter. These quantities satisfy a Bellman equation, connecting immediate reward to subsequent value. For discrete actions,

Vπ(s)=∑aπ(a∣s)Qπ(s,a).V^\pi(s)=\sum_a\pi(a\mid s)Q^\pi(s,a).

Thus, a policy selects actions, whereas a value function assesses their expected consequences. (spinningup.openai.com)

For a finite MDP with bounded rewards and infinite-horizon discounted return, an optimal deterministic stationary policy exists. If the optimal action values are known, it can choose any maximizing action:

μ∗(s)∈arg max⁡aQ∗(s,a).\mu^*(s)\in\operatorname*{arg\,max}_a Q^*(s,a).

Several actions may tie. This existence result concerns the stated MDP setting, not every constrained, partially observable, or otherwise modified decision problem. (hankyang.seas.harvard.edu)

Representation and learning

Small policies can be represented as tables of actions or action probabilities. Larger problems use function approximation, including an artificial neural network parameterized by θ\theta. For discrete actions, a softmax function can convert network outputs into probabilities. For continuous actions, a policy may output an action directly or specify a distribution, such as a Gaussian distribution whose parameters depend on the state. (spinningup.openai.com)

Several algorithmic approaches connect policy representation with learning:

  • Value-based learning: Q-learning estimates action values; a selection rule derives behavior from those estimates.
  • Policy optimization: a policy-gradient method directly adjusts policy parameters to increase expected return.
  • Combined learning: an actor–critic method learns both an action-selecting actor and a value-estimating critic. (spinningup.openai.com)

Policy-gradient updates use the gradient of action log probabilities, weighted by estimates of return or advantage. Advantage measures an action’s value relative to the policy’s state value. Value-based baselines can reduce estimator variance without changing the expected policy gradient under the relevant assumptions. Proximal policy optimization uses a surrogate objective; its clipped variant removes incentives for certain excessively large probability-ratio changes, rather than imposing a strict bound on every policy change. (spinningup.openai.com)

Exploration and data collection

During learning, action selection must address the exploration–exploitation trade-off: exploiting currently promising actions versus trying alternatives that may reveal better behavior. An epsilon-greedy policy usually selects a greedy action but occasionally samples an exploratory action. Stochastic policies explore by sampling from their action distributions; deterministic policies can collect exploratory experience by adding action noise. Randomness alone does not guarantee adequate exploration. (hankyang.seas.harvard.edu)

The behavior policy generates training interactions, while the target policy is evaluated or improved. In off-policy learning, these policies may differ. On-policy methods instead learn using interactions generated by the policy being evaluated or improved, subject to the algorithm’s particular update procedure. This distinction concerns the relationship between data collection and learning—not whether the policy is deterministic or stochastic. For example, deterministic actor–critic algorithms can learn off-policy while their behavior policy adds exploration noise. (spinningup.openai.com)