States & Actions (S, A)
The situations the agent can be in, and the choices available in each.
Transition P(s′|s,a)
The probability of landing in state s′ after taking action a in state s.
Reward R & Discount γ
Immediate feedback for each step; γ (0–1) weights future vs. present reward.
Policy π & Value V
π maps states→actions; V(s) is the expected total reward from state s onward.