1. Function Approximation
A neural net maps state → Q-value for every action. Generalizes to unseen states
instead of needing a lookup entry for each one.
2. Experience Replay
Store transitions (s, a, r, s′) in a buffer; train on random minibatches.
Breaks temporal correlation and reuses data efficiently.
3. Target Network
A periodically-frozen copy of the Q-net supplies the TD target. Fixing it stops the
target from moving every step, which would otherwise diverge.
4. ε-Greedy Exploration
With probability ε act randomly, else greedily. ε decays over time —
explore early, exploit once Q-estimates sharpen.