Perceptron Loss, Weight Update, SGD, 그리고 SVM¶
Perceptron은 binary classification을 위한 linear classifier 임.
- Input 자체는 binary일 필요가 없으며,
- class label은 0/1 또는 -1/+1로 표현할 수 있음.
주의할 점은, Label notation에 따라 loss와 weight update의 표현도 달라진다.
1. 0/1 notation¶
Target과 prediction은 다음과 같음: $$ y,\hat y\in \{0,1\} $$
Perceptron의 weighted sum은 다음과 같음: $$ z=\mathbf{w}^{\top}\mathbf{x}+b $$
Prediction은 step activation을 통해 결정됨: $$ \hat y= \begin{cases} 1, & z\geq0 \\ 0, & z<0 \end{cases} $$
Perceptron loss는 다음과 같이 표현할 수 있음: $$ \begin{aligned}L_{\mathrm{Perceptron}} &= \max\left(0,-(2y-1)z\right) \\ &=\max\left(0,-tz)\right)\end{aligned} $$
- \(t=2y-1\)
- 즉, 정답을 맞추는 경우에 \(tz\)는 양수가 됨.
Target별 표현은 다음과 같음: $$ L_{\mathrm{Perceptron}} = \begin{cases} \max(0,-z), & y=1 \\ \max(0,z), & y=0 \end{cases} $$
- correct classification: loss가 0.
- misclassification: decision boundary의 잘못된 쪽으로 이동할수록 loss가 증가.
- misclassified sample에 대해서만 penalty가 발생.
Weight update는 다음과 같음: $$ \boxed{ w_i^{(\mathrm{next})} = w_i+\eta(y-\hat y)x_i } $$
각 경우의 update는 다음과 같음:
- correct classification: no update.
- false negative: positive direction으로 update.
- false positive: negative direction으로 update.
False negative의 경우: $$ y=1,\qquad \hat y=0 $$
Weight update는 다음과 같음: $$ w_i^{(\mathrm{next})} = w_i+\eta x_i $$
False positive의 경우: $$ y=0,\qquad \hat y=1 $$
Weight update는 다음과 같음: $$ w_i^{(\mathrm{next})} = w_i-\eta x_i $$
Bias update는 다음과 같음: $$ \boxed{ b^{(\mathrm{next})} = b+\eta(y-\hat y) } $$
즉, 0/1 notation에서는 prediction error를 직접 이용하는 형태로 update를 표현할 수 있음.
2. -1/+1 notation¶
Perceptron을 수학적으로 표현할 때는 -1/+1 notation이 더 간단함.
Target은 다음과 같음: $$ t\in \{-1,+1\} $$
앞서 살펴본 \(t=2y-1\)의 관계는 \(y\in \{0,1\}\)인 경우임. 여기서 라벨값이 \(t\)로 처리한 경우.
Weighted sum은 동일함: $$ z=\mathbf{w}^{\top}\mathbf{x}+b $$
Prediction은 sign function으로 표현됨: $$ \hat t=\operatorname{sign}(z) $$
Classification의 correctness는 다음과 같은 signed margin 형태로 표현할 수 있음: $$ tz $$
- positive: correct classification.
- negative: misclassification.
- magnitude: decision boundary로부터 떨어진 정도.
Perceptron loss는 다음과 같음: $$ \boxed{ L_{\mathrm{Perceptron}} = \max(0,-tz) } $$
- 즉, correct classification인 경우, \(-tz\)는 무조건 negative이므로, \(L\)은 0이 됨.
- 틀린 경우엔 \(tz<0\)이고, \(-tz\)는 positive이므로 \(L>0\)이 됨.
Correct classification에서는 loss가 0이고, misclassification에서만 loss가 발생함.
Misclassified sample에 대한 weight update는 다음과 같음: $$ \boxed{ \mathbf{w}^{(\mathrm{next})} = \mathbf{w}+\eta t\mathbf{x} } $$
Bias update는 다음과 같음: $$ \boxed{ b^{(\mathrm{next})} = b+\eta t } $$
실제로 각 경우를 살펴보자.
Positive sample의 misclassification: $$ t=+1 $$
Positive sample에서 misclassification이라면 Weight update는 다음과 같음: $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w}+\eta\mathbf{x} $$
Negative sample의 misclassification: $$ t=-1 $$
Negative sample에서 misclassification이라면 Weight update는 다음과 같음: $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w}-\eta\mathbf{x} $$
따라서 0/1 notation과 -1/+1 notation은 표현 방식만 다를 뿐 동일한 Perceptron learning rule을 나타냄.
두 notation의 mapping은 다음과 같음: $$ t=2y-1 $$
따라서 다음과 같이 대응됨: $$ y=0 \rightarrow t=-1 \\ y=1 \rightarrow t=+1 $$
3. Perceptron과 Stochastic Gradient Descent¶
Perceptron learning은
- sample 단위로 parameter를 update한다 는 점에서
- Stochastic Gradient Descent와 직접 연결할 수 있음.
일반적인 single-sample SGD update는 다음과 같음: $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w} - \eta \nabla_{\mathbf w}L_i $$
위의 graident를 구한 실제 업데이트는 다음과 같음: $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w} - \eta (\hat y_i - y_i) \mathbf{x} $$
앞서 살펴본 0/1 notation에서의 classical Perceptron update는 다음과 같음: $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w} + \eta(y_i-\hat y_i)\mathbf{x} $$
결국, SGD와 Perceptron 모두 동일한 식임.
Perceptron loss를 signed label notation으로 표현 하면 다음과 같음: $$ L_i = \max(0,-t_i z_i) $$
Misclassified sample (\(L_i>0\) 인 경우)에서 i번째 샘플로 구해진 gradient는 다음과 같음: $$ \nabla_{\mathbf w}L_i = -t_i\mathbf{x}_i $$
Misclassified sample (\(L_i>0\) 인 경우)에서 i번째 샘플에서의 SGD update는 다음과 같음: $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w} - \eta (-t_i\mathbf{x}_i) \quad \text{ if } t\hat y < 0 $$
- \(L_i>0\)인 경우에만 update가 일어난다는 것임.
- \(\hat y_i = \mathbf{w}^\top \mathbf{x}_i + b\)
이 업데이트 및 loss는 classical Perceptron의 -1/+1 notation update와 동일함. $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w}+\eta t_i\mathbf{x}_i $$
결국 두 경우 모두 label coding만 다를 뿐 같은 learning rule을 표현함.
참고로,
scikit-learn의 Perceptron은
다음 설정의 SGDClassifier와 equivalent하게 볼 수 있음:
각 설정의 의미는 다음과 같음.
loss="perceptron": Perceptron loss 사용.learning_rate="constant": learning rate를 iteration에 따라 감소시키지 않음.eta0=1: constant learning rate를 1로 설정.penalty=None: regularization을 사용하지 않음.
Learning rate는 다음과 같음: $$ \eta=1 $$
Misclassified sample에 대한 update 는 다음과 같음: $$ \mathbf{w}^{(\mathrm{next})} = \mathbf{w} + t_i\mathbf{x}_i $$
즉, SGDClassifier에 Perceptron loss, constant learning rate, unit learning rate, no regularization을 적용하면 classical Perceptron update와 연결됨.
4. Perceptron Loss와 Hinge Loss¶
Perceptron loss와 Hinge loss의 관계는 -1/+1 notation에서 가장 명확 하게 나타남.
Perceptron loss는 다음과 같음: $$ L_{\mathrm{Perceptron}} = \max(0,-tz) $$
Hinge loss는 다음과 같음: $$ \boxed{ L_{\mathrm{hinge}} = \max(0,1-tz) } $$
- 위 그래프의 x축이 \(tz\)임.
두 loss의 핵심적인 차이는 margin 의 유무(1이 더해짐)에 있음.
Perceptron에서는 다음과 같이 correct side에 위치하면 loss가 0: $$ tz>0 $$
Hinge loss에서는 일정한 margin까지 확보(1이상이 되어야)해야 loss가 0: $$ tz\geq1 $$
따라서 sample은 다음과 같이 구분할 수 있음.
- misclassification: Perceptron loss와 Hinge loss 모두 penalty 발생.
- correct classification with small margin: Perceptron loss는 0, Hinge loss는 penalty 발생.
- correct classification with sufficient margin: 두 loss 모두 0.
5. Hinge Loss와 SVM¶
Hinge loss를 사용하는 대표적인 linear classifier가 Support Vector Machine, SVM임.
Linear SVM의 objective function은 다음과 같음:
- \(t_i\)는 i번째 -1/1 notation 라벨값.
첫 번째 term은 regularization term (Hard SVM에서 margin maximization. 단, margin constraint를 만족해야함):
두 번째 term은 Hinge loss (Soft SVM 의 slack variable 을 unconstrained objective로 도입한 결과):
SVM은 단순한 correct classification뿐 아니라 large margin을 추구함.
Perceptron과 SVM의 차이는 loss의 threshold에서 명확하게 나타남.
Perceptron:
- correct side에 있으면 loss가 0.
- misclassified sample에 대해서만 penalty.
- margin의 크기는 직접적으로 요구하지 않음.
SVM:
- correct classification만으로는 충분하지 않음.
- margin 내부에 위치한 sample에도 penalty.
- regularization과 함께 large-margin decision boundary를 추구함.
6. Perceptron, SGD, SVM의 연결¶
전체 관계를 정리하면 다음과 같음.
Perceptron:
- misclassification 중심.
- sample 단위 update.
- SGD 형태로 optimize 가능.
- classical Perceptron은 constant learning rate와 no regularization의 특수한 형태로 볼 수 있음.
Classical Perceptron과 equivalent한 SGDClassifier 설정:
Hinge loss:
- misclassification뿐 아니라 insufficient margin에도 penalty.
- correct classification 이후에도 margin 확보를 계속 요구.
SVM:
- Hinge loss를 기반으로 margin을 확보.
- regularization을 통해 large-margin decision boundary를 추구(margin maximization).
- SGD를 이용해서도 optimize 가능.
결국 SGDClassifier는 Perceptron과 linear SVM을 동일한 stochastic optimization framework 안에서 이해하는 데 유용함.
loss="perceptron": Perceptron으로 연결.loss="hinge": linear SVM의 Hinge loss로 연결.- learning rate와 regularization 설정에 따라 실제 optimization behavior가 달라짐.