데이터셋 구성
사용할 데이터셋은 대출 승인 결과를 예측하는 이진 분류 문제입니다.
| 컬럼명 | 설명 |
|---|---|
| Loan_ID | 대출 고유 식별자 |
| Gender | 성별 |
| Married | 결혼 여부 |
| Dependents | 부양 가족 수 |
| Education | 학력 수준(대졸/비대졸) |
| Self_Employed | 자영업 여부 |
| ApplicantIncome | 신청자 소득 |
| CoapplicantIncome | 공동 신청자 소득 |
| LoanAmount | 대출 금액(천 단위) |
| Loan_Amount_Term | 대출 기간(월) |
| Credit_History | 신용 이력 충족 여부 |
| Property_Area | 부동산 위치(도시/준도시/농촌) |
| Loan_Status | 대출 승인 상태(타겟 변수) |
학습 데이터는 614개 샘플과 13개 특성을 포함하고, 테스트 데이터는 367개 샘플에 타겟 변수가 제외된 12개 특성으로 구성됩니다.
1. 데이터 불러오기
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# 데이터 로드
raw_train = pd.read_csv("data/train_LoanPrediction.csv")
raw_test = pd.read_csv("data/test_LoanPrediction.csv")
2. 데이터 전처리
클래스 불균형 처리
타겟 클래스의 분포가 고르지 않을 경우 모델이 편향될 수 있으므로 샘플링을 통해 균형을 맞춥니다.
approved = raw_train.Loan_Status == 'Y'
rejected = raw_train.Loan_Status == 'N'
approved_samples = raw_train[approved].sample(n=192, random_state=42)
rejected_samples = raw_train[rejected]
balanced_train = pd.concat([approved_samples, rejected_samples])
인코딩 변환
범주형 데이터를 수치형으로 변환하여 알고리즘이 처리할 수 있도록 합니다.
from sklearn.preprocessing import LabelEncoder
encoder = LabelEncoder()
balanced_train['Loan_Status'] = encoder.fit_transform(balanced_train['Loan_Status'])
balanced_train['Education'] = encoder.fit_transform(balanced_train['Education'])
raw_test['Education'] = encoder.transform(raw_test['Education'])
balanced_train = pd.get_dummies(balanced_train, columns=['Property_Area'], drop_first=True)
raw_test = pd.get_dummies(raw_test, columns=['Property_Area'], drop_first=True)
3. 특성 선택
selected_features = ['Education', 'Credit_History', 'Property_Area_Semiurban', 'Property_Area_Urban']
X_train = balanced_train[selected_features]
y_train = balanced_train['Loan_Status']
X_test = raw_test[selected_features]
4. 모델 학습 및 평가
여러 분류 알고리즘을 비교하고 교차 검증과 하이퍼파라미터 튜닝을 수행합니다.
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.svm import SVC
from sklearn.model_selection import StratifiedKFold, cross_val_score, GridSearchCV
# 후보 모델 정의
candidates = [
LogisticRegression(random_state=42),
RandomForestClassifier(random_state=42),
SVC(random_state=42)
]
# 5-폴드 교차 검증
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for candidate in candidates:
scores = cross_val_score(candidate, X_train, y_train, cv=skf, scoring='accuracy')
print(f"{candidate.__class__.__name__} - 평균 정확도: {scores.mean():.4f}, 표준편차: {scores.std():.4f}")
# 최적 파라미터 탐색
param_candidates = {'solver': ['newton-cg', 'lbfgs', 'liblinear', 'saga']}
optimizer = GridSearchCV(
LogisticRegression(random_state=42, max_iter=1000),
param_candidates,
scoring="accuracy",
cv=skf,
refit=True
)
optimizer.fit(X_train, y_train)
5. 특성 중요도 해석
final_model = optimizer.best_estimator_
importance_df = pd.DataFrame({
'feature': selected_features,
'weight': final_model.coef_[0]
}).sort_values('weight', ascending=True)
importance_df.plot.barh(x='feature', y='weight', legend=False, figsize=(10, 6))
plt.xlabel('계수 값')
plt.title('로지스틱 회귀 특성 가중치')
plt.tight_layout()
plt.show()
6. 모델 배포 및 예측
import joblib
# 모델 영속화
model_path = 'model/loan_approval_model.pkl'
joblib.dump(optimizer, model_path)
# 모델 복원
deployed_model = joblib.load(model_path)
# 신규 샘플 예측
new_application = np.array([[0, 1, 0, 0]])
prediction = deployed_model.predict(new_application)
probability = deployed_model.predict_proba(new_application)
print(f"예측 결과: {'승인' if prediction[0] == 1 else '거절'}")
print(f"승인 확률: {probability[0][1]:.4f}")
핵심 개념 정리
분류 문제의 특성
Loan_Status가 이산적인 두 클래스(Y/N)로 구분되므로 지도학습의 분류 유형에 해당합니다. 회귀와 달리 출력값이 연속적이지 않습니다.
클래스 불균형의 영향
학습 데이터에서 한 클래스가 과도하게 많으면 모델이 다수 클래스로 편향되어 예측합니다. 이는 소수 클래스의 중요한 패턴을 놓치게 만들어 실제 성능과 검증 성능의 괴리를 발생시킵니다.
인코딩 방식 선택
순서가 있는 범주(학력 수준: 고졸-대졸-석사)에는 레이블 인코딩을, 순서가 없는 범주(지역 유형)에는 원-핫 인코딩을 적용합니다. 잘못된 인코딩은 알고리즘이 가짜 순서 관계를 학습하게 합니다.
교차 검증의 필요성
데이터를 여러 번 분할해 평가하면 특정 분할에 대한 과적합을 방지하고 모델의 일반화 성능을 신뢰할 수 있습니다. StratifiedKFold는 클래스 비율을 유지하며 분할합니다.
정확도의 한계
불균형 데이터에서는 높은 정확도가 오해를 불러올 수 있습니다. 정밀도(예측한 승인 중 실제 승인 비율), 재현율(실제 승인 중 예측 성공 비율), F1-점수(둘의 조화평균)를 함께 확인해야 합니다.
ROC 곡선 해석
임계값을 변화시켜 얻는 참긍정률과 거짓긍정률의 관계를 시각화합니다. AUC가 0.5면 무작위 예측, 1.0이면 완벽 분류를 의미합니다. 다양한 임계값에서의 성능 변화를 한눈에 파악할 수 있습니다.
과적합과 과소적합
훈련 정확도는 높으나 검증 정확도가 낮으면 과적합(모델이 훈련 데이터의 노이즈까지 학습), 둘 다 낮으면 과소적합(모델이 데이터의 패턴을 충분히 학습하지 못함)을 의심합니다. 학습 곡선을 통해 진단합니다.