Python 기반 대출 승인 예측 분류 모델 구축

이번 튜토리얼에서는 대출 승인 여부를 예측하는 분류 문제를 다룹니다. 금융 데이터를 활용해 머신러닝 모델을 구축하고 평가하는 전체 과정을 살펴봅니다.

데이터셋 구성

사용할 데이터셋은 대출 승인 결과를 예측하는 이진 분류 문제입니다.

컬럼명설명
Loan_ID대출 고유 식별자
Gender성별
Married결혼 여부
Dependents부양 가족 수
Education학력 수준(대졸/비대졸)
Self_Employed자영업 여부
ApplicantIncome신청자 소득
CoapplicantIncome공동 신청자 소득
LoanAmount대출 금액(천 단위)
Loan_Amount_Term대출 기간(월)
Credit_History신용 이력 충족 여부
Property_Area부동산 위치(도시/준도시/농촌)
Loan_Status대출 승인 상태(타겟 변수)

학습 데이터는 614개 샘플과 13개 특성을 포함하고, 테스트 데이터는 367개 샘플에 타겟 변수가 제외된 12개 특성으로 구성됩니다.

1. 데이터 불러오기

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# 데이터 로드
raw_train = pd.read_csv("data/train_LoanPrediction.csv")
raw_test = pd.read_csv("data/test_LoanPrediction.csv")

2. 데이터 전처리

클래스 불균형 처리

타겟 클래스의 분포가 고르지 않을 경우 모델이 편향될 수 있으므로 샘플링을 통해 균형을 맞춥니다.

approved = raw_train.Loan_Status == 'Y'
rejected = raw_train.Loan_Status == 'N'

approved_samples = raw_train[approved].sample(n=192, random_state=42)
rejected_samples = raw_train[rejected]

balanced_train = pd.concat([approved_samples, rejected_samples])

인코딩 변환

범주형 데이터를 수치형으로 변환하여 알고리즘이 처리할 수 있도록 합니다.

from sklearn.preprocessing import LabelEncoder

encoder = LabelEncoder()
balanced_train['Loan_Status'] = encoder.fit_transform(balanced_train['Loan_Status'])
balanced_train['Education'] = encoder.fit_transform(balanced_train['Education'])
raw_test['Education'] = encoder.transform(raw_test['Education'])

balanced_train = pd.get_dummies(balanced_train, columns=['Property_Area'], drop_first=True)
raw_test = pd.get_dummies(raw_test, columns=['Property_Area'], drop_first=True)

3. 특성 선택

selected_features = ['Education', 'Credit_History', 'Property_Area_Semiurban', 'Property_Area_Urban']

X_train = balanced_train[selected_features]
y_train = balanced_train['Loan_Status']

X_test = raw_test[selected_features]

4. 모델 학습 및 평가

여러 분류 알고리즘을 비교하고 교차 검증과 하이퍼파라미터 튜닝을 수행합니다.

from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.svm import SVC
from sklearn.model_selection import StratifiedKFold, cross_val_score, GridSearchCV

# 후보 모델 정의
candidates = [
    LogisticRegression(random_state=42),
    RandomForestClassifier(random_state=42),
    SVC(random_state=42)
]

# 5-폴드 교차 검증
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for candidate in candidates:
    scores = cross_val_score(candidate, X_train, y_train, cv=skf, scoring='accuracy')
    print(f"{candidate.__class__.__name__} - 평균 정확도: {scores.mean():.4f}, 표준편차: {scores.std():.4f}")

# 최적 파라미터 탐색
param_candidates = {'solver': ['newton-cg', 'lbfgs', 'liblinear', 'saga']}
optimizer = GridSearchCV(
    LogisticRegression(random_state=42, max_iter=1000),
    param_candidates,
    scoring="accuracy",
    cv=skf,
    refit=True
)
optimizer.fit(X_train, y_train)

5. 특성 중요도 해석

final_model = optimizer.best_estimator_

importance_df = pd.DataFrame({
    'feature': selected_features,
    'weight': final_model.coef_[0]
}).sort_values('weight', ascending=True)

importance_df.plot.barh(x='feature', y='weight', legend=False, figsize=(10, 6))
plt.xlabel('계수 값')
plt.title('로지스틱 회귀 특성 가중치')
plt.tight_layout()
plt.show()

6. 모델 배포 및 예측

import joblib

# 모델 영속화
model_path = 'model/loan_approval_model.pkl'
joblib.dump(optimizer, model_path)

# 모델 복원
deployed_model = joblib.load(model_path)

# 신규 샘플 예측
new_application = np.array([[0, 1, 0, 0]])
prediction = deployed_model.predict(new_application)
probability = deployed_model.predict_proba(new_application)

print(f"예측 결과: {'승인' if prediction[0] == 1 else '거절'}")
print(f"승인 확률: {probability[0][1]:.4f}")

핵심 개념 정리

분류 문제의 특성

Loan_Status가 이산적인 두 클래스(Y/N)로 구분되므로 지도학습의 분류 유형에 해당합니다. 회귀와 달리 출력값이 연속적이지 않습니다.

클래스 불균형의 영향

학습 데이터에서 한 클래스가 과도하게 많으면 모델이 다수 클래스로 편향되어 예측합니다. 이는 소수 클래스의 중요한 패턴을 놓치게 만들어 실제 성능과 검증 성능의 괴리를 발생시킵니다.

인코딩 방식 선택

순서가 있는 범주(학력 수준: 고졸-대졸-석사)에는 레이블 인코딩을, 순서가 없는 범주(지역 유형)에는 원-핫 인코딩을 적용합니다. 잘못된 인코딩은 알고리즘이 가짜 순서 관계를 학습하게 합니다.

교차 검증의 필요성

데이터를 여러 번 분할해 평가하면 특정 분할에 대한 과적합을 방지하고 모델의 일반화 성능을 신뢰할 수 있습니다. StratifiedKFold는 클래스 비율을 유지하며 분할합니다.

정확도의 한계

불균형 데이터에서는 높은 정확도가 오해를 불러올 수 있습니다. 정밀도(예측한 승인 중 실제 승인 비율), 재현율(실제 승인 중 예측 성공 비율), F1-점수(둘의 조화평균)를 함께 확인해야 합니다.

ROC 곡선 해석

임계값을 변화시켜 얻는 참긍정률과 거짓긍정률의 관계를 시각화합니다. AUC가 0.5면 무작위 예측, 1.0이면 완벽 분류를 의미합니다. 다양한 임계값에서의 성능 변화를 한눈에 파악할 수 있습니다.

과적합과 과소적합

훈련 정확도는 높으나 검증 정확도가 낮으면 과적합(모델이 훈련 데이터의 노이즈까지 학습), 둘 다 낮으면 과소적합(모델이 데이터의 패턴을 충분히 학습하지 못함)을 의심합니다. 학습 곡선을 통해 진단합니다.

태그: scikit-learn LogisticRegression RandomForestClassifier classification GridSearchCV

10월 1일 17:48에 게시됨