학습 목표:
1. YOLOv5 v6.0 알고리즘 프레임워크 이해
2. 날씨 장면 인식을 위한 코드 구현
모델 개요:
YOLOv5 v6.0은 세 가지 주요 구성요소인 백본(Backbone), 넥(Neck), 예측(Prediction)으로 이루어져 있다. 백본은 입력 이미지를 단계적으로 서브샘플링하면서 저수준의 질감 정보에서 고수준의 의미론적 정보까지 다양한 스케일의 특성을 추출한다. 넥은 업샘플링, 특성 연결, 다시 한번 서브샘플링을 통해 서로 다른 스케일의 특성 간 양방향 융합을 실현한다. 예측부는 세 가지 스케일에서 객체 클래스, 객체 신뢰도, 바운딩 박스 위치를 출력한다.
본 문서의 코드는 날씨 장면 분류 작업에 맞게 구조를 조정하였다. Specifically, the feature extraction structure corresponding to layers 0-9 (Conv-C3-SPPF) in the figure is retained, while layers 10-23 (Neck and three detection branches) are removed. After flattening the backbone output, it is passed through a two-layer fully connected network to output predictions for four classes: cloudy, rain, shine, and sunrise.
코드 구현:
필요 라이브러리 임포트 및 GPU 설정:
이전 과정에서 VGG-16 모델을 학습하였다. 해당 모델은 卷積層과 풀링層을 연속으로 쌓아 특성을 추출하는 간단하고 직관적인 구조를 가지고 있으나, 파라미터 수가 많고跨層 연결이 부재하여 과대적합 문제가 발생하기 쉬운 단점이 있다. 그 다음으로는 ResNet34 모델을 학습하였는데, 이 모델은 잔차 연결을 도입하여深层 네트워크에서의 기울기 전파 문제를 해소하였다. VGG-16에 비해 파라미터 활용 효율이 높고 훈련이 안정적이다. 이번 주에는 YOLOv5의 백본 네트워크를 기반으로 한 분류 모델을 학습하였다. Conv, C3, SPPF 모듈을 활용하여 잔차 학습의基础上分支融合 및 다중 스케일 특성 추출 능력을 갖추고 있어, 날씨 이미지에서 구름, 조명, 하늘 색상 등의 특성을 추출하기에 더욱 적합하다.
세 가지 모델 모두 卷積을 통해 이미지 특성을 추출하지만, 설계 철학은 점차 진화하고 있다. VGG-16의 卷積堆積에서 ResNet34의 잔차 학습, 그리고 YOLOv5의跨단계 특성 융합과 다중 스케일 특성 추출로 발전해왔으며, 이를 통해 卷積 신경망 구조의演进와 다양한 모듈의 역할을 더욱 깊이 이해하게 되었다.
1. YOLOv5 v6.0 알고리즘 프레임워크 이해
2. 날씨 장면 인식을 위한 코드 구현
모델 개요:
YOLOv5 v6.0은 세 가지 주요 구성요소인 백본(Backbone), 넥(Neck), 예측(Prediction)으로 이루어져 있다. 백본은 입력 이미지를 단계적으로 서브샘플링하면서 저수준의 질감 정보에서 고수준의 의미론적 정보까지 다양한 스케일의 특성을 추출한다. 넥은 업샘플링, 특성 연결, 다시 한번 서브샘플링을 통해 서로 다른 스케일의 특성 간 양방향 융합을 실현한다. 예측부는 세 가지 스케일에서 객체 클래스, 객체 신뢰도, 바운딩 박스 위치를 출력한다.
본 문서의 코드는 날씨 장면 분류 작업에 맞게 구조를 조정하였다. Specifically, the feature extraction structure corresponding to layers 0-9 (Conv-C3-SPPF) in the figure is retained, while layers 10-23 (Neck and three detection branches) are removed. After flattening the backbone output, it is passed through a two-layer fully connected network to output predictions for four classes: cloudy, rain, shine, and sunrise.
코드 구현:
필요 라이브러리 임포트 및 GPU 설정:
import torch
import torch.nn as nn
import torchvision.transforms as transforms
import torchvision
from torchvision import transforms, datasets
import os, pathlib, warnings
warnings.filterwarnings("ignore")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"사용 디바이스: {device}")
데이터셋 로드:data_path = './weather-data/'
data_dir = pathlib.Path(data_path)
image_paths = list(data_dir.glob('*'))
class_labels = [str(p).split("\\")[1] for p in image_paths]
print(f"클래스 목록: {class_labels}")
데이터 전처리 설정:train_transform = transforms.Compose([
transforms.Resize([224, 224]),
transforms.ToTensor(),
transforms.Normalize(
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225])
])
val_transform = transforms.Compose([
transforms.Resize([224, 224]),
transforms.ToTensor(),
transforms.Normalize(
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225])
])
full_dataset = datasets.ImageFolder("./weather-data/", transform=train_transform)
print(f"전체 데이터셋 크기: {len(full_dataset)}")
훈련/검증 데이터 분할:train_samples = int(0.8 * len(full_dataset))
val_samples = len(full_dataset) - train_samples
train_data, val_data = torch.utils.data.random_split(full_dataset, [train_samples, val_samples])
print(f"훈련 데이터: {len(train_data)}, 검증 데이터: {len(val_data)}")
데이터로더 생성:batch_size = 8
train_loader = torch.utils.data.DataLoader(
train_data,
batch_size=batch_size,
shuffle=True,
num_workers=1
)
val_loader = torch.utils.data.DataLoader(
val_data,
batch_size=batch_size,
shuffle=True,
num_workers=1
)
배치 데이터 형태 확인:for batch_x, batch_y in val_loader:
print(f"배치 형태 [N, C, H, W]: {batch_x.shape}")
print(f"레이블 형태: {batch_y.shape}, 타입: {batch_y.dtype}")
break
모델 아키텍처 구축:import torch.nn.functional as F
def auto_pad(kernel_size, padding=None):
if padding is None:
padding = kernel_size // 2 if isinstance(kernel_size, int) else [x // 2 for x in kernel_size]
return padding
class ConvBlock(nn.Module):
def __init__(self, in_channels, out_channels, kernel=1, stride=1, padding=None, groups=1, activation=True):
super().__init__()
self.conv = nn.Conv2d(in_channels, out_channels, kernel, stride, auto_pad(kernel, padding), groups=groups, bias=False)
self.bn = nn.BatchNorm2d(out_channels)
self.act = nn.SiLU() if activation is True else (activation if isinstance(activation, nn.Module) else nn.Identity())
def forward(self, x):
return self.act(self.bn(self.conv(x)))
class ResidualBlock(nn.Module):
def __init__(self, in_channels, out_channels, use_shortcut=True, groups=1, expansion=0.5):
super().__init__()
hidden_channels = int(out_channels * expansion)
self.block1 = ConvBlock(in_channels, hidden_channels, 1, 1)
self.block2 = ConvBlock(hidden_channels, out_channels, 3, 1, groups=groups)
self.use_shortcut = use_shortcut and in_channels == out_channels
def forward(self, x):
return x + self.block2(self.block1(x)) if self.use_shortcut else self.block2(self.block1(x))
class CSPModule(nn.Module):
def __init__(self, in_channels, out_channels, num_blocks=1, use_shortcut=True, groups=1, expansion=0.5):
super().__init__()
hidden_channels = int(out_channels * expansion)
self.conv1 = ConvBlock(in_channels, hidden_channels, 1, 1)
self.conv2 = ConvBlock(in_channels, hidden_channels, 1, 1)
self.conv3 = ConvBlock(2 * hidden_channels, out_channels, 1)
self.blocks = nn.Sequential(*(
ResidualBlock(hidden_channels, hidden_channels, use_shortcut, groups, 1.0)
for _ in range(num_blocks)
))
def forward(self, x):
return self.conv3(torch.cat((self.blocks(self.conv1(x)), self.conv2(x)), dim=1))
class SpatialPoolFast(nn.Module):
def __init__(self, in_channels, out_channels, kernel_size=5):
super().__init__()
hidden_channels = in_channels // 2
self.conv1 = ConvBlock(in_channels, hidden_channels, 1, 1)
self.conv2 = ConvBlock(hidden_channels * 4, out_channels, 1, 1)
self.pool = nn.MaxPool2d(kernel_size=kernel_size, stride=1, padding=kernel_size // 2)
def forward(self, x):
x = self.conv1(x)
with warnings.catch_warnings():
warnings.simplefilter('ignore')
y1 = self.pool(x)
y2 = self.pool(y1)
return self.conv2(torch.cat([x, y1, y2, self.pool(y2)], 1))
class YOLOv5Backbone(nn.Module):
def __init__(self, num_classes=4):
super(YOLOv5Backbone, self).__init__()
self.layer1 = ConvBlock(3, 64, 3, 2, 2)
self.layer2 = ConvBlock(64, 128, 3, 2)
self.layer3 = CSPModule(128, 128)
self.layer4 = ConvBlock(128, 256, 3, 2)
self.layer5 = CSPModule(256, 256)
self.layer6 = ConvBlock(256, 512, 3, 2)
self.layer7 = CSPModule(512, 512)
self.layer8 = ConvBlock(512, 1024, 3, 2)
self.layer9 = CSPModule(1024, 1024)
self.layer10 = SpatialPoolFast(1024, 1024, 5)
self.classifier = nn.Sequential(
nn.Linear(in_features=65536, out_features=256),
nn.ReLU(),
nn.Dropout(0.5),
nn.Linear(in_features=256, out_features=num_classes)
)
def forward(self, x):
x = self.layer1(x)
x = self.layer2(x)
x = self.layer3(x)
x = self.layer4(x)
x = self.layer5(x)
x = self.layer6(x)
x = self.layer7(x)
x = self.layer8(x)
x = self.layer9(x)
x = self.layer10(x)
x = torch.flatten(x, start_dim=1)
x = self.classifier(x)
return x
model = YOLOv5Backbone(num_classes=4).to(device)
print(model)
모델 파라미터 확인:import torchsummary as torch_summary
torch_summary.summary(model, (3, 224, 224))
훈련 함수 정의:def train_epoch(dataloader, network, criterion, optimizer):
dataset_size = len(dataloader.dataset)
num_batches = len(dataloader)
epoch_loss = 0.0
epoch_correct = 0
for inputs, targets in dataloader:
inputs = inputs.to(device)
targets = targets.to(device)
predictions = network(inputs)
loss = criterion(predictions, targets)
optimizer.zero_grad()
loss.backward()
optimizer.step()
epoch_correct += (predictions.argmax(1) == targets).type(torch.float).sum().item()
epoch_loss += loss.item()
accuracy = epoch_correct / dataset_size
avg_loss = epoch_loss / num_batches
return accuracy, avg_loss
평가 함수 정의:def evaluate(dataloader, network, criterion):
dataset_size = len(dataloader.dataset)
num_batches = len(dataloader)
eval_loss = 0.0
eval_correct = 0
with torch.no_grad():
for inputs, targets in dataloader:
inputs = inputs.to(device)
targets = targets.to(device)
predictions = network(inputs)
loss = criterion(predictions, targets)
eval_correct += (predictions.argmax(1) == targets).type(torch.float).sum().item()
eval_loss += loss.item()
accuracy = eval_correct / dataset_size
avg_loss = eval_loss / num_batches
return accuracy, avg_loss
모델 훈련 실행:import copy
optimizer = torch.optim.Adam(model.parameters(), lr=1e-4)
criterion = nn.CrossEntropyLoss()
num_epochs = 50
history_train_acc = []
history_train_loss = []
history_val_acc = []
history_val_loss = []
best_accuracy = 0.0
best_model_state = None
for epoch in range(num_epochs):
model.train()
train_acc, train_loss = train_epoch(train_loader, model, criterion, optimizer)
model.eval()
val_acc, val_loss = evaluate(val_loader, model, criterion)
if val_acc > best_accuracy:
best_accuracy = val_acc
best_model_state = copy.deepcopy(model.state_dict())
history_train_acc.append(train_acc)
history_train_loss.append(train_loss)
history_val_acc.append(val_acc)
history_val_loss.append(val_loss)
current_lr = optimizer.state_dict()['param_groups'][0]['lr']
print(f"Epoch {epoch+1:02d} | Train Acc: {train_acc*100:.2f}% | Train Loss: {train_loss:.4f} | Val Acc: {val_acc*100:.2f}% | Val Loss: {val_loss:.4f} | LR: {current_lr:.2E}")
torch.save(best_model_state, './best_weather_model.pth')
print("훈련 완료!")
결과 시각화:import matplotlib.pyplot as plt
import warnings
warnings.filterwarnings("ignore")
plt.rcParams['font.sans-serif'] = ['SimHei']
plt.rcParams['axes.unicode_minus'] = False
plt.rcParams['figure.dpi'] = 100
epoch_range = range(num_epochs)
fig, axes = plt.subplots(1, 2, figsize=(14, 4))
axes[0].plot(epoch_range, history_train_acc, label='훈련 정확도')
axes[0].plot(epoch_range, history_val_acc, label='검증 정확도')
axes[0].legend(loc='lower right')
axes[0].set_title('훈련 및 검증 정확도')
axes[0].set_xlabel('Epoch')
axes[0].set_ylabel('Accuracy')
axes[1].plot(epoch_range, history_train_loss, label='훈련 손실')
axes[1].plot(epoch_range, history_val_loss, label='검증 손실')
axes[1].legend(loc='upper right')
axes[1].set_title('훈련 및 검증 손실')
axes[1].set_xlabel('Epoch')
axes[1].set_ylabel('Loss')
plt.tight_layout()
plt.show()
최종 모델 성능 평가:final_model = YOLOv5Backbone(num_classes=4).to(device)
final_model.load_state_dict(torch.load('./best_weather_model.pth', map_location=device))
final_acc, final_loss = evaluate(val_loader, final_model, criterion)
print(f"최종 검증 정확도: {final_acc*100:.2f}%")
print(f"최종 검증 손실: {final_loss:.4f}")
학습 내용 정리:이전 과정에서 VGG-16 모델을 학습하였다. 해당 모델은 卷積層과 풀링層을 연속으로 쌓아 특성을 추출하는 간단하고 직관적인 구조를 가지고 있으나, 파라미터 수가 많고跨層 연결이 부재하여 과대적합 문제가 발생하기 쉬운 단점이 있다. 그 다음으로는 ResNet34 모델을 학습하였는데, 이 모델은 잔차 연결을 도입하여深层 네트워크에서의 기울기 전파 문제를 해소하였다. VGG-16에 비해 파라미터 활용 효율이 높고 훈련이 안정적이다. 이번 주에는 YOLOv5의 백본 네트워크를 기반으로 한 분류 모델을 학습하였다. Conv, C3, SPPF 모듈을 활용하여 잔차 학습의基础上分支融合 및 다중 스케일 특성 추출 능력을 갖추고 있어, 날씨 이미지에서 구름, 조명, 하늘 색상 등의 특성을 추출하기에 더욱 적합하다.
세 가지 모델 모두 卷積을 통해 이미지 특성을 추출하지만, 설계 철학은 점차 진화하고 있다. VGG-16의 卷積堆積에서 ResNet34의 잔차 학습, 그리고 YOLOv5의跨단계 특성 융합과 다중 스케일 특성 추출로 발전해왔으며, 이를 통해 卷積 신경망 구조의演进와 다양한 모듈의 역할을 더욱 깊이 이해하게 되었다.