목차
- 설치
- 기본 사용법
- 데이터 타입 3.1. object 타입 3.2. string 타입 3.3. numeric 타입 3.4. array 타입
본 문서에서는 JSON Schema의 기본 문법을 다루며, 필요에 따라 Python 코드를 통해 검증 과정을 확인한다.
JSON Schema는 JSON 데이터의 구조, 형식 및 제약 조건을 정의하기 위한 표준이다. JSON Schema를 활용하면 JSON 데이터의 유효성을 검증하고 문서화하여 데이터의 정확성과 무결성을 보장할 수 있다.
- 설치
pip install jsonschema
- 기본 사용법
JSON Schema를 사용하려면 검증 대상 JSON 데이터와 해당 데이터의 구조를 정의한 스키마를 준비해야 한다. validate 함수를 통해 인스턴스가 스키마에 부합하는지 확인할 수 있다.
from jsonschema.validators import validate
sample_data = {
"username": "kimseong",
"userAge": 24,
"isActive": True,
"interest": "programming"
}
validation_schema = {
"type": "object",
"properties": {
"username": {
"type": "string"
},
"userAge": {
"type": "number"
},
"isActive": {
"type": "boolean"
},
"interest": {
"type": "string"
}
}
}
def validate_user_data():
validate(sample_data, validation_schema)
API 응답 검증 예시
외부 API를 호출한 후 반환되는 응답 데이터의 구조를 검증해야 하는 경우가 많다. 다음은 API 응답의 필드 타입을 확인하는 예시이다.
import requests
response_schema = {
"type": "object",
"properties": {
"statusCode": {
"type": "string"
},
"message": {
"type": "string"
},
"payload": {
"type": "array",
"items": {
"type": "object",
"properties": {
"idx": {
"type": "number"
},
"subject": {
"type": "string"
},
"body": {
"type": "string"
},
"authorIdx": {
"type": "number"
},
"removed": {
"type": "number"
},
"createdAt": {
"type": "string"
},
"modifiedAt": {
"type": "string"
},
"currentUser": {
"type": "boolean"
}
}
}
}
}
}
api_endpoint = ("https://api.example.com/articles/list")
auth_headers = {
"authorization": "Bearer eyJhbGciOiJIUzI1NiJ9.eyJ1c2VySWQiOjEsInVzZXJuYW1lIjoiYW5vbnltIiwiZXhwIjoxNzU0ODI3OTU5fQ.AAAA-BBBB-CCCC-DDDD"
}
def check_api_response():
response = requests.get(url=api_endpoint, headers=auth_headers)
validate(response.json(), response_schema)
- 데이터 타입
이전 섹션에서 type 키워드를 object로 설정하여 JSON 객체를 검증하는 방법을 살펴보았다. 그러나 type 키워드는 다른 값도 받을 수 있으며, 각 데이터 타입마다 고유한 제약 키워드가 존재한다.
type 키워드가 지원하는 값은 다음과 같다: object, string, array, integer, number, boolean, null. 본 문서에서는 boolean과 null에 대해서는 별도로 다루지 않는다.
JSON Schema는 중첩 구조를 갖는 형태로, 여러 키워드의 값으로 또 다른 JSON Schema를 사용할 수 있다.
주요 데이터 타입:
| 타입 | 설명 |
|---|---|
| string | 문자열 타입, 텍스트 데이터에 사용 |
| number | 숫자 타입, 부동소수점 수를 표현 |
| integer | 정수 타입, 정수 값을 표현 |
| boolean | 布尔 타입, true 또는 false 값 |
| object | 객체 타입, 중첩된 JSON 객체에 사용 |
| array | 배열 타입, 리스트나 집합에 사용 |
| null | null 값 타입 |
3.1. object 타입
object 타입은 Python의 딕셔너리와 유사하며, 다음과 같은 검증 키워드를 제공한다:
| 키워드 | 설명 |
|---|---|
| properties | 검증할 JSON 객체의 각 속성 이름을 키로 가지며, 값은 해당 속성을 검증할 스키마이다. 일치하지 않는 속성은 무시된다. |
| patternProperties | 속성 이름이 정규식과 일치할 경우 해당 패턴 스키마를 적용하여 검증 |
| additionalProperties | true 또는 false 값을 가지며, properties에 정의되지 않은 추가 속성을 허용할지 결정. 기본값은 true |
| required | 필수로 존재해야 하는 속성들의 목록을 지정 |
| minProperties | 객체가 가져야 하는 최소 속성 개수 |
| maxProperties | 객체가 가질 수 있는 최대 속성 개수 |
| propertyNames | 속성 이름이 만족해야 하는 패턴. pattern 키워드와 함께 사용 가능 |
적용 예시:
from jsonschema import validate
schema_definition = {
"type": "object",
"properties": {
"username": {
"type": "string"
},
"userAge": {
"type": "number"
},
"contactEmail": {
"type": "string"
}
},
"required": ["username", "contactEmail"],
"additionalProperties": False,
"minProperties": 2,
"maxProperties": 5
}
user_info = {
"username": "kimseong",
"userAge": 28,
"contactEmail": "user@example.com"
}
def verify_object_schema():
validate(user_info, schema_definition)
3.2. string 타입
string 타입에는 다음과 같은 검증 키워드가 적용된다:
| 키워드 | 설명 |
|---|---|
| minLength / maxLength | 문자열의 최소 길이와 최대 길이를 지정 |
| pattern | 문자열이 만족해야 하는 정규식 패턴 |
| format | 날짜와 시간 형식: "date-time", "time", "date", "duration"이메일 형식: "email", "idn-email"도메인 형식: "hostname", "idn-hostname"IP 주소 형식: "ipv4", "ipv6"리소스 식별자 형식: "uuid", "uri", "uri-reference", "iri", "iri-reference"URI 템플릿: "uri-template"JSON 포인터: "json-pointer", "relative-json-pointer"정규식: "regex" |
길이 제한 예시:
def check_string_constraints():
string_schema = {
"type": "object",
"properties": {
"fullName": {
"type": "string",
"maxLength": 30,
"minLength": 2
},
"registeredDate": {
"type": "string",
"format": "date"
}
}
}
sample_record = {
"fullName": "김성민",
"registeredDate": "2025-08-07"
}
validate(sample_record, string_schema)
정규식 패턴 예시:
def validate_with_regex():
test_record = {
"fullName": "leejiwon",
"score": 85
}
regex_schema = {
"type": "object",
"properties": {
"fullName": {
"type": "string",
"pattern": "^[a-zA-Z]{2,10}$"
},
"score": {
"type": "number"
}
}
}
validate(test_record, regex_schema)
3.3. numeric 타입
JSON Schema의 numeric 타입은 두 가지로 구분된다. number는 Python의 float과 int 모두를 포함하며, integer는 정수만 다룬다. 여기서는 number 타입의 주요 검증 키워드를 살펴본다:
| 키워드 | 설명 |
|---|---|
| multipleOf | 양의 정수. 해당 값의 배수여야 함 |
| minimum | 숫자가 이 값보다 크거나 같아야 함 |
| maximum | 숫자가 이 값보다 작거나 같아야 함 |
| exclusiveMinimum | 숫자가 이 값보다 커야 함 |
| exclusiveMaximum | 숫자가 이 값보다 작아야 함 |
범위 검증 예시:
def numeric_range_validation():
numeric_schema = {
"type": "object",
"properties": {
"userAge": {
"type": "number",
"minimum": 18,
"maximum": 65,
"exclusiveMinimum": False,
"exclusiveMaximum": False
}
}
}
person_data = {
"userAge": 30
}
validate(person_data, numeric_schema)
3.4. array 타입
array 타입은 Python의 리스트나 튜플과 유사하며, 다음과 같은 검증 키워드를 사용한다:
| 키워드 | 설명 |
|---|---|
| items | 배열의 모든 요소를 검증할 스키마 |
| prefixItems | 튜플 검증 시 인덱스별로 각 요소를 검증. items를 false로 설정하면 추가 요소 허용 불가 |
| contains | 배열에 최소 하나 이상 포함해야 하는 요소의 스키마 |
| unevaluatedItems | true/false, items, prefixItems, contains에서 지정하지 않은 요소의 존재 허용 여부 |
| minItems | 최소 요소 개수 |
| maxItems | 최대 요소 개수 |
| uniqueItems | true/false, 모든 요소가 고유해야 하는지 여부 |
배열 검증 예시:
def array_schema_validation():
test_array = {
"values": [100, "greeting", 3.14]
}
array_schema = {
"type": "object",
"properties": {
"values": {
"type": "array",
"prefixItems": [
{"type": "number"},
{"type": "string"}
],
"unevaluatedItems": True
}
}
}
validate(test_array, array_schema)