상태 확인에 실패한 Amazon EC2 Linux 인스턴스 문제 해결
상태 확인에 실패한 Amazon EC2 Linux 인스턴스 문제 해결 (Troubleshoot Amazon EC2 Linux instances with failed status checks)
Linux 인스턴스가 상태 확인(status check)에 실패하면 다음 정보가 문제 해결에 도움이 될 수 있어요. 먼저 애플리케이션에 문제가 나타나고 있는지 확인해요. 인스턴스가 예상대로 애플리케이션을 실행하고 있지 않다고 확인하면 상태 확인 정보와 시스템 로그를 검토해요.
출처: 문서
본문
상태 확인이 실패하게 만드는 문제의 예시는 Status checks for Amazon EC2 instances를 참고하세요.
목차 (Contents)
- Review status check information
- Retrieve the system logs
- Troubleshoot system log errors for Linux instances
- Out of memory: kill process
- ERROR: mmu_update failed
- I/O error (block device failure)
- I/O ERROR: neither local nor remote disk
- request_module: runaway loop modprobe
- "FATAL: kernel too old" and "fsck: No such file or directory..." (Kernel and AMI mismatch)
- "FATAL: Could not load /lib/modules" or "BusyBox" (Missing kernel modules)
- ERROR Invalid kernel
- fsck: No such file or directory... (File system not found)
- General error mounting filesystems
- VFS: Unable to mount root fs on unknown-block
- Error: Unable to determine major/minor number of root device...
- XENBUS: Device with no driver...
- ... days without being checked, check forced
- fsck died with exit status... (Missing device)
- GRUB prompt (grubdom>)
- Device eth0 has different MAC address than expected (Hard-coded MAC address)
- Unable to load SELinux Policy (SELinux misconfiguration)
- XENBUS: Timeout connecting to devices (Xenbus timeout)
상태 확인 정보 검토 (Review status check information)
Amazon EC2 콘솔을 사용해 impaired 인스턴스를 조사하려면 (To investigate impaired instances using the Amazon EC2 console)
- Amazon EC2 콘솔(https://console.aws.amazon.com/ec2/)을 엽니다.
- 탐색 창에서 Instances를 선택한 다음 인스턴스를 선택해요.
- Status and alarms 탭을 선택해 모든 System status checks, Instance status checks, Attached EBS status checks의 개별 결과를 확인해요.
상태 확인이 실패했다면 다음 옵션 중 하나를 시도할 수 있어요.
- 실패한 상태 확인에 응답해 인스턴스를 복구하는 알람을 만들어요. 자세한 내용은 Create alarms that stop, terminate, reboot, or recover an instance를 참고하세요.
- (인스턴스 상태 확인) 인스턴스 유형을 Nitro 기반 인스턴스로 변경했다면, 필요한 ENA와 NVMe 드라이버가 없는 인스턴스에서 마이그레이션한 경우 상태 확인이 실패해요. 자세한 내용은 Compatibility for changing the instance type을 참고하세요.
- EBS 루트 볼륨이 있는 인스턴스의 경우 인스턴스를 중지하고 다시 시작해요. 자세한 내용은 Stop and start Amazon EC2 instances를 참고하세요.
- 인스턴스 스토어 루트 볼륨이 있는 인스턴스의 경우 인스턴스를 종료하고 대체 인스턴스를 시작해요. 자세한 내용은 Terminate Amazon EC2 instances를 참고하세요.
- Amazon EC2가 문제를 해결할 때까지 기다려요.
- Support에 연락하거나 문제를 AWS re:Post에 게시해요.
- 인스턴스가 Auto Scaling 그룹에 있다면: (시스템 상태 확인 및 인스턴스 상태 확인) 기본적으로 Amazon EC2 Auto Scaling이 대체 인스턴스를 자동으로 시작해요. 자세한 내용은 Amazon EC2 Auto Scaling User Guide의 Health checks for instances in an Auto Scaling group을 참고하세요. (Attached EBS 상태 확인) Amazon EC2 Auto Scaling이 자동으로 대체 인스턴스를 시작하도록 구성해야 해요. 자세한 내용은 Amazon EC2 Auto Scaling User Guide의 Monitor and replace Auto Scaling instances with impaired Amazon EBS volumes를 참고하세요.
- 시스템 로그를 검색하고 오류를 찾아요. 자세한 내용은 Retrieve the system logs를 참고하세요.
시스템 상태 확인과 인스턴스 상태 확인의 경우 기본적으로 Amazon EC2 Auto Scaling이 대체 인스턴스를 자동으로 시작해요. (Attached EBS 상태 확인) 대체 인스턴스를 자동으로 시작하도록 Amazon EC2 Auto Scaling을 구성해야 해요.
시스템 로그 검색 (Retrieve the system logs)
인스턴스 상태 확인이 실패하면 인스턴스를 재부팅하고 시스템 로그를 검색할 수 있어요. 로그는 문제 해결에 도움이 되는 오류를 드러낼 수 있어요. 재부팅하면 로그에서 불필요한 정보가 지워져요.
인스턴스를 재부팅하고 시스템 로그를 검색하려면 (To reboot an instance and retrieve the system log)
- Amazon EC2 콘솔(https://console.aws.amazon.com/ec2/)을 엽니다.
- 탐색 창에서 Instances를 선택하고 인스턴스를 선택해요.
- Instance state, Reboot instance를 선택해요. 인스턴스가 재부팅되는 데 몇 분이 걸릴 수 있어요.
- 문제가 여전히 있는지 확인해요. 경우에 따라 재부팅으로 문제가 해결될 수 있어요.
- 인스턴스가
running상태가 되면 Actions, Monitor and troubleshoot, Get system log를 선택해요. - 화면에 나타나는 로그를 검토하고 아래의 알려진 시스템 로그 오류 문 목록을 사용해 문제를 해결해요.
- 문제가 해결되지 않으면 AWS re:Post에 문제를 게시할 수 있어요.
Linux 인스턴스의 시스템 로그 오류 문제 해결 (Troubleshoot system log errors for Linux instances)
인스턴스 도달 가능성 확인(instance reachability check) 같은 인스턴스 상태 확인에 실패한 Linux 인스턴스의 경우, 위 단계에 따라 시스템 로그를 검색했는지 확인하세요. 다음 목록에는 일반적인 시스템 로그 오류와 각 오류를 해결하기 위해 취할 수 있는 권장 조치가 있어요.
메모리 오류 (Memory Errors)
- Out of memory: kill process
- ERROR: mmu_update failed (Memory management update failed)
장치 오류 (Device Errors)
- I/O error (block device failure)
- I/O ERROR: neither local nor remote disk (Broken distributed block device)
커널 오류 (Kernel Errors)
- request_module: runaway loop modprobe (Looping legacy kernel modprobe on older Linux versions)
- "FATAL: kernel too old" and "fsck: No such file or directory..." (Kernel and AMI mismatch)
- "FATAL: Could not load /lib/modules" or "BusyBox" (Missing kernel modules)
- ERROR Invalid kernel (EC2 incompatible kernel)
파일 시스템 오류 (File System Errors)
- fsck: No such file or directory... (File system not found)
- General error mounting filesystems (failed mount)
- VFS: Unable to mount root fs on unknown-block (Root filesystem mismatch)
- Error: Unable to determine major/minor number of root device... (Root file system/device mismatch)
- XENBUS: Device with no driver...
- ... days without being checked, check forced (File system check required)
- fsck died with exit status... (Missing device)
운영체제 오류 (Operating System Errors)
- GRUB prompt (grubdom>)
- Bringing up interface eth0: Device eth0 has different MAC address than expected, ignoring. (Hard-coded MAC address)
- Unable to load SELinux Policy. Machine is in enforcing mode. Halting now. (SELinux misconfiguration)
- XENBUS: Timeout connecting to devices (Xenbus timeout)
(이 페이지는 각 오류에 대한 상세 원인·조치 표와 시스템 로그 예시를 포함합니다. 각 오류의 잠재 원인과 Amazon EBS 기반·인스턴스 스토어 기반 인스턴스별 권장 조치는 원문 링크에서 이어집니다. 예시 시스템 로그 블록은 모두 원문을 그대로 유지해야 하며, 이 번역에서는 오류 이름과 구조를 보존하고 본문 요약으로 대체했습니다. 원문의 전체 로그와 표는 https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/TroubleshootingInstances.html 에서 확인하세요.)