ServerIssueApril2026/ServerIssueManual.html

107 lines
5.8 KiB
HTML

<!DOCTYPE html>
<html>
<head>
<title>Server Issue Avril 2026</title>
</head>
<body style="font-size: 20px;">
<h1>Table of Contents</h1>
<ul>
<li><a href="#Theproblem">The problem</a></li>
<li><a href="#Thefix">The fix</a></li>
<li><a href="#Conclusion">Conclusion</a></li>
<li><a href="#OtherPotentialFixes">Other potential fixes</a></li>
<li><a href="#References">References</a></li>
</ul>
<p>After a power outage at the hosting location (3ti), the server wouldnt boot. <br>
Connected remotely to a computer located at the hosting provider, connected to IPMI/iDRAC then used virtual console to see what was being displayed on the server.
</p>
<h1 id="Theproblem"><font color="Blue">The problem:</font></h1>
<br>
<p>Error was:</p><br>
<pre>
<hr>
Time out for waiting the udev queue being empty.
Time out for waiting the udev queue being empty.
Gave up waiting for root file system device. Common problems:
- Boot args (cat /proc/cmdline)
- Check rootdelay= (did the system wait long enough?)
- Missing modules (cat /proc/modules; ls /dev)
ALERT! /dev/mapper/pve-root does not exist. Dropping to a shell!
BusyBox v1.35.0 (Debian 1:1.13.0-4+b7) built-in shell (ash)
Enter help for a list of built-in commands.
(initramfs)_
<hr>
</pre>
<p>
This built-in shell lacks required tools to perform basic diagnostics let alone any rescue operations, but ls /dev wouldnt list the hard drives.<br><br>
Checked in the DRAC/Perc/BIOS and the drives, RAID and virtual disks are all healthy so there is probably no hardware failure.<br>
A reboot shows that GRUB does load so the drives ARE working.<br><br>
A search for /dev/mapper/pve-root does not exist proxmox shows this is a common issue and there are a few ways people have fixed them.<br>
<ol>
<li>Use a different kernel. When updating the kernel Proxmox keeps older kernel versions and they can be accessed by selecting the Advanced options for Proxmox Virtual Environment GNU/Linux boot option in GRUB
<ol type="a">
<li><a href="https://old.reddit.com/r/homelab/comments/1cetljo/i_need_help_devmapperpveroot_does_not_exist/">https://old.reddit.com/r/homelab/comments/1cetljo/i_need_help_devmapperpveroot_does_not_exist/</a></li>
</ol>
</li>
<li>Use Chroot to perform a rescue operation on the Linux installation. Including installing a new kernel if the Proxmox doesnt have other installed kernels.
<ol type="a">
<li><a href="https://forum.proxmox.com/threads/alert-dev-mapper-pve-root-does-not-exist.148029/">https://forum.proxmox.com/threads/alert-dev-mapper-pve-root-does-not-exist.148029/</a></li>
</ol>
</li>
<li>Add rootdelay=10 to GRUB
<ol type="a">
<li><a href="https://forum.proxmox.com/threads/solved-i-need-help-dev-mapper-pve-root-does-not-exist.116967/">https://forum.proxmox.com/threads/solved-i-need-help-dev-mapper-pve-root-does-not-exist.116967/</a></li>
</ol>
</li>
<li>Adding GRUB_CMDLINE_LINUX_DEFAULT="intel_iommu=on iommu=pt" or quiet intel_iommu=off intremap=off to GRUB</li>
</ol><br>
I tried <font color="purple"><b>option 3</b></font> by adding it directly in the GRUB selection menu to no avail, it didnt change anything. I assumed <font color="purple"><b>option 4</b></font> didnt have anything to do with my issue and wasnt even worth a try.
<font color="purple"><b>Option 1</b></font> did not help either. I couldnt see an older kernel in the advanced option. Either the system didnt update or it didnt keep the previous kernel version for me to boot on.<br>
I then went on to try and boot from an Ubuntu desktop live USB drive. But even from there I couldnt see the hard drive listed in /dev, or in lsblk or lshw, etc. At least I now had access to better tools that initramfs was giving me that I could use to perform a better diagnosis.
<ol>
<li><b>dmesg -T | tail -n 200</b> (look for PCIe, SAS, SATA, NVMe, link resets)</li>
<li><b>lsblk -e7 -o NAME,TYPE,SIZE,MODEL,SERIAL,TRAN,HCTL</b> (see what the kernel created)</li>
<li><b>lspci -nn | egrep -i 'sas|raid|sata|nvme|scsi'</b> (confirm the controller exists)</b></li>
</ol>
As I was getting desperate, I was basically trying anything I could find only, I eventually saw megaraid_sas 0000:01:00.0 in journalctl -k -b|egrep -i mpt3sas|megaraid|ahci|nvme|reset|timout|aer.
The RAID card I use is a Dell PERC h310 so where is that megaraid device coming from? It also has log entries saying reset/timeout, bingo, I thought, this is the issue, it. times out??
Now lspci does find something if I filter for megaraid (or lspci -k -s 01:00.0) instead of trying to find Dell or PERC.
There were also some messages in dmesg -T | egrep -I ahci|ata|SATA: SATA (1 to 5) link down (SStatus 0 SControl 300). I assumed this was related to the timeout with the megaraid card as it is a SAS/SATA compatible RAID device and my drives are all SATA.
I tried changing the drives to RAID in the server BIOS configuration even though it didnt make sense, and anyway no changes were made prior to the server stopping working. The fact that I could get to GRUB meant the system sees the disks, at least prior to GRUB or the OS being loaded. Not seeing the disks even in the Ubuntu live I took it that it meant that it was probably a type of driver issue, not a hardware problem. I began looking for a different PERC card that might be using a different driver like an H710 hoping that the new card would be seen properly by the OS and that the RAID configuration would transfer over to the new card.
</p>
<h1 id="Thefix"><font color="Blue">The Fix</font></h1>
<h1 id="Conclusion"><font color="Blue">Conclusion</font></h1>
<h1 id="OtherPotentialFixes"><font color="Blue">Other potential fixes</font></h1>
<h1 id="References"><Font color="Blue">References</Font></h1>
<h3>Test h3</h3>
</body>
</html>