Hardware and update recovery
Hardware and update recovery
OPEMOS targets hardware outside Valve’s currently validated Steam Deck set. Ethernet, Wi-Fi, hybrid graphics, suspend, display routing, and firmware must therefore be verified for each hardware profile. A successful image build does not certify every device in a PC.
If an update leaves a black screen
Do not reinstall or wipe the system immediately. A black display does not by itself prove that the update failed: the graphical session may have failed while the machine, another virtual terminal, and networking remain alive.
Try Ctrl+Alt+F3. If a login prompt appears, sign in and collect the
following without changing the boot configuration:
uname -r
cat /etc/os-release
cat /proc/cmdline
steamos-bootconf this-image 2>&1
modinfo -F version nvidia 2>&1
modinfo -F vermagic nvidia 2>&1
modinfo nvidia 2>&1 | sed -n '1,30p'
lsmod | grep -E 'nvidia|nouveau'
journalctl -b -p warning..alert --no-pager
journalctl -b | grep -iE 'nvidia|nouveau|gamescope|drm|firmware'
If SSH was already enabled, the same read-only collection can be run over
Ethernet. Preserve /var/log/steamos-nvidia-repatch.log if the legacy
installer’s self-heal path is present. Do not disable signature checking,
install a nearest-version module, or edit the inactive-slot boot files until
the running kernel, active slot, NVIDIA module vermagic, userspace version,
firmware, and failure log have been compared.
As a temporary diagnostic route, the legacy installer documents
steamos-session-select plasma from a TTY. That may recover a desktop when
Gaming Mode alone failed, but it does not repair a mismatched kernel module.
The first OPEMOS hardware update test observed the sharper failure case: the
newly activated rootfs-A contained kernel 6.16.12-valve24.5 but no
nvidia.ko, while rootfs-B retained an NVIDIA module. The firmware displayed
neither SteamOS’s automatic rollback menu nor its forced steamcl-menu UI on
that machine. This confirms that a structurally valid recovery image is not an
installed-update guarantee and that a visible bootloader menu cannot be the
only recovery path.
Generated development media now includes Roll Back Last SteamOS Update on the recovery desktop. The action lists only unmounted disks with one complete SteamOS A/B layout, requires explicit disk selection, offers only slots whose sole kernel has matching NVIDIA vermagic and GSP firmware, revalidates the partition identity, and requires an exact confirmation phrase. It then asks Valve’s disk-scoped boot tools to select that slot; it never invokes the broad repair/re-image workflow. This remains a development recovery action until its bootconf result and cancellation behavior pass sacrificial-disk and real- hardware testing.
Why a SteamOS update can invalidate NVIDIA
SteamOS stages operating-system updates into an inactive A/B root slot. NVIDIA contains a kernel-specific interface layer, so a module built for the old kernel must not be treated as compatible with the new slot. The matching userspace libraries and GSP firmware must agree with the module release too.
The legacy steamos-nvidia-installer demonstrates a useful transaction shape:
after Valve stages the inactive slot, it builds the driver for that slot and
marks the slot invalid if preparation fails. OPEMOS must not copy its fallback
that retries package installation with SigLevel = Never. HTTPS transport is
not a replacement for the reviewed signatures, hashes, locks, and provenance
required by the OPEMOS support contract.
OPEMOS update guardian
Generated installation media now embeds the exact pinned support guardian and the Open OPEMOS installer stages it into both target A/B slots after Valve’s installation completes. Its console-safe fallback, delayed-network repair, and machine-readable state are implemented. Physical A/B update, graphical-failure, power-loss, and rollback coverage remain certification gates.
The installed-system updater should use a persistent, machine-readable state transaction:
- Detect the exact staged SteamOS version, kernel, slot, partition identities, and selected NVIDIA source policy.
- Resolve or build only the exact compatible NVIDIA artifact and complete authenticated userspace closure.
- Apply it to the inactive slot and verify modules, vermagic, userspace, firmware, initramfs, boot arguments, and package database independently.
- Permit the boot-slot switch only after every verification succeeds.
- On failure or cancellation, mark the candidate slot invalid, keep the current slot selected, retain bounded diagnostics, and offer retry.
- On first boot, count attempts and automatically return to the last verified slot if the graphical health check does not complete.
The graphical updater should show Downloading update, Preparing NVIDIA for
the new kernel, Validating next boot, and Ready to restart. Because a
graphics update can terminate the compositor, the same status must also be
written to a persistent log and shown through a console-safe fallback. A
progress window alone is not a recovery mechanism.
Stable shell, updateable services
The installed graphical application should not need to be replaced for every resolver, recovery, or transaction fix. Its visible shell and its backend are separate compatibility domains:
- The unprivileged shell owns presentation, accessibility, and a bounded versioned request protocol. It remains open while services update.
- A separately signed backend generation owns network resolution, update state, validation, and calls into the fixed privileged support-helper protocol. A downloaded backend is never loaded into the GUI process.
- The current backend continues serving the shell while a candidate downloads and starts. The shell changes channels only after the candidate proves its signed identity, compatible protocol range, and bounded startup health.
- Failure before health acknowledgement restores the last-known-good backend. The existing signed A/B desktop-generation manager supplies the initial content-addressed staging, activation, health-deadline, and rollback foundation; it still needs a shell/backend protocol split before this can be described as seamless.
- Offline and delayed-network startup always uses the last-known-good backend. Internet access is an update opportunity, not a prerequisite for opening the UI, reading local status, or invoking recovery.
A backend manifest must bind the exact backend version, supported shell protocol interval, target OS and architecture, executable hash, support revision, guardian schema, release channel, and reviewed signer. If its protocol does not overlap the installed shell, the update is a normal full-application release—not a hot swap. Release discovery and TLS never substitute for manifest signature verification.
Activation must also wait for a safe transaction boundary. An update may not take ownership midway through package installation, initramfs creation, slot selection, device writing, or another destructive operation. Once healthy, new requests move to the candidate and the previous backend is retained for bounded rollback; the graphical window does not reload or disappear.
Boot interstitial status
The pinned support snapshot now includes a no-input DRM/KMS interstitial for
showing bounded guardian progress before SteamOS starts Gaming or Desktop Mode.
Generated media carries its launcher, progress writer, executable validator,
and systemd service contract, and the welcome installer verifies those files in
both installed A/B slots. The service is conditionally inactive unless an exact
authenticated x86_64 interstitial executable and its hash binding are present.
That executable is intentionally not taken from an unsigned CI artifact. The visual boot status becomes an end-user feature only after OPEMOS publishes it under the reviewed production desktop-update signer policy and physical SteamOS testing proves DRM release, watchdog behavior, display handoff, and continued boot after renderer failure. The guardian remains authoritative with or without the cosmetic display.
OPEMOS should additionally install an explicit recovery entry that can reach a text/rescue environment without starting Gaming Mode. The exact SteamOS boot entry and rollback edits remain hardware-test gates; they must not be inferred from generic systemd behavior or applied to an unrecognized Valve layout.
Recovery graphics tiers
A fallback should be prepared before an update, not downloaded after graphics have already failed:
- On hybrid systems, prefer the already installed Intel or AMD in-kernel driver and Mesa stack when that GPU can own a display.
- Keep a console-only entry that avoids the NVIDIA modules and Gaming Mode while retaining the firmware-provided framebuffer where the machine supports it. This is intended for status, diagnostics, and rollback—not accelerated gaming.
- Offer Nouveau only for GPU/kernel combinations that pass a hardware profile test. Its kernel component follows the installed kernel, but modern NVIDIA generations may still depend on GSP firmware and compatible Mesa userspace.
Nouveau and the NVIDIA open modules must never race to bind the same GPU. A Nouveau recovery entry therefore needs its own validated command line and initramfs policy that disables NVIDIA and does not inherit the normal boot’s Nouveau blacklist. Failure of that entry must still leave recovery from the known-good USB image available.
Wi-Fi diagnosis on non-Deck hardware
Start by identifying the controller and its current kernel binding. Run:
lspci -nnk | grep -A3 -iE 'network|wireless'
lsusb
rfkill list
nmcli radio
nmcli device status
nmcli -f GENERAL.DEVICE,GENERAL.TYPE,GENERAL.DRIVER,GENERAL.DRIVER-VERSION,GENERAL.FIRMWARE-VERSION,GENERAL.FIRMWARE-MISSING,GENERAL.STATE,GENERAL.REASON device show
journalctl -b -u NetworkManager --no-pager
dmesg | grep -iE 'firmware|wifi|wlan|iwlwifi|ath|rtw|brcm|mt76'
These results separate common causes:
| Evidence | Likely class of problem |
|---|---|
| No PCI/USB device | Firmware/UEFI setting, hardware, or bus enumeration |
| Device exists but no kernel driver | Unsupported or omitted kernel module |
FIRMWARE-MISSING: yes or firmware load errors |
Required firmware blob absent |
Hardware/software blocked in rfkill or nmcli radio |
Radio-kill state |
| Driver and firmware load, but activation fails | Credentials, security mode, regulatory domain, DHCP, or NetworkManager policy |
Do not include passwords or full connection profiles in a diagnostic report. The PCI/USB vendor and device IDs, driver name, firmware filename/version, NetworkManager state reason, and bounded kernel errors are sufficient for the hardware compatibility record.
Hardware-profile policy
Future image builds should accept a detected or user-selected target hardware profile, then verify before export that every required in-kernel module and firmware file exists for the image’s exact kernel. Profiles may add reviewed Wi-Fi firmware or modules, but must never fetch an arbitrary out-of-tree driver at first boot. Both A/B slots and the update guardian must preserve the same profile, and unsupported devices must remain clearly labeled rather than being silently treated as compatible.
References
- NVIDIA open kernel-module build and matching-component requirements
- Linux kernel Nouveau documentation
- Nouveau project status and GSP support
- Linux kernel firmware API
- NetworkManager device state and missing-firmware reporting
- NetworkManager command-line reference
- Legacy SteamOS NVIDIA installer’s update-repair implementation
- systemd rescue and emergency boot guidance