Server memory errors and diagnostics
Memory faults show up as three things: correctable ECC errors logged by the BMC, uncorrectable errors that crash the OS with a machine check, and DIMMs that disappear from the inventory after a reboot. Each has a different response.
Modern servers hide a lot of this from you. Dell's ADDDC, Lenovo's memory sparing and page retirement in Linux keep a marginal DIMM in service. That is good for uptime but it also means the alert is the only early warning you will get.
At a glance
| Dell | iDRAC Lifecycle Log MEM0701/MEM0702 (correctable), MEM0001 (uncorrectable), amber LED and DIMM slot LED |
|---|---|
| Lenovo | XCC event log, light path diagnostics, FRU callouts |
| Supermicro | IPMI SEL, sensor readings, optional MemTest86 boot |
| OS side | Windows WHEA-Logger event 19/47, Linux mcelog / rasdaemon, ESXi vmkernel MCE |
| Fix | Reseat once, then replace; never swap slots to hide the error |
What RackLedge does
- Reading BMC logs and mapping errors to a physical slot
- Warranty and out-of-warranty DIMM replacement with matched parts
- Post-replacement burn-in and error-rate monitoring
- Root cause when it is not the DIMM: CPU memory controller, board, or a dusty heatsink
How we work
A memory engagement starts with evidence, not a parts list. We pull utilization from your monitoring or hypervisor, confirm memory is the real constraint, and only then choose modules that keep every channel populated at full speed.
Firmware comes first. BIOS and BMC updates fix memory training bugs and unlock newer module types, so we baseline them before opening the chassis. Installation is scheduled in a maintenance window and validated with a memory test and an application check before the server goes back into rotation.
Parts are new or tested refurbished with a warranty, matched to the vendor's qualified list or an equivalent specification, and we keep a record of exactly what went where so the next upgrade starts from documentation rather than guesswork.
Related services
More on server memory
Frequently asked questions
The error moved to a different slot after I reseated the DIMM. What now?
If the error follows the DIMM, replace the DIMM. If it stays with the slot, the problem is the board or the CPU's memory controller and the server needs a service call.
Can memory errors be caused by heat?
Yes. A failed fan or a blocked rear exhaust raises DIMM temperature and correctable-error rates. Check thermal logs before buying memory.
Need a hand with this?
Tell us what you are running and what is slowing you down. You get a straight assessment and a plan, with no obligation. Support desk is staffed 24/7.