This is now published on the Imunify360 Status Page:
https://imunify360.statuspage.io/incidents/mplmcsfby1kq
We’ll continue to provide updates until a fixed version is released and updated to.
Many thanks for your continued patience as we work with Imunify360 engineers to resolve this critical bug.
Imunify developers have reproduced and confirmed the root cause. References: DEF-49872, DEF-50315
Official article now available: https://cloudlinux.zendesk.com/hc/en-us/articles/28982664003868-Imunify360-RustBolit-duplicate-scan-processes-cause-high-CPU-RAM-and-I-O-usage
A workaround has been applied to impacted machines.
The missing .rapid-scan-db directory is not the cause.
“Since imunify360-firewall 8.13.6-1, the agent’s systemd services run in a hardened mount namespace (ProtectSystem=true) with no supplementary groups. On hosts that hide /proc (hidepid), the agent can no longer see its own scanner processes, so its liveness check wrongly concludes the scan is gone and re-launches it on every poll - producing the pile-up of concurrent RustBolt processes you observed.
A permanent fix on the agent side (making liveness checks immune to proc-hiding, plus adopt-or-kill of any existing scan before spawning) is in progress. We will follow up on this ticket with the specific package version once the fixed release is scheduled. Thank you for your continued patience.”
G’day,
Following our earlier note about the Imunify360 bug, we now have clarity and movement. This is important as the bug has already taken down 2x client servers requiring hypervisor-level intervention to bring back up.
CL’s Imunify team have confirmed the root cause of the I/O runaway bug in the Imunify360 software.
“The root cause is an Imunify360 scanner process entering an unrecoverable loop targeting an account. The scanner is being relaunched every 10~ minutes because an internal state file it relies on to track scan completion is missing from disk — without it, the scan never registers as finished, and Imunify360 keeps re-spawning it. On a large account, each individual scan run is also far too large to complete before the next one is triggered, compounding the problem.
The missing scan state directory is /home/.rapid-scan-db/ on your servers. Without it, RustBolt cannot record scan completion, so the scheduler continuously re-launches the job with no concurrency limit.
This has been escalated to our engineering team as DEF-50315, which is approved for development.
We are also seeking guidance on a safe interim mitigation to relieve the server load while the fix is in progress, and will follow up as soon as we have that confirmed.”
We’ll continue to work with Imunify’s team towards patch and resolution. Our prompt referral has led to early discovery of this critical bug in their software, which had managed to get through their delayed rollout process.
Cheers,
Merlot Digital
Network: AS138521