2023 · IEEE International Conference on Big Data

AltOOM: A Data-driven Out of Memory Root Cause Identification Strategy

Pranjal Chakraborty | Naser Ezzati-Jivan | Vahid Azhari | François Tetreault

Evidence basis: metadata-only · Review status: catalog-reviewed; paper-author approval pending

resource-analysis root-cause-analysis system-tracing predictive-monitoring

out-of-memory OOM diagnosis data-driven RCA resource analysis memory pressure forecasting process-level profiling Perf sar

Core contribution: AltOOM combines early memory-pressure forecasting with selective process-level profiling to identify the process most responsible for an impending out-of-memory event.

Problem and motivation

Linux's reactive OOM killer can terminate a high-memory process even when a lower-memory process is the one whose memory usage is growing; resource-constrained systems need earlier warning and lower-overhead process attribution.

Method and contribution

AltOOM samples 34 system-level signals at 0.5-second intervals - 28 virtual-memory statistics, three memory-related system calls (brk, sbrk, and mmap), and three kernel events (kmalloc, mm_page_alloc, and vmscan) - using perf and sar. It labels pressure at %memused >= 85%, compares SVM, vanilla DNN, and bidirectional-LSTM predictors, filters 34 features to 15, then uses burst-collected process-level allocation signals and moving-average growth ranking after an alert.

Findings and evidence

In the reported evaluation, the feature-filtered DNN reaches 0.82 accuracy for (n,k)=(4,3) and 0.81 for (3,3); the abstract reports 85% memory-pressure forecasting accuracy. AltOOM process identification reaches 0.56-0.83 as the burst count increases from 3 to 7, versus 0.42 for the Linux OOM killer. A Firefox PDF-preview case reports 0.83 and 0.79 forecasting accuracy for (4,3) and (3,3), respectively.

Limitations and future directions

Limitations: The method can miss gradual memory buildup when only three timestamps (1.5 seconds) are observed. Fixed-rate monitoring creates overhead, and the evaluation centers on generated pressure scenarios plus a Firefox case rather than broad production-device coverage.

Future work: Evaluate adaptive sampling of rates, metrics, and metric groups, and implement actions such as controlling or adjusting the responsible process, terminating it, or restarting the system.

Sources and identifiers

When to cite this paper

Cite this paper when studying proactive memory-pressure forecasting or process-level out-of-memory root-cause identification.

Citation

BibTeX
@inproceedings{ezzatiJivan2023altooma,
  author = {Pranjal Chakraborty and Naser Ezzati-Jivan and Vahid Azhari and François Tetreault},
  title = {AltOOM: A Data-driven Out of Memory Root Cause Identification Strategy},
  year = {2023},
  booktitle = {IEEE International Conference on Big Data},
  pages = {1637-1646},
  publisher = {IEEE},
  doi = {10.1109/bigdata59044.2023.10386937},
  url = {https://doi.org/10.1109/bigdata59044.2023.10386937}
}
Other citation formats for Word and reference managers
APA 7
Chakraborty, P., Ezzati-Jivan, N., Azhari, V., & Tetreault, F. (2023). AltOOM: A Data-driven Out of Memory Root Cause Identification Strategy. In IEEE International Conference on Big Data (pp. 1637-1646). https://doi.org/10.1109/bigdata59044.2023.10386937
IEEE
P. Chakraborty, N. Ezzati-Jivan, V. Azhari, and F. Tetreault, "AltOOM: A Data-driven Out of Memory Root Cause Identification Strategy," in IEEE International Conference on Big Data, pp. 1637-1646, 2023, doi: 10.1109/bigdata59044.2023.10386937

Readable Markdown record · JSON record · Download RIS