Product Overview

Predictive storage failure detection for SRE teams.

From raw drive telemetry to a risk-ranked list of drives to swap - without writing a single query.

Fleet Health Monitor - 5 drives
Drive ID Model Health Days to Risk Status
WD-7A2F WD Blue 4TB 97 - Healthy
ST-B91C Seagate IronWolf 81 18 Watch
NV-04E2 Samsung 980 Pro 74 7 At Risk
WD-C3A1 WD Red Plus 8TB 94 - Healthy
KX-55D7 Kioxia CD6-R 88 - Healthy

What Crest reads.

Crest's agent reads all available health telemetry from every drive in your fleet - including the signals that SMART alone does not surface.

  • SMART attribute deltas (all 255 attributes)
  • NVMe health log entries (SMART/Health Information Log)
  • Vendor-specific wear indicator counters
  • Reallocated sector delta tracking over time
  • Temperature trends and power cycle counts
  • IOPS and throughput degradation patterns
Drive SMART + NVMe Drive Vendor Telemetry Crest Agent TLS transport Risk Model Survival analysis Alert Poll every 5 min Per-model baseline

How the failure model works.

1
Baseline normalisation
Each drive model has its own wear curve. Crest normalises attribute values against a per-model baseline before analysis - a Samsung 980 Pro's reallocated sector count means something different from a Seagate IronWolf's.
2
Anomaly detection against trajectory
For each drive, Crest tracks how the attribute values move over time relative to the model's expected trajectory. A reallocated sector count climbing faster than baseline for that model is a signal. A stable count at a high absolute value may not be.
3
Survival analysis output
The model outputs a probability distribution over the next N days. You see a risk score, a predicted failure window, and a confidence level - not just a binary healthy/unhealthy flag. This is what makes the alert actionable: swap before day 18, not just "something is wrong".

Fits inside the runbooks you already have.

PagerDuty
Drive risk alert opens a PagerDuty incident with drive ID, host, predicted failure window, and recommended action - not a generic critical event.
See integration details
Prometheus
crest_drive_risk_score gauge metric per drive, ready to scrape with your existing Prometheus setup and Grafana dashboards.
See integration details
Slack
Direct channel message with at-risk drive summary on your schedule. Configure the channel and notification threshold in the Crest dashboard.
See integration details
See all integrations

One view for your whole fleet.

Fleet Dashboard - all hosts
Sort by risk Filter by host Export CSV
Drive ID Host Model Days Running Health Score Risk Level Predicted Window
NV-04E2 storage-02 Samsung 980 Pro 2TB 483 74 Critical 7-12 days
ST-B91C storage-01 Seagate IronWolf 6TB 721 81 Watch 14-21 days
WD-7A2F storage-01 WD Blue 4TB 312 97 Healthy N/A
WD-C3A1 storage-03 WD Red Plus 8TB 198 94 Healthy N/A
KX-55D7 storage-02 Kioxia CD6-R 3.84TB 567 88 Healthy N/A

Start with the drives you have.

The Crest agent reads existing SMART data on your current drives. No hardware changes, no additional sensors.