Most organizations discover an application-layer DDoS attack when the helpdesk starts receiving user complaints. By that point, the attack has been running for minutes, the origin is already degraded, and the team is making decisions under pressure without baseline data. The metrics that would have detected the attack early were never wired to an alert.
This post covers the specific signals that distinguish an L7 attack from a legitimate traffic spike, where each metric lives in a typical stack, and what threshold to alert on. The goal is to instrument before the attack, not during it.
Request Rate per Endpoint
Total request rate is a poor L7 DDoS signal because it includes cached traffic and varies with marketing campaigns. Request rate per specific endpoint is the signal that matters. An endpoint that normally receives 200 requests per minute jumping to 8,000 requests per minute is a meaningful anomaly; a 40x increase in total traffic might just be a product launch. Instrument per-endpoint request rates in your WAF or API gateway, not in your application layer, so the metric is available before the origin is saturated.
Alert threshold: 5x the 7-day P95 per-endpoint rate, sustained for 60 seconds. Use a short window to catch fast-onset attacks rather than a rolling average that smooths the signal.
Origin Response Time Percentiles
P50 latency degrades last. Instrument P95 and P99 latency per endpoint at the edge, not at the application server. When P99 crosses 3x the baseline, the origin is approaching saturation. When P50 crosses 2x baseline, saturation is complete and users are experiencing failures. These two thresholds give you a warning signal and a critical signal in sequence.
Measure at the edge, not the origin
Application-level latency metrics (from APM tools) miss the queue time in front of your workers. Edge latency metrics capture the full user experience including connection wait time, which rises first during L7 attacks.
WAF Block Rate
A sudden increase in WAF blocks -- even if the blocks are working correctly -- is an early attack signal. If your WAF blocks 500 requests per hour normally and jumps to 50,000 blocks per hour, you are under attack regardless of whether users are impacted yet. Wire an alert on the WAF block rate time series. Many teams disable this alert because it fires during rule tuning, but it is one of the fastest attack indicators available.
Database Connection Pool Saturation
For L7 attacks targeting compute-heavy endpoints, database connection pool utilization is the resource that saturates first. Most application stacks expose this metric via their connection pool library. Alert when the active connection count exceeds 80% of the pool maximum, sustained for more than 30 seconds. At that point, new requests are queuing waiting for a connection, and latency spikes are imminent even if individual query times are still normal.
TLS Handshake Rate
A TLS connection flood (distinct from an HTTP flood) shows up as a spike in TLS handshakes that do not complete to HTTP requests. The ratio of completed TLS sessions to issued HTTP requests drops significantly. This metric is available at the load balancer or edge layer. A TLS handshake rate that is 10x the completed request rate indicates a connection-exhaustion attack that bypasses HTTP-level rate limits entirely.
Wiring the Alerts
For Cloudflare, the metrics above are available in Cloudflare Logs (R2 export) or Cloudflare Analytics via API. For AWS CloudFront with WAF, use CloudWatch metrics: RequestCount, BlockedRequests, OriginLatency. For Nginx, expose via ngx_http_stub_status_module and scrape with Prometheus. The key is that every metric in this list has a data source that is available before an attack begins -- the work is configuration and threshold selection, not new infrastructure.
Verify Your Observability Gaps
DDactic's assessment identifies which L7 metrics are missing from your monitoring stack and which endpoints lack the alerting needed for early detection.
Run a Free Scan