One of the most persistent misconceptions in API security is that requiring authentication on an endpoint prevents DDoS. It does not. An attacker with a single valid API key, or a botnet with thousands of stolen credentials, can exhaust your origin just as effectively as an unauthenticated flood -- and often more so, because authenticated requests bypass bot-score filters at the edge.
Authentication answers the question "who is this?" Rate limiting answers the question "how much is too much?" These are separate questions with separate answers. Conflating them is a configuration gap that shows up repeatedly in DDactic assessments: endpoints that require a Bearer token but have no per-token or per-IP ceiling on request frequency.
Why Authenticated Endpoints Are Vulnerable
When a request arrives at your API with a valid JWT, the Cloudflare bot score is typically high (the request looks legitimate), the WAF signature rules do not fire (the payload is valid JSON), and the request proceeds directly to your origin. The authentication check your application performs costs time and compute before the rate limit would have been checked -- if one even exists.
Credential stuffing attacks exploit this directly. They send authentication requests with breached username/password pairs at rates of 10,000 per hour per IP cluster. Even if 99% fail, the 1% that succeed provide valid tokens that can then be used to hammer expensive authenticated endpoints. The authentication layer not only fails to stop the attack; it provides the attacker with attack tokens.
The authenticated API key amplification problem
API keys issued to developers or partners often have no rate limit configured. A compromised key, or a developer running a test script incorrectly, can saturate an endpoint. Key-level rate limits are distinct from IP-level limits and must be configured separately.
What Rate Limiting Provides That Auth Does Not
A correctly configured rate limit enforces a ceiling on request frequency independent of identity validity. It fires before the authentication check in properly designed gateway architectures, which means invalid and valid credentials both count against the limit. It measures behavior, not identity. An authenticated user sending 1,000 requests per minute is just as dangerous to availability as an unauthenticated bot sending the same rate -- authentication tells you which user it is, not whether the rate is acceptable.
The correct model is to apply rate limits at three levels simultaneously: per source IP (catches unauthenticated floods and distributed botnets), per API key or user ID (catches authenticated abuse and compromised credential attacks), and per endpoint globally (catches coordinated attacks that stay below per-IP thresholds by spreading across many sources).
Where to Apply Each Control
Authentication validation belongs at the origin or at an API gateway that terminates the session. Rate limiting belongs at the edge -- as close to the attacker as possible, before compute is spent. Applying rate limiting at the application layer (in Django middleware, Spring filters, or Express middleware) means the request has already consumed a connection, a worker thread, and TLS handshake overhead before being rejected. Edge-level rate limiting eliminates that overhead for blocked requests.
For endpoints that must be public before authentication (login, password reset, OAuth callback), authentication-layer protection is impossible by definition. These endpoints are fully exposed to unauthenticated traffic and require the most aggressive rate limits. A login endpoint that allows 60 attempts per minute per IP is effectively unprotected against distributed credential stuffing using a /24 subnet.
"We had OAuth on every endpoint and assumed that was sufficient. The assessment showed our search API had no rate limit at all. A single valid token could have taken us down."
Implementing Both Correctly
The order of operations in a well-configured API gateway: TLS termination, then connection rate limit check, then IP-level request rate limit check, then authentication validation, then per-user rate limit check, then routing to origin. Each check earlier in the chain prevents compute from being spent on the next. Putting auth before rate limits inverts this and ensures your most expensive check runs before your cheapest one.
Check Whether Your APIs Have Both Controls
DDactic tests for the presence and effectiveness of rate limits independently from authentication requirements, and reports endpoints that have one without the other.
Run a Free Scan