Skip to main content
Subscribe
Backend Infrastructure

API Rate Limiting in 2026: A Step-by-Step Redis and Lua Guide

I once watched a clean PostgreSQL node choke to death on 45,000 unthrottled requests in six seconds.

It wasn’t an organized nation-state cyberattack. It was just an enthusiastic junior developer on a partner team who wrote an infinite loop inside a client-side polling function. They forgot to add a sleep block. By the time our monitoring alert woke me up, our database pool was entirely depleted, and the application was throwing 500 errors across three continents.

If you expose an HTTP endpoint to the internet without a strict rate-limiting barrier, you haven’t built an API. You’ve built a denial-of-service vector.

As a backend engineer, you cannot rely on the good behavior of clients. Whether it’s a buggy frontend script, a scrap-happy AI bot trying to train its next model, or a targeted application-layer attack, your backend needs defensive armor.

But how do you build a system that shields your databases without introducing debilitating latency for legitimate human users?

Here’s the architecture I use for production-grade traffic control in 2026.

TL;DR: The Core Reality

Vulnerable APIs and automated bot attacks cost businesses up to $186 billion a year globally (Imperva, The Economic Impact of API and Bot Attacks, 2024). Defensive engineering demands that rate limiting occur as close to the edge as possible, using a distributed store like Redis to handle dynamic traffic shaping without bottlenecking your database layer.


Why Is API Rate Limiting Critical in 2026?

API traffic accounts for 57% of all dynamic internet traffic (Cloudflare 2024 API Security and Management Report, 2024). That share means your server endpoints are the primary target for automated internet scanners, malicious scraping tools, and aggressive API integrations.

+------------------------------------------------------------+
| Dynamic Internet Traffic Composition (2024)                |
+------------------------------------------------------------+
| [██████████████████████████████████████] 57% API Traffic    |
| [████████████████████████████] 43% Standard Web Page Traffic |
+------------------------------------------------------------+

When you leave an endpoint unprotected, you aren’t just risking a minor performance hit. You’re giving third-party clients direct control over your compute costs and database connection pools.

Consider the modern bot ecosystem. Imperva’s 2026 Bad Bot Report says bots now account for over 53% of all web traffic, “with 40% classified as malicious,” and that AI-driven bot attacks surged 12.5x (Imperva 2026 Bad Bot Report, 2026). For a year-over-year baseline, the 2025 edition had bad bots at 37% of all internet traffic (Imperva 2025 Bad Bot Report, 2025). These bots scrape pricing information, test stolen credentials against authentication endpoints, and exhaust computational power.

But there’s a more insidious threat: Layer 7 application DDoS attacks. HTTP DDoS attacks grew 129% year-over-year in Q2 2025 (Cloudflare DDoS Threat Report for 2025 Q2, 2025), and by the midpoint of 2026 Cloudflare had already mitigated 29.64 trillion HTTP DDoS requests for the year (Cloudflare DDoS Threat Report H1 2026, 2026). These attacks don’t try to saturate your network bandwidth; instead, they target resource-intensive API operations, like PDF generation or complex SQL queries, and knock your systems offline with a surprisingly small volume of requests.

API Security Citation Capsule
“According to Cloudflare’s 2024 API Security and Management Report, API traffic now makes up 57% of all dynamic internet traffic. This rapid expansion makes centralized backend rate limiting a baseline requirement for preventing resource exhaustion and mitigating automated application-layer attacks.”

If you are currently exposing custom GPTs, MCP setups, or agent workflows without rate limiting, you are begging for a massive cloud invoice. If you are building AI-driven backends, check out implementing Model Context Protocol backends securely to understand how this risk profile changes in agentic environments.

Donut chart showing dynamic internet traffic breakdown

Source: Cloudflare, 2024


Which Rate Limiting Algorithm Fits Your System?

78% of API attack attempts use one or more OWASP API Top 10 methods (Salt Security State of AI and API Security Report, 2026), and rate limiting is the direct control for one of them: API4, “Unrestricted Resource Consumption.” Selecting the wrong rate-limiting algorithm can lead to false positives that lock out your best customers, or false negatives that allow spikes to bypass your defenses.

Production backends use four main algorithms.

Rate Limiting Algorithms:
├── Fixed Window (Simple, but vulnerable to boundary bursting)
├── Sliding Window Log (Precise, but high memory footprint)
├── Token Bucket (Supports bursts, highly space-efficient)
└── Leaky Bucket (Constant output flow, smooths traffic spikes)

1. Token Bucket

In a Token Bucket design, your system stores a maximum number of tokens for each client. Every incoming request consumes a token. If the bucket is empty, the request is rejected. Tokens replenish at a constant, defined rate over time.

The primary advantage? It handles bursty traffic smoothly. If a developer loads an app and fires off ten parallel API calls to populate a dashboard, the Token Bucket lets them through. But once those tokens are spent, the client is forced to fall back to the refill rate.

2. Leaky Bucket

The Leaky Bucket algorithm processes incoming requests through a FIFO (First-In, First-Out) queue. Requests enter the bucket at arbitrary rates and leak out at a constant, steady speed. If the incoming request rate exceeds the queue’s capacity, the excess requests overflow and are dropped immediately.

This algorithm is exceptional for smooth traffic shaping. However, it can introduce latency for legitimate user actions because every request must wait its turn to leak out of the queue.

3. Fixed Window Counter

This is the simplest approach. The system divides time into static windows (e.g., 1-minute blocks). Each user has a counter for that specific window.

But here’s the flaw: what happens if a client sends their entire limit at the very end of Window A, and another batch at the very beginning of Window B? They successfully double their rate limit right at the boundary. Your backend just absorbed twice the allowed load within a sub-second interval.

4. Sliding Window Counter (and its cousin, the Sliding Window Log)

This hybrid approach tracks the current window counter and the previous window counter. Using the exact timestamp of the request, the system computes a weighted average of the request volume. It delivers most of the accuracy of a sliding log without the memory overhead of saving every single timestamp to Redis. The sliding log, which the implementation below uses, stores one sorted-set entry per request: exact, and fine at 100 requests a minute per client, but costly if your limits run into the tens of thousands.

Current Window Estimation:
Request Count = (Previous Window Count * (Remaining Ratio of Prev Window)) + Current Window Count

If you are using tools like AI agents or sub-agent orchestration platforms, your traffic patterns will be incredibly bursty. If you run complex workflows like managing sub-agent processes, a Token Bucket approach is highly recommended to accommodate parent-child task bursts without throwing unexpected errors.

Algorithm Selection Citation Capsule
“According to Salt Security’s H1 2026 State of AI and API Security Report, 78% of API attack attempts use one or more OWASP API Top 10 methods. Rate limiting is the primary control for API4, Unrestricted Resource Consumption, and a Token Bucket or Sliding Window algorithm prevents resource exhaustion while reducing false-positive blocks for legitimate users.”


How Do You Design a Distributed Rate Limiting Architecture?

Nearly all (99%) of the attack attempts Salt Labs analyzed came from authenticated sources (Salt Security State of AI and API Security Report, 2026). That’s the problem with edge-only rules: a WAF that throttles anonymous IPs doesn’t help when the abuse arrives with a valid token. You need limits keyed on identity, with state shared across every instance. To build an enterprise-ready system, you must deploy your rate limiter as a high-speed, state-sharing layer using Redis.

If you have a stateless backend replicated across five load-balanced containers, you cannot track rate limits in local memory. If Client A hits Container 1, and then hits Container 2, those instances don’t know about each other’s state. You need a centralized, ultra-low latency data store.

An abstract, fluid 3D render representing smooth data flow control and rate-limiting bucket algorithms.

But aren’t we just introducing another bottleneck? If every API request requires a network round-trip to Redis before doing any work, we risk bloating our response times.

To mitigate this, we write our rate-limiting logic in Lua scripts and execute them directly inside Redis. This guarantees three critical advantages:

  1. Atomicity: Redis runs Lua scripts sequentially on a single thread. There are no race conditions between reading a key and writing its new value.
  2. Reduced Latency: Instead of executing three separate commands (GET, INCR, EXPIRE) across three network roundtrips, the backend performs one single roundtrip to execute the Lua script.
  3. Isolation: No other client sees the key in a half-updated state while the script runs. Note that Redis doesn’t roll back writes if a script errors partway through, so keep scripts short and validate arguments before the first write.

Why bother with identity-aware middleware when a gateway exists? Look at what organizations actually found in their production APIs over the past year:

Horizontal bar chart of security problems organizations found in production APIs: sensitive data exposure 44%, vulnerabilities 43%, authentication problems 41%, credential stuffing or brute force attacks 27%

Source: Salt Security, H1 2026

Credential stuffing and brute force showed up at 27% of organizations, and authentication problems at 41% (Salt Security, 2026). Those are attacks a per-account limit in your own middleware catches and a per-IP gateway rule often doesn’t. Standard gateways still handle basic perimeter protection; application-level limits give you user identity awareness. For databases linked directly to your AI stacks, read connecting databases to AI workflows safely to understand how custom rate-limiting middleware serves as your last line of defense.

Distributed Architecture Citation Capsule
“Salt Security’s H1 2026 report found that 99% of the API attack attempts it analyzed originated from authenticated sources. That makes identity-keyed, state-sharing rate limits in Redis, coupled with context-aware middleware, necessary to prevent API abuse that IP-based edge rules never see.”


What Is the Step-by-Step Backend Implementation Flow?

32% of organizations experienced an API security incident in the past year, and 47% delayed an application deployment over API security concerns (Salt Security State of AI and API Security Report, 2026). To keep your infrastructure out of that first number, implement a Redis-backed Sliding Window Log.

Here’s the exact request lifecycle.

Client Request ──> API Gateway/Middleware ──> Execute Lua Script in Redis
                                                     │
                                           ┌─────────┴─────────┐
                                      [Limit Exceeded]   [Within Limit]
                                           │                   │
                                   Reject (HTTP 429)     Allow & Pass to Backend

Step 1: Write the Atomic Lua Script

This script tracks hits in a sliding window. We pass two arguments: the client’s rate-limiting key and the maximum threshold within the window size.

-- KEYS[1]: Rate limit key (e.g., "rate_limit:user_12345")
-- ARGV[1]: Current UNIX timestamp in seconds
-- ARGV[2]: Window size in seconds (e.g., 60)
-- ARGV[3]: Max requests allowed in that window
-- ARGV[4]: Unique request ID (so two requests in the same second both count)

local key = KEYS[1]
local now = tonumber(ARGV[1])
local window = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])
local request_id = ARGV[4]

local clear_before = now - window

-- Remove old logs outside of our sliding window
redis.call('ZREMRANGEBYSCORE', key, 0, clear_before)

-- Count the remaining active timestamps in this window
local current_requests = redis.call('ZCARD', key)

if current_requests < limit then
    -- Score is the timestamp; member must be unique per request
    redis.call('ZADD', key, now, request_id)
    -- Extend key TTL so it cleans up automatically if inactive
    redis.call('EXPIRE', key, window)
    return {1, limit - current_requests - 1} -- Allowed (True), Remaining Tokens
else
    return {0, 0} -- Denied (False), 0 Tokens Remaining
end

Step 2: Integrate Node.js Backend Middleware

Now, write the middleware that intercepts requests, communicates with Redis via our Lua script, and appends the standard HTTP headers.

import Redis from 'ioredis';

const redis = new Redis(process.env.REDIS_URL);

// Load our Lua script into Redis memory for high-performance execution
const rateLimitLua = `
local key = KEYS[1]
local now = tonumber(ARGV[1])
local window = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])
local request_id = ARGV[4]
local clear_before = now - window
redis.call('ZREMRANGEBYSCORE', key, 0, clear_before)
local current_requests = redis.call('ZCARD', key)
if current_requests < limit then
    redis.call('ZADD', key, now, request_id)
    redis.call('EXPIRE', key, window)
    return {1, limit - current_requests - 1}
else
    return {0, 0}
end
`;

// Register the custom command with ioredis
redis.defineCommand('evaluateRateLimit', {
  numberOfKeys: 1,
  lua: rateLimitLua,
});

export async function rateLimiterMiddleware(req, res, next) {
  // Use unique identifiers: API key first, fall back to IP
  const clientId = req.headers['x-api-key'] || req.ip;
  const key = `rate_limit:${clientId}`;
  
  const now = Math.floor(Date.now() / 1000);
  const windowSizeSeconds = 60;
  const maxRequests = 100;

  try {
    // Run atomic check in Redis
    const [allowed, remaining] = await redis.evaluateRateLimit(
      key,
      now,
      windowSizeSeconds,
      maxRequests,
      `${Date.now()}-${Math.random().toString(36).slice(2)}`
    );

    // Set standard dynamic API headers
    res.setHeader('X-RateLimit-Limit', maxRequests);
    res.setHeader('X-RateLimit-Remaining', remaining);
    res.setHeader('X-RateLimit-Reset', now + windowSizeSeconds);

    if (allowed === 0) {
      res.setHeader('Retry-After', windowSizeSeconds);
      return res.status(429).json({
        error: 'Too Many Requests',
        message: 'Rate limit exceeded. Please back off and retry later.',
        resetAt: now + windowSizeSeconds
      });
    }

    next();
  } catch (err) {
    // Fail-open strategy to protect client experience if Redis is down
    console.error('Rate limiting system failure:', err);
    next();
  }
}

Do not let a redis crash destroy your core user experience. Designing a fallback mechanism to fail-open ensures that if Redis goes down, your users can still complete transactions while your engineering team repairs the state store.

An illuminated, futuristic architectural gateway representing an API gateway enforcing traffic controls.

If you are running complex automation scripts, custom CLI tools, or background processes, you will want a robust integration pattern. For example, if you are setting up local developer tool integrations, explore setting up local CLI integration patterns to ensure limits match developer workflows.

Implementation Security Citation Capsule
“According to Salt Security’s H1 2026 State of AI and API Security Report, 32% of organizations had an API security incident in the past year and 27% found credential stuffing or brute force attacks against production APIs. Server-side atomic evaluation with sliding-window sets is the standard defense against rapid-fire credential stuffing.”


How Do You Prevent Sophisticated Bot-Driven Bypass Techniques?

Bots account for over 53% of all web traffic in 2026, and Imperva says 27% of attacks now focus on APIs and identity systems (Imperva 2026 Bad Bot Report, 2026). Attackers don’t just write a single curl script and run it forever. They rototill their target vectors. They rotate IP addresses, exploit public proxies, mask user agents, and systematically bypass simple, single-dimensional rate limiters.

                  ┌───────────────┐
                  │ Client Request│
                  └───────┬───────┘
                          │
          ┌───────────────┴───────────────┐
          ▼                               ▼
   [Validate Headers]             [Analyze Fingerprint]
   - Is API Key valid?            - Match TLS Fingerprint (JA4)
   - Is IP in proxy pool?         - Check HTTP Header ordering
          │                               │
          └───────────────┬───────────────┘
                          ▼
             ┌─────────────────────────┐
             │ Multi-Dimensional Score │
             └─────────────────────────┘

If you only use client IP addresses to determine limits, you are leaving your system wide open. Attackers can spin up residential proxies for pennies, giving them access to thousands of dynamic IPs.

How do we defend against this? We enforce multi-dimensional checks:

1. TLS Fingerprinting (JA4)

Every client handshake sends a list of supported SSL/TLS extensions and cipher suites in a specific order. Traditional browsers have very specific ciphers. Python scripts, Go binaries, and Node runtimes have entirely different signatures.

Using JA4 TLS fingerprinting, your gateway can identify if an incoming request claiming to be a current Chrome build is actually a generic python-requests script. If the fingerprint doesn’t match the claimed User-Agent, you drop their rate limit threshold by 95%.

2. Multi-Tiered Identification Keys

Do not evaluate rate limits on a single attribute. Instead, construct your keys hierarchically:

Tier 1 Check: API Key (Strict Limit)
Tier 2 Check: JWT Claim / Account ID (Group Limit)
Tier 3 Check: IP address + Route path (Default Limit)

If a client makes a series of unauthenticated requests, throttle them strictly by IP. Once authenticated, transition them immediately to account-based limits. This isolates resource exhaustion issues to single malicious accounts rather than whole networks behind shared NAT setups.

If you are developing complex AI-driven application backends in common frameworks, we can apply these multi-tiered identifiers directly in your route middleware. For PHP and modern framework designs, read designing robust agent middleware patterns for specific patterns.

Bot-Prevention Citation Capsule
“Imperva’s 2026 Bad Bot Report found that bots now account for over 53% of all web traffic, with 40% classified as malicious, and that AI-driven bot attacks surged 12.5x. To mitigate these dynamic threats, system architects must abandon IP-only tracking in favor of cryptographic signatures, multi-tiered accounts, and TLS JA4 finger-print validations.”


How Do You Monitor Rate Limiting Performance Without Killing Your Database?

In Postman’s 2025 State of the API Report, 51% of developers said they worry about “unauthorized or excessive API calls from AI agents,” making it their number one security concern (Postman 2025 State of the API Report, 2025). You can’t manage excessive calls you can’t see. Often, the metrics systems designed to monitor performance end up introducing more performance degradation than the actual traffic spikes.

Logging Flow:
API Gateway Middleware ──> Process Request ──> Push Event to Memory Ring Buffer
                                                               │
                                                       (Asynchronous Batch)
                                                               ▼
                                                     Metrics DB (Prometheus)

If your rate limiter writes a new log row to PostgreSQL every time a request is dropped, your database will collapse during an active DDoS attack.

To monitor rate limiting correctly, you must rely on decoupled, asynchronous metrics collection.

  1. Use Prometheus Counters: Maintain high-performance counters directly in your application process memory or fetch aggregate metrics from Redis keys.
  2. Batch Logs via Vector/Fluentbit: Output structured JSON to stdout from your backend servers. Let a background sidecar process like Vector collect, bundle, and push those logs to aggregate systems like Grafana Loki or Datadog.
  3. Audit Shadow APIs: F5’s 2024 API security research found the average organization manages 421 different APIs (SecurityMEA, reporting F5’s findings, 2024). At that scale, some routes always escape the inventory. Monitor your gateway routing statistics to identify unregistered endpoints that are bypassing rate-limiting middleware entirely.
Shadow API Threat:
[Dynamic Client Traffic] 
   ├──> Protected Routes (/api/v1/checkout)  ──> (Enforces Rate Limiting)
   └──> Unlisted Routes (/temp/debug-export) ──> (Bypasses Limiters - Shadow API)

To systematically eliminate shadow routes, ensure your API Gateway registers and checks all incoming paths against a central routing manifest. Any unregistered route should automatically fall back to the most restrictive rate-limiting tier.

Grouped bar chart: web app and API attacks up 49% (Akamai, 2024) and Layer 7 DDoS up 129% (Cloudflare, Q2 2025)

Source: Akamai, 2024 & Cloudflare, 2025

Akamai’s 2024 figure covers web attacks against both applications and APIs, up 49% in a year, with 108 billion API attacks recorded between January 2023 and June 2024 (Akamai, 2024). Both numbers are a year or more old, but the 2026 volumes above show the direction hasn’t changed. Monitoring these attack profiles is critical. For instance, if you run distributed agents or external developer CLI tools, read analyzing agent developer environments to learn how different tool suites communicate and impact underlying system APIs.

Telemetry Performance Citation Capsule
“With F5’s 2024 research putting the average organization at 421 APIs, some routes inevitably escape the inventory and run without rate limiting. Light, decoupled, in-memory metrics plus a central routing manifest are how you find and fix unsecured endpoints without adding load to your database.”


API Rate Limiting Frequently Asked Questions

A high-tech monitoring dashboard showing real-time graphs and analytics for API performance and rate limit triggers.

Why does basic IP-based rate limiting fail in multi-tenant environments?

Basic IP rate limiting fails because large enterprise customers route all corporate user traffic through single, shared outbound NAT proxies or cloud gateways.

If you throttle based on raw IP addresses, you run a high risk of locking out an entire corporate customer office because one of their internal testing developers triggered a low rate-limiting rule.

Additionally, bots account for over 53% of all web traffic in 2026 (Imperva 2026 Bad Bot Report, 2026), and the malicious ones use distributed proxy configurations to easily rotate past IP blocks. This is why using application API keys or JWT tokens is critical for accurate client identification.

What HTTP status code and headers should rate limiters return?

Your backend rate-limiting middleware must respond with HTTP 429 (Too Many Requests). Avoid returning general errors like 400 or 403, which can confuse clients.

Always include these standardized headers to help developer clients back off dynamically:

Proper implementation of these headers helps avoid security incidents and unnecessary downtime. That matters at scale: vulnerable APIs and bot attacks cost businesses up to $186 billion a year globally (Imperva, 2024).

How do you protect Redis from becoming a single point of failure?

If your Redis cluster crashes, your backend must not fail. When configuring your middleware, use a fail-open policy wrapped in a try/catch block.

If the Redis connection times out (set a strict connection timeout limit of 15-20 milliseconds), log the outage to a tracking service and fall back to allowing the API request through.

This design pattern is essential for high availability, especially as HTTP application-layer DDoS attacks (which directly exploit resources) grew 129% year-over-year in Q2 2025 (Cloudflare DDoS Threat Report, 2025). Protecting core availability takes priority over strict throttling when internal systems fail.

What is OWASP API4 and why should backend developers care?

OWASP API4 represents “Unrestricted Resource Consumption,” which covers systems that lack protection against infinite loops, oversized file uploads, or missing rate-limiting configurations.

It matters because attackers lean on the OWASP list: 78% of API attack attempts use one or more OWASP API Top 10 methods (Salt Security State of AI and API Security Report, 2026).

Implementing sliding-window filters at your application layer directly addresses this vulnerability, shielding downstream servers and database pools from resource starvation.

How do you handle client SDK retries during rate limit spikes?

When your API throws an HTTP 429, client applications should not immediately retry their requests. Instead, design your client-side SDKs to use exponential backoff with jitter.

Exponential backoff multiplies the wait time with each consecutive failure, while jitter adds random noise to stagger client retries.

Without jitter, thousands of blocked clients will retry at the exact same millisecond interval, creating a secondary “retry storm” that mimics an active Layer 7 DDoS attack.

With 32% of organizations reporting an API security incident in the past year (Salt Security State of AI and API Security Report, 2026), implementing proper client retry mechanics is critical for overall platform stability.

How do you identify and rate-limit shadow APIs?

Shadow APIs are unmapped, unmonitored routes that bypass security controls. They’re hard to rule out when the average organization runs 421 APIs (SecurityMEA, reporting F5’s findings, 2024).

To secure them, configure your main API gateway to act as a strict routing whitelist. Any path not explicitly registered in your central gateway config must be rejected with an HTTP 404 or fall back to an extremely restrictive global rate limit.

Running automated route discovery scans against your staging and production environments can help you identify and secure these hidden entry points.


Defensive Engineering Over Passive Optimization

Building modern backend infrastructure is about preparing for worst-case scenarios. Optimizing database queries and scaling server instances doesn’t help if an unthrottled loop or a scraper bot can easily exhaust your database connection pools.

High-Stability Backend Defense Checklist:
├── Deploy stateless rate-limiting checks close to the edge
├── Run atomic Lua evaluations in Redis
├── Enforce multi-dimensional client keys (Keys, JWTs, and IPs)
├── Design decoupled, asynchronous metric streams
└── Standardize 429 responses with accurate header info

Stop treating traffic control as an afterthought. Build your rate limiters into your core system design, use Redis for high-speed state sharing, and protect your downstream resources before a buggy client update forces you to spend your weekend debugging production crashes.

If you are currently deploying AI engines, LLM integrations, or custom development environments, read setting up pre-commit production hooks to ensure developer actions are secure, monitored, and bounded before they reach production servers. Ensure your backend has the defensive shielding it needs.

Written by Nishil Bhave

Builder, maker, and tech writer at MakeToCreate.

Never miss a post

Get the latest tech insights delivered to your inbox. No spam, unsubscribe anytime.

Related Posts