> ## Documentation Index
> Fetch the complete documentation index at: https://docs.guardway.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Sizing

> VM, disk and retention for a self-hosted Guardway gateway at 100k, 1M and 10M requests per day, with the settings that keep the local database bounded.

## What this is for

This page sizes a dedicated virtual machine for the Docker Compose bundle from the **Register Gateway** dialog (see [Deploy](/guardway-gateway/deploy)) at **100k, 1M and 10M requests per day**. The numbers come from load measurements (see [How these numbers were measured](#how-these-numbers-were-measured)) with planning headroom on top.

They apply to **gateway v0.5.46 or later**. Earlier builds saturate at roughly 25–30 requests per second per gateway; do not size a proof of concept on them. [Requirements](/guardway-gateway/requirements) covers the software and network prerequisites.

## Quick reference

| Daily requests | Peak to plan for | VM (single host, Compose) | Disk (SSD) | Local retention |
| - | - | - | - | - |
| **100k / day** | \~10 req/s (bursts to 100 req/s) | **4 vCPU, 8 GB RAM** | **200 GB** | 30 days |
| **1M / day** | \~60 req/s (bursts to 300 req/s) | **8 vCPU, 16 GB RAM** | **1 TB NVMe** | 14–30 days |
| **10M / day** | \~600 req/s (bursts to 1,500 req/s) | **2 gateway VMs: 8 vCPU / 16 GB each**, behind a load balancer, **plus a Postgres VM: 8 vCPU / 32 GB** (or managed Postgres) | Gateways: 100 GB each. Postgres: **2 TB NVMe** | 7 days, with `GUARDWAY_LOG_BODY_MAX_BYTES=4096` |

<Warning>
  **Set retention on day one.** Local retention and the stored-body cap are opt-in (see [Configuration for each tier](#configuration-for-each-tier)). Without them the gateway's database grows forever.
</Warning>

**Compute is cheap and storage is what you size for.** The gateway stores each request's prompt and response locally, so the database grows with your *prompt size* × volume × retention. The disk column assumes typical chat traffic (\~20 KB stored per request). See [Database size](#database-size) for short-prompt workloads (about 10× less) and the two settings that bound it: `GUARDWAY_RETENTION_DAYS` and `GUARDWAY_LOG_BODY_MAX_BYTES`.

Add to the VM above if you enable these optional components:

| Optional component | Add |
| - | - |
| ML PII detection (the `risk-engine` service) on every request | See [ML PII detection](#ml-pii-detection). At 100k/day: +8 vCPU and +8 GB RAM. At 1M/day and above: one NVIDIA GPU (T4/L4 class), or restrict ML PII to selected applications. |
| Live prompt-injection inspection (the `inspection` service, LLM-based) | No significant VM CPU. It makes one extra LLM call per inspected request, so size your LLM quota for it. |
| Observability profile (Jaeger `tracing`) | +1 vCPU, +2 GB RAM |

Use x86-64 or arm64 with SSD storage. Network-attached HDD-class storage is not suitable for the database. Leave 30 GB of the disk for container images (the ML PII image alone is about 10 GB unpacked).

## What drives the size

Measured per proxied request, with \~20 regex/keyword guardrail policies on request and response:

| Resource | Cost per request | What it means |
| - | - | - |
| Gateway CPU | \~2 ms | One vCPU saturates near 500 req/s. Plan for 150 req/s per vCPU, leaving room for larger prompts, streaming and TLS. |
| Postgres CPU | \~1–1.5 ms | One vCPU saturates near 600 req/s. Plan for 200 req/s per vCPU. |
| Database disk | \~2 KB + stored bodies | Fixed overhead per request is \~2 KB: the `requests` row plus `audit_logs` (which keeps up to 10 KB of the masked request body). On top of that, the request and response bodies are stored, compressed, each capped at `GUARDWAY_LOG_BODY_MAX_BYTES` (64 KiB recommended). Measured on short load-test prompts: \~2.1 KB in total. On a real gateway's chat traffic: \~21 KB average, up to \~100 KB (agents and coding tools resend their whole context every turn). |
| Container logs | \~0.45 KB | Rotated by the bundle at 5 × 50 MB per service, so bounded. |
| Postgres WAL | bounded | Capped by `max_wal_size` (bundle default 2 GB). |
| Gateway memory | \~150–300 MB base, plus \~50–100 KB per in-flight request | In-flight requests ≈ request rate × LLM latency. Example: 300 req/s × 5 s = 1,500 in flight ≈ 150 MB. |

Request rate is rarely flat. A day's traffic usually lands in about 8 business hours, with short bursts on top (batch jobs, load tests, agent fan-out). The quick-reference table plans peak = 5 × the daily average and lets bursts reach several times that. The gateway itself stays fast during bursts. Your LLM provider's rate limits are usually the first limit hit.

<Note>
  Audit logs are **local-only**: the `audit_logs` table lives in the gateway's own database and is never sent to the Guardway platform. Retention settings on this page are what bound it.
</Note>

### Database size

Plan stored bytes per request from your workload, then multiply: **DB size ≈ daily requests × bytes per request × retention days × 1.5**. The 1.5× covers indexes and vacuum headroom.

| Workload | Stored per request | Notes |
| - | - | - |
| Short prompts (classification, extraction, load tests) | \~3 KB | |
| Typical chat / RAG | \~20 KB | Real-gateway average |
| Agents / coding tools (long context) | 30–130 KB | Capped by `GUARDWAY_LOG_BODY_MAX_BYTES` (64 KiB per body recommended) |
| Any workload with `GUARDWAY_LOG_BODY_MAX_BYTES=4096` | ≤ \~10 KB | Keeps a 4 KB preview of each body |

Typical chat (\~20 KB per request):

| Daily requests | 7 days | 14 days | 30 days |
| - | - | - | - |
| 100k | 21 GB | 42 GB | 90 GB |
| 1M | 210 GB | 420 GB | 900 GB |
| 10M | 2.1 TB | 4.2 TB | — (use a body cap) |

Short prompts (\~3 KB per request), or typical chat with `GUARDWAY_LOG_BODY_MAX_BYTES=4096`:

| Daily requests | 7 days | 14 days | 30 days |
| - | - | - | - |
| 100k | 3 GB | 6 GB | 14 GB |
| 1M | 32 GB | 63 GB | 135 GB |
| 10M | 315 GB | 630 GB | 1.4 TB |

Measure your real number after a day of traffic. From the gateway host, with the bundle's `database` service running:

```bash theme={null}
docker compose exec database psql -U gateway -d gateway -c "
SELECT pg_size_pretty(
         (pg_total_relation_size('requests') + pg_total_relation_size('audit_logs'))
         / GREATEST((SELECT count(*) FROM requests), 1)
       ) AS stored_per_request,
       (SELECT count(*) FROM requests) AS requests;"
```

## Configuration for each tier

All settings go in the bundle's `.env`. Defaults are shown in parentheses. The full variable reference is on [Environment](/guardway-gateway/environment).

```bash theme={null}
# Local request history (requests, audit logs, webhook deliveries…).
# These three are OFF by default (0), so upgrading a gateway never deletes or
# truncates history it already holds. Set them on every new VM.
GUARDWAY_RETENTION_DAYS=30            # (0 = keep forever)  recommended 30
GUARDWAY_WEBHOOK_RETENTION_DAYS=7     # (0)                 recommended 7
GUARDWAY_LOG_BODY_MAX_BYTES=65536     # (0 = no cap)        65536; 4096 at high volume

# Postgres (bundle defaults shown)
POSTGRES_SHARED_BUFFERS=256MB         # 100k/day: 256MB · 1M/day: 2GB · 10M/day: 8GB
POSTGRES_MAX_CONNECTIONS=200
POSTGRES_MAX_WAL_SIZE=2GB             # 10M/day: 8GB

# Gateway
DB_MAX_OPEN_CONNS=25                  # 10M/day: 50 per gateway (keep gateways × value < max_connections)
GUARDWAY_UPSTREAM_MAX_CONNS_PER_HOST=1024

# Container log rotation (per service)
GUARDWAY_LOG_MAX_SIZE=50m
GUARDWAY_LOG_MAX_FILES=5
```

<Note>
  Two more capacity variables exist on the container, `GUARDWAY_MAX_INFLIGHT` (default `4096`, see [Load shedding and bursts](#load-shedding-and-bursts)) and `DB_MAX_IDLE_CONNS` (default `10`; at 10M/day raise it to 25–50 per gateway to avoid reconnect churn). The onboarding bundle's `docker-compose.yml` does not pass these two through from `.env` yet. If you bring your own compose file, set them in the gateway service's `environment:` block.
</Note>

**100k/day (POC).** One VM with the bundle defaults. Everything, Postgres included, runs on the same host.

**1M/day.** One VM. Set `POSTGRES_SHARED_BUFFERS=2GB`. With typical chat traffic, 14-day retention fits the 1 TB disk with headroom. For 30 days, set `GUARDWAY_LOG_BODY_MAX_BYTES=16384` or add disk. Back up the `database_data` volume daily.

**10M/day.** Run Postgres on its own VM (or a managed Postgres such as RDS, Azure Database for PostgreSQL or Cloud SQL, 8 vCPU / 32 GB) with NVMe storage. Run two or more gateway VMs behind a load balancer for capacity and high availability. Point each gateway at the shared Postgres (`DB_HOST`, `DB_PASSWORD`, …) and set `DB_MAX_OPEN_CONNS=50`. Use 7-day retention with `GUARDWAY_LOG_BODY_MAX_BYTES=4096` (about 315 GB). If you need full prompts kept longer at this volume, push them to your SIEM or data lake through [Settings → Notifications](/platform/settings/notifications) rather than the gateway's operational database.

## ML PII detection

The `risk-engine` service runs a GLiNER transformer model for PII entity detection. It runs only for guardrail policies that use the ML/SLM provider (see [SLM Guardrails](/platform/configuration/security#slm-guardrails)). Regex-based PII, keyword and content-filter policies are cheap and are already included in the base numbers above.

Measured on CPU: **about 2 vCPU-seconds per short prompt**. Four workers on 4 vCPU complete about 2 scans per second, and each worker holds about 1 GB of RAM. That makes ML PII on *every* request practical only at low volume:

* **100k/day:** average \~1.2 req/s, peaks \~6 req/s. Add 8 vCPU / 8 GB and set `WEB_CONCURRENCY=4` on the `risk-engine` service (one model per worker).
* **1M/day and above:** use a GPU host for `risk-engine`. One T4/L4-class GPU is expected to cover this tier, but that is an estimate (GPU throughput was not measured); validate with your own prompts before committing. Otherwise apply ML PII only to the applications that need it, and use regex PII elsewhere.

## Load shedding and bursts

When more than `GUARDWAY_MAX_INFLIGHT` requests (default 4096) are in progress, the gateway answers `503 gateway_overloaded` with `Retry-After: 1` instead of queueing. Client SDKs (OpenAI, Anthropic, aiohttp with retries) back off and retry. Sustained 503s mean the VM is undersized for the offered load: add vCPUs or a second gateway.

Rate limits are separate from capacity. An org limit from [Settings → Traffic → Rate Limits](/platform/settings/traffic) or a per-key limit returns `429 rate_limit_exceeded` by design. Check those first when a load test sees 429s.

## Disk growth: diagnosing and reclaiming

On the host, `docker system df -v` shows image, volume and container-log usage. Inside the database, check the largest tables, dead rows and how far back request history goes:

```bash theme={null}
docker compose exec database psql -U gateway -d gateway -c "
SELECT relname, pg_size_pretty(pg_total_relation_size(relid)) AS size, n_live_tup, n_dead_tup
FROM pg_stat_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 8;"
docker compose exec database psql -U gateway -d gateway -c "SELECT min(created_at) FROM requests;"
```

Common causes and fixes:

| Symptom | Cause | Fix |
| - | - | - |
| `requests` / `audit_logs` large, old `min(created_at)` | No retention configured (always the case before v0.5.46; opt-in since) | Upgrade to v0.5.46 or later and set `GUARDWAY_RETENTION_DAYS`. The pruner runs within 2 minutes of start, then hourly. |
| `requests` large even with few rows | Large prompt/response bodies (agents, long context) | Lower `GUARDWAY_LOG_BODY_MAX_BYTES` (applies to new rows) and/or `GUARDWAY_RETENTION_DAYS`. |
| Container log files of several GB | Docker json-file logs never rotated (bundles generated before gateway v0.5.46) | Download a fresh onboarding bundle, or add the same `logging` rotation to your compose file, and recreate the services with `docker compose up -d`. |
| `docker system df` shows large images | Old image versions after upgrades | `docker image prune -a` (removes images no container uses). |
| Many dead rows | A long-open transaction blocked vacuum | The bundle sets `idle_in_transaction_session_timeout=10min`. Then run `VACUUM (ANALYZE)` on the table. |
| Large WAL | Archiving or an inactive replication slot | Disable `archive_mode`, or drop the stale slot. |

Deleted rows free space *inside* Postgres for reuse; the files on disk do not shrink. To return space to the OS after a large one-time prune, run `VACUUM FULL <table>` in a maintenance window. It locks the table while it runs.

## How these numbers were measured

A load rig run on 2026-10-07:

* **Setup:** a 2-vCPU / 1 GB gateway container and a 1-vCPU / 512 MB Postgres; an OpenAI-compatible mock upstream with a fixed 800 ms latency; an open-loop load generator; \~23 customer-shaped content-filter and keyword/regex guardrail policies on request and response.
* **Throughput:** 580 req/s sustained at p95 814 ms, with the gateway at 1.1 vCPU and Postgres at 0.6 vCPU. At 1,200 req/s offered the gateway served 937 req/s and shed the rest with 503s, with no crash.
* **Disk:** 100,200 requests at 300 req/s, all 200, p95 980 ms. The database grew by 209 MB (\~2.1 KB per request) with short prompts. The same rig on v0.5.45 saturated at 24 req/s. The typical-chat figure (\~20 KB) comes from a long-running cloud-connected QA gateway: average compressed request body 21 KB over 8,086 chat completions.
* **ML PII:** measured separately, with the `risk-engine` image v0.1.0 and 155-character Portuguese prompts.

These measurements ran on Apple-silicon cores. Typical server x86-64 cores perform within the planning headroom used above.

## Related

* [Requirements](/guardway-gateway/requirements) — software, network and baseline hardware.
* [Deploy](/guardway-gateway/deploy) — the onboarding bundle and `docker compose up -d`.
* [Environment](/guardway-gateway/environment) — every variable the gateway reads, including the capacity and retention ones.
* [Settings → Traffic](/platform/settings/traffic) — org-wide rate limits, which return 429s independent of capacity.
* [Settings → Notifications](/platform/settings/notifications) — push events to a SIEM instead of keeping full prompts in the gateway database.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.