What this is for
This page sizes a dedicated virtual machine for the Docker Compose bundle from the Register Gateway dialog (see Deploy) at 100k, 1M and 10M requests per day. The numbers come from load measurements (see How these numbers were measured) with planning headroom on top. They apply to gateway v0.5.46 or later. Earlier builds saturate at roughly 25–30 requests per second per gateway; do not size a proof of concept on them. Requirements covers the software and network prerequisites.Quick reference
Compute is cheap and storage is what you size for. The gateway stores each request’s prompt and response locally, so the database grows with your prompt size × volume × retention. The disk column assumes typical chat traffic (~20 KB stored per request). See Database size for short-prompt workloads (about 10× less) and the two settings that bound it:
GUARDWAY_RETENTION_DAYS and GUARDWAY_LOG_BODY_MAX_BYTES.
Add to the VM above if you enable these optional components:
Use x86-64 or arm64 with SSD storage. Network-attached HDD-class storage is not suitable for the database. Leave 30 GB of the disk for container images (the ML PII image alone is about 10 GB unpacked).
What drives the size
Measured per proxied request, with ~20 regex/keyword guardrail policies on request and response:
Request rate is rarely flat. A day’s traffic usually lands in about 8 business hours, with short bursts on top (batch jobs, load tests, agent fan-out). The quick-reference table plans peak = 5 × the daily average and lets bursts reach several times that. The gateway itself stays fast during bursts. Your LLM provider’s rate limits are usually the first limit hit.
Audit logs are local-only: the
audit_logs table lives in the gateway’s own database and is never sent to the Guardway platform. Retention settings on this page are what bound it.Database size
Plan stored bytes per request from your workload, then multiply: DB size ≈ daily requests × bytes per request × retention days × 1.5. The 1.5× covers indexes and vacuum headroom.
Typical chat (~20 KB per request):
Short prompts (~3 KB per request), or typical chat with
GUARDWAY_LOG_BODY_MAX_BYTES=4096:
Measure your real number after a day of traffic. From the gateway host, with the bundle’s
database service running:
Configuration for each tier
All settings go in the bundle’s.env. Defaults are shown in parentheses. The full variable reference is on Environment.
Two more capacity variables exist on the container,
GUARDWAY_MAX_INFLIGHT (default 4096, see Load shedding and bursts) and DB_MAX_IDLE_CONNS (default 10; at 10M/day raise it to 25–50 per gateway to avoid reconnect churn). The onboarding bundle’s docker-compose.yml does not pass these two through from .env yet. If you bring your own compose file, set them in the gateway service’s environment: block.POSTGRES_SHARED_BUFFERS=2GB. With typical chat traffic, 14-day retention fits the 1 TB disk with headroom. For 30 days, set GUARDWAY_LOG_BODY_MAX_BYTES=16384 or add disk. Back up the database_data volume daily.
10M/day. Run Postgres on its own VM (or a managed Postgres such as RDS, Azure Database for PostgreSQL or Cloud SQL, 8 vCPU / 32 GB) with NVMe storage. Run two or more gateway VMs behind a load balancer for capacity and high availability. Point each gateway at the shared Postgres (DB_HOST, DB_PASSWORD, …) and set DB_MAX_OPEN_CONNS=50. Use 7-day retention with GUARDWAY_LOG_BODY_MAX_BYTES=4096 (about 315 GB). If you need full prompts kept longer at this volume, push them to your SIEM or data lake through Settings → Notifications rather than the gateway’s operational database.
ML PII detection
Therisk-engine service runs a GLiNER transformer model for PII entity detection. It runs only for guardrail policies that use the ML/SLM provider (see SLM Guardrails). Regex-based PII, keyword and content-filter policies are cheap and are already included in the base numbers above.
Measured on CPU: about 2 vCPU-seconds per short prompt. Four workers on 4 vCPU complete about 2 scans per second, and each worker holds about 1 GB of RAM. That makes ML PII on every request practical only at low volume:
- 100k/day: average ~1.2 req/s, peaks ~6 req/s. Add 8 vCPU / 8 GB and set
WEB_CONCURRENCY=4on therisk-engineservice (one model per worker). - 1M/day and above: use a GPU host for
risk-engine. One T4/L4-class GPU is expected to cover this tier, but that is an estimate (GPU throughput was not measured); validate with your own prompts before committing. Otherwise apply ML PII only to the applications that need it, and use regex PII elsewhere.
Load shedding and bursts
When more thanGUARDWAY_MAX_INFLIGHT requests (default 4096) are in progress, the gateway answers 503 gateway_overloaded with Retry-After: 1 instead of queueing. Client SDKs (OpenAI, Anthropic, aiohttp with retries) back off and retry. Sustained 503s mean the VM is undersized for the offered load: add vCPUs or a second gateway.
Rate limits are separate from capacity. An org limit from Settings → Traffic → Rate Limits or a per-key limit returns 429 rate_limit_exceeded by design. Check those first when a load test sees 429s.
Disk growth: diagnosing and reclaiming
On the host,docker system df -v shows image, volume and container-log usage. Inside the database, check the largest tables, dead rows and how far back request history goes:
Deleted rows free space inside Postgres for reuse; the files on disk do not shrink. To return space to the OS after a large one-time prune, run
VACUUM FULL <table> in a maintenance window. It locks the table while it runs.
How these numbers were measured
A load rig run on 2026-10-07:- Setup: a 2-vCPU / 1 GB gateway container and a 1-vCPU / 512 MB Postgres; an OpenAI-compatible mock upstream with a fixed 800 ms latency; an open-loop load generator; ~23 customer-shaped content-filter and keyword/regex guardrail policies on request and response.
- Throughput: 580 req/s sustained at p95 814 ms, with the gateway at 1.1 vCPU and Postgres at 0.6 vCPU. At 1,200 req/s offered the gateway served 937 req/s and shed the rest with 503s, with no crash.
- Disk: 100,200 requests at 300 req/s, all 200, p95 980 ms. The database grew by 209 MB (~2.1 KB per request) with short prompts. The same rig on v0.5.45 saturated at 24 req/s. The typical-chat figure (~20 KB) comes from a long-running cloud-connected QA gateway: average compressed request body 21 KB over 8,086 chat completions.
- ML PII: measured separately, with the
risk-engineimage v0.1.0 and 155-character Portuguese prompts.
Related
- Requirements — software, network and baseline hardware.
- Deploy — the onboarding bundle and
docker compose up -d. - Environment — every variable the gateway reads, including the capacity and retention ones.
- Settings → Traffic — org-wide rate limits, which return 429s independent of capacity.
- Settings → Notifications — push events to a SIEM instead of keeping full prompts in the gateway database.