Hardware requirements
Sizing a Face Matcher host is driven almost entirely by how many video streams it processes and where the processing happens. Server-side RTSP puts the whole decoding, detection and extraction load on the host, so demand grows with every camera. Edge streams move detection and extraction onto the camera or AI box, leaving the server a fraction of the work per camera. API-only identification has no video at all and is sized by request rate and watchlist size. Under-sizing shows up as latency, timeouts and dropped frames, so evaluate on adequate hardware from the start.
Minimum requirements
- x86_64 CPU with the AVX2 instruction set (Intel Haswell or AMD Zen and newer), or an ARM CPU (AWS Graviton, NVIDIA Jetson).
- At least 4 physical CPU cores.
- 16 GB RAM.
- 80 GB of storage; SSD recommended. Stored images grow with the retention period and the number of cameras, see Data retention.
- Linux with Docker Engine and Docker Compose.
- Optional: a supported NVIDIA GPU for the detection, extraction and liveness services, see GPU acceleration.
Evaluation hardware
Machines used by Innovatrics for demonstrations, with Full HD RTSP streams and watchlists under 10,000 members:
| Hardware | RTSP cameras |
|---|---|
| NUC with Intel Core i7-1360P, 16 GB RAM | up to 3 |
| Lenovo Legion 9 with Intel Core i9-14900HX, 64 GB RAM | up to 18 |
| Dell PowerEdge R250 with Intel Xeon E-2336, 32 GB RAM | up to 4 |
Production sizing
The figures assume Full HD RTSP streams and fewer than 10,000 watchlist members unless stated otherwise.
Server-side RTSP processing
| Hardware | RTSP cameras |
|---|---|
| Dell PowerEdge R250 with Intel Xeon E-2336, 32 GB RAM | up to 4 |
| Dell PowerEdge R660 with Intel Xeon Gold 5416S | up to 14 |
| Dell PowerEdge R760 with 2x Intel Xeon Platinum 8470 | up to 100 |
API-only identification
Search requests per second with a response latency under 1,000 ms; each request carries one selfie-style image with a single face.
| Hardware | Search requests | Watchlist size |
|---|---|---|
| NUC with Intel Core i7-1360P, 16 GB RAM | 4 per second | up to 100,000 |
| Amazon EC2 c7g.2xlarge | 2 per second | up to 1 million |
| Amazon EC2 c7g.4xlarge | 14 per second | up to 1 million |
| 2x Amazon EC2 c7g.12xlarge | 120 per second | up to 1 million |
Edge streams
With detection and extraction on the devices, the same server hardware handles many more cameras:
| Hardware | Edge devices | Watchlist size |
|---|---|---|
| NUC with Intel Core i7-1360P, 16 GB RAM | up to 150 | up to 10,000 |
| Dell PowerEdge R250 with Intel Xeon E-2336, 32 GB RAM | up to 150 | up to 10,000 |
| Dell PowerEdge R760 with 2x Intel Xeon Platinum 8470 | up to 400 | up to 10 million |
Supported devices are IP cameras with an Ambarella SoC, AI boxes with a Hailo accelerator and NVIDIA Jetson devices; see Supported devices.
Service instances by traffic
Beyond the host itself, the number of instances of each service decides how much traffic a given camera count absorbs. The samples below assume Full HD RTSP cameras, a 500 ms re-detection interval and a 10,000-member watchlist; they are illustrative, and a project-specific sizing is agreed with Innovatrics before procurement. How to run several instances is described in Scaling.
| Scenario | Traffic | Camera | Detector | Extractor | Matcher | Liveness |
|---|---|---|---|---|---|---|
| 2 cameras, demo or proof of concept | 3 people constantly in view | 2 | 1 | 2 | 1 | 1 |
| 50 cameras, low traffic | 4 people per minute, 3 s in view | 50 | 1 | 1 | 1 | 0 |
| 50 cameras, medium traffic | 20 people per minute, 3 s in view | 50 | 1 | 4 | 1 | 0 |
| 50 cameras, high traffic | 120 people per minute, 5 s in view | 50 | 1 | 34 | 4 | 0 |
| 50 cameras, extreme traffic | 300 people per minute, 5 s in view | 50 | 1 | 84 | 8 | 0 |
What changes the math
- Camera count per host scales roughly linearly. When a host saturates, add a host rather than overloading one; see Multi-server deployment.
- Processing type: moving cameras to edge devices reclaims most of the server capacity they used, see Edge streams.
- Stored images and retention: face crops and full frames go to S3 storage; size the disk for your retention period and per-camera save strategy.
- Watchlist size raises matching cost and memory; watchlists in the millions call for the Accurate extractor and more matcher instances.
Share the intended camera layout, expected throughput and camera shortlist with Innovatrics early; a short sizing check is cheaper than re-procurement.