Skip to main content

Monitoring and Logs

Face Matcher services publish metrics and traces through OpenTelemetry and write structured JSON logs to standard output. The package ships no monitoring stack; you connect your own Prometheus, OTLP collector and log shipper. Configuration uses the standard OpenTelemetry SDK environment variables in section 2.8 of .env, plus a few Face Matcher settings.

Metrics​

Every engine service runs a Prometheus exporter on port 4318 inside the container:

OTEL_METRICS_EXPORTER=prometheus
OTEL_EXPORTER_PROMETHEUS_HOST=*
OTEL_EXPORTER_PROMETHEUS_PORT=4318
Metrics__RequestsProcessingTimeHistogramBands="100, 200, 500, 1000, 2500, 5000"

The port is exposed on face-matcher-network (not published on the host), and the services carry the label scrapeMetrics: "true". A Prometheus container on the same host joins the network and scrapes <service>:4318 (for example extractor:4318, cam-1:4318, api:4318), or discovers targets through Docker service discovery filtered on that label. Metrics__RequestsProcessingTimeHistogramBands sets the histogram buckets (milliseconds) for request processing time on the RPC services, base, matcher and the face search service.

To read the metrics of one service by hand, publish its port in docker-compose.override.yml with a distinct host port per service, apply with docker compose up -d and open http://localhost:4321/metrics:

services:
cam-1:
ports:
- 4321:4318

Each service also serves an HTTP health check on port 6060 (HealthCheck__Port); api and graphql-api additionally check the database and RabbitMQ (HealthCheck__Tags__db, HealthCheck__Tags__rmq). RabbitMQ ships with the rabbitmq_prometheus plugin enabled and its management UI on localhost:15672.

Tracing​

Tracing is off by default. Point it at any OTLP receiver (OpenTelemetry Collector, Jaeger, Grafana Tempo) and enable it:

Tracing__Enabled=true
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=0.1

The shipped default http://jaeger:4317 assumes a receiver container named jaeger on face-matcher-network; none is included, so run your own or change the address. OTEL_TRACES_SAMPLER_ARG=1.0 traces every request and is fine for debugging; lower it in production. Apply with docker compose up -d.

Logs​

Services log JSON to standard output (AppSettings__Log_JsonConsole_Enabled=true); rolling files are off. Levels follow the Serilog__MinimumLevel__* settings in section 2.7 of .env; raise Serilog__MinimumLevel__Default to Debug temporarily when you investigate, then apply with docker compose up -d. Read logs from the deployment directory:

docker compose ps # container names and state
docker compose logs -f extractor # follow one service
docker compose logs --tail=0 -f # everything, live
docker compose -f dependencies/docker-compose.yml logs -f rabbitmq
docker logs -f fm-station # Station by container name

The services carry the label logging: promtail, which you can use as a selector in Promtail or another Docker log shipper to forward the logs to Loki or a similar system.

Log rotation​

Docker's json-file driver keeps container logs without a limit by default, and a busy camera container writes a lot. Check what the logs occupy:

sudo du -h $(docker inspect --format='{{.LogPath}}' $(docker ps -qa))

Set a limit per service in docker-compose.override.yml, or host-wide in /etc/docker/daemon.json ("log-opts": {"max-size": "10m", "max-file": "5"} under "log-driver": "json-file", then restart Docker). Either way the limit applies to containers created afterwards, so run docker compose up -d --force-recreate:

services:
extractor:
logging:
driver: json-file
options:
max-size: 10m
max-file: "5"