Skip to main content

GStreamer input

On AI boxes (Hailo-8, NVIDIA Jetson, x86) the Embedded Stream Processor gets its frames from a GStreamer pipeline configured in the gstreamer_input solver. The pipeline can read a video file, an RTSP camera or a USB camera; smart cameras use their vendor input solver instead and skip this page. The solver appends its own app sink to whatever you write, so your pipeline describes only the source and the conversion to raw BGR frames at a known resolution.

Pipeline rules​

  • The pipeline must end with frames in video/x-raw,format=BGR; in most cases the last elements are videoconvert ! video/x-raw,format=BGR.
  • gst_width and gst_height must match the resolution the pipeline actually outputs. Add videoscale and a caps filter with width=...,height=... if the source resolution differs.
  • Never add a sink element (appsink, autovideosink, xvimagesink, ...); the solver adds its own.
  • Limit the frame rate with videorate and framerate=10/1 — ten frames per second is enough for walk-through recognition and keeps the accelerator free for detection.
  • The face-detection solver on some accelerators expects a fixed input shape (for example w1280h720 in its file name); the pipeline resolution must match it.

The pipeline goes into settings.yaml as the gst_pipeline parameter:

solvers:
frame_input:
solver: ./solver/gstreamer_input.cpu.solver
parameters:
- name: gst_pipeline
value: "<your pipeline>"
- name: gst_width
value: 1280
- name: gst_height
value: 720

Validate a pipeline first​

Test every pipeline with gst-launch-1.0 before giving it to the stream processor. Take the pipeline as written for settings.yaml, prepend gst-launch-1.0, and append a conversion and a display sink so you can see the frames:

$ gst-launch-1.0 <pipeline from settings.yaml> ! videoconvert ! xvimagesink

If the window shows video at the expected resolution and rate, remove the two trailing elements and paste the pipeline into settings.yaml. gst-device-monitor-1.0 lists the available capture devices with their supported formats and resolutions, which is the quickest way to find the right device= path and caps for a USB camera:

$ gst-device-monitor-1.0
Device found:
name : Integrated_Webcam_HD
class : Video/Source
caps : video/x-raw, format=YUY2, width=640, height=480, framerate=30/1
image/jpeg, width=1280, height=720, framerate=30/1
properties:
device.path = /dev/video0

Example pipelines​

MP4 file (software decoding):

filesrc location=test.mp4 ! qtdemux ! h264parse ! avdec_h264 ! videorate drop-only=true ! videoconvert ! video/x-raw,format=BGR,framerate=10/1

MP4 file with Intel VA-API hardware decoding (Axiomtek RSC101 and other Intel-based boxes; needs gstreamer1.0-vaapi, see install):

filesrc location=test.mp4 ! qtdemux ! vaapidecodebin ! queue leaky=no max-size-buffers=5 max-size-bytes=0 max-size-time=0 ! vaapipostproc ! videorate ! video/x-raw,framerate=10/1,width=1280,height=720 ! queue leaky=no max-size-buffers=5 max-size-bytes=0 max-size-time=0 ! videoconvert ! video/x-raw,format=BGR

RTSP camera (H.264):

rtspsrc location=rtsp://user:pass@192.0.2.40:554/stream latency=30 ! queue ! rtph264depay ! h264parse ! avdec_h264 ! videorate ! videoconvert ! videoscale ! video/x-raw,format=BGR,width=1280,height=720,framerate=10/1

Replace avdec_h264 with the platform's hardware decoder where available (for example vaapih264dec on Intel or nvv4l2decoder on Jetson) and with avdec_h265 / rtph265depay / h265parse for H.265 streams. Camera URL formats are covered in RTSP URLs and masking.

USB camera:

v4l2src device=/dev/video0 ! video/x-raw,framerate=10/1,width=1280,height=720 ! videoconvert ! video/x-raw,format=BGR

If the camera only offers MJPEG at the wanted resolution, decode it first: v4l2src device=/dev/video0 ! image/jpeg,width=1280,height=720,framerate=10/1 ! jpegdec ! videoconvert ! video/x-raw,format=BGR.

Several streams on one box​

One sfe_stream_processor process handles one stream per settings file, and you can pass several files to a single process. Duplicate settings.yaml, give each copy its own connection.client_id (each stream is a separate edge stream in Station) and its own frame_input pipeline, and start:

$ ./bin/sfe_stream_processor settings/lobby.yaml settings/entrance.yaml settings/test-video.yaml

For example, settings/lobby.yaml reads a USB camera (v4l2src ...), settings/entrance.yaml an RTSP camera (rtspsrc ...) and settings/test-video.yaml a file (filesrc ...), each with gst_width/gst_height matching its pipeline. All other sections can stay identical. Size the box for the sum of the streams: each one runs detection at its own frame rate on the shared accelerator, so reduce framerate or resolution per stream if detection latency grows.