Pasture Monitoring

This installation demonstrates the application of the Scelaris approach to livestock monitoring, using camera-based detection and vision-language models (VLM) to provide operational context for animal movement and behavior.

Task

Detect and distinguish individual equines visible in camera/event images and determine whether the event contains relevant information.

Input / Acquisition

A dual-sensor MOBOTIX camera provides event images or live imagery. Images enter the processing workflow through camera-generated alerts, email, or direct file-server transfer.

Processing

The camera captures/detects the event. Server-side processing uses a vision-language model (VLM) to analyze the image and identify/assign the visible animals. The analysis result is then converted into an operational action.

Action

Relevant alerts can be enriched with the analysis and forwarded for human attention. Non-relevant alerts can be discarded. Image archiving is independent of whether an analyzed alert is forwarded.

Dual-sensor MOBOTIX camera view of the monitored pasture with equines
Real-world monitoring environment: Dual-sensor MOBOTIX camera view covering the monitored pasture.
MOBOTIX camera interface showing configured event detection regions
Camera-side event detection: Existing MOBOTIX detection regions provide event input for downstream analysis.
Original pasture camera image before VLM analysis
Camera input: Original event image before AI analysis.
VLM analysis result identifying horses, pony and donkey in pasture image
VLM analysis: The vision-language model identifies and distinguishes the visible equines and returns structured results for further processing.
Parallel Monitoring

Zabbix independently monitors the MOBOTIX camera and relevant operational parameters, including system/vital parameters and illumination-related values.

Independent operational monitoring: Zabbix continuously monitors camera availability, operating parameters and sensor values independently of the AI image-analysis pipeline.
Result

The human decision-maker receives contextualized information rather than only an unexplained camera alarm or raw image, supporting a quick and informed decision.