R
Projects
Research·2026★ Featured

Project Ariel — Low-Altitude Aerial Detection

ML Engineering · Group-One · THD × THI × Fraunhofer IVI · April – July 2026

91.8% mAP@0.5Overall
4.2 msPer frame
YOLOv8nModel
4,776Annotations

YOLOv8n · 2× speed · conf=0.50

Problem

Existing aerial object-detection datasets — VisDrone, UAVDT — are designed for flight altitudes above 30 metres. At 3–9 metres the geometry changes entirely: objects fill the frame, perspective distortion dominates, and ground shadows frequently exceed the physical footprint of the objects they cast. Standard detectors miss detections and generate high false-positive rates on shadows and ground texture.

The project's goal was to close that gap by building a fully annotated dataset from real low-altitude drone footage and fine-tuning a nano-scale model suitable for onboard edge deployment.

4-Stage Pipeline

01

Frame Extraction

30 FPS drone footage filtered with a pixel-wise intensity threshold (≥10% delta per frame). Reduces 2,374 raw frames to ~237 per video sequence without sacrificing scene diversity.

02

Annotation

Collaborative labeling in Roboflow: tight bounding boxes, shadow exclusion (label the physical silhouette only), ≥20% body-visibility rule for truncated objects, ~15% negative-sample ratio.

03

Training

YOLOv8n on Google Colab T4 GPU for 150 epochs. Cosine LR decay, early stopping (patience=15), tuned loss weights (box=7.5, cls=1.5), batch=16, imgsz=640.

04

Evaluation

Locked 10% test split — 231 images, 501 instances. Metrics: mAP@0.5, mAP@0.5:0.95, per-class precision and recall, and confusion matrix analysis.

Dataset

4,776
Annotations
2,227
Frames labeled
10
Video sequences
4
Distinct scenes

Classes: Human 3,642 (76.2%) · Bicycle 597 (12.5%) · Vehicle 537 (11.2%) · Scenes: DK_backyard, DK_parking, THI_Bikepark, THI_Grass · Split: 80 / 10 / 10

Results

ClassPrecisionRecallmAP@0.5mAP@0.5:0.95
All0.9560.9260.9180.850
Human0.9490.9580.9510.845
Bicycle0.9180.8180.8070.713
Vehicle1.0001.0000.9950.993

Inference Latency · T4 GPU

0.9 ms
Preprocess
2.3 ms
Inference
1.0 ms
Postprocess
4.2 ms
Total

Key Finding

Confusion matrix analysis revealed the primary failure mode: 83% of background false positives were classified as Human. The model occasionally hallucinates people on empty terrain — a direct consequence of Human being the majority class (76.2%). Future iterations should expand the negative-sample pool and apply hard-negative mining to address this bias.

Team

Sample Predictions

THI_Bikepark and THI_Grass scene predictions

THI_Bikepark · THI_Grass + Mixed Scenes — model predictions

Ground Truth Labels vs Model Predictions

Ground Truth Labels vs Model Predictions — THI_Bikepark