Skip to content

Latest commit

Β 

History

22 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🎯 Detecto β€” Real-Time Person Detection & Counting System

A full-stack application that detects and counts people in images using YOLOv8, served via a FastAPI backend with a React + Vite frontend.

Detecto Banner


πŸ“– Table of Contents

  1. Project Purpose
  2. Architecture
  3. Features
  4. Tech Stack
  5. Repository Structure
  6. Setup & Run Instructions
  7. API Reference
  8. Testing Methodology
  9. Results
  10. Screenshots
  11. What Worked Well / What Could Improve
  12. Bonus Features
  13. Team

🎯 Project Purpose

Detecto automates person detection and counting in images for safety and occupancy monitoring. It provides:

  • Real-time visual feedback β€” bounding boxes with confidence scores drawn directly on uploaded images.
  • Analytics β€” a history of every detection with timestamp, count, average confidence, and inference time.
  • A clean API β€” so the detection engine can be reused by other apps, scripts, or dashboards.

The system is designed for operators who need fast, clear feedback without touching any ML code.


πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                       FRONTEND (React + Vite)                β”‚
β”‚                                                              β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚
β”‚   β”‚  DetectionView   β”‚         β”‚     HistoryView       β”‚     β”‚
β”‚   β”‚  - upload image  β”‚         β”‚  - list detections    β”‚     β”‚
β”‚   β”‚  - show boxes    β”‚         β”‚  - filter by conf.    β”‚     β”‚
β”‚   β”‚  - show stats    β”‚         β”‚  - reset history      β”‚     β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚  POST /detect                 β”‚  GET /history
             β”‚  (multipart/form-data)        β”‚  POST /reset
             β–Ό                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     BACKEND (FastAPI)                        β”‚
β”‚                                                              β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
β”‚   β”‚ detect.py  │──►│ preprocessing │──►│  YOLOv8 model  β”‚    β”‚
β”‚   β”‚            β”‚   β”‚    .py        β”‚   β”‚  (ultralytics) β”‚    β”‚
β”‚   β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚
β”‚         β”‚ save                                               β”‚
β”‚         β–Ό                                                    β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”          β”‚
β”‚   β”‚  SQLite    │◄──│  history.py (read + reset)   β”‚          β”‚
β”‚   β”‚ detections β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜          β”‚
β”‚   β”‚    .db     β”‚                                             β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                                             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Data flow for a detection request:

  1. User selects an image in the browser.
  2. Frontend POSTs the image to /detect as multipart/form-data.
  3. Backend decodes the image, runs YOLOv8 inference, filters for the person class (COCO class 0), and draws bounding boxes.
  4. Backend returns JSON with count, boxes, confidences, inference time, and a base64-encoded annotated image.
  5. Backend persists the detection (timestamp, count, avg confidence, inference time) to SQLite.
  6. Frontend renders the annotated image plus summary statistics.

✨ Features

Detection View

  • Upload JPEG or PNG images
  • Draws bounding boxes on all detected people
  • Displays:
    • Total person count
    • Per-person confidence scores
    • Average confidence
    • Inference time in milliseconds
  • Handles invalid uploads with clear error messages

History View

  • Table of all past detections: timestamp, count, avg confidence, inference time
  • Filter by minimum confidence threshold
  • Reset entire detection history
  • Sorted newest-first

Backend API

  • POST /detect β€” run inference on an uploaded image
  • GET /history β€” retrieve past detections
  • POST /reset β€” clear all stored detections
  • Auto-generated API docs at /docs (Swagger UI)

πŸ› οΈ Tech Stack

Layer Technology
Detection YOLOv8n (Ultralytics), OpenCV, NumPy, Pillow
Backend FastAPI, Uvicorn, SQLAlchemy, Pydantic
Database SQLite
Frontend React 19, Vite, Axios, React Router
Styling Tailwind CSS
Config python-dotenv (root .env), Vite env (.env)

πŸ“ Repository Structure

detecto/
β”œβ”€β”€ .env.example              # MODEL_PATH, CONFIDENCE_THRESHOLD, DB_PATH
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ main.py
β”‚   β”œβ”€β”€ config.py             # env loading + path resolution
β”‚   β”œβ”€β”€ routes/
β”‚   β”‚   β”œβ”€β”€ detect.py
β”‚   β”‚   └── history.py
β”‚   β”œβ”€β”€ models/
β”‚   β”‚   β”œβ”€β”€ record.py         # SQLAlchemy model + engine
β”‚   β”‚   └── schemas.py        # Pydantic DTOs
β”‚   β”œβ”€β”€ utils/
β”‚   β”‚   └── preprocessing.py
β”‚   β”œβ”€β”€ samples/
β”‚   β”‚   β”œβ”€β”€ frame1.jpg
β”‚   β”‚   └── ...
β”‚   β”œβ”€β”€ requirements.txt
β”‚   └── yolov8n.pt            # model weights
β”‚
β”œβ”€β”€ frontend/
β”‚   β”œβ”€β”€ public/
β”‚   β”‚   └── samples/
β”‚   β”‚       β”œβ”€β”€ frame1.jpg
β”‚   β”‚       β”œβ”€β”€ frame2.jpg
β”‚   β”‚       └── frame3.jpg
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ components/
β”‚   β”‚   β”‚   β”œβ”€β”€ Navbar.jsx
β”‚   β”‚   β”‚   └── StatsCard.jsx
β”‚   β”‚   β”œβ”€β”€ pages/
β”‚   β”‚   β”‚   β”œβ”€β”€ DetectionView.jsx
β”‚   β”‚   β”‚   └── HistoryView.jsx
β”‚   β”‚   β”œβ”€β”€ App.jsx
β”‚   β”‚   └── main.jsx
β”‚   β”œβ”€β”€ .env.example          # VITE_API_URL
β”‚   β”œβ”€β”€ package.json
β”‚   └── vite.config.js
β”‚
β”œβ”€β”€ detections.db             # SQLite (created at runtime)
β”œβ”€β”€ README.md
└── .gitignore

βš™οΈ Setup & Run Instructions

Prerequisites

  • Python 3.10+
  • Node.js 18+ and npm
  • (Optional) NVIDIA GPU with CUDA for faster inference
  • ~500 MB free disk space for the YOLO weights and dependencies

1. Clone the repository

git clone <your-repo-url> detecto
cd detecto

2. Backend Setup

Run from the repository root (backend/ is a Python package):

# Create and activate virtual environment (one-time)
python -m venv backend/venv
source backend/venv/bin/activate          # macOS / Linux
# backend\venv\Scripts\activate           # Windows

# Install dependencies
pip install -r backend/requirements.txt

# Create environment file (see .env.example)
cp .env.example .env

# Run the API
uvicorn backend.main:app --reload --port 8000

The backend will be available at http://localhost:8000 Interactive docs: http://localhost:8000/docs

πŸ’‘ On first run, the YOLOv8n model (~6 MB) is downloaded automatically, or place your own yolov8n.pt in backend/ to skip the download.


3. Frontend Setup

Open a new terminal:

cd frontend

# Install dependencies
npm install

# Create environment file (see .env.example)
cp .env.example .env

# Run the dev server
npm run dev

The frontend will be available at http://localhost:5173


4. Verify Everything Works

  1. Open http://localhost:5173
  2. Go to Detection View
  3. Upload any image from frontend/public/samples/
  4. You should see bounding boxes, count, avg confidence, and inference time
  5. Go to History View β€” your detection should appear in the table

backend/requirements.txt

fastapi
uvicorn[standard]
python-multipart
ultralytics
opencv-python
numpy
pillow
sqlalchemy
python-dotenv

πŸ”Œ API Reference

POST /detect

Upload an image and receive detection results.

Request: multipart/form-data with field file (JPEG or PNG)

Response:

{
  "count": 3,
  "detections": [
    { "x1": 120.4, "y1": 55.1, "x2": 210.7, "y2": 340.2, "confidence": 0.91 },
    { "x1": 300.2, "y1": 60.5, "x2": 390.9, "y2": 345.0, "confidence": 0.87 },
    { "x1": 480.1, "y1": 58.3, "x2": 570.4, "y2": 342.8, "confidence": 0.83 }
  ],
  "inference_time_ms": 412.5,
  "avg_confidence": 0.87,
  "annotated_image_b64": "/9j/4AAQ..."
}

The annotated image is a bare base64-encoded JPEG β€” the frontend prepends the data:image/jpeg;base64, prefix when rendering it in an <img> tag.

Errors:

  • 400 β€” unsupported file type or empty upload
  • 500 β€” model inference failure

GET /history

Retrieve past detections.

Query parameters:

Param Type Default Description
min_confidence float 0.0 Filter by minimum avg confidence
since ISO str null Only return detections after this
limit int 100 Max number of rows

Response:

[
  {
    "id": 12,
    "timestamp": "2025-01-15T14:32:11",
    "count": 3,
    "avg_confidence": 0.87,
    "inference_time_ms": 412.5
  }
]

POST /reset

Clears all stored detections.

Response: { "status": "cleared" }


πŸ§ͺ Testing Methodology

Test Set

  • 10+ images taken from frontend/public/samples/
  • Mix of: single person, small groups (2–5), larger crowds (6+)
  • Include some challenging cases (occlusion, partial visibility, low light)

Procedure

  1. For each test image, manually count the visible people and record the number as ground truth.
  2. Run the image through /detect and record:
    • Model's person count
    • Average confidence of valid detections
    • Inference time (ms)
  3. Compare counts and calculate accuracy.
  4. Note any false positives (boxes on non-persons).

Metrics & Formulas

Metric Formula Target
Detection Accuracy (Correct detections) Γ· (Total visible persons) Γ— 100% β‰₯ 85%
False Positive Rate (Non-person boxes) Γ· (Total boxes) Γ— 100% ≀ 10%
Average Inference Time Mean of inference_time_ms across all test images ≀ 1.5 s
Average Confidence Mean confidence of valid person detections β‰₯ 0.7
System Reliability (# images processed without crash) Γ· (# total images) Γ—100 100%

Hardware used: [e.g., Intel i7-1165G7 CPU, no GPU]


πŸ“Š Results

Per-Image Results

⚠️ Replace the numbers below with your actual test data.

# Image Visible Detected Correct False Pos. Avg Conf Time (ms)
1 frame1.jpg 1 1 1 0 0.92 380
2 frame2.jpg 3 3 3 0 0.88 410
3 frame3.jpg 5 5 5 0 0.85 425
4 frame4.jpg 2 2 2 0 0.90 395
5 frame5.jpg 7 6 6 0 0.81 440
6 frame6.jpg 4 4 4 1 0.79 415
7 frame7.jpg 1 1 1 0 0.94 370
8 frame8.jpg 8 8 8 0 0.83 460
9 frame9.jpg 3 3 3 0 0.86 405
10 frame10.jpg 6 6 6 0 0.84 435
11 frame11.jpg 2 2 2 0 0.91 390
12 frame12.jpg 5 5 5 1 0.80 430

Aggregate Summary

Metric Target Achieved
Detection Accuracy β‰₯ 85% 96.6%
False Positive Rate ≀ 10% 5.4%
Average Inference Time ≀ 1.5 s 0.41 s
Average Confidence β‰₯ 0.7 0.86
System Reliability 100% 100%

Observations

  • Worked well: Single and small-group images (1–4 people) β€” near-perfect accuracy with high confidence (> 0.85).
  • Challenges:
    • Dense crowds (>6 people): one missed detection in frame5.jpg β€” a partially occluded person behind others.
    • False positives: two cases where background objects (mannequin, tall backpack) were classified as persons with confidence ~0.55–0.60.
  • Performance: CPU-only inference averaged ~410 ms per 640Γ—640 image. Resizing larger inputs to 640Γ—640 reduced time without hurting accuracy.

πŸ“Έ Screenshots

1. Detection View β€” Single Person

Single person detection

High-confidence detection (0.94) with bounding box drawn.

2. Detection View β€” Group of People

Group detection

5 people detected with average confidence 0.85.

3. History View

History table

Detection history with timestamps, counts, and confidence filtering.

πŸ“Œ To add screenshots: create a docs/ folder in the repo root, place your PNGs there, and reference them above.


βœ… What Worked Well / πŸ”§ What Could Improve

What worked well

  • FastAPI + Ultralytics integration β€” clean, minimal code; the /docs page made testing easy without a frontend.
  • Base64 annotated images β€” simplified the frontend; no canvas math needed.
  • SQLite persistence β€” zero-config, perfect for a demo.
  • Separation of concerns β€” preprocessing, model, routes, and DB models are in separate files, making the code easy to test and extend.

What could improve

  • Real-time video streaming β€” currently image-only; adding frame-by-frame webcam support would make it a true monitoring tool.
  • Tracking across frames β€” YOLO detects per frame; adding ByteTrack or DeepSORT would give stable person IDs.
  • Region-based alerts β€” allow operators to draw a zone and alert when too many people are inside.
  • GPU acceleration β€” inference time would drop from ~400 ms to <30 ms on a CUDA-enabled device.
  • Model accuracy on crowds β€” fine-tuning on CrowdHuman would reduce missed detections in dense scenes.
  • Better error UI β€” currently relies on browser alerts for some errors.

🎁 Bonus Features

(Mark which ones you implemented.)

  • Real-time webcam feed with detection overlays
  • Region-based alerts (restricted zone count)
  • Heatmap / tracking lines
  • Statistics panel (average crowd size per hour)
  • CSV / Excel export of history
  • Confidence-threshold filtering in History View

πŸ‘₯ Team

  • Andrew Kihara β€” [#akihara]
  • Benjamin Koimett β€” [#bkoimett]

πŸ“š Resources


πŸ“„ License

This project was built for educational purposes as part of a bootcamp assignment.

About

Detecto is a real-time web app that counts people in images. It uses a YOLOv8 AI model with a FastAPI backend and a React frontend. Users can upload photos to see detected people in boxes, track processing speed, and view past uploads.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages