A full-stack application that detects and counts people in images using YOLOv8, served via a FastAPI backend with a React + Vite frontend.
- Project Purpose
- Architecture
- Features
- Tech Stack
- Repository Structure
- Setup & Run Instructions
- API Reference
- Testing Methodology
- Results
- Screenshots
- What Worked Well / What Could Improve
- Bonus Features
- Team
Detecto automates person detection and counting in images for safety and occupancy monitoring. It provides:
- Real-time visual feedback β bounding boxes with confidence scores drawn directly on uploaded images.
- Analytics β a history of every detection with timestamp, count, average confidence, and inference time.
- A clean API β so the detection engine can be reused by other apps, scripts, or dashboards.
The system is designed for operators who need fast, clear feedback without touching any ML code.
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β FRONTEND (React + Vite) β
β β
β ββββββββββββββββββββ βββββββββββββββββββββββββ β
β β DetectionView β β HistoryView β β
β β - upload image β β - list detections β β
β β - show boxes β β - filter by conf. β β
β β - show stats β β - reset history β β
β ββββββββββ¬ββββββββββ βββββββββββββ¬ββββββββββββ β
ββββββββββββββΌββββββββββββββββββββββββββββββββΌββββββββββββββββββ
β POST /detect β GET /history
β (multipart/form-data) β POST /reset
βΌ βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β BACKEND (FastAPI) β
β β
β ββββββββββββββ βββββββββββββββββ ββββββββββββββββββ β
β β detect.py ββββΊβ preprocessing ββββΊβ YOLOv8 model β β
β β β β .py β β (ultralytics) β β
β βββββββ¬βββββββ βββββββββββββββββ ββββββββββββββββββ β
β β save β
β βΌ β
β ββββββββββββββ ββββββββββββββββββββββββββββββββ β
β β SQLite βββββ history.py (read + reset) β β
β β detections β ββββββββββββββββββββββββββββββββ β
β β .db β β
β ββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Data flow for a detection request:
- User selects an image in the browser.
- Frontend POSTs the image to
/detectasmultipart/form-data. - Backend decodes the image, runs YOLOv8 inference, filters for the
personclass (COCO class 0), and draws bounding boxes. - Backend returns JSON with count, boxes, confidences, inference time, and a base64-encoded annotated image.
- Backend persists the detection (timestamp, count, avg confidence, inference time) to SQLite.
- Frontend renders the annotated image plus summary statistics.
- Upload JPEG or PNG images
- Draws bounding boxes on all detected people
- Displays:
- Total person count
- Per-person confidence scores
- Average confidence
- Inference time in milliseconds
- Handles invalid uploads with clear error messages
- Table of all past detections: timestamp, count, avg confidence, inference time
- Filter by minimum confidence threshold
- Reset entire detection history
- Sorted newest-first
POST /detectβ run inference on an uploaded imageGET /historyβ retrieve past detectionsPOST /resetβ clear all stored detections- Auto-generated API docs at
/docs(Swagger UI)
| Layer | Technology |
|---|---|
| Detection | YOLOv8n (Ultralytics), OpenCV, NumPy, Pillow |
| Backend | FastAPI, Uvicorn, SQLAlchemy, Pydantic |
| Database | SQLite |
| Frontend | React 19, Vite, Axios, React Router |
| Styling | Tailwind CSS |
| Config | python-dotenv (root .env), Vite env (.env) |
detecto/
βββ .env.example # MODEL_PATH, CONFIDENCE_THRESHOLD, DB_PATH
βββ backend/
β βββ main.py
β βββ config.py # env loading + path resolution
β βββ routes/
β β βββ detect.py
β β βββ history.py
β βββ models/
β β βββ record.py # SQLAlchemy model + engine
β β βββ schemas.py # Pydantic DTOs
β βββ utils/
β β βββ preprocessing.py
β βββ samples/
β β βββ frame1.jpg
β β βββ ...
β βββ requirements.txt
β βββ yolov8n.pt # model weights
β
βββ frontend/
β βββ public/
β β βββ samples/
β β βββ frame1.jpg
β β βββ frame2.jpg
β β βββ frame3.jpg
β βββ src/
β β βββ components/
β β β βββ Navbar.jsx
β β β βββ StatsCard.jsx
β β βββ pages/
β β β βββ DetectionView.jsx
β β β βββ HistoryView.jsx
β β βββ App.jsx
β β βββ main.jsx
β βββ .env.example # VITE_API_URL
β βββ package.json
β βββ vite.config.js
β
βββ detections.db # SQLite (created at runtime)
βββ README.md
βββ .gitignore
- Python 3.10+
- Node.js 18+ and npm
- (Optional) NVIDIA GPU with CUDA for faster inference
- ~500 MB free disk space for the YOLO weights and dependencies
git clone <your-repo-url> detecto
cd detectoRun from the repository root (backend/ is a Python package):
# Create and activate virtual environment (one-time)
python -m venv backend/venv
source backend/venv/bin/activate # macOS / Linux
# backend\venv\Scripts\activate # Windows
# Install dependencies
pip install -r backend/requirements.txt
# Create environment file (see .env.example)
cp .env.example .env
# Run the API
uvicorn backend.main:app --reload --port 8000The backend will be available at http://localhost:8000 Interactive docs: http://localhost:8000/docs
π‘ On first run, the YOLOv8n model (~6 MB) is downloaded automatically, or place your own
yolov8n.ptinbackend/to skip the download.
Open a new terminal:
cd frontend
# Install dependencies
npm install
# Create environment file (see .env.example)
cp .env.example .env
# Run the dev server
npm run devThe frontend will be available at http://localhost:5173
- Open http://localhost:5173
- Go to Detection View
- Upload any image from
frontend/public/samples/ - You should see bounding boxes, count, avg confidence, and inference time
- Go to History View β your detection should appear in the table
fastapi
uvicorn[standard]
python-multipart
ultralytics
opencv-python
numpy
pillow
sqlalchemy
python-dotenvUpload an image and receive detection results.
Request: multipart/form-data with field file (JPEG or PNG)
Response:
{
"count": 3,
"detections": [
{ "x1": 120.4, "y1": 55.1, "x2": 210.7, "y2": 340.2, "confidence": 0.91 },
{ "x1": 300.2, "y1": 60.5, "x2": 390.9, "y2": 345.0, "confidence": 0.87 },
{ "x1": 480.1, "y1": 58.3, "x2": 570.4, "y2": 342.8, "confidence": 0.83 }
],
"inference_time_ms": 412.5,
"avg_confidence": 0.87,
"annotated_image_b64": "/9j/4AAQ..."
}The annotated image is a bare base64-encoded JPEG β the frontend prepends
the data:image/jpeg;base64, prefix when rendering it in an <img> tag.
Errors:
400β unsupported file type or empty upload500β model inference failure
Retrieve past detections.
Query parameters:
| Param | Type | Default | Description |
|---|---|---|---|
min_confidence |
float | 0.0 | Filter by minimum avg confidence |
since |
ISO str | null | Only return detections after this |
limit |
int | 100 | Max number of rows |
Response:
[
{
"id": 12,
"timestamp": "2025-01-15T14:32:11",
"count": 3,
"avg_confidence": 0.87,
"inference_time_ms": 412.5
}
]Clears all stored detections.
Response: { "status": "cleared" }
- 10+ images taken from
frontend/public/samples/ - Mix of: single person, small groups (2β5), larger crowds (6+)
- Include some challenging cases (occlusion, partial visibility, low light)
- For each test image, manually count the visible people and record the number as ground truth.
- Run the image through
/detectand record:- Model's person count
- Average confidence of valid detections
- Inference time (ms)
- Compare counts and calculate accuracy.
- Note any false positives (boxes on non-persons).
| Metric | Formula | Target |
|---|---|---|
| Detection Accuracy | (Correct detections) Γ· (Total visible persons) Γ 100% | β₯ 85% |
| False Positive Rate | (Non-person boxes) Γ· (Total boxes) Γ 100% | β€ 10% |
| Average Inference Time | Mean of inference_time_ms across all test images |
β€ 1.5 s |
| Average Confidence | Mean confidence of valid person detections | β₯ 0.7 |
| System Reliability | (# images processed without crash) Γ· (# total images) Γ100 | 100% |
Hardware used: [e.g., Intel i7-1165G7 CPU, no GPU]
β οΈ Replace the numbers below with your actual test data.
| # | Image | Visible | Detected | Correct | False Pos. | Avg Conf | Time (ms) |
|---|---|---|---|---|---|---|---|
| 1 | frame1.jpg | 1 | 1 | 1 | 0 | 0.92 | 380 |
| 2 | frame2.jpg | 3 | 3 | 3 | 0 | 0.88 | 410 |
| 3 | frame3.jpg | 5 | 5 | 5 | 0 | 0.85 | 425 |
| 4 | frame4.jpg | 2 | 2 | 2 | 0 | 0.90 | 395 |
| 5 | frame5.jpg | 7 | 6 | 6 | 0 | 0.81 | 440 |
| 6 | frame6.jpg | 4 | 4 | 4 | 1 | 0.79 | 415 |
| 7 | frame7.jpg | 1 | 1 | 1 | 0 | 0.94 | 370 |
| 8 | frame8.jpg | 8 | 8 | 8 | 0 | 0.83 | 460 |
| 9 | frame9.jpg | 3 | 3 | 3 | 0 | 0.86 | 405 |
| 10 | frame10.jpg | 6 | 6 | 6 | 0 | 0.84 | 435 |
| 11 | frame11.jpg | 2 | 2 | 2 | 0 | 0.91 | 390 |
| 12 | frame12.jpg | 5 | 5 | 5 | 1 | 0.80 | 430 |
| Metric | Target | Achieved |
|---|---|---|
| Detection Accuracy | β₯ 85% | 96.6% |
| False Positive Rate | β€ 10% | 5.4% |
| Average Inference Time | β€ 1.5 s | 0.41 s |
| Average Confidence | β₯ 0.7 | 0.86 |
| System Reliability | 100% | 100% |
- Worked well: Single and small-group images (1β4 people) β near-perfect accuracy with high confidence (> 0.85).
- Challenges:
- Dense crowds (>6 people): one missed detection in
frame5.jpgβ a partially occluded person behind others. - False positives: two cases where background objects (mannequin, tall backpack) were classified as persons with confidence ~0.55β0.60.
- Dense crowds (>6 people): one missed detection in
- Performance: CPU-only inference averaged ~410 ms per 640Γ640 image. Resizing larger inputs to 640Γ640 reduced time without hurting accuracy.
High-confidence detection (0.94) with bounding box drawn.
5 people detected with average confidence 0.85.
Detection history with timestamps, counts, and confidence filtering.
π To add screenshots: create a
docs/folder in the repo root, place your PNGs there, and reference them above.
- FastAPI + Ultralytics integration β clean, minimal code; the
/docspage made testing easy without a frontend. - Base64 annotated images β simplified the frontend; no canvas math needed.
- SQLite persistence β zero-config, perfect for a demo.
- Separation of concerns β preprocessing, model, routes, and DB models are in separate files, making the code easy to test and extend.
- Real-time video streaming β currently image-only; adding frame-by-frame webcam support would make it a true monitoring tool.
- Tracking across frames β YOLO detects per frame; adding ByteTrack or DeepSORT would give stable person IDs.
- Region-based alerts β allow operators to draw a zone and alert when too many people are inside.
- GPU acceleration β inference time would drop from ~400 ms to <30 ms on a CUDA-enabled device.
- Model accuracy on crowds β fine-tuning on CrowdHuman would reduce missed detections in dense scenes.
- Better error UI β currently relies on browser alerts for some errors.
(Mark which ones you implemented.)
- Real-time webcam feed with detection overlays
- Region-based alerts (restricted zone count)
- Heatmap / tracking lines
- Statistics panel (average crowd size per hour)
- CSV / Excel export of history
- Confidence-threshold filtering in History View
- Andrew Kihara β [#akihara]
- Benjamin Koimett β [#bkoimett]
This project was built for educational purposes as part of a bootcamp assignment.



