Skip to content

About

Smart farming decisions, powered by AI & real data

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

31 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

🌾 FarmPulse

Smart farming decisions, powered by AI & real data.
An intelligent agronomy platform providing Indian farmers with real-time weather telemetry, live mandi commodity prices, soil health diagnostics, and AI-powered crop advisory.


🌐 Live Deployments


🏆 Hackathon Context

Built by Team BlackBox for the Into the Scrape-Verse Hackathon:

"You write a scraper, it works, and a week later the site changes its layout and everything breaks quietly. Build one that repairs itself instead, run it from your coding agent, and spend the week turning the data into something real."

FarmPulse leverages Bright Data Web Unlocker & Scraping APIs alongside Google Gemini AI to scrape dynamic, multi-source agricultural datasets (weather, mandi prices across 28 states, government schemes, and soil research) and synthesize them into real-time, actionable insights for farmers.


🚀 Key Features

  • 🌤️ Live Weather & Microclimate Telemetry: Real-time rainfall, temperature trends, and forecast alerts with ambient weather responsiveness.
  • 📊 Mandi Price Tracker: Live commodity pricing and market trends from 1,000+ mandis across 28 states.
  • 🌱 Soil Health & Agronomy Advisories: N-P-K nutrient diagnostics, soil moisture tracking, and actionable crop protection guidance.
  • 🏛️ Government Schemes & Subsidies: Curated agricultural schemes, loan waivers, and subsidies with deadlines and application links.
  • ✦ AI Farm Advisor (Gemini AI): Context-aware interactive assistant providing instant agronomic and market recommendations tailored to the farmer's exact crop and location.

🛠️ Tech Stack

Layer Technologies
Frontend React 19, Vite, Vanilla CSS Modules, Recharts
Backend FastAPI, Python 3.11, uv, Pydantic
Data & AI Bright Data Scraping APIs, Google Gemini AI (google-genai)
Deployment Vercel (Frontend), Render (Backend)

⚡ Quick Start

1. Clone the Repository

git clone https://github.com/Aayushi10/FarmPulse.git
cd FarmPulse

2. Backend Setup (FastAPI + uv)

cd backend

# Install dependencies with uv
pip install uv && uv sync

# Create .env file
cp .env.example .env  # or create .env and add your API keys

# Start the backend server
uv run uvicorn src.main:app --reload --port 8000

Backend runs at: http://localhost:8000


3. Frontend Setup (React + Vite)

cd ../frontend

# Install dependencies
pnpm install  # or npm install

# Start the dev server
pnpm dev      # or npm run dev

Frontend runs at: http://localhost:5173


🔑 Environment Variables

Backend (backend/.env)

GEMINI_API_KEY=your_gemini_api_key
BRIGHT_DATA_API_KEY=your_bright_data_api_key
BRIGHT_DATA_ZONE=your_bright_data_zone

Frontend (frontend/.env)

VITE_API_BASE_URL=http://localhost:8000

📝 Hackathon Submission Q&A

Question 1: What does your project do?

FarmPulse is an AI-powered agricultural intelligence platform that unifies fragmented web data into a single, personalized command center for farmers.

A farmer specifies their location, crop, and growth stage. FarmPulse automatically collects real-time weather forecasts, regional crop advisories, pest alerts, live mandi market prices, cultivation practices, and government schemes. Structured facts are validated and displayed directly, while Gemini AI synthesizes the data into prioritized, source-backed insights, identifies cross-source conflicts (e.g., rainfall vs. irrigation advice), and answers follow-up questions in natural language.

FarmPulse also features continuous data quality validation, leveraging Bright Data Scraper Studio's self-healing capabilities to automatically repair collectors and recover data when source websites change their layouts.


Question 2: How did you use Scraper Studio in your project?

Bright Data Scraper Studio serves as the backbone data acquisition layer of FarmPulse, powering 8 distinct collectors across agricultural portals, research institutes (IARI, Vikaspedia), market mandis, and government welfare sites.

Our backend manages these collectors with a trigger-and-cache architecture:

  • Data Normalization & Validation: Scraper outputs (JSON & NDJSON) are validated deterministically for data integrity (e.g., realistic temperature ranges, valid mandi price boundaries, non-empty advisory fields).
  • Self-Healing Resilience: If a target website updates its DOM and extraction degrades, our data quality engine detects the anomaly and triggers Bright Data's self-healing workflow to repair the collector and re-fetch clean data without manual code maintenance.
  • Controlled Demo Verification: For demonstration, we host a dedicated mock source (FarmDetailsWebsite). Intentionally breaking its HTML structure causes the collector to detect degradation and self-heal in real time, demonstrating continuous pipeline reliability.

Validated facts are fed into our UI and provided to Gemini as grounded context, ensuring all AI recommendations remain strictly factual and source-attributed.


Question 3: What was the most frustrating thing you hit while building with Scraper Studio or the CLI?

The most friction came from the workspace and permission model around Scraper Studio collectors. We initially assumed collector IDs could be shared across team members using individual API keys. However, collectors created under a specific workspace required matching account credentials to trigger.

Once we centralized ownership and aligned API credentials across the team, it was smooth sailing, but having explicit documentation or team-sharing features for hackathon collaboration would reduce initial setup friction.


Question 4: Where did you get stuck for the longest, and what got you unstuck?

The biggest challenge was moving from generic web scraping to identifying data sources with genuine agricultural utility. We evaluated dozens of portals (weather bulletins, district advisory PDFs, mandi price tables, and state scheme databases) that varied wildly in structure and update frequency.

We got unstuck by shifting our mindset from "What websites can we scrape?" to "What exact decisions does a farmer make on a given morning?" This clarity led us to select high-impact sources (IARI weather recaps, live mandi price feeds, and pest advisories) and design a unified schema normalization layer that standardizes different formats before feeding them to the AI advisor.


Question 5: How was the overall developer experience? What would you change, and anything else you want the Bright Data team to hear?

Overall, the developer experience was exceptional. Being able to define scrapers from natural language prompts and test them rapidly via the CLI drastically cut down the time spent writing custom CSS/XPath selectors. The proxy rotation and headless execution worked out of the box without IP blocks.

Feedback for the Bright Data Team:

  1. Team/Workspace Collaboration: Allow invite-based collector sharing or team API tokens so hackathon teams can collaborate without sharing primary credentials.
  2. Webhook / Completion Callbacks: An optional webhook callback when long-running collector jobs finish would simplify asynchronous architectures compared to polling.

Overall, abstracting away proxy management and scraper maintenance let us focus entirely on building a high-value product for farmers.


👥 Team

BlackBox

About

Smart farming decisions, powered by AI & real data

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages